Compare commits

..

31 Commits

Author SHA1 Message Date
Niklas Ye 7efd1bbba7 Set the chart's placeholder version to 0.37.0
CI / chart (push) Successful in 1s
CI / test (push) Successful in 7s
CI / security (push) Successful in 12s
Release / test (push) Successful in 6s
Release / chart (push) Successful in 1s
Release / binaries (push) Successful in 21s
Release / image (push) Successful in 1m3s
Release / scan-image (push) Successful in 1s
2026-10-03 16:17:20 +02:00
Niklas Ye 3bf94a5d7f Default to 2 replicas and RollingUpdate now that the singleton jobs are locked
CI / chart (push) Successful in 0s
CI / test (push) Successful in 7s
CI / security (push) Successful in 12s
replicas and strategy: Recreate were the chart's only guard against the
archiver, notifier and migration races; v0.36.0 closed all three with
advisory locks and a conflict-resolving incident insert, which made
that guard redundant rather than load-bearing. Expose replicaCount
(new, no values.yaml key existed before) defaulting to 2, and switch
to strategy: RollingUpdate with no explicit maxUnavailable/maxSurge --
the 25%/25% default rounds to 0/1 at replicaCount: 2, which is already
zero-downtime.

The chart does not gate this on image.tag, so pointing it at a
pre-v0.36.0 image with the new default is a foot-gun by omission --
noted in both the values.yaml comment and the deployment.yaml comment,
not guarded in code, same as the chart does for every other
version-coupled assumption today.

Co-authored-by: Claude <noreply@anthropic.com>
2026-10-03 12:30:27 +02:00
Niklas Ye 5c4e0bdd0e Set the chart's placeholder version to 0.36.0
CI / chart (push) Successful in 1s
CI / security (push) Successful in 16s
CI / test (push) Successful in 4m50s
Release / test (push) Successful in 7s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 48s
Release / image (push) Successful in 1m12s
Release / scan-image (push) Successful in 2s
2026-10-03 12:03:28 +02:00
Niklas Ye 9e5b085d8b Resolve the new-incident insert conflict instead of dropping the payload
CI / chart (push) Successful in 1s
CI / security (push) Successful in 20s
CI / test (push) Has been cancelled
openIncident's INSERT had no ON CONFLICT clause, relying entirely on
incidentForGroup's earlier SELECT to avoid a duplicate. On more than
one replica, two webhook deliveries for the very first occurrence of
a brand-new groupKey can both pass that SELECT before either INSERTs;
the loser then hit incidents_open_group_key_idx's unique violation,
which rolled back its whole transaction — including that payload's
alert upserts, done earlier in the same transaction. ingest's error is
only logged and receiveWebhook answers 200 regardless, so nothing
retried it: the loser's alerts silently never existed.

Add ON CONFLICT (team_id, group_key) WHERE resolved_at IS NULL DO
NOTHING to the INSERT, matching the partial unique index. Postgres
only resolves that conflict after the winning transaction commits (or
rolls back), so by the time RETURNING comes back empty,
existingOpenIncident's follow-up SELECT is guaranteed to see the
winner's row. The loser attaches to it instead of failing outright,
and the rest of its payload commits normally. Covers both callers,
since the dead man's switch sweeper shares this same function.

New test (package api_test, fires N webhook deliveries for one
groupKey from a synchronized start with distinct fingerprints, so
they aren't accidentally serialized by upsertAlerts' own per-
fingerprint lock) confirmed meaningful: with the ON CONFLICT clause
reverted, it fails 10/10 on a missing alert fingerprint; restored, 0/10.
Note while building it: "exactly one incident" alone cannot
distinguish fixed from broken, since the DB's own unique index already
guarantees that either way — the real signal is the loser's payload
surviving.

Chart comment updated: all three of the chart's original reasons for
Recreate are now addressed in code, though replicas stays at 1 and the
strategy stays Recreate pending a deliberate decision to raise it.

Co-authored-by: Claude <noreply@anthropic.com>
2026-10-03 11:58:39 +02:00
Niklas Ye 0050738ca0 Advisory-lock DB migrations against concurrent replica startup
Migrate's check-then-apply loop against schema_migrations had no
locking: two replicas booting at once against a fresh or
partially-migrated database could both pass the "not yet applied"
check for the same file and race applying it, crashing whichever lost
the duplicate-key insert (confirmed: reverting the lock fails the new
test 10/10 on a duplicate-key violation, racing as early as the
CREATE TABLE IF NOT EXISTS schema_migrations statement itself).

Hold a Postgres advisory lock for Migrate's whole run, on a dedicated
connection reserved via db.Conn so lock and unlock happen on the same
session. Blocking (pg_advisory_lock), unlike the archiver/notifier's
pg_try_advisory_lock: on boot there's no later tick to defer to, so a
second replica should wait for the first to finish migrating rather
than skip ahead.

Adds internal/db's first test file, exercising two concurrent Migrate
calls against a fresh schema.

Still open: the new-incident-insert race on a webhook for a brand-new
groupKey, noted in the chart's updated comment. Login rate limiting
staying in-process, diluted across replicas, is an accepted tradeoff.

Co-authored-by: Claude <noreply@anthropic.com>
2026-10-03 11:43:09 +02:00
Niklas Ye 42180948d1 Guard the archiver and notifier passes with a Postgres advisory lock
Both background loops run unconditionally on every instance with no
coordination between them, which the chart's replicas: 1 + strategy:
Recreate exists specifically to paper over: with more than one replica,
every one of them would sweep and deliver notifications independently,
and two overlapping during a rollout would both page for the same
incident.

Add withAdvisoryLock, which takes a Postgres advisory lock on a
dedicated connection and runs a pass only if it gets the lock,
otherwise skipping until the next tick. Wire StartArchiver and
StartNotifier through it with their own lock keys, so Sweep and
NotifySweep themselves are untouched and every existing test calling
them directly keeps working unchanged.

This also closes the notifier's double-delivery race in passing: two
replicas can no longer both be inside deliverPending at once, since
only one can hold notifierLockKey at a time.

Deliberately not addressed here, and still blocking a replica count
above 1: the in-memory login rate limiter, the unlocked migration
runner, and the new-incident-insert race on a webhook for a brand-new
groupKey. Noted in the updated chart comment.

Co-authored-by: Claude <noreply@anthropic.com>
2026-10-03 11:36:48 +02:00
Niklas Ye c83c7c2a8b Set the chart's placeholder version to 0.35.1
CI / chart (push) Successful in 2s
CI / security (push) Successful in 20s
CI / test (push) Successful in 4m45s
Release / test (push) Successful in 3s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 50s
Release / image (push) Successful in 1m21s
Release / scan-image (push) Successful in 3s
Cosmetic, as before (497086c, 4358e84): make helm-package passes
--version and --app-version from the tag, so this decides nothing
about what gets published. Kept in step so the tree heading for
v0.35.1 doesn't say 0.35.0 to a reader who hasn't yet seen the tag.

🤖 Generated with [Claude Code](https://claude.ai/code)

Co-Authored-By: Claude <noreply@anthropic.com>
2026-10-03 09:39:01 +02:00
niklas 43beda9a30 Merge pull request 'Fix the queue chip fade falling short of the right edge' (#30) from fix-chips-fade-edge into main
CI / chart (push) Successful in 1s
CI / security (push) Successful in 21s
CI / test (push) Has been cancelled
2026-10-03 07:36:12 +00:00
Niklas Ye 710521a73c Fix the queue chip fade falling short of the right edge
CI / chart (pull_request) Successful in 1s
CI / security (pull_request) Successful in 16s
CI / test (pull_request) Successful in 5m13s
position: sticky, as a flex item of the row it's pinning itself
against, interacted with that row's gap and its own negative margin in
a way that landed it short of the true edge -- visibly, a sliver of
the next chip stayed poking out past where the fade should have
covered it, which is the "ends before the screen edge" bug reported
against the release.

Replaced with the simpler, better-supported pattern for this: an
absolutely positioned overlay against a position:relative, overflow:
auto parent. Unlike a sticky descendant, an absolutely positioned one
is resolved against the parent's own (non-scrolling) box, so it stays
flush with the real edge regardless of scroll position, without the
flex-gap/margin interaction that caused this.

Verified visually: a standalone reproduction of both versions,
screenshotted with Chromium's headless_shell (no browser automation
tool available in this environment, but the binary's right there) --
the old version shows the next chip's edge peeking past the fade, the
new one doesn't.

🤖 Generated with [Claude Code](https://claude.ai/code)

Co-Authored-By: Claude <noreply@anthropic.com>
2026-10-03 09:33:33 +02:00
niklas d9492913ed Merge pull request 'Fix Stats/Admin/Account showing in the bottom bar too' (#29) from fix-nav-secondary-cascade into main
CI / chart (push) Successful in 0s
CI / test (push) Successful in 8s
CI / security (push) Successful in 14s
Reviewed-on: #29
2026-10-03 07:30:04 +00:00
Niklas Ye 1770e5d945 Fix Stats/Admin/Account showing in the bottom bar too
CI / chart (pull_request) Successful in 1s
CI / security (pull_request) Successful in 16s
CI / test (pull_request) Successful in 5m2s
.nav-link-secondary's display:none sat before .nav-link's own
display:flex in the file. Both are single-class selectors, so they tie
on specificity, and a tie is broken by which one comes later in the
file -- not by which class the element happens to carry. .nav-link's
declaration, being later, won for every element wearing both classes,
so Stats/Admin/Account rendered as three extra tabs on the phone bar
instead of folding into "More" as intended.

Moved the rule below .nav-link instead of changing either declaration,
since nothing about the values was wrong -- only their order was.

🤖 Generated with [Claude Code](https://claude.ai/code)

Co-Authored-By: Claude <noreply@anthropic.com>
2026-10-03 09:22:48 +02:00
Niklas Ye 497086cb51 Set the chart's placeholder version to 0.35.0
CI / test (push) Successful in 4s
CI / chart (push) Successful in 0s
CI / security (push) Successful in 12s
Release / test (push) Successful in 4s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 51s
Release / image (push) Successful in 1m5s
Release / scan-image (push) Successful in 23s
Cosmetic: make helm-package passes --version and --app-version from the
tag, so these two fields decide nothing about what gets published. Kept
in step anyway, the same as 4358e84 and fd26fef before it, so the tree
heading for v0.35.0 does not say 0.34.0 to a reader who hasn't yet seen
the tag.

🤖 Generated with [Claude Code](https://claude.ai/code)

Co-Authored-By: Claude <noreply@anthropic.com>
2026-10-03 09:10:30 +02:00
niklas e3090d2779 Merge pull request 'Web: visual design pass (issue #26)' (#28) from design-polish-issue-26 into main
CI / test (push) Successful in 4s
CI / chart (push) Successful in 1s
CI / security (push) Successful in 13s
Reviewed-on: #28
2026-10-03 07:08:19 +00:00
Niklas Ye 91f03c21e8 Web: visual design pass (issue #26)
CI / chart (pull_request) Successful in 2s
CI / security (pull_request) Successful in 18s
CI / test (pull_request) Successful in 4m49s
Addresses the screenshot-review feedback in #26. No framework or build
step added — all of this stays within the existing plain HTML/CSS/
vanilla-JS + go:embed architecture.

- Nav: re-enable the bottom tab bar that was already built and
  switched off (Queue/On-call/Alerts/Team + a "More" sheet for
  Stats/Admin/Account), replacing the hamburger on phone width.
- Queue: chip counts, a scroll fade on the filter row, a
  "Triggered Xh ago" + severity label per row, a chevron on the team
  switcher so it reads as a dropdown.
- On-call: collapse repeated same-person days into shift bars (week
  view and "your shifts" both), show the week as a date range with
  the ISO week number as secondary text, split "Current shift" out
  from "Next shifts" with "ends in Nd", a pill badge + row highlight
  for "you".
- Incident detail: fix the actual bug behind the duplicate
  "acknowledged" timeline entries (acknowledgeIncident's UPDATE had no
  guard on the incident's current status, so acknowledging an
  already-acknowledged incident silently re-logged the event — now
  idempotent, with regression tests on both the authenticated route
  and the ntfy ack-button route). Relabel escalation re-pages so they
  don't look like the same page landing twice. Copy the primary action
  up near the top. Label the "···" button. Group the timeline by
  phase (triggered/acknowledged/resolved). Add an "at a glance"
  summary row (duration/severity/responsible) and collapse the group
  labels by default.
- Team overview: reword the vague copy ("One owner." etc.) into plain
  labels.
- Empty states: fill in missing icons/one-liners across queue,
  alerts, stats and the incident timeline.
- CSS: fix card padding bugs, verify link contrast already passes AA,
  introduce a --fs-* type-scale token set and migrate the few
  genuinely isolated cases onto it (left sizes tied to a fixed shape,
  a deliberately prominent display, or a non-negotiable constraint
  like the iOS-zoom-prevention input size as documented exceptions
  rather than guess at a render this change can't see).

Verified with the full fmt/lint/test/helm-lint gate, plus a live
instance against the test DB with seeded incidents and schedule data
to trace the on-call grouping and timeline phase-splitting logic
against real API responses.

🤖 Generated with [Claude Code](https://claude.ai/code)

Co-Authored-By: Claude <noreply@anthropic.com>
2026-10-03 09:01:54 +02:00
Niklas Ye 4358e84b24 Set the chart's placeholder version to 0.34.0
CI / test (push) Successful in 4s
CI / chart (push) Successful in 1s
CI / security (push) Successful in 12s
Release / test (push) Successful in 4s
Release / chart (push) Successful in 7s
Release / image (push) Successful in 2m27s
Release / binaries (push) Successful in 2m35s
Release / scan-image (push) Successful in 3s
2026-10-02 22:08:35 +02:00
niklas f45dc2f925 Merge pull request 'internal/api: unify human/service-account authz into one Caller type' (#24) from unify-caller-authz into main
CI / chart (push) Successful in 1s
CI / security (push) Successful in 22s
CI / test (push) Successful in 27s
2026-10-02 20:06:37 +00:00
Niklas Ye 774fdfcaa8 internal/api: unify human/service-account authz into one Caller type
CI / chart (pull_request) Successful in 1s
CI / security (pull_request) Successful in 15s
CI / test (pull_request) Successful in 5m21s
ctxUser/ctxTeams (human) and ctxServiceAccount (+ a synthetic ctxTeams
entry, service account) used to be two parallel, un-unified context
representations -- every authz predicate had to remember which one(s) it
needed, and every place that forgot either wrongly 403'd a service account
(terdut-server#23, terdut-operator#3), crashed on an unchecked zero-value
user id, or silently no-op'd. New internal/api/caller.go collapses both
into one Caller, stored under one ctxCaller key by serveAs/serveAsServiceAccount;
every existing predicate (userFromContext, callerTeamIDs, callerRole,
callerIsAdmin, isInstanceServiceAccount, AdminOnly, requireSelfOrAdmin,
requireTeamOwner, OperatorModeBlock) now reads through it, with identical
behavior for every untouched call site (alerts.go, incidents.go,
schedule.go, stats.go, etc.) -- confirmed by the full existing suite
passing unchanged.

Four real fixes land alongside the refactor, not just the restructuring:

1. callerMayManageServiceAccount gains the one load-bearing branch this
   exists for: an instance-scoped service account may now manage (mint or
   revoke a key on) any team-scoped account, not only a human admin, that
   team's human owner, or the account itself. handleCreateServiceAccount
   already let an instance-scoped caller *create* a team-scoped account for
   any team; adopting or rotating one it didn't just create in the same
   call -- terdut-operator's own documented crash-window recovery -- had no
   equivalent permission and 403'd forever. Closes terdut-operator#3.

2. handleCreateInvite wrote a service-account caller's zero-value user id
   straight into invites.created_by (nullable, but never passed as nil),
   which foreign-key-violates against users(id) -- a 500, not success, for
   any team-scoped service account minting an invite. Fixed the same way
   handleCreateServiceAccount already handles the analogous case. Found
   live while verifying this change, not filed separately since it's fixed
   in the same place it was found.

3. handleMe and handleTestNotification 500'd for a service-account caller
   (fetchUser/the ntfy_topic lookup against a zero-value user id that
   matches no row); handleDismissOnboarding silently no-op'd (UPDATE ...
   WHERE id = 0). All three now call Caller.AsHuman() and return an
   explicit 403 ("this endpoint is for human accounts only").

4. Ratifies, rather than further narrows, two capabilities a team-scoped
   service account already had by construction and this document's own
   text once called "a gap acknowledged rather than closed": owner-equivalent
   reach over membership/invites, and minting another service account for
   its own team. terdut-operator's new TerdutTeam invite-minting feature is
   about to depend on the first one, so this makes it documented, tested,
   intentional behavior instead of an accident nobody was supposed to rely
   on.

AdminOnly/requireSelfOrAdmin are unchanged in effect: still human-only,
forever, for every scope of service account -- confirmed by
TestAdminOnly_RefusesEveryServiceAccountScope. terdut-server#23's named
routes (POST /api/users, PUT /api/admin/settings) were never the right
thing to widen; its real fix is the terdut-operator invite feature,
recorded in SERVICE-ACCOUNTS.md's "What this unblocks" and closing that
issue once it ships.

SERVICE-ACCOUNTS.md amended in place (not a new file, its own established
convention) to describe the as-built Caller model, correct its own
aspirational claim about AdminOnly that TEAM-LOOKUP.md had already flagged
as not matching shipped code, and record all of the above.
2026-10-02 21:52:43 +02:00
Niklas Ye fd26fef1ba Set the chart's placeholder version to 0.33.2
CI / test (push) Successful in 4s
CI / chart (push) Successful in 1s
CI / security (push) Successful in 12s
Release / test (push) Successful in 3s
Release / chart (push) Successful in 3s
Release / binaries (push) Successful in 16s
Release / image (push) Successful in 1m8s
Release / scan-image (push) Successful in 1s
make helm-package passes --version and --app-version from the tag, so
these fields decide nothing about what is published -- but a tree heading
for v0.33.2 that still says 0.33.1 tells its reader something false.
Same as 4b15079 and 6f8499f before it.
2026-10-02 09:42:50 +02:00
Niklas Ye a9d788cc83 Wait for Postgres to accept connections before the main container starts
A Deployment created before Postgres has finished its very first boot --
initdb plus Patroni leader election, on a from-scratch postgres-operator
cluster -- crash-looped a few times. db.Open()'s own ping-retry budget
(pingAttempts/pingRetryDelay, internal/db/db.go) is sized for a much
shorter, different race -- NetworkPolicy propagation, a few seconds -- not
for genuine first-time cluster creation, which routinely takes longer, so
it exhausted and the process exited before ever binding its HTTP port. A
startupProbe cannot fix that: the crash happens before there is anything
to probe.

Added a wait-for-postgres init container instead: it loops pg_isready
against database.dsn until Postgres actually answers, before the main
container's own, unchanged retry budget gets a chance to run out.
pg_isready needs no credentials -- it reports PQPING_OK on anything that
amounts to a Postgres backend answering, including an auth challenge --
so no PGPASSWORD is wired into it.

Chart-only; no Go code changed. database.waitForPostgres.enabled defaults
to true and can be turned off if something else already guarantees
Postgres is reachable before this Deployment is created.
2026-10-02 09:42:42 +02:00
Niklas Ye 4b15079ac2 Set the chart's placeholder version to 0.33.1
CI / chart (push) Successful in 1s
CI / security (push) Successful in 19s
CI / test (push) Successful in 4m42s
Release / test (push) Successful in 3s
Release / chart (push) Successful in 2s
Release / image (push) Successful in 2m11s
Release / scan-image (push) Successful in 3s
Release / binaries (push) Successful in 2m47s
Cosmetic: `make helm-package` passes --version/--app-version from the
tag, so this field decides nothing about what gets published. Still done
so the tree doesn't say 0.33.0 while heading for a v0.33.1 release.
Cites fc9f47c, the fix this version actually is.
2026-10-02 09:18:28 +02:00
Niklas Ye fc9f47cc8d Refuse an OIDC sign-in from creating the very first user
Closes the race terdut-operator#1 found: /api/bootstrap and OIDC
auto-provisioning both key off the same signal (SELECT COUNT(*) FROM
users) with no coordination between them, so an otherwise-ordinary OIDC
sign-in against a freshly-created, not-yet-bootstrapped install could
create user #1 itself and take the one slot /api/bootstrap expects to win
uncontested (terdut-operator's own design, DESIGN.md §1/§6, assumes it is
the only caller). The operator has no way to recover from losing that
race -- it never gets a credential, and nothing it owns can clear the
occupying user row.

resolveSSOUser now checks the same gate handleBootstrap already does,
right where it's about to create a brand-new user (an identity nobody has
linked yet, that also matches no existing local account by email) -- not
anywhere else, since every other sign-in on an already-bootstrapped
install is unaffected. New sso_error code `not_bootstrapped`: the person
sees "this install is still setting up, try again in a moment" and a
second attempt once something has actually bootstrapped succeeds
normally, same as any other first sign-in.

Does not fix the other half of that issue (BootstrapStateLost's own
"delete and recreate" instructions still don't work once something has
occupied the slot some other way) -- this closes the specific race, not
every path to that state.
2026-10-02 09:12:32 +02:00
Niklas Ye bc9f793f1f Add GET /api/teams?name= (TEAM-LOOKUP.md)
CI / chart (push) Successful in 1s
CI / security (push) Successful in 1m15s
CI / test (push) Successful in 5m48s
Resolves the gap TEAM-LOOKUP.md raised: an instance-scoped service
account had no way to recover a team's id after a 409 on POST
/api/teams, unlike the already-solved equivalent for service accounts
themselves (GET /api/service-accounts?name=).

Same route, extended the same way handleListServiceAccounts already
branches on ?name=: unset behaves exactly as before (the caller's own
teams via team_members); set looks up one team by exact name, open to
any authenticated caller -- not gated by isInstanceServiceAccount or
AdminOnly, since what it discloses (a name is taken, nothing about who's
in it) is the same low sensitivity that lookup already accepts for
service-account names.

Tests cover the exact motivating scenario (create, 409 on a retry,
recover the id via ?name=), the empty-array-not-an-error case, that no
role/source is reported for a non-member match, and that a caller who
isn't a member of the matched team still gets it.
2026-10-01 10:49:11 +02:00
Niklas Ye 871274a3a0 Request: team lookup for service accounts (TEAM-LOOKUP.md)
Raised by terdut-operator's TerdutTeam controller (ROADMAP.md Stage 2):
an instance-scoped service account has no way to recover a team's id
after a 409 on POST /api/teams, unlike the equivalent, already-solved
case for service accounts themselves (GET /api/service-accounts?name=).
Also corrects a claim in SERVICE-ACCOUNTS.md's own text that doesn't
match AdminOnly's actual code -- confirmed against source, not assumed.
2026-10-01 10:43:09 +02:00
Niklas Ye 6f8499fa42 Set the chart's placeholder version to 0.33.0
CI / chart (push) Successful in 1s
CI / security (push) Successful in 18s
CI / test (push) Successful in 4m26s
Release / test (push) Successful in 9s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 29s
Release / image (push) Successful in 1m10s
Release / scan-image (push) Successful in 31s
Cosmetic: make helm-package passes --version and --app-version from the
tag, so these two fields decide nothing about what gets published. Kept
in step anyway, the same as 0ee576f (0.32.0) and 949d659 (0.31.0) before
it, so a tree heading for v0.33.0 does not read as still being on 0.32.0.
2026-09-29 21:33:15 +02:00
Niklas Ye ef731e85c5 Document service accounts and operator mode in the README
The last commit (a4dd60f) shipped the code with no README update, against
this repo's own convention of documenting the whole API surface there.
Adds the Service accounts section and table, the Authentication and Teams
prose covering the new principal and TERDUT_OPERATOR_MODE, and the
/api/version row.

Corrects one thing along the way: a first draft claimed team-scoped
accounts are refused on team membership/invite endpoints. Checked against
teams.go and that is false — nothing server-side carves those two out,
only convention (no sane operator would call them) keeps them out of
automation's hands. README and SERVICE-ACCOUNTS.md now say that plainly
instead of the stronger, incorrect claim.
2026-09-29 21:33:04 +02:00
Niklas Ye a4dd60f6b8 Add service accounts and operator mode
Service accounts (SERVICE-ACCOUNTS.md) are a scoped, non-human credential:
not a users row, so they never touch OIDC sync, login or the is_admin flag.
Instance scope can create a team and mint a team-scoped account for it;
team scope is owner-equivalent for that one team and nothing else. This is
what unblocks terdut-operator's DESIGN.md §6 — no more impersonating a
human admin, and a real rotation story instead of the unworkable
delete-and-re-bootstrap /api/bootstrap can't actually do.

- migration 014: service_accounts + service_account_keys
- POST /api/service-accounts, POST/DELETE .../keys, GET ?name= self-lookup
- AuthMiddleware resolves a tdsa_-prefixed key to a distinct principal;
  a team-scoped account gets a synthetic single membership so
  requireTeamMember/requireTeamOwner work on it unmodified
- handleCreateTeam accepts an instance-scoped caller; the team it creates
  has no human owner, which is the expected shape for one an operator is
  about to hand a team-scoped credential to

Operator mode (TERDUT_OPERATOR_MODE / values.operatorMode) declares an
install gitops-managed: session and user-API-key writes to teams,
escalation policies, dead man's switches and integrations get 403
reason=operator_managed, while a service account's writes still go
through. Team membership/invites and the schedule are deliberately left
out — never gitops-managed by design, and still human day-to-day work.
/api/auth/config reports operator_mode so the web UI can grey these
sections out from the start rather than only after a write fails.

Also: GET /api/version (both terdut-tui and terdut-operator currently
detect server capability by route-probing; this gives them a real answer),
and a PUT for dead man's switches so a reconciler can update one in place
instead of deleting and recreating it.
2026-09-29 21:25:17 +02:00
Niklas Ye b5573fbca2 Add design note: scoped service-account/token type
Proposes a non-human credential type — service_accounts +
service_account_keys, instance- or team-scoped, distinct from both
user API keys (always tied to a human's full rights) and integration
keys (narrow, one-way webhook auth only). Directly unblocks
terdut-operator's DESIGN.md §6, whose bootstrap/rotation plan doesn't
work against /api/bootstrap's actual single-shot-per-install behavior.

Design note only, no implementation yet.
2026-09-29 20:35:28 +02:00
Niklas Ye 0ee576f793 Set the chart's placeholder version to 0.32.0
CI / chart (push) Successful in 1s
CI / security (push) Successful in 16s
CI / test (push) Successful in 4m2s
Release / test (push) Successful in 4s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 31s
Release / image (push) Successful in 1m10s
Release / scan-image (push) Successful in 3s
Cosmetic: make helm-package sets the published version and appVersion
from the tag, so these two fields decide nothing (see the comment
above them). Kept in step anyway, same as 949d659 and e5b4df7, so a
tree heading for v0.32.0 doesn't say 0.31.0.
2026-09-27 22:05:26 +02:00
Niklas Ye b610b1817a Fold the account page's ntfy and password forms behind disclosures
Both sat open by default, competing with the rest of the page for
attention on every visit even though most visits need neither. Team's
rota already has the same problem for its bulk-assign form and solves
it with a native <details>/<summary> disclosure, styled generically in
app.css; this reuses that idiom rather than inventing a JS toggle.

Each section now shows a one-line status (the topic, or whether a
password is set) with the actual form folded under a summary naming
the action ("Set a topic" / "Change topic", "Set a password" /
"Change password"). A successful save closes the fold and confirms
with a toast, since the point of folding is that a saved form goes
back to being just a status line; a validation or API error keeps the
fold open and shows inline, next to the field it's about.

The password section's heading no longer says "Change password" or
"Set a password" itself, since that verb now lives on the summary; it
just says "Password", matching the existing SSO-off case.
2026-09-27 22:04:59 +02:00
Niklas Ye 949d6595ba Set the chart's placeholder version to 0.31.0
CI / chart (push) Successful in 1s
CI / security (push) Successful in 17s
CI / test (push) Successful in 3m48s
Release / test (push) Successful in 5s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 41s
Release / image (push) Successful in 1m3s
Release / scan-image (push) Successful in 5s
Cosmetic: make helm-package sets the published version and appVersion
from the tag, so these two fields decide nothing (see the comment
above them). Kept in step anyway, same as e5b4df7 and 97a4814, so a
tree heading for v0.31.0 doesn't say 0.30.0.
2026-09-27 18:23:43 +02:00
Niklas Ye 33356ca978 Add a global, colour-coded team selector to the nav
state.js's currentTeam() was hard-coded to teams[0] and never really meant
"the team currently selected" — team.js's settings page and queue.js's
filter chips each kept their own separate, unsynchronized notion of "which
team" instead, so picking one on one page had no effect on the other.

Replaces both with a single state.selectedTeamID, set only through the new
setSelectedTeam (persisted in localStorage, unlike the queue's old per-tab
sessionStorage filter) and broadcast to listeners via onTeamChange. A new
teamselector.js control — a coloured dot plus the team's name, or "All
teams" — sits at the top of both the desktop sidebar and the mobile topbar,
opening the existing bottom-sheet menu to switch. Shown only once someone is
in more than one team, matching every other team-aware control in this app.

Colours come from a new teamColorClass() in format.js, hashing a team's id
into the six-colour rc1..rc6 palette already used for the rota's per-person
chips, so no schema or API change is needed. The queue's team filter chips
pick up the same colours.
2026-09-27 18:14:54 +02:00
51 changed files with 3292 additions and 293 deletions
+70 -10
View File
@@ -334,6 +334,7 @@ over an administrator's edit.
| `TERDUT_PUBLIC_URL` | — | Base URL a phone uses to reach this server: the notification's link into the web UI, its Acknowledge button, and whether the session cookie is `Secure` |
| `TERDUT_NOTIFY_REPEAT` | `15m` | **seed.** How long an incident may sit unacknowledged before it is paged again. `0` notifies once and never repeats |
| `TERDUT_PASSWORD_LOGIN` | `true` | `false` refuses password login and password sign-up (`403`), leaving single sign-on the only way in. Refused at startup unless SSO is configured |
| `TERDUT_OPERATOR_MODE` | `false` | Declares this install gitops-managed: a session's or a user's own API key's writes to teams, escalation policies, dead man's switches and integrations are refused (`403 reason:"operator_managed"`); a [service account](#service-accounts)'s are not. Team membership and the schedule stay editable regardless |
| `TERDUT_OIDC_ISSUER` | — | Turns single sign-on on. The provider's issuer URL; discovery is read from `<issuer>/.well-known/openid-configuration`. See [Single sign-on](#single-sign-on-oidc) |
| `TERDUT_OIDC_CLIENT_ID` / `TERDUT_OIDC_CLIENT_SECRET` | — | **Required with an issuer.** The confidential client registered at the provider. Keep the secret in a Secret, not in values |
| `TERDUT_OIDC_NAME` | `SSO` | What the sign-in button calls the provider |
@@ -350,7 +351,7 @@ Note that `TERDUT_STALE_AFTER` and `TERDUT_DEADMAN_TIMEOUT` point in opposite di
is a generous grace period around a `repeat_interval` you do not control; a dead man's switch is a
deadline you set deliberately, and the heartbeat's route is configured to beat faster than it.
In the Helm chart the two sweeper durations are set via `sweeper.staleAfter` and `sweeper.archiveAfter`, dead man's switches via the `deadman.*` values, notifications via the `notify.*` values, and single sign-on via `oidc.*` and `passwordLogin`.
In the Helm chart the two sweeper durations are set via `sweeper.staleAfter` and `sweeper.archiveAfter`, dead man's switches via the `deadman.*` values, notifications via the `notify.*` values, single sign-on via `oidc.*` and `passwordLogin`, and operator mode via `operatorMode`.
---
@@ -705,8 +706,8 @@ of the last heartbeat, and the heartbeat's labels are on the incident's
All endpoints except `/api/bootstrap`, `/api/integrations/{key}/alertmanager`,
`/api/notify/ack/{token}`, `/api/login`, `/api/logout`, `/api/auth/config`,
`/api/oidc/login`, `/api/oidc/callback`, `/api/oidc/device` and `/api/oidc/device/token`
require either an API key:
`/api/version`, `/api/oidc/login`, `/api/oidc/callback`, `/api/oidc/device` and
`/api/oidc/device/token` require either an API key:
```
Authorization: Bearer <api-key>
@@ -721,6 +722,13 @@ granting the flag itself. Everybody else works incidents — acknowledging,
assigning, snoozing, resolving, noting — and manages their own account and
nobody else's. An API key carries exactly the rights of the user it belongs to.
A third principal, the **service account**, exists for automation (a
Kubernetes operator, most likely) that needs to manage teams, escalation
policies, dead man's switches and integrations without impersonating a human.
It is not a user — it never signs in, never appears in a team's member list,
and never holds the administrator flag — and its key is prefixed `tdsa_` so it
reads as one at a glance in a log line. See [Service accounts](#service-accounts).
**Getting an account.** The first one comes from `/api/bootstrap`. After that
it depends on `signup_mode`, an administrator setting:
@@ -764,16 +772,28 @@ the shape of a team, not about reading other people's incidents.
Anything belonging to a team you are not in answers `404`, not `403`: whether an
incident exists is itself something only its team should learn.
**Operator mode** (`TERDUT_OPERATOR_MODE`, see [Configuration](#configuration))
declares this install gitops-managed. When it is on, a session or a user's own
API key gets `403 {"error": "...", "reason": "operator_managed"}` on every
write this README marks **owner**-gated under Teams below (creating, renaming
or deleting a team; its OIDC group binding; its escalation ladder; its dead
man's switches; its integrations) — a service account's writes are unaffected.
Team membership and invites are deliberately excluded: they are never
gitops-managed, in operator mode or out of it. `GET /api/auth/config` reports
`operator_mode` so a client can grey those sections out before a write is ever
attempted.
| Method | Path | Description |
|---|---|---|
| `GET` | `/api/auth/config` | How to sign in: `{"password_login", "oidc": {"enabled","name"}, "device_login"}`. No session needed |
| `GET` | `/api/auth/config` | How to sign in: `{"password_login", "oidc": {"enabled","name"}, "device_login", "operator_mode"}`. No session needed |
| `GET` | `/api/version` | `{"version"}` — this build's version string. No session needed, the same as `/healthz` |
| `POST` | `/api/login` | `{"username","password"}` → sets the session cookie, returns `{user, has_password}`. `429` after too many failures; `403` when `TERDUT_PASSWORD_LOGIN=false` |
| `GET` | `/api/oidc/login` | Starts a single sign-on sign-in: redirects the browser to the provider. `?next=/path` is where to land afterwards; only a path on this server is honoured. Only exists when SSO is configured |
| `POST` | `/api/oidc/device` | Starts a device login: returns `{device_code, user_code, verification_url, interval, expires_in}`. Only exists when SSO is configured |
| `POST` | `/api/oidc/device/token` | `{"device_code"}` → `202 {"status":"pending"}`, then `200` with the session cookie once approved (once only). `410` with `{"error":"expired"}` or `{"error":"denied"}`; `429 {"error":"slow_down"}` if polled faster than `interval` |
| `POST` | `/api/oidc/device/approve` | **session** — `{"user_code"}`. Approves a pending device login as the caller. `403` for an API key; `404` for an unknown, expired or already decided code |
| `POST` | `/api/oidc/device/deny` | **session** — `{"user_code"}`. Refuses it |
| `GET` | `/api/oidc/callback` | Where the provider sends the browser back. Sets the session cookie and redirects to `/`, or to `/?sso_error=<code>` — one of `denied`, `expired`, `failed`, `unavailable`, `not_allowed`, `no_email`, `email_conflict`, `disabled` |
| `GET` | `/api/oidc/callback` | Where the provider sends the browser back. Sets the session cookie and redirects to `/`, or to `/?sso_error=<code>` — one of `denied`, `expired`, `failed`, `unavailable`, `not_allowed`, `no_email`, `email_conflict`, `disabled`, `not_bootstrapped` (no user exists on this install yet — sign in again once something has called `/api/bootstrap`) |
| `POST` | `/api/logout` | Ends the session and clears the cookie |
| `GET` | `/api/me` | The caller: `{user, has_password}` |
@@ -810,6 +830,41 @@ on anybody's.
| `GET` | `/api/admin/settings` | **admin** | The editable settings with their bounds, plus the environment-configured ones, read-only. Never credentials |
| `PUT` | `/api/admin/settings` | **admin** | Change one or more `{"key": seconds}`, or `{"signup_mode": "open"\|"invite_only"}`. `400` for an unknown key or a value outside its bounds |
### Service accounts
A service account is a scoped, non-human credential for automation — not a
`users` row, so it never signs in, is never a team member, and never carries
the administrator flag. Two scopes:
- **instance** — the same reach system administration has over teams: create
one, and mint a **team**-scoped account against any of them. There is no
cap on how many instance-scoped accounts exist, but ordinarily there is one,
belonging to whatever is provisioning this install end to end.
- **team** — owner-equivalent for that one team, and nothing else: every
**owner**-gated endpoint under [Teams](#teams), membership and invites
included. Nothing narrower is enforced server-side; what actually keeps
membership out of automation's hands is that no operator built against this
scope should ever call those two endpoints — see
[operator mode](#authentication) and `SERVICE-ACCOUNTS.md`'s note on this.
A key is shown once, at creation or rotation, and only its hash is stored —
the same handling as a user's API key. Losing it means minting a new one;
there is no way to recover a raw key from the server.
| Method | Path | Who | Description |
|---|---|---|---|
| `GET` | `/api/service-accounts` | **admin** | Every service account. Pass `?name=` instead to look one up by its exact name — open to **any** authenticated caller (human or service account), since it returns no key material and is how an account finds its own id |
| `POST` | `/api/service-accounts` | owner\* | Create one and mint its first key `{"name","scope","team_id"?}` (`team_id` required for `scope:"team"`, absent for `scope:"instance"`). Returns `{"service_account", "key"}` — `key.key` shown once |
| `POST` | `/api/service-accounts/{id}/keys` | owner\* | Mint an additional key `{"name"}` — rotation without recreating the account. Shown once |
| `DELETE` | `/api/service-accounts/{id}/keys/{keyID}` | owner\* | Revoke one key |
\* For an **instance**-scoped account: a system administrator only. For a
**team**-scoped account: a system administrator, that team's own human owner,
an instance-scoped service account (minting a narrower credential for a team
it just created), or — for the two key endpoints only — the account rotating
or revoking its own key, which is not a privilege escalation, the same
reasoning a user's own API keys rest on.
### Alert ingestion
Alerts arrive on a team's integration key. The key is both the credential and the
@@ -827,15 +882,19 @@ and was removed in v0.13.0 once senders had moved onto keys.
### Teams
**owner** below means an owner of that team *or* a system administrator, who
passes every one of these without being a member — see
[Authentication](#authentication). **member** means membership and nothing else: an
administrator who is not in the team gets the same `404` as anybody else.
**owner** below means an owner of that team, a system administrator (who
passes every one of these without being a member), or that team's own
team-scoped [service account](#service-accounts) — including membership and
invites, technically, though no automation this scope was designed for
(a Kubernetes operator's CRDs, see `SERVICE-ACCOUNTS.md`) ever models team
membership or would call those two. See [Authentication](#authentication).
**member** means membership and nothing else: an administrator who is not in
the team gets the same `404` as anybody else.
| Method | Path | Who | Description |
|---|---|---|---|
| `GET` | `/api/teams` | any | The caller's own teams, each with their role |
| `POST` | `/api/teams` | any | Create a team `{"name"}`; the creator becomes its first owner |
| `POST` | `/api/teams` | any | Create a team `{"name"}`; a human creator becomes its first owner. An instance-scoped [service account](#service-accounts) may also create one, and it gets no owner at all — expected for a team an operator is about to hand a team-scoped credential to, not an orphaned team a human made |
| `PUT` | `/api/teams/{teamID}` | **owner** | Rename it `{"name"}`. `409` if the name is taken |
| `DELETE` | `/api/teams/{teamID}` | **owner** | Delete a team and everything under it. `409` while it has open incidents |
| `GET` | `/api/teams/{teamID}/members` | member | Who is in the team, with `status` (`oncall` if the rota has them today, `unpageable` when a page to them would go nowhere — even if they are on call — else `reachable`), `on_call`, `next_shift` (first rota day after today), `pageable` and `problem` (`has no ntfy topic` / `account is disabled`; never the topic itself) and `last_active_at` (their newest session or API-key use). Every member sees the same list |
@@ -854,6 +913,7 @@ administrator who is not in the team gets the same `404` as anybody else.
| `PUT` | `/api/teams/{teamID}/escalation` | **owner** | Replace it wholesale. `400` for a level with no targets or no timeout — a rung that pages nobody is a silence with a number on it |
| `GET` | `/api/teams/{teamID}/deadman/switches` | member | The team's [dead man's switches](#dead-mans-switch), each `{id, name, matcher, timeout_seconds, severity, status, last_heartbeat_at, last_triggered_at, open_incident_id, sources[]}`. `status` is `healthy`, `dead` or `dormant`; `sources` has one entry per heartbeat fingerprint. Empty when the team watches nothing |
| `POST` | `/api/teams/{teamID}/deadman/switches` | **owner** | Add one: `{name?, matcher, timeout_seconds, severity?}`. `400` when the matcher names no `alertname` or holds several, or the timeout is not positive — a switch that silently watches nothing is the failure this feature exists to prevent |
| `PUT` | `/api/teams/{teamID}/deadman/switches/{switchID}` | **owner** | Replace one in place, same body and validation as create. Its id is unchanged — for an automated caller reconciling a spec change, unlike delete-and-recreate |
| `DELETE` | `/api/teams/{teamID}/deadman/switches/{switchID}` | **owner** | Stop watching. An incident it opened stays open. `404` for a switch of another team |
### Notifications
+263
View File
@@ -0,0 +1,263 @@
# Service accounts: a scoped, non-human credential type
This is a design note for a feature, not an implementation plan — it exists to
propose the shape before writing code. It's raised directly by `terdut-operator`
(a separate repo, no shared code — see its `DESIGN.md` §6, §9, §13), which needs
a credential for unattended, repeatable API access and currently has no good one
available. Anything automating terdut-server long-term (this operator, CI, future
integrations) hits the same gap, so this is written as a general primitive, not
operator-specific.
## The problem
terdut-server has two credential types today, and neither fits "an unattended
process that manages teams/schedules/policies on someone's behalf":
- **User API keys** (`api_keys`, `internal/api/users.go`) are always tied to a
real `users` row and carry that user's full rights — every team they're a
member of, their admin flag if set. There's no `kind`/`service` marker
distinguishing "a human's personal automation key" from "a login session," and
no way to mint one scoped to less than the full user.
- **Integration keys** (`integrations`, `internal/api/*teams*.go`) are team-scoped,
but narrowly: they authenticate exactly one inbound Alertmanager webhook call
(`POST /api/integrations/{key}/alertmanager`) and nothing else. They're not a
general management-API credential and shouldn't become one — overloading a
narrow, one-way ingestion credential with broad read/write access would weaken
the one property that makes it safe to embed in an Alertmanager config today.
The result: any automation that needs to create teams, set escalation policies,
manage dead-man switches, or rotate integration keys has to hold a real human
admin's or team owner's API key. That key is exactly as powerful as that person
logging in — full team access, and full instance access if they're an admin.
`terdut-operator`'s design ran directly into this (its DESIGN.md §6): its
described bootstrap/rotation flow assumed a repeatable, identity-scoped way to
get a credential, and `/api/bootstrap`'s actual behavior (single-shot per
install, gated on `COUNT(*) FROM users`, confirmed via `internal/api/users.go`
and `charts/terdut-server/templates/bootstrap-job.yaml`) doesn't provide one —
it mints exactly one founding admin, once, ever.
## Goals
- A credential type that isn't a human: doesn't touch OIDC group sync, login,
session, or the `is_admin`/account-management semantics that come with a real
`users` row.
- Two scopes matching the two shapes automation actually needs: instance-wide
(create/list teams — what a server-owning controller needs) and team-scoped
(manage one team's escalation policy, dead-man switches, integrations,
schedule, OIDC group bindings — what a per-team controller or integration
needs).
- Repeatable issuance and rotation — unlike `/api/bootstrap`, callable more than
once, by anything that already holds admin rights, without destroying and
recreating state to get a fresh credential.
- Visibly distinct from a human in every place identity shows up (audit trails,
timeline entries, UI attribution) — a service account acting on a team should
never be indistinguishable from a person.
## Non-goals
- Not a general OAuth2/OIDC client-credentials flow — this is a bearer-token
primitive matching the shape `api_keys` already uses (SHA-256 hash stored,
raw key shown once at creation), not a new auth protocol.
- Not replacing integration keys — those stay as the narrow, one-way webhook
credential they are today.
- Not modeling per-endpoint or per-verb permissions within a scope — `instance`
and `team` are the only two scopes for now; finer-grained scoping is future
work if a real need shows up.
## Proposed shape
### Schema
```sql
CREATE TABLE service_accounts (
id BIGSERIAL PRIMARY KEY,
name TEXT NOT NULL UNIQUE, -- e.g. "terdut-operator"
scope TEXT NOT NULL CHECK (scope IN ('instance', 'team')),
team_id BIGINT REFERENCES teams(id) ON DELETE CASCADE,
-- team_id required iff scope = 'team'; NULL iff scope = 'instance'
created_by BIGINT REFERENCES users(id),
created_at TIMESTAMPTZ NOT NULL DEFAULT now()
);
CREATE TABLE service_account_keys (
id BIGSERIAL PRIMARY KEY,
service_account_id BIGINT NOT NULL REFERENCES service_accounts(id) ON DELETE CASCADE,
key_hash TEXT NOT NULL UNIQUE,
name TEXT NOT NULL, -- e.g. "initial", "2026-Q4-rotation"
created_at TIMESTAMPTZ NOT NULL DEFAULT now(),
last_used_at TIMESTAMPTZ
);
```
Deliberately not a `users` row: no `password_hash`, no `is_admin`, no
`user_identities` linkage, so it's structurally impossible for a service account
to be pulled into OIDC group sync or password login. Multiple keys per account
(mirroring `api_keys`' existing one-user-many-keys shape) so rotation is "mint a
new key, revoke the old one," not "recreate the account."
### Endpoints
- `POST /api/service-accounts` — instance-scope/admin-only. Body:
`{"name": ..., "scope": "instance"|"team", "teamID": ... }` (teamID required
iff scope=team, and caller must be that team's owner or a system admin).
Returns the account plus its first raw key (shown once, same pattern as
`POST /api/users/{id}/api-keys`). Safe to call again with the same `name` —
see "idempotent lookup" below — unlike `/api/bootstrap`, which is inherently
one-shot by design (it's answering "does any user exist yet," a question with
no analogue once one already does).
- `POST /api/service-accounts/{id}/keys` — mint an additional key on an existing
account (self-service-equivalent: instance admin for `instance` scope, team
owner or system admin for `team` scope). Enables rotation without recreating
the account or losing its identity/audit history.
- `DELETE /api/service-accounts/{id}/keys/{keyID}` — revoke one key, mirroring
`DELETE /api/users/{id}/api-keys/{keyID}`.
- `GET /api/service-accounts?name=` — look up an existing account by name.
This is what turns "I tried to create my account and got a conflict" into a
normal flow instead of an error: a controller that expects to have already
registered itself calls this first, and only falls through to `POST` if
nothing comes back.
### Auth middleware
**Revised** (this section originally described an aspiration that didn't
match what shipped — `TEAM-LOOKUP.md` already caught one instance of that,
and a fuller audit found three more; this is the corrected, as-built
description, not the original proposal).
`internal/api/middleware.go`'s dual resolution (`Authorization: Bearer` →
`apiKeyUser()`, or session cookie → `sessionUser()`) and the service-account
path (`serviceAccountFor()`) both resolve into one `Caller` type
(`internal/api/caller.go`), not two parallel, un-unified context
representations the way an earlier version of this server kept them. Every
authorization predicate reads `Caller`'s methods:
- `Caller.IsAdmin()` — true **only** for a human system administrator, never
for a service account of either scope, under any circumstance. `AdminOnly`
and `requireSelfOrAdmin` key on this alone — user management
(`POST /api/users`, `PUT /api/users/{id}/admin`, etc.) and
`GET/PUT /api/admin/settings` stay human-only, forever. The original text
here claimed an instance-scoped service account satisfies `AdminOnly` "for
team-creation/listing purposes" — that was never true of the shipped code
(`TEAM-LOOKUP.md` caught the listing half; the creation half was always a
separate, bespoke check in `handleCreateTeam`, not `AdminOnly` itself) and
is not being made true now. Don't widen `AdminOnly`: every time this has
come up, the fix has been a narrower, purpose-built capability instead
(`?name=` lookups for teams and service accounts; now
`terdut-operator`'s own invite-minting feature for the one real gap this
boundary left — how a human ever gets a first login on a no-OIDC,
operator-managed install. See the bottom of "What this unblocks.")
- `Caller.IsInstanceServiceAccount()` — true only for an instance-scoped
service account, never for a human (including a human admin).
`handleCreateTeam` uses exactly this: a human creates a team by being a
human (and becomes its owner); an instance-scoped service account creates
one with no human owner at all. The two paths are not interchangeable, so
this predicate deliberately does not also admit a human admin.
- `Caller.Role(teamID)`/`TeamIDs()` — a human's real `team_members` rows, or
a team-scoped service account's single synthetic owner membership
(`serveAsServiceAccount`). This is what makes `requireTeamMember`/
`requireTeamOwner` treat a team-scoped service account as owner-equivalent
for that one team, with no separate branch needed in either function.
- `Caller.ServiceAccountID()` — used by `OperatorModeBlock` ("any service
account passes") and by `callerMayManageServiceAccount`'s self-rotation
check.
- `Caller.AsHuman()` — the accessor every handler that needs a real
`user_id` to act on behalf of must call and check, instead of reading a
user off context unconditionally. Before the `Caller` type existed, four
handlers did the latter and silently misbehaved for a service-account
caller: `handleMe` and `handleTestNotification` 500'd (a zero-value user id
that matches no row), `handleDismissOnboarding` silently no-op'd (`UPDATE
... WHERE id = 0` affects nothing, still returns 204), and
`handleCreateInvite` wrote that same zero value into `invites.created_by`
— a real foreign-key violation, not just a wrong answer, since that column
is nullable but was never passed as `nil`. All four now call `AsHuman()`
and return an explicit 403 ("this endpoint is for human accounts only")
or, for the invite case, leave `created_by` `NULL` the same way
`handleCreateServiceAccount` already did for the analogous situation.
**Team scope is owner-equivalent for every `requireTeamOwner` endpoint,
membership and invites included — by design, not by an unclosed gap.** An
earlier version of this document flagged this as "acknowledged rather than
closed," kept in check only by the social convention that nobody *builds*
automation against those two routes. That convention is retired:
`terdut-operator`'s `TerdutTeam` controller now mints and revokes its own
team's invite link through exactly this capability (its existing
team-scoped credential, `POST`/`DELETE /api/teams/{teamID}/invites`), which
is the real fix for the human-onboarding gap below — not a narrower
carve-out of this capability. `service_accounts_test.go`'s
`TestServiceAccount_TeamScopeManagesItsOwnInvites` pins it.
**A team-scoped account can also mint another service account scoped to its
own team** (`handleCreateServiceAccount`'s `callerOwnsTeam` branch, which a
team-scoped caller already satisfies for its own team via the synthetic
membership above). Kept, not restricted, for the same reason: a team-scoped
credential is that team's owner's reach, full stop — carving this one
capability out while leaving membership/invites alone would be an arbitrary
asymmetry. Pinned by
`TestServiceAccount_TeamScopeCanMintAnotherAccountForItsOwnTeam`.
**`callerMayManageServiceAccount` gained the one load-bearing fix this
redesign exists for:** an instance-scoped service account may manage
(mint/revoke a key on) *any* team-scoped account, not only one admin, that
team's human owner, or the account itself. `handleCreateServiceAccount`
already let an instance-scoped caller *create* a team-scoped account for
any team; this closes the gap where adopting or rotating one it didn't just
create in the same call — exactly `terdut-operator`'s documented
adopt-on-409 crash-window recovery (its own `DESIGN.md` §5) — 403'd forever
instead of succeeding (`terdut-operator#3`). Pinned by
`TestServiceAccount_InstanceScopeAdoptsAnExistingTeamScopedAccountsKey`.
Anywhere identity is recorded for a human (incident timeline
`acknowledged_by`/`assigned_to`, audit-relevant fields), a service-account
caller is still coerced into a bare `user_id` of `0` today — `Caller`'s new
`Identity()` accessor exists for exactly this follow-up, but wiring it in
needs a schema migration (an actor-attribution column distinct from
`user_id`) and is deliberately out of scope here. Tracked separately, not by
this document.
## What this unblocks
Directly resolves `terdut-operator` DESIGN.md §6's two broken assumptions:
1. **Bootstrap becomes single-purpose again.** `/api/bootstrap` mints exactly
the founding human admin, once. The operator's actual first-reconcile flow:
call `/api/bootstrap` only on a genuinely empty install; otherwise (or
immediately after, if it won the bootstrap race) call
`GET /api/service-accounts?name=terdut-operator`, and `POST` one if it
doesn't exist yet. From then on the operator never touches `/api/bootstrap`
again.
2. **Rotation becomes real.** `POST /api/service-accounts/{id}/keys` + revoke the
old one — no destructive DB-level workaround, no re-triggering a single-shot
endpoint that can't fire twice.
3. **Cross-namespace credential mirroring is no longer needed at all.**
`terdut-operator`'s current design holds every credential — instance- and
team-scoped alike — privately in the operator's own namespace, never in
the namespace of the CR each one authenticates for; reconciliation happens
entirely inside the operator's controller loop, so no CR owner ever needs
read access to a terdut-server credential regardless of same- or
cross-namespace `serverRef`. Team scoping is still what bounds the blast
radius of any individual credential: a leaked team-scoped key exposes
exactly one team's resources, never the whole server, which is what makes
holding many credentials in one place (the operator's namespace) an
acceptable trade rather than reintroducing the mirrored design's
server-admin-equivalent-everywhere problem.
4. **A human can get a first login on a no-OIDC, operator-managed install —
without ever touching `AdminOnly` or `/api/admin/settings`.** This was
filed as `terdut-server#23` ("no API path to create a human login after
bootstrap") and diagnosed, at the time, as this server needing to let a
service account through `AdminOnly`. It doesn't: the fix lives entirely
in `terdut-operator`, because a team-scoped credential was *already*
owner-equivalent for `POST /api/teams/{teamID}/invites`, and invite
redemption (`POST /api/signup` with an `invite` token) bypasses
`signup_mode` entirely — `terdut-operator` just never grew a feature to
use either fact. Its `TerdutTeam` controller now mints and surfaces one
via its own existing team-scoped credential (`spec.invite`,
`status.inviteSecretRef`, see that repo's own docs), so a human joins a
CRD-managed team by a real invite link, the same way anyone else would.
`terdut-server#23` is closed with this note once that feature ships — its
named routes stay human-only, correctly, not a gap.
## Suggested sequencing
Land this before `terdut-operator` implements any bootstrap/credential-handling
code — that code would otherwise be written against the current one-shot,
user-only credential model as a known-temporary workaround, which is wasted
effort on a repo that currently has zero implementation to begin with.
+73
View File
@@ -0,0 +1,73 @@
# Team lookup for service accounts: closing terdut-operator's create-path crash window
This is a design note for a feature, not an implementation plan — same posture as
`SERVICE-ACCOUNTS.md`, and raised for the same reason: `terdut-operator`'s `TerdutTeam`
controller (ROADMAP.md Stage 2, a separate repo, no shared code) hit a gap this server has
no answer for yet.
## The problem
`POST /api/teams` (`handleCreateTeam`, confirmed against `internal/api/teams.go`) lets an
instance-scoped service account create a team — it has its own explicit
`isInstanceServiceAccount(...)` branch alongside the human-user path, not gated by
`AdminOnly`. If that call succeeds server-side but the caller (`TerdutTeam`'s controller)
crashes before persisting the resulting team ID locally, a retry's `POST` 409s on the name's
unique constraint (confirmed: the `isUniqueViolation` branch in the same handler).
Recovering from that 409 means looking the team up by name, and nothing today permits that
for a service account:
- `GET /api/teams` (`handleListTeams`) answers "what teams does the *caller* belong to", via
a `team_members` join keyed on `userFromContext`'s `caller.ID` — confirmed against source.
A service account is never a member of anything, so this always returns empty for one,
regardless of what exists.
- `GET /api/admin/teams` (`handleAdminListTeams`) is gated by `AdminOnly`, and `AdminOnly`'s
actual code (`internal/api/middleware.go`) checks only `userFromContext(...).IsAdmin` — no
branch for a service account at all, confirmed against source. This contradicts
`SERVICE-ACCOUNTS.md`'s own text, which claims "an instance-scoped [service account
satisfies] `AdminOnly` for team-creation/listing purposes" — that claim doesn't match this
endpoint's actual, shipped code. (Team *creation* is fine: `handleCreateTeam` isn't behind
`AdminOnly` at all, it has its own check. Only the listing half of that sentence is wrong.)
This is exactly the shape of gap `SERVICE-ACCOUNTS.md`'s own `GET /api/service-accounts?name=`
closed for service accounts themselves (confirmed: that endpoint's own comment —
"the name lookup is open to any authenticated caller... what lets a service account find its
own account on the 403 that follows a second POST"). Teams never got the equivalent, because
nothing needed it until an operator started creating them unattended.
## Goals
- A service-account-accessible way to look up one team by exact name, mirroring
`GET /api/service-accounts?name=` as closely as possible — same shape, same reasoning,
same low sensitivity of what it discloses.
- No change to today's behavior for an empty/no-name request.
## Proposed shape
Extend `GET /api/teams` itself, the same way `handleListServiceAccounts` already branches on
a `?name=` query param, rather than adding a new route:
- `name` unset (today's behavior, unchanged): the caller's own teams, via `team_members`.
- `name=<value>` set: look up that one team by exact name — a one-or-zero-length array, not
an error on no match, mirroring `GET /api/service-accounts?name=`'s own response shape and
status codes exactly. Deliberately **not** gated by `isInstanceServiceAccount` or
`AdminOnly`: a human caller who's already a member sees this same information in their own
team list regardless, and a non-member learning only that a name is taken — not who's in
the team, not any of its data — is the same low-sensitivity disclosure
`GET /api/service-accounts?name=` already accepts for service-account names.
## What this unblocks
Directly resolves the crash-window gap in `terdut-operator`'s `TerdutTeam` controller: on a
409 from `POST /api/teams`, `GET /api/teams?name=<the same name>` — authenticated with the
same instance-scoped credential that just got the 409 — finds the id, and the controller
proceeds as if its own create had returned it directly. The same adopt-on-409 pattern already
proven for service accounts (that repo's `DESIGN.md` §6 point 1, §5's general rule), not a
new one.
## Suggested sequencing
Land this before `TerdutTeam`'s create path is implemented — the same reasoning
`SERVICE-ACCOUNTS.md` gave for its own sequencing: writing that code against today's gap as a
"known-temporary workaround" is wasted effort when the fix is this small and this
well-precedented.
+2 -2
View File
@@ -15,5 +15,5 @@ type: application
# appVersion and image.tag in values.yaml no longer agree, and that is not an oversight:
# image.tag stays "latest", which is what a local install actually pulls. appVersion is
# metadata and drives nothing.
version: 0.30.0
appVersion: "v0.30.0"
version: 0.37.0
appVersion: "v0.37.0"
+28 -5
View File
@@ -6,21 +6,42 @@ metadata:
labels:
{{- include "terdut-server.labels" . | nindent 4 }}
spec:
replicas: 1
replicas: {{ .Values.replicaCount }}
selector:
matchLabels:
{{- include "terdut-server.selectorLabels" . | nindent 6 }}
# Recreate, not RollingUpdate, even though the PVC that forced it is gone: the
# sweeper and the notifier are unsynchronised singletons, and two replicas
# overlapping during a rollout would both page for the same incident.
# RollingUpdate, not Recreate: the sweeper, notifier and migration runner
# each take a Postgres advisory lock around their own pass, and new-incident
# creation on the first webhook for a brand-new groupKey resolves its own
# insert conflict -- so two replicas overlapping during a rollout no longer
# double-page, race a migration, or drop a webhook payload (v0.36.0). No
# explicit maxUnavailable/maxSurge: the 25%/25% default rounds to 0/1 at
# replicaCount: 2, which is zero-downtime already.
strategy:
type: Recreate
type: RollingUpdate
template:
metadata:
labels:
{{- include "terdut-server.selectorLabels" . | nindent 8 }}
spec:
enableServiceLinks: false
{{- if .Values.database.waitForPostgres.enabled }}
initContainers:
- name: wait-for-postgres
image: "{{ .Values.database.waitForPostgres.image.repository }}:{{ .Values.database.waitForPostgres.image.tag }}"
imagePullPolicy: {{ .Values.database.waitForPostgres.image.pullPolicy }}
env:
- name: TERDUT_DB_DSN
value: {{ required "database.dsn is required" .Values.database.dsn | quote }}
command:
- sh
- -c
- |
until pg_isready -d "$TERDUT_DB_DSN"; do
echo "wait-for-postgres: not ready yet, retrying in 2s"
sleep 2
done
{{- end }}
containers:
- name: terdut-server
image: "{{ .Values.image.repository }}:{{ .Values.image.tag }}"
@@ -75,6 +96,8 @@ spec:
value: "{{ .Values.notify.publicUrl | default (printf "https://%s" .Values.networking.hostname) }}"
- name: TERDUT_PASSWORD_LOGIN
value: {{ .Values.passwordLogin | quote }}
- name: TERDUT_OPERATOR_MODE
value: {{ .Values.operatorMode | quote }}
{{- if .Values.oidc.enabled }}
- name: TERDUT_OIDC_ISSUER
value: {{ required "oidc.issuer is required when oidc.enabled" .Values.oidc.issuer | quote }}
+35
View File
@@ -1,3 +1,10 @@
# Safe above 1 since v0.36.0: the sweeper, notifier and migration runner each
# take a Postgres advisory lock around their own pass, and a webhook that
# loses the race to open a brand-new incident attaches to the winner's row
# instead of dropping its payload. An image older than v0.36.0 does not have
# these guards -- do not raise this against one.
replicaCount: 2
networking:
hostname: "terdut.example.com"
servicePort: 8080
@@ -33,6 +40,26 @@ database:
passwordSecret:
name: ""
key: password
# Blocks the main container from starting until Postgres accepts
# connections. Without this, a Deployment created before Postgres has
# finished its very first boot -- initdb plus Patroni leader election, on a
# from-scratch postgres-operator cluster -- crash-loops a few times: the
# app's own ping-retry budget on startup (pingAttempts/pingRetryDelay in
# internal/db/db.go) is sized for a much shorter, different race --
# NetworkPolicy propagation, a few seconds -- not for genuine first-time
# cluster creation, which routinely takes longer, so it exhausts and the
# process exits before ever binding its HTTP port. A startupProbe cannot
# help here: the crash happens before there is anything to probe.
#
# pg_isready needs no credentials -- it reports PQPING_OK on anything that
# amounts to "a Postgres backend answered", including an auth challenge --
# so no PGPASSWORD is wired into this container.
waitForPostgres:
enabled: true
image:
repository: postgres
tag: "17-alpine"
pullPolicy: IfNotPresent
service:
type: ClusterIP
@@ -112,6 +139,14 @@ notify:
# redeploy) if the identity provider is down and somebody has to get in.
passwordLogin: true
# Declares this install gitops-managed: writes to teams, escalation policies,
# dead man's switches and integrations from a session or a user's own API key
# are refused, while a service account's (see SERVICE-ACCOUNTS.md) are not.
# Off by default — turning it on is a statement that something like
# terdut-operator, not a person in the web UI, owns this install's
# configuration from here on.
operatorMode: false
# Single sign-on through an OpenID Connect provider such as Authentik.
#
# At the provider, create an OAuth2/OpenID application whose redirect URI is
+1 -1
View File
@@ -54,7 +54,7 @@ func main() {
log.Fatalf("seed settings: %v", err)
}
router := api.NewRouter(database, notify, cfg)
router := api.NewRouter(database, notify, cfg, version)
srv := &http.Server{
Addr: cfg.Addr,
+122
View File
@@ -0,0 +1,122 @@
package api
// This file is internal (package api, not api_test) because withAdvisoryLock is
// unexported and these tests exercise its locking semantics directly rather than
// through the full StartArchiver/StartNotifier loop, which would make the "does
// not run while held" case timing-dependent instead of deterministic. It opens a
// plain connection to TERDUT_TEST_DSN rather than reusing testdb_test.go's
// newTestDB, since that helper lives in the separate, already-compiled
// api_test package and a Postgres advisory lock needs no schema or migration
// to exercise.
import (
"context"
"database/sql"
"os"
"testing"
_ "github.com/jackc/pgx/v5/stdlib"
)
// advisoryTestDB opens a plain, unmigrated connection to the test database. An
// unset DSN fails rather than skips, matching testdb_test.go's rationale: a
// suite that quietly tests nothing is worse than one that does not run.
func advisoryTestDB(t *testing.T) *sql.DB {
t.Helper()
dsn := os.Getenv("TERDUT_TEST_DSN")
if dsn == "" {
t.Fatalf("TERDUT_TEST_DSN is not set: these tests need Postgres.\n" +
"Run `make test-db` for a local one, then\n" +
" export TERDUT_TEST_DSN=postgres://terdut:terdut@localhost:5432/terdut_test?sslmode=disable")
}
db, err := sql.Open("pgx", dsn)
if err != nil {
t.Fatalf("connect to TERDUT_TEST_DSN: %v", err)
}
t.Cleanup(func() { db.Close() })
return db
}
func TestWithAdvisoryLock_RunsWhenFree(t *testing.T) {
db := advisoryTestDB(t)
ctx := context.Background()
ran := false
withAdvisoryLock(ctx, db, archiverLockKey, "test", func() { ran = true })
if !ran {
t.Fatal("fn did not run although the lock was free")
}
}
func TestWithAdvisoryLock_SkipsWhileHeldElsewhere(t *testing.T) {
db := advisoryTestDB(t)
ctx := context.Background()
// Hold the lock on a connection of our own, standing in for another
// replica mid-pass.
holder, err := db.Conn(ctx)
if err != nil {
t.Fatalf("acquire holder connection: %v", err)
}
defer holder.Close()
if _, err := holder.ExecContext(ctx, "SELECT pg_advisory_lock($1)", archiverLockKey); err != nil {
t.Fatalf("pre-acquire lock: %v", err)
}
ran := false
withAdvisoryLock(ctx, db, archiverLockKey, "test", func() { ran = true })
if ran {
t.Fatal("fn ran although another connection already held the lock")
}
if _, err := holder.ExecContext(ctx, "SELECT pg_advisory_unlock($1)", archiverLockKey); err != nil {
t.Fatalf("release held lock: %v", err)
}
// Now that the holder released it, the next caller should get it.
ran = false
withAdvisoryLock(ctx, db, archiverLockKey, "test", func() { ran = true })
if !ran {
t.Fatal("fn did not run after the other connection released the lock")
}
}
func TestWithAdvisoryLock_ReleasesAfterFnReturns(t *testing.T) {
db := advisoryTestDB(t)
ctx := context.Background()
withAdvisoryLock(ctx, db, notifierLockKey, "test", func() {})
// If the first call had leaked the lock, this one would see it held and
// skip, leaving ran false.
ran := false
withAdvisoryLock(ctx, db, notifierLockKey, "test", func() { ran = true })
if !ran {
t.Fatal("fn did not run on a later call: the earlier call leaked its lock")
}
}
func TestWithAdvisoryLock_KeysAreIndependent(t *testing.T) {
db := advisoryTestDB(t)
ctx := context.Background()
holder, err := db.Conn(ctx)
if err != nil {
t.Fatalf("acquire holder connection: %v", err)
}
defer holder.Close()
if _, err := holder.ExecContext(ctx, "SELECT pg_advisory_lock($1)", archiverLockKey); err != nil {
t.Fatalf("pre-acquire archiver lock: %v", err)
}
defer holder.ExecContext(ctx, "SELECT pg_advisory_unlock($1)", archiverLockKey)
// Holding archiverLockKey must not block notifierLockKey.
ran := false
withAdvisoryLock(ctx, db, notifierLockKey, "test", func() { ran = true })
if !ran {
t.Fatal("fn did not run under a different key although only archiverLockKey was held")
}
}
+35 -1
View File
@@ -376,6 +376,17 @@ func incidentForGroup(ctx context.Context, tx *sql.Tx, notify NotifyConfig, team
// its own. Hence the querier rather than a *sql.Tx. A nil severity leaves the
// column for refreshSeverity to fill from the member alerts; the sweeper passes
// one because its incidents have no members to derive it from.
//
// Both callers get here only after their own SELECT found no open incident for
// this group_key — but on more than one replica, two webhook deliveries for the
// very first occurrence of a brand-new group_key can both pass that SELECT
// before either INSERTs. ON CONFLICT DO NOTHING against
// incidents_open_group_key_idx is what makes the loser's INSERT a no-op instead
// of a unique-violation error that would otherwise roll back its entire
// payload; existingOpenIncident then hands it the winner's row. Postgres
// resolves that conflict only once the winner's transaction has committed (or
// rolled back), so by the time this RETURNING comes back empty, the winner's
// row is guaranteed visible to that follow-up SELECT.
func openIncident(ctx context.Context, q querier, notify NotifyConfig, teamID int64, groupKey, title string, groupLabels map[string]string, severity *string) (int64, error) {
onCall, err := currentOnCall(ctx, q, teamID)
if err != nil {
@@ -391,10 +402,18 @@ func openIncident(ctx context.Context, q querier, notify NotifyConfig, teamID in
err = q.QueryRowContext(ctx, `
INSERT INTO incidents (team_id, group_key, title, group_labels, signature, status, severity, triggered_at, assigned_to)
VALUES ($1, $2, $3, $4::jsonb, $5, 'triggered', $6, $7, $8)
ON CONFLICT (team_id, group_key) WHERE resolved_at IS NULL DO NOTHING
RETURNING id`,
teamID, groupKey, title, string(labelsJSON), incidentSignature(groupLabels, title), severity,
time.Now().Unix(), onCall).Scan(&id)
if err != nil {
switch {
case err == sql.ErrNoRows:
// Lost the race: someone else's incident for this group_key exists now.
// Everything below — the trigger event, assignment, page, escalation
// clock — already happened for that row when it was created; attach to
// it rather than fail this call (and the whole payload) outright.
return existingOpenIncident(ctx, q, teamID, groupKey)
case err != nil:
return 0, err
}
@@ -423,6 +442,21 @@ func openIncident(ctx context.Context, q querier, notify NotifyConfig, teamID in
return id, nil
}
// existingOpenIncident looks up the open incident openIncident's own INSERT just
// lost a conflict against — the same lookup incidentForGroup does before ever
// calling openIncident, repeated here for the caller that arrived second.
func existingOpenIncident(ctx context.Context, q querier, teamID int64, groupKey string) (int64, error) {
var id int64
err := q.QueryRowContext(ctx,
"SELECT id FROM incidents WHERE team_id = $1 AND group_key = $2 AND resolved_at IS NULL",
teamID, groupKey,
).Scan(&id)
if err != nil {
return 0, err
}
return id, nil
}
// linkAlert adds an alert to an incident, emitting a timeline entry only the
// first time. Re-sends of an already-linked alert are silent.
func linkAlert(ctx context.Context, tx *sql.Tx, incidentID, alertID int64) error {
+112
View File
@@ -0,0 +1,112 @@
package api_test
import (
"bytes"
"encoding/json"
"fmt"
"net/http"
"sync"
"testing"
)
// TestWebhook_ConcurrentFirstOccurrenceOpensOneIncident reproduces two
// replicas racing the very first webhook delivery for a brand-new group_key:
// both see no open incident yet (incidentForGroup's own SELECT finds
// nothing) and race openIncident's INSERT.
//
// The DB's own unique index already guarantees at most one incident either
// way, with or without this fix — so "exactly one incident" alone cannot
// tell a fixed run from a broken one. What ON CONFLICT handling actually
// changes is what happens to the *loser*: before it, the loser's INSERT hit
// incidents_open_group_key_idx's unique violation, which — since
// upsertAlerts ran earlier in that same transaction — rolled back its whole
// payload, alert insert included. ingest's error is only logged and
// receiveWebhook answers 200 regardless, so nothing ever retried it: the
// loser's alert silently never existed. That is the regression signal this
// test checks — every caller's fingerprint must show up in /api/alerts, not
// just the winner's.
func TestWebhook_ConcurrentFirstOccurrenceOpensOneIncident(t *testing.T) {
s := newTS(t)
const callers = 8
const groupKey = "race-group"
// Every caller needs its own fingerprint. A shared one would serialize all
// of them at upsertAlerts' own ON CONFLICT (team_id, fingerprint) row lock,
// long before any of them reached incidentForGroup — which would hide the
// very race this test exists to force.
bodies := make([][]byte, callers)
for i := range callers {
payload := map[string]any{
"version": "4",
"status": "firing",
"groupKey": groupKey,
"groupLabels": map[string]string{"alertname": "RaceAlert"},
"alerts": []map[string]any{amAlert(fmt.Sprintf("fp-race-%d", i), "RaceAlert", "firing",
"2026-05-20T10:00:00Z", "0001-01-01T00:00:00Z", nil)},
}
bodies[i], _ = json.Marshal(payload)
}
// A start line, so every request is fired as close to simultaneously as
// goroutine scheduling allows, rather than trickling out one dial at a
// time — the race window is the gap between incidentForGroup's SELECT and
// openIncident's INSERT, which a staggered start could easily miss.
var ready sync.WaitGroup
start := make(chan struct{})
statuses := make([]int, callers)
var wg sync.WaitGroup
for i := range callers {
ready.Add(1)
wg.Add(1)
go func(i int) {
defer wg.Done()
ready.Done()
<-start
resp, err := http.Post(s.URL+"/api/integrations/"+s.ingestKey+"/alertmanager",
"application/json", bytes.NewReader(bodies[i]))
if err != nil {
t.Errorf("post webhook #%d: %v", i, err)
return
}
defer resp.Body.Close()
statuses[i] = resp.StatusCode
}(i)
}
ready.Wait()
close(start)
wg.Wait()
for i, code := range statuses {
if code != http.StatusOK {
t.Errorf("webhook #%d returned %d, want 200", i, code)
}
}
var matched []any
for _, inc := range listIncidents(t, s, "") {
if inc["group_key"] == groupKey {
matched = append(matched, inc["id"])
}
}
if len(matched) != 1 {
t.Fatalf("expected exactly 1 incident for group_key %q after %d concurrent deliveries, got %d: %v",
groupKey, callers, len(matched), matched)
}
var alerts []map[string]any
decode(t, s.req(t, http.MethodGet, "/api/alerts", nil), &alerts)
seen := map[string]bool{}
for _, a := range alerts {
if fp, ok := a["fingerprint"].(string); ok {
seen[fp] = true
}
}
for i := range callers {
fp := fmt.Sprintf("fp-race-%d", i)
if !seen[fp] {
t.Errorf("alert %q is missing: its delivery's whole payload was silently rolled back "+
"when it lost the race for the incident", fp)
}
}
}
+1 -1
View File
@@ -60,7 +60,7 @@ func newTSWith(t *testing.T, deadman api.DeadmanConfig, cfg api.NotifyConfig, co
t.Helper()
database := newTestDB(t)
srv := httptest.NewServer(api.NewRouter(database, cfg, conf))
srv := httptest.NewServer(api.NewRouter(database, cfg, conf, "test"))
t.Cleanup(srv.Close)
body, _ := json.Marshal(map[string]string{"username": "admin", "email": "admin@test.com"})
+18 -2
View File
@@ -15,6 +15,12 @@ const (
// expiryGrace absorbs clock skew and notification latency before an alert
// whose ends_at watermark has passed is treated as stale.
expiryGrace = 5 * time.Minute
// archiverLockKey is the Postgres advisory lock the sweeper takes for the
// duration of each pass, so that running more than one replica does not run
// the sweep concurrently on all of them. Its value has no meaning beyond
// being distinct from notifierLockKey.
archiverLockKey int64 = 7265_0001
)
// StartArchiver runs the alert sweeper until ctx is cancelled, starting with an
@@ -23,15 +29,25 @@ const (
// the fallback, not the setting: each pass reads the current value from the
// settings table, so an administrator's change takes effect on the next tick
// instead of at the next restart.
//
// Each pass runs under archiverLockKey (see withAdvisoryLock), so that on more
// than one replica only whichever instance's tick takes the lock first actually
// sweeps; the rest skip that tick rather than racing the same pass.
func StartArchiver(ctx context.Context, db *sql.DB, archiveAfter, staleAfter time.Duration, notify NotifyConfig) {
ticker := time.NewTicker(sweepInterval)
defer ticker.Stop()
Sweep(ctx, db, archiveAfter, staleAfter, notify)
sweep := func() {
withAdvisoryLock(ctx, db, archiverLockKey, "sweeper", func() {
Sweep(ctx, db, archiveAfter, staleAfter, notify)
})
}
sweep()
for {
select {
case <-ticker.C:
Sweep(ctx, db, archiveAfter, staleAfter, notify)
sweep()
case <-ctx.Done():
return
}
+5 -1
View File
@@ -279,7 +279,11 @@ type meResponse struct {
// between the login form and the app, since it cannot read its own cookie.
func handleMe(db *sql.DB) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
caller, _ := userFromContext(r.Context())
caller, ok := userFromContext(r.Context())
if !ok {
respond(w, http.StatusForbidden, errResp("this endpoint is for human accounts only"))
return
}
user, err := fetchUser(r.Context(), db, caller.ID)
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
+1 -1
View File
@@ -301,7 +301,7 @@ func TestSetPassword_EndsOtherSessionsButNotThisOne(t *testing.T) {
func TestBootstrap_WithPassword(t *testing.T) {
database := newTestDB(t)
srv := httptest.NewServer(api.NewRouter(database, api.NotifyConfig{}, testConfig()))
srv := httptest.NewServer(api.NewRouter(database, api.NotifyConfig{}, testConfig(), "test"))
t.Cleanup(srv.Close)
body := `{"username":"admin","email":"a@test.com","password":"` + adminPassword + `"}`
+123
View File
@@ -0,0 +1,123 @@
package api
import (
"context"
"fmt"
"git.ryuvia.com/niklas/terdut-server/internal/models"
)
// Caller is the one principal type every authorization predicate in this
// package reads from. Before this, a human (ctxUser + ctxTeams) and a
// service account (ctxServiceAccount + a synthetic ctxTeams entry) were two
// parallel, un-unified context representations — every predicate had to
// remember which one(s) it needed to check, and the ones that forgot either
// 403'd a service account that should have been let through (terdut-server#23,
// terdut-operator#3), crashed on an unchecked zero-value user ID (handleMe,
// handleTestNotification), or silently no-op'd (handleDismissOnboarding).
// serveAs and serveAsServiceAccount now both build exactly one Caller and
// store it under one context key; everything else in this file is a read
// of one of its methods.
type Caller struct {
// user is set for a human caller (session cookie or a user's own API
// key), nil for a service account of either scope.
user *models.User
// sa is set for a service-account caller, nil for a human.
sa *serviceAccountPrincipal
// memberships is the caller's real team_members rows for a human, or —
// for a team-scoped service account — the single synthetic owner
// membership serveAsServiceAccount injects (see its own comment for
// why). Always nil for an instance-scoped service account: it acts on
// teams by id, not by belonging to one.
memberships []membership
}
// AsHuman returns the real user behind this caller, or false for a service
// account of either scope. Every handler that needs a real user_id to act
// on behalf of — not just "is this caller sufficiently privileged" — calls
// this and handles the false case explicitly, replacing the unchecked
// userFromContext(ctx) zero-value reads that used to silently misbehave for
// a service-account caller.
func (c Caller) AsHuman() (models.User, bool) {
if c.user == nil {
return models.User{}, false
}
return *c.user, true
}
// IsAdmin is true only for a human system administrator — never for a
// service account, of either scope, under any circumstance. AdminOnly and
// requireSelfOrAdmin key on this and nothing else: user management and
// /api/admin/settings stay human-only forever, by design (SERVICE-ACCOUNTS.md).
func (c Caller) IsAdmin() bool {
return c.user != nil && c.user.IsAdmin
}
// IsInstanceServiceAccount reports whether this caller is specifically an
// instance-scoped service account — never true for a human, including a
// human admin. handleCreateTeam needs exactly this: a human creates a team
// by being a human (and becomes its owner as a side effect), an
// instance-scoped service account creates one with no human owner at all;
// the two paths are not interchangeable, so this predicate must not also
// admit a human admin the way MayActAsInstanceAdmin deliberately does.
func (c Caller) IsInstanceServiceAccount() bool {
return c.sa != nil && c.sa.scope == models.ServiceAccountScopeInstance
}
// Role reports the caller's role in teamID, and whether they belong to it
// at all.
func (c Caller) Role(teamID int64) (string, bool) {
for _, m := range c.memberships {
if m.teamID == teamID {
return m.role, true
}
}
return "", false
}
// TeamIDs lists every team this caller belongs to: a human's real
// memberships, or a team-scoped service account's own single team. Always
// empty for an instance-scoped service account.
func (c Caller) TeamIDs() []int64 {
ids := make([]int64, 0, len(c.memberships))
for _, m := range c.memberships {
ids = append(ids, m.teamID)
}
return ids
}
// ServiceAccountID reports this caller's own service-account id, for the
// "may manage/rotate its own credential" self-check in
// callerMayManageServiceAccount, and for OperatorModeBlock's "any service
// account passes" rule.
func (c Caller) ServiceAccountID() (int64, bool) {
if c.sa == nil {
return 0, false
}
return c.sa.id, true
}
// Identity is a stable, log/audit-facing string distinguishing a human
// caller from a service account — "user:42" or "service-account:7". Not
// wired into any database column today (incidents.go's acknowledged_by/
// assigned_to/user_id are explicitly out of scope for this change — that
// needs its own schema migration, tracked separately), but this is the one
// place in the request path that already knows which kind of caller this
// is, and that follow-up will want exactly this accessor.
func (c Caller) Identity() string {
switch {
case c.user != nil:
return fmt.Sprintf("user:%d", c.user.ID)
case c.sa != nil:
return fmt.Sprintf("service-account:%d", c.sa.id)
default:
return "unknown"
}
}
func callerFromContext(ctx context.Context) (Caller, bool) {
c, ok := ctx.Value(ctxCaller).(Caller)
return c, ok
}
+39
View File
@@ -1,8 +1,11 @@
package api
import (
"context"
"database/sql"
"encoding/json"
"errors"
"log"
"net/http"
"strconv"
"strings"
@@ -77,3 +80,39 @@ func decodeJSON(r *http.Request, v any) error {
func errResp(msg string) map[string]string {
return map[string]string{"error": msg}
}
// withAdvisoryLock runs fn only if it can take the named Postgres advisory lock on a
// dedicated connection, and skips fn otherwise. This is what keeps the archiver and
// notifier safe to run on more than one replica: whichever instance's tick gets there
// first does the work; the rest see the lock held and simply wait for their next tick
// instead of running the same pass concurrently.
//
// pg_try_advisory_lock is session-scoped, so taking and releasing it must happen on the
// same connection, reserved via db.Conn rather than borrowed from the pool's shared
// connections fn itself may use — and released (unlocked, then closed) before returning,
// since a session lock otherwise outlives this call and leaks onto whatever reuses the
// pooled connection next.
func withAdvisoryLock(ctx context.Context, db *sql.DB, key int64, name string, fn func()) {
conn, err := db.Conn(ctx)
if err != nil {
log.Printf("%s: advisory lock: acquire connection: %v", name, err)
return
}
defer conn.Close()
var locked bool
if err := conn.QueryRowContext(ctx, "SELECT pg_try_advisory_lock($1)", key).Scan(&locked); err != nil {
log.Printf("%s: advisory lock: %v", name, err)
return
}
if !locked {
return // another replica is already running this pass
}
defer func() {
if _, err := conn.ExecContext(ctx, "SELECT pg_advisory_unlock($1)", key); err != nil {
log.Printf("%s: advisory unlock: %v", name, err)
}
}()
fn()
}
+7 -4
View File
@@ -261,14 +261,17 @@ func resolveIfSettled(ctx context.Context, q querier, incidentID int64) (bool, e
}
// acknowledgeIncident records that userID has picked an incident up, and reports
// whether it changed anything — an already-resolved incident is left alone.
// Shared by the authenticated handler and the Acknowledge button in a push
// notification, so both write the same state and the same timeline entry.
// whether it changed anything — an already-resolved or already-acknowledged
// incident is left alone, so a second acknowledge (a retried request, or a
// stale push notification tapped after the web UI already acked it) is a
// no-op rather than a second "acknowledged" timeline entry. Shared by the
// authenticated handler and the Acknowledge button in a push notification,
// so both write the same state and the same timeline entry.
func acknowledgeIncident(ctx context.Context, q querier, incidentID, userID int64) (bool, error) {
res, err := q.ExecContext(ctx, `
UPDATE incidents
SET status = 'acknowledged', acknowledged_by = $1, acknowledged_at = $2
WHERE id = $3 AND resolved_at IS NULL`,
WHERE id = $3 AND status = 'triggered'`,
userID, time.Now().Unix(), incidentID)
if err != nil {
return false, err
+12 -2
View File
@@ -197,10 +197,20 @@ func handleIncidentAcknowledge(db *sql.DB) http.HandlerFunc {
return
}
if !acked {
if !incidentExists(w, r, db, id) {
// incidentIDParam above already confirmed the incident exists, so this
// is either resolved, or already acknowledged — the latter is now a
// no-op rather than an error, since the caller's desired state
// (acknowledged) already holds.
inc, err := fetchIncident(r.Context(), db, id)
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
respond(w, http.StatusConflict, errResp("incident is resolved"))
if inc.Status == "resolved" {
respond(w, http.StatusConflict, errResp("incident is resolved"))
return
}
respond(w, http.StatusOK, inc)
return
}
respondIncident(w, r, db, id)
+42
View File
@@ -338,6 +338,48 @@ func TestIncident_Acknowledge(t *testing.T) {
}
}
// A second acknowledge — a retried request, or a stale push notification
// tapped after the web UI already acked it — must be a no-op: same state,
// no second "acknowledged" timeline entry. Regression test for the bug
// described in issue #26 ("two acknowledged entries look like a bug").
func TestIncident_AcknowledgeTwiceIsIdempotent(t *testing.T) {
s := newTS(t)
postWebhook(t, s, []map[string]any{
amAlert("fp-ack2", "Y", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
})
resp := s.req(t, http.MethodPost, "/api/incidents/1/acknowledge", nil)
if resp.StatusCode != http.StatusOK {
t.Fatalf("first acknowledge returned %d", resp.StatusCode)
}
var first map[string]any
decode(t, resp, &first)
resp = s.req(t, http.MethodPost, "/api/incidents/1/acknowledge", nil)
if resp.StatusCode != http.StatusOK {
t.Fatalf("second acknowledge returned %d, want 200 (idempotent)", resp.StatusCode)
}
var second map[string]any
decode(t, resp, &second)
if second["status"] != "acknowledged" {
t.Errorf("expected status still acknowledged, got %v", second["status"])
}
if second["acknowledged_by"] != first["acknowledged_by"] {
t.Errorf("expected the same acknowledged_by, got %v then %v", first["acknowledged_by"], second["acknowledged_by"])
}
types := eventTypes(timeline(t, s, 1))
n := 0
for _, ty := range types {
if ty == "acknowledged" {
n++
}
}
if n != 1 {
t.Errorf("expected exactly one acknowledged event, got %d in %v", n, types)
}
}
func TestIncident_ManualResolveIsTerminal(t *testing.T) {
s := newTS(t)
postWebhook(t, s, []map[string]any{
+131 -23
View File
@@ -9,15 +9,20 @@ import (
"strings"
"time"
"git.ryuvia.com/niklas/terdut-server/internal/config"
"git.ryuvia.com/niklas/terdut-server/internal/models"
)
type contextKey string
const (
ctxUser contextKey = "user"
// ctxCaller holds the one Caller (see caller.go) every authorization
// predicate in this package reads from — a human and a service account
// used to be two parallel, un-unified context keys (ctxUser/ctxTeams vs.
// ctxServiceAccount); this is why that was a mistake, not a smaller
// version of the same idea.
ctxCaller contextKey = "caller"
ctxSession contextKey = "session"
ctxTeams contextKey = "teams"
)
// AuthMiddleware accepts either of the two credentials the server issues: an
@@ -39,12 +44,19 @@ func AuthMiddleware(db *sql.DB) func(http.Handler) http.Handler {
respond(w, http.StatusUnauthorized, errResp("unauthorized"))
return
}
userID, ok := apiKeyUser(r.Context(), db, token)
if !ok {
respond(w, http.StatusUnauthorized, errResp("unauthorized"))
if userID, ok := apiKeyUser(r.Context(), db, token); ok {
serveAs(w, r, next, db, userID, 0)
return
}
serveAs(w, r, next, db, userID, 0)
// Tried second, not first: a user API key is the common case,
// and a service-account key is visibly prefixed (tdsa_) so this
// second lookup is rarely reached on a request that was going
// to fail anyway.
if sa, ok := serviceAccountFor(r.Context(), db, token); ok {
serveAsServiceAccount(w, r, next, sa)
return
}
respond(w, http.StatusUnauthorized, errResp("unauthorized"))
return
}
@@ -170,8 +182,7 @@ func serveAs(w http.ResponseWriter, r *http.Request, next http.Handler, db *sql.
return
}
ctx := context.WithValue(r.Context(), ctxTeams, teams)
ctx = context.WithValue(ctx, ctxUser, u)
ctx := context.WithValue(r.Context(), ctxCaller, Caller{user: &u, memberships: teams})
if sessionID != 0 {
ctx = context.WithValue(ctx, ctxSession, sessionID)
}
@@ -183,9 +194,115 @@ func hashToken(token string) string {
return hex.EncodeToString(h[:])
}
// userFromContext is a thin compatibility wrapper over Caller.AsHuman(), so
// every call site written before the Caller abstraction (alerts.go,
// incidents.go, schedule.go, stats.go, and more) needs no change and keeps
// its exact existing behavior.
func userFromContext(ctx context.Context) (models.User, bool) {
u, ok := ctx.Value(ctxUser).(models.User)
return u, ok
c, _ := callerFromContext(ctx)
return c.AsHuman()
}
// serviceAccountPrincipal is a service account as resolved from its key:
// enough to authorize requests, never the key itself.
type serviceAccountPrincipal struct {
id int64
name string
scope string
teamID int64 // meaningless (zero) for instance scope
}
// serviceAccountFor resolves a service-account key to its account and stamps
// its last use, the same shape apiKeyUser has for a user's own key.
func serviceAccountFor(ctx context.Context, db *sql.DB, token string) (serviceAccountPrincipal, bool) {
var sa serviceAccountPrincipal
var keyID int64
var teamID sql.NullInt64
err := db.QueryRowContext(ctx, `
SELECT k.id, a.id, a.name, a.scope, a.team_id
FROM service_account_keys k
JOIN service_accounts a ON a.id = k.service_account_id
WHERE k.key_hash = $1`, hashToken(token),
).Scan(&keyID, &sa.id, &sa.name, &sa.scope, &teamID)
if err != nil {
return serviceAccountPrincipal{}, false
}
if teamID.Valid {
sa.teamID = teamID.Int64
}
// best-effort; don't fail the request if this update fails
db.ExecContext(ctx,
"UPDATE service_account_keys SET last_used_at = $1 WHERE id = $2",
time.Now().Unix(), keyID)
return sa, true
}
// serveAsServiceAccount hands the request on with a service account's
// identity in context. A team-scoped account gets a single synthetic
// membership — owner of its own team, nothing else — which is what makes it
// satisfy requireTeamMember/requireTeamOwner exactly as a real owner would,
// without teaching either function about a second kind of caller. An
// instance-scoped account gets no memberships at all: it acts on teams by id,
// not by belonging to one.
//
// No CSRF check, for the same reason an API key needs none: a service-account
// key is only ever set by the client that holds it, never attached by a
// browser to a request another site makes.
func serveAsServiceAccount(w http.ResponseWriter, r *http.Request, next http.Handler, sa serviceAccountPrincipal) {
var memberships []membership
if sa.scope == models.ServiceAccountScopeTeam {
memberships = []membership{{teamID: sa.teamID, role: models.RoleOwner}}
}
ctx := context.WithValue(r.Context(), ctxCaller, Caller{sa: &sa, memberships: memberships})
next.ServeHTTP(w, r.WithContext(ctx))
}
// isInstanceServiceAccount is a thin compatibility wrapper over
// Caller.IsInstanceServiceAccount(), for call sites outside this package's
// core predicates (handleCreateTeam, handleCreateServiceAccount) that
// needed this exact, narrow check before the Caller abstraction existed.
func isInstanceServiceAccount(ctx context.Context) bool {
c, _ := callerFromContext(ctx)
return c.IsInstanceServiceAccount()
}
// operatorReason marks a write that operator mode refused as such, distinct
// from every other 403 this server returns, so a client — the web UI or
// terdut-tui — can tell "you may not" from "this is managed elsewhere" and
// show the right message instead of a bare "forbidden".
const operatorReason = "operator_managed"
// OperatorModeBlock refuses a human write (session or a user's own API key)
// on a route it wraps, while letting a service account through. That is the
// whole point of operator mode: automation holding a service-account key
// (terdut-operator, most likely) keeps reconciling these resources, and a
// person in the web UI or terdut-tui gets a clear "edit this through your
// GitOps source instead" rather than a write that the next resync would only
// undo.
//
// Checked after AuthMiddleware, the same way AdminOnly is: by the time a
// request reaches here the caller is already known to be a service account
// or not. A router that never enables operator mode pays nothing for this —
// it hands back next unchanged rather than wrapping it in a check that would
// always pass.
func OperatorModeBlock(cfg config.Config) func(http.Handler) http.Handler {
return func(next http.Handler) http.Handler {
if !cfg.OperatorMode {
return next
}
return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
caller, _ := callerFromContext(r.Context())
if _, ok := caller.ServiceAccountID(); ok {
next.ServeHTTP(w, r)
return
}
respond(w, http.StatusForbidden, map[string]string{
"error": "this server is in operator mode; edit this through your GitOps source instead of the web UI or API",
"reason": operatorReason,
})
})
}
}
// membership is the caller's role in one team.
@@ -218,24 +335,15 @@ func callerMemberships(ctx context.Context, db *sql.DB, userID int64) ([]members
// administration is about accounts, not about reading other people's incidents,
// and an admin who needs to see a team's queue can add themselves to it.
func callerTeamIDs(ctx context.Context) []int64 {
ms, _ := ctx.Value(ctxTeams).([]membership)
ids := make([]int64, 0, len(ms))
for _, m := range ms {
ids = append(ids, m.teamID)
}
return ids
c, _ := callerFromContext(ctx)
return c.TeamIDs()
}
// callerRole reports the caller's role in one team, and whether they are in it
// at all.
func callerRole(ctx context.Context, teamID int64) (string, bool) {
ms, _ := ctx.Value(ctxTeams).([]membership)
for _, m := range ms {
if m.teamID == teamID {
return m.role, true
}
}
return "", false
c, _ := callerFromContext(ctx)
return c.Role(teamID)
}
// requireTeamMember answers the request and reports false unless the caller
+19 -2
View File
@@ -38,6 +38,13 @@ const (
ackTokenTTL = 24 * time.Hour
)
// notifierLockKey is the Postgres advisory lock the notifier takes for the
// duration of each pass, so that running more than one replica does not
// deliver (or double-deliver) the same notification from more than one of
// them at once. Its value has no meaning beyond being distinct from
// archiverLockKey.
const notifierLockKey int64 = 7265_0002
// Notification kinds, recording why a push was sent.
const (
notifyTriggered = "triggered"
@@ -97,6 +104,10 @@ var notifyClient = &http.Client{Timeout: 10 * time.Second}
// StartNotifier delivers queued notifications until ctx is cancelled, starting
// with an immediate pass so a restart flushes whatever the last one left behind.
//
// Each pass runs under notifierLockKey (see withAdvisoryLock), so that on more
// than one replica only whichever instance's tick takes the lock first actually
// delivers; the rest skip that tick rather than racing the same pass.
func StartNotifier(ctx context.Context, db *sql.DB, cfg NotifyConfig) {
if !cfg.enabled() {
log.Print("notifier: disabled (no ntfy URL configured)")
@@ -107,11 +118,17 @@ func StartNotifier(ctx context.Context, db *sql.DB, cfg NotifyConfig) {
ticker := time.NewTicker(notifyInterval)
defer ticker.Stop()
NotifySweep(ctx, db, cfg)
sweep := func() {
withAdvisoryLock(ctx, db, notifierLockKey, "notifier", func() {
NotifySweep(ctx, db, cfg)
})
}
sweep()
for {
select {
case <-ticker.C:
NotifySweep(ctx, db, cfg)
sweep()
case <-ctx.Done():
return
}
+11 -4
View File
@@ -66,12 +66,19 @@ func handleNotifyAck(db *sql.DB) http.HandlerFunc {
return
}
if !acked {
// The incident closed between the page and the tap. Nothing to do,
// and nothing the responder did wrong — report the state, not an error,
// so ntfy shows a success toast rather than a failure.
// Either the incident closed between the page and the tap, or it was
// already acknowledged (e.g. from the web UI, or an earlier tap of
// the same button) — either way nothing the responder did wrong, so
// report the actual state rather than assuming "resolved", and let
// ntfy show a success toast rather than a failure.
inc, err := fetchIncident(r.Context(), db, incidentID)
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
respond(w, http.StatusOK, map[string]any{
"incident_id": incidentID,
"status": "resolved",
"status": inc.Status,
})
return
}
+41
View File
@@ -408,6 +408,47 @@ func TestNotify_AckButtonAcknowledgesIncident(t *testing.T) {
}
}
// The ack token isn't single-use (it stays valid for a day, in case the
// first tap never reaches the server), so tapping the same notification's
// Acknowledge button twice is a real scenario, not just a retried request.
// It must report the incident's actual state, not assume "resolved" —
// see handleNotifyAck's !acked branch — and must not log a second
// "acknowledged" event.
func TestNotify_AckButtonTwiceIsIdempotent(t *testing.T) {
s, f := notifyTS(t, api.NotifyConfig{PublicURL: "https://terdut.example.com"})
fireCritical(t, s)
s.sweepNotify(t)
ackURL := f.messages()[0].Actions[0].URL
path := ackURL[strings.Index(ackURL, "/api/notify/ack/"):]
for i := range 2 {
resp, err := http.Post(s.URL+path, "application/json", nil)
if err != nil {
t.Fatalf("ack %d: %v", i+1, err)
}
var body map[string]any
decode(t, resp, &body)
if resp.StatusCode != http.StatusOK {
t.Fatalf("ack %d returned %d", i+1, resp.StatusCode)
}
if body["status"] != "acknowledged" {
t.Errorf("ack %d: expected status acknowledged, got %v", i+1, body["status"])
}
}
n := 0
for _, ty := range eventTypes(timeline(t, s, 1)) {
if ty == "acknowledged" {
n++
}
}
if n != 1 {
t.Errorf("expected exactly one acknowledged event after two taps, got %d", n)
}
}
func TestNotify_AckRejectsUnknownToken(t *testing.T) {
s, _ := notifyTS(t, api.NotifyConfig{PublicURL: "https://terdut.example.com"})
fireCritical(t, s)
+35 -1
View File
@@ -48,6 +48,17 @@ const (
ssoNoEmail ssoError = "no_email" // the provider sent no email address
ssoEmailConflict ssoError = "email_conflict" // a local account has this email and cannot be linked
ssoDisabled ssoError = "disabled" // the linked account is disabled
// ssoNotBootstrapped: this identity has no existing account, and no user
// exists on this install yet either -- creating one here would race
// POST /api/bootstrap for the one gitops-managed installs expect to win
// it (terdut-operator's own DESIGN.md §1, §6), which has no way to
// recover if it loses. The person sees this for at most as long as it
// takes whatever is bootstrapping this install to finish; signing in
// again afterward hits the ordinary first-sign-in path. Found by
// terdut-operator#1: nothing stopped an otherwise-ordinary OIDC sign-in
// from quietly winning this race against an operator that assumed it
// was the only caller.
ssoNotBootstrapped ssoError = "not_bootstrapped"
)
// handleAuthConfig says how this server can be signed in to, so the login form
@@ -66,9 +77,18 @@ func handleAuthConfig(cfg config.Config) http.HandlerFunc {
// DeviceLogin is whether a client that cannot open a browser (the TUI)
// can sign in by showing a code, through /api/oidc/device.
DeviceLogin bool `json:"device_login"`
// OperatorMode is whether this install is gitops-managed: writes to
// teams, escalation policies, dead man's switches and integrations
// from a session or a user's own API key are refused (OperatorModeBlock),
// though a service account's are not. The web UI reads this before
// anybody signs in, the same way it reads PasswordLogin/OIDC, so it can
// show those sections read-only from the start rather than only after
// a write fails.
OperatorMode bool `json:"operator_mode"`
}
return func(w http.ResponseWriter, r *http.Request) {
resp := response{PasswordLogin: !cfg.DisablePasswordLogin}
resp := response{PasswordLogin: !cfg.DisablePasswordLogin, OperatorMode: cfg.OperatorMode}
if cfg.OIDC.Enabled() {
resp.OIDC = oidcInfo{Enabled: true, Name: cfg.OIDC.Name}
resp.DeviceLogin = true
@@ -357,6 +377,20 @@ func resolveSSOUser(ctx context.Context, tx *sql.Tx, cfg config.OIDC, id *oidc.I
return 0, ssoEmailConflict
}
case errors.Is(err, sql.ErrNoRows):
// Creating the very first user is /api/bootstrap's own job (same
// gate, same table: SELECT COUNT(*) FROM users in handleBootstrap).
// An identity nobody has linked yet, on an install with no users at
// all, is exactly the race terdut-operator#1 found: whoever gets
// here first wins a slot the other side has no way to recover from
// losing. Refusing it here costs an otherwise-ordinary sign-in
// nothing but a retry once bootstrap has actually run.
var userCount int
if err := tx.QueryRowContext(ctx, "SELECT COUNT(*) FROM users").Scan(&userCount); err != nil {
return 0, err
}
if userCount == 0 {
return 0, ssoNotBootstrapped
}
userID, err = createSSOUser(ctx, tx, id)
if err != nil {
return 0, err
+39
View File
@@ -543,6 +543,45 @@ func TestSSO_DisabledUserIsRefused(t *testing.T) {
}
}
// TestSSO_FirstUserIsRefusedUntilBootstrap is terdut-operator#1: an
// otherwise-ordinary OIDC sign-in against a brand-new, not-yet-bootstrapped
// install must not be allowed to create the first user and win the race
// POST /api/bootstrap expects to win uncontested. Built directly over
// api.NewRouter rather than newSSOTS/newTS, both of which bootstrap before
// a test body ever runs -- exactly the state this test needs to not have yet.
func TestSSO_FirstUserIsRefusedUntilBootstrap(t *testing.T) {
idp := newFakeIdP(t)
database := newTestDB(t)
srv := httptest.NewServer(api.NewRouter(database, api.NotifyConfig{PublicURL: "http://terdut.test"}, ssoConfig(idp), "test"))
t.Cleanup(srv.Close)
first := newBrowser(t, srv.URL)
first.CheckRedirect = func(*http.Request, []*http.Request) error { return http.ErrUseLastResponse }
if loc := signInSSO(t, idp, first, alice); loc != "/?sso_error=not_bootstrapped" {
t.Fatalf("sent to %q, want not_bootstrapped", loc)
}
// Bootstrap the install for real, the way terdut-operator's own
// reconcileBootstrap does.
resp, err := http.Post(srv.URL+"/api/bootstrap", "application/json",
strings.NewReader(`{"username":"admin","email":"admin@test.com"}`))
if err != nil {
t.Fatal(err)
}
resp.Body.Close()
if resp.StatusCode != http.StatusCreated {
t.Fatalf("bootstrap: %d", resp.StatusCode)
}
// The same identity, signing in again, is this install's ordinary first
// SSO user now -- no longer refused.
second := newBrowser(t, srv.URL)
second.CheckRedirect = func(*http.Request, []*http.Request) error { return http.ErrUseLastResponse }
if loc := signInSSO(t, idp, second, alice); loc != "/" {
t.Errorf("sent to %q after bootstrap, want success", loc)
}
}
func TestSSO_NoEmailIsRefused(t *testing.T) {
idp := newFakeIdP(t)
s := newSSOTS(t, idp)
+38 -13
View File
@@ -14,8 +14,12 @@ import (
// NewRouter builds the HTTP surface. notify is passed through to the webhook,
// the only handler that has to decide where a new incident's page goes; a zero
// notify disables notifications. Dead man's switches are per team and read from
// the database, so nothing about them is wired in here.
func NewRouter(db *sql.DB, notify NotifyConfig, cfg config.Config) http.Handler {
// the database, so nothing about them is wired in here. version is reported
// verbatim by GET /api/version, unauthenticated like /healthz: a client
// deciding whether it can talk to this server — terdut-tui, terdut-operator —
// needs to ask before it holds a credential for it, and the version is not a
// secret.
func NewRouter(db *sql.DB, notify NotifyConfig, cfg config.Config, version string) http.Handler {
// One limiter each, both process-wide for the life of the router: login
// counts failed passwords, sign-up counts account creation, and mixing the
// two would let a burst of sign-ups lock somebody out of logging in.
@@ -30,6 +34,9 @@ func NewRouter(db *sql.DB, notify NotifyConfig, cfg config.Config) http.Handler
r.Get("/healthz", func(w http.ResponseWriter, r *http.Request) {
respond(w, http.StatusOK, map[string]string{"status": "ok"})
})
r.Get("/api/version", func(w http.ResponseWriter, r *http.Request) {
respond(w, http.StatusOK, map[string]string{"version": version})
})
// Unauthenticated: bootstrap, the Alertmanager webhook receiver, and the
// Acknowledge button in a push notification. The last one is authorised by
@@ -151,11 +158,26 @@ func NewRouter(db *sql.DB, notify NotifyConfig, cfg config.Config) http.Handler
r.Post("/api/incidents/{id}/notes", handleCreateNote(db))
r.Delete("/api/incidents/{id}/notes/{eventID}", handleDeleteNote(db))
// Service accounts: a scoped, non-human credential for automation
// (terdut-operator, most likely) that needs to manage the resources
// below without impersonating a human user. See SERVICE-ACCOUNTS.md.
r.Get("/api/service-accounts", handleListServiceAccounts(db))
r.Post("/api/service-accounts", handleCreateServiceAccount(db))
r.Post("/api/service-accounts/{id}/keys", handleCreateServiceAccountKey(db))
r.Delete("/api/service-accounts/{id}/keys/{keyID}", handleDeleteServiceAccountKey(db))
// Operator mode (TERDUT_OPERATOR_MODE) makes every write below refuse a
// human caller (a session or a user's own API key) while still letting
// a service account through — see OperatorModeBlock. opMode is a no-op
// wrapper when the flag is off, so this costs nothing on a server that
// never sets it.
opMode := OperatorModeBlock(cfg)
// Teams. A user sees the teams they belong to; an owner configures one.
r.Get("/api/teams", handleListTeams(db))
r.Post("/api/teams", handleCreateTeam(db))
r.Put("/api/teams/{teamID}", handleRenameTeam(db))
r.Delete("/api/teams/{teamID}", handleDeleteTeam(db))
r.With(opMode).Post("/api/teams", handleCreateTeam(db))
r.With(opMode).Put("/api/teams/{teamID}", handleRenameTeam(db))
r.With(opMode).Delete("/api/teams/{teamID}", handleDeleteTeam(db))
r.Get("/api/teams/{teamID}/members", handleListTeamMembers(db))
r.Post("/api/teams/{teamID}/members", handleAddTeamMember(db))
r.Delete("/api/teams/{teamID}/members/{userID}", handleRemoveTeamMember(db))
@@ -163,28 +185,31 @@ func NewRouter(db *sql.DB, notify NotifyConfig, cfg config.Config) http.Handler
// A team's own OIDC group binding: which provider groups grant member
// and owner access to it.
r.Get("/api/teams/{teamID}/oidc-groups", handleGetTeamOIDCGroups(db))
r.Put("/api/teams/{teamID}/oidc-groups", handleSetTeamOIDCGroups(db))
r.With(opMode).Put("/api/teams/{teamID}/oidc-groups", handleSetTeamOIDCGroups(db))
// Invite links into this team.
// Invite links into this team. Not operator-mode-gated: membership is
// deliberately never gitops-managed (see terdut-operator's DESIGN.md
// §4.2), so it stays editable regardless of this flag.
r.Get("/api/teams/{teamID}/invites", handleListInvites(db))
r.Post("/api/teams/{teamID}/invites", handleCreateInvite(db, notify.PublicURL))
r.Delete("/api/teams/{teamID}/invites/{inviteID}", handleRevokeInvite(db))
// A team's escalation ladder: who is paged when nobody answers.
r.Get("/api/teams/{teamID}/escalation", handleGetEscalation(db))
r.Put("/api/teams/{teamID}/escalation", handleSetEscalation(db))
r.With(opMode).Put("/api/teams/{teamID}/escalation", handleSetEscalation(db))
// A team's own dead man's switches: which of its alerts are heartbeats,
// and how long a silence has to last before somebody is paged.
r.Get("/api/teams/{teamID}/deadman/switches", handleListTeamDeadman(db))
r.Post("/api/teams/{teamID}/deadman/switches", handleCreateTeamDeadman(db))
r.Delete("/api/teams/{teamID}/deadman/switches/{switchID}", handleDeleteTeamDeadman(db))
r.With(opMode).Post("/api/teams/{teamID}/deadman/switches", handleCreateTeamDeadman(db))
r.With(opMode).Put("/api/teams/{teamID}/deadman/switches/{switchID}", handleUpdateTeamDeadman(db))
r.With(opMode).Delete("/api/teams/{teamID}/deadman/switches/{switchID}", handleDeleteTeamDeadman(db))
// Integrations: where a team's alerts come in, and the key that says so.
r.Get("/api/teams/{teamID}/integrations", handleListIntegrations(db))
r.Post("/api/teams/{teamID}/integrations", handleCreateIntegration(db, notify.PublicURL))
r.Patch("/api/teams/{teamID}/integrations/{integrationID}", handleRenameIntegration(db))
r.Delete("/api/teams/{teamID}/integrations/{integrationID}", handleDeleteIntegration(db))
r.With(opMode).Post("/api/teams/{teamID}/integrations", handleCreateIntegration(db, notify.PublicURL))
r.With(opMode).Patch("/api/teams/{teamID}/integrations/{integrationID}", handleRenameIntegration(db))
r.With(opMode).Delete("/api/teams/{teamID}/integrations/{integrationID}", handleDeleteIntegration(db))
// The rota is per team. /api/schedule/current is the exception: it
// answers across every team the caller is in, which is what somebody on
+350
View File
@@ -0,0 +1,350 @@
package api
import (
"context"
"database/sql"
"errors"
"net/http"
"strconv"
"strings"
"time"
"git.ryuvia.com/niklas/terdut-server/internal/models"
"github.com/go-chi/chi/v5"
)
// serviceAccountKeyPrefix marks a service-account key visibly, in logs and at
// a glance, distinct from a user's own personal API key. It carries no
// meaning to the server itself — the hash is looked up the same way either
// kind of key is — it exists entirely for whoever is reading a log line or an
// audit trail.
const serviceAccountKeyPrefix = "tdsa_"
// randomServiceAccountToken is randomToken with serviceAccountKeyPrefix on the
// raw value, hashed as a whole: the prefix is not a fixed header stripped
// before hashing, it is part of the secret, the same as if it had been
// generated that long to begin with.
func randomServiceAccountToken() (raw, hash string, err error) {
body, _, err := randomToken()
if err != nil {
return "", "", err
}
raw = serviceAccountKeyPrefix + body
return raw, hashToken(raw), nil
}
// callerIsAdmin reports whether the caller is a signed-in human system
// administrator. A service account never is, by design (SERVICE-ACCOUNTS.md):
// account and user management stays human-only, service accounts included.
func callerIsAdmin(ctx context.Context) bool {
u, ok := userFromContext(ctx)
return ok && u.IsAdmin
}
// callerOwnsTeam reports whether the caller is owner-equivalent for teamID:
// a human owner, or that team's own team-scoped service account (its single
// synthetic membership, serveAsServiceAccount — ratified in
// SERVICE-ACCOUNTS.md as intentional, not an accident: a team-scoped
// credential is that team's owner's reach, full stop, membership and
// invites included). Built on callerRole like requireTeamOwner, but without
// writing a response: callers here need to combine it with other ways of
// being allowed, not stop at the first no.
func callerOwnsTeam(ctx context.Context, teamID int64) bool {
role, ok := callerRole(ctx, teamID)
return ok && role == models.RoleOwner
}
// handleCreateServiceAccount creates a service account and mints its first
// key. Who may do this depends on scope: an instance-scoped account (which
// can in turn create a team and a team-scoped account for it) is system
// administration's own reach extended to automation, so only a human admin
// grants one. A team-scoped account is that team's owner's reach, so a human
// admin, the target team's own human owner, or an existing instance-scoped
// service account (minting itself a narrower credential for a team it just
// created) may create one.
func handleCreateServiceAccount(db *sql.DB) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
var req struct {
Name string `json:"name"`
Scope string `json:"scope"`
TeamID int64 `json:"team_id"`
}
if err := decodeJSON(r, &req); err != nil {
respond(w, http.StatusBadRequest, errResp("invalid request body"))
return
}
req.Name = strings.TrimSpace(req.Name)
if req.Name == "" {
respond(w, http.StatusBadRequest, errResp("name is required"))
return
}
if req.Scope != models.ServiceAccountScopeInstance && req.Scope != models.ServiceAccountScopeTeam {
respond(w, http.StatusBadRequest, errResp("scope must be instance or team"))
return
}
if req.Scope == models.ServiceAccountScopeTeam && req.TeamID == 0 {
respond(w, http.StatusBadRequest, errResp("team_id is required for a team-scoped account"))
return
}
if req.Scope == models.ServiceAccountScopeInstance && req.TeamID != 0 {
respond(w, http.StatusBadRequest, errResp("team_id must not be set for an instance-scoped account"))
return
}
allowed := callerIsAdmin(r.Context())
if !allowed && req.Scope == models.ServiceAccountScopeTeam {
allowed = callerOwnsTeam(r.Context(), req.TeamID) || isInstanceServiceAccount(r.Context())
}
if !allowed {
respond(w, http.StatusForbidden, errResp("team owner, system administrator, or instance-scoped service account access required"))
return
}
var callerUserID *int64
if u, ok := userFromContext(r.Context()); ok {
id := u.ID
callerUserID = &id
}
var teamID *int64
if req.Scope == models.ServiceAccountScopeTeam {
teamID = &req.TeamID
}
var sa models.ServiceAccount
var created int64
if err := db.QueryRowContext(r.Context(), `
INSERT INTO service_accounts (name, scope, team_id, created_by)
VALUES ($1, $2, $3, $4)
RETURNING id, name, scope, team_id, created_by, created_at`,
req.Name, req.Scope, teamID, callerUserID,
).Scan(&sa.ID, &sa.Name, &sa.Scope, &sa.TeamID, &sa.CreatedBy, &created); err != nil {
if isUniqueViolation(err) {
respond(w, http.StatusConflict, errResp("a service account with that name already exists"))
return
}
// The only foreign key that can fail here is team_id: an
// instance-scoped caller is not otherwise checked against it
// (callerOwnsTeam already proved it exists for a human owner).
respond(w, http.StatusBadRequest, errResp("unknown team_id"))
return
}
sa.CreatedAt = time.Unix(created, 0).UTC()
key, err := mintServiceAccountKey(r.Context(), db, sa.ID, "initial")
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
respond(w, http.StatusCreated, map[string]any{"service_account": sa, "key": key})
}
}
// mintServiceAccountKey inserts one key for an existing account and returns
// it with its raw value populated — the one moment that value exists outside
// the request that generated it.
func mintServiceAccountKey(ctx context.Context, db *sql.DB, serviceAccountID int64, name string) (models.ServiceAccountKey, error) {
raw, hash, err := randomServiceAccountToken()
if err != nil {
return models.ServiceAccountKey{}, err
}
var key models.ServiceAccountKey
var created int64
if err := db.QueryRowContext(ctx, `
INSERT INTO service_account_keys (service_account_id, key_hash, name)
VALUES ($1, $2, $3)
RETURNING id, service_account_id, name, created_at`,
serviceAccountID, hash, name,
).Scan(&key.ID, &key.ServiceAccountID, &key.Name, &created); err != nil {
return models.ServiceAccountKey{}, err
}
key.CreatedAt = time.Unix(created, 0).UTC()
key.Key = raw
return key, nil
}
func fetchServiceAccount(ctx context.Context, db *sql.DB, id int64) (models.ServiceAccount, error) {
var sa models.ServiceAccount
var created int64
err := db.QueryRowContext(ctx,
"SELECT id, name, scope, team_id, created_by, created_at FROM service_accounts WHERE id = $1", id,
).Scan(&sa.ID, &sa.Name, &sa.Scope, &sa.TeamID, &sa.CreatedBy, &created)
if err != nil {
return sa, err
}
sa.CreatedAt = time.Unix(created, 0).UTC()
return sa, nil
}
// callerMayManageServiceAccount reports whether the caller may mint or revoke
// a key on sa: a system administrator, that team-scoped account's own human
// owner, the account rotating its own credential (not a privilege
// escalation, the same reasoning requireSelfOrAdmin already rests on for a
// user's own API keys) — or, new, an instance-scoped service account
// managing any team-scoped account.
//
// That last branch closes terdut-operator#3: handleCreateServiceAccount
// already lets an instance-scoped caller *create* a team-scoped account for
// any team (the branch below it, isInstanceServiceAccount(ctx)) — this
// account didn't have an equivalent reach to *adopt or rotate* one it
// didn't just create in the same call, which is exactly the recovery path
// terdut-operator's own documented crash-window handling depends on
// (DESIGN.md §5's general adopt-on-conflict rule): a reconcile that creates
// the account successfully but crashes before persisting its credential
// locally retries into a 409, and without this branch the only available
// recovery — minting a fresh key on the now-existing account — 403'd
// forever, with no way out. Granting it here is not a new power: it
// mirrors the create-time reach this scope already has, just extended to
// the retry path DESIGN.md's own crash-window reasoning requires.
func callerMayManageServiceAccount(ctx context.Context, sa models.ServiceAccount) bool {
if callerIsAdmin(ctx) {
return true
}
if sa.TeamID != nil && callerOwnsTeam(ctx, *sa.TeamID) {
return true
}
caller, _ := callerFromContext(ctx)
if id, ok := caller.ServiceAccountID(); ok && id == sa.ID {
return true
}
if sa.TeamID != nil && caller.IsInstanceServiceAccount() {
return true
}
return false
}
func serviceAccountParam(w http.ResponseWriter, r *http.Request) (int64, bool) {
id, err := strconv.ParseInt(chi.URLParam(r, "id"), 10, 64)
if err != nil {
respond(w, http.StatusBadRequest, errResp("invalid service account id"))
return 0, false
}
return id, true
}
func handleCreateServiceAccountKey(db *sql.DB) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
id, ok := serviceAccountParam(w, r)
if !ok {
return
}
sa, err := fetchServiceAccount(r.Context(), db, id)
if errors.Is(err, sql.ErrNoRows) {
respond(w, http.StatusNotFound, errResp("service account not found"))
return
}
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
if !callerMayManageServiceAccount(r.Context(), sa) {
respond(w, http.StatusForbidden, errResp("team owner, system administrator, or the account itself may rotate its key"))
return
}
var req struct {
Name string `json:"name"`
}
if err := decodeJSON(r, &req); err != nil {
respond(w, http.StatusBadRequest, errResp("invalid request body"))
return
}
if req.Name == "" {
respond(w, http.StatusBadRequest, errResp("name is required"))
return
}
key, err := mintServiceAccountKey(r.Context(), db, sa.ID, req.Name)
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
respond(w, http.StatusCreated, key)
}
}
func handleDeleteServiceAccountKey(db *sql.DB) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
id, ok := serviceAccountParam(w, r)
if !ok {
return
}
sa, err := fetchServiceAccount(r.Context(), db, id)
if errors.Is(err, sql.ErrNoRows) {
respond(w, http.StatusNotFound, errResp("service account not found"))
return
}
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
if !callerMayManageServiceAccount(r.Context(), sa) {
respond(w, http.StatusForbidden, errResp("team owner, system administrator, or the account itself may revoke its key"))
return
}
keyID, err := strconv.ParseInt(chi.URLParam(r, "keyID"), 10, 64)
if err != nil {
respond(w, http.StatusBadRequest, errResp("invalid key id"))
return
}
res, err := db.ExecContext(r.Context(),
"DELETE FROM service_account_keys WHERE id = $1 AND service_account_id = $2", keyID, sa.ID)
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
if n, _ := res.RowsAffected(); n == 0 {
respond(w, http.StatusNotFound, errResp("key not found"))
return
}
w.WriteHeader(http.StatusNoContent)
}
}
// handleListServiceAccounts lists every service account, or looks one up by
// its exact name with ?name=. The name lookup is open to any authenticated
// caller, human or service account: it returns no key material, and it is
// what lets a service account find its own account on the 403 that follows a
// second POST — the self-registration pattern SERVICE-ACCOUNTS.md describes.
// Listing everything, with no filter, stays administrator-only.
func handleListServiceAccounts(db *sql.DB) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
name := strings.TrimSpace(r.URL.Query().Get("name"))
if name == "" && !callerIsAdmin(r.Context()) {
respond(w, http.StatusForbidden, errResp("administrator access required to list every service account; pass ?name= to look up one by name"))
return
}
query := "SELECT id, name, scope, team_id, created_by, created_at FROM service_accounts"
var args []any
if name != "" {
query += " WHERE name = $1"
args = append(args, name)
}
query += " ORDER BY id"
rows, err := db.QueryContext(r.Context(), query, args...)
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
defer rows.Close()
accounts := []models.ServiceAccount{}
for rows.Next() {
var sa models.ServiceAccount
var created int64
if err := rows.Scan(&sa.ID, &sa.Name, &sa.Scope, &sa.TeamID, &sa.CreatedBy, &created); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
sa.CreatedAt = time.Unix(created, 0).UTC()
accounts = append(accounts, sa)
}
if err := rows.Err(); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
respond(w, http.StatusOK, accounts)
}
}
+492
View File
@@ -0,0 +1,492 @@
package api_test
import (
"bytes"
"encoding/json"
"io"
"net/http"
"net/http/httptest"
"strings"
"testing"
"git.ryuvia.com/niklas/terdut-server/internal/api"
"git.ryuvia.com/niklas/terdut-server/internal/models"
)
// reqAs is s.req with an arbitrary bearer credential in place of the admin's
// own key, for exercising a service account's or another user's key.
func (s *ts) reqAs(t *testing.T, key, method, path string, body any) *http.Response {
t.Helper()
var r io.Reader
if body != nil {
data, _ := json.Marshal(body)
r = bytes.NewReader(data)
}
req, _ := http.NewRequest(method, s.URL+path, r)
req.Header.Set("Authorization", "Bearer "+key)
if body != nil {
req.Header.Set("Content-Type", "application/json")
}
resp, err := http.DefaultClient.Do(req)
if err != nil {
t.Fatalf("%s %s: %v", method, path, err)
}
return resp
}
// createServiceAccount creates a service account as callerKey and returns its
// freshly minted raw key.
func createServiceAccount(t *testing.T, s *ts, callerKey, name, scope string, teamID int64) string {
t.Helper()
body := map[string]any{"name": name, "scope": scope}
if teamID != 0 {
body["team_id"] = teamID
}
resp := s.reqAs(t, callerKey, http.MethodPost, "/api/service-accounts", body)
if resp.StatusCode != http.StatusCreated {
resp.Body.Close()
t.Fatalf("create service account %s: %d", name, resp.StatusCode)
}
var result struct {
Key struct {
Key string `json:"key"`
} `json:"key"`
}
decode(t, resp, &result)
if result.Key.Key == "" {
t.Fatalf("create service account %s: no key returned", name)
}
return result.Key.Key
}
// createTeamAs creates a team as callerKey and returns its id.
func createTeamAs(t *testing.T, s *ts, callerKey, name string) int64 {
t.Helper()
resp := s.reqAs(t, callerKey, http.MethodPost, "/api/teams", map[string]string{"name": name})
if resp.StatusCode != http.StatusCreated {
resp.Body.Close()
t.Fatalf("create team %s: %d", name, resp.StatusCode)
}
var team struct {
ID int64 `json:"id"`
}
decode(t, resp, &team)
return team.ID
}
// ---------------------------------------------------------------------------
// Instance scope
// ---------------------------------------------------------------------------
func TestServiceAccount_InstanceScopeCreatesTeamWithNoHumanOwner(t *testing.T) {
s := newTS(t)
instanceKey := createServiceAccount(t, s, s.key, "terdut-operator", models.ServiceAccountScopeInstance, 0)
if !strings.HasPrefix(instanceKey, "tdsa_") {
t.Errorf("expected a service-account key to carry the tdsa_ prefix, got %q", instanceKey)
}
resp := s.reqAs(t, instanceKey, http.MethodPost, "/api/teams", map[string]string{"name": "provisioned"})
if resp.StatusCode != http.StatusCreated {
t.Fatalf("instance-scoped account creating a team: %d", resp.StatusCode)
}
var team struct {
ID int64 `json:"id"`
Role string `json:"role"`
}
decode(t, resp, &team)
if team.Role != "" {
t.Errorf("expected no role on a team a service account created (no human owner), got %q", team.Role)
}
// It still exists, visible to an administrator, even with no member.
var admin []map[string]any
decode(t, s.req(t, http.MethodGet, "/api/admin/teams", nil), &admin)
found := false
for _, tm := range admin {
if int64(tm["id"].(float64)) == team.ID {
found = true
}
}
if !found {
t.Errorf("expected the service-account-created team to appear in /api/admin/teams")
}
}
func TestServiceAccount_TeamScopeCannotCreateTeam(t *testing.T) {
s := newTS(t)
instanceKey := createServiceAccount(t, s, s.key, "terdut-operator", models.ServiceAccountScopeInstance, 0)
teamA := createTeamAs(t, s, instanceKey, "team-a")
keyA := createServiceAccount(t, s, instanceKey, "team-a-sa", models.ServiceAccountScopeTeam, teamA)
resp := s.reqAs(t, keyA, http.MethodPost, "/api/teams", map[string]string{"name": "should-fail"})
if resp.StatusCode != http.StatusForbidden {
t.Errorf("expected 403, a team-scoped account creating a team, got %d", resp.StatusCode)
}
resp.Body.Close()
}
// ---------------------------------------------------------------------------
// Team scope
// ---------------------------------------------------------------------------
// The whole point of team scope: bound to its own team, refused everywhere
// else, the same as an instance-scoped account minting a key per TerdutTeam
// rather than sharing one server-admin-equivalent credential would need.
func TestServiceAccount_TeamScopeIsBoundToItsOwnTeam(t *testing.T) {
s := newTS(t)
instanceKey := createServiceAccount(t, s, s.key, "terdut-operator", models.ServiceAccountScopeInstance, 0)
teamA := createTeamAs(t, s, instanceKey, "team-a")
teamB := createTeamAs(t, s, instanceKey, "team-b")
keyA := createServiceAccount(t, s, instanceKey, "team-a-sa", models.ServiceAccountScopeTeam, teamA)
policy := map[string]any{"repeat_count": 0, "fallback_topic": "", "levels": []any{}}
resp := s.reqAs(t, keyA, http.MethodPut, "/api/teams/"+id64(teamA)+"/escalation", policy)
if resp.StatusCode != http.StatusOK {
t.Fatalf("team-a's own key setting its escalation: %d", resp.StatusCode)
}
resp.Body.Close()
// 404, not 403: the same "does this exist" refusal a human non-member
// gets from requireTeamMember, not a distinguishable "you may not".
resp2 := s.reqAs(t, keyA, http.MethodPut, "/api/teams/"+id64(teamB)+"/escalation", policy)
if resp2.StatusCode != http.StatusNotFound {
t.Errorf("expected 404 reaching into another team, got %d", resp2.StatusCode)
}
resp2.Body.Close()
}
// Team scope is owner-equivalent broadly (SERVICE-ACCOUNTS.md), not limited to
// one endpoint: escalation, dead man's switches and integrations all work.
func TestServiceAccount_TeamScopeManagesItsResources(t *testing.T) {
s := newTS(t)
instanceKey := createServiceAccount(t, s, s.key, "terdut-operator", models.ServiceAccountScopeInstance, 0)
teamA := createTeamAs(t, s, instanceKey, "team-a")
keyA := createServiceAccount(t, s, instanceKey, "team-a-sa", models.ServiceAccountScopeTeam, teamA)
resp := s.reqAs(t, keyA, http.MethodPost, "/api/teams/"+id64(teamA)+"/deadman/switches",
map[string]any{"matcher": "alertname=Watchdog", "timeout_seconds": 900, "severity": "critical"})
if resp.StatusCode != http.StatusCreated {
t.Errorf("team-scoped account creating a dead man's switch: %d", resp.StatusCode)
}
resp.Body.Close()
resp2 := s.reqAs(t, keyA, http.MethodPost, "/api/teams/"+id64(teamA)+"/integrations",
map[string]string{"name": "prod"})
if resp2.StatusCode != http.StatusCreated {
t.Errorf("team-scoped account creating an integration: %d", resp2.StatusCode)
}
resp2.Body.Close()
}
// Documents the capability already granted at create time (handleCreateServiceAccount's
// own callerOwnsTeam branch) also applies here: a team-scoped account is that
// team's owner's reach, membership and further accounts included, not just
// the handful of endpoints exercised above. Kept, not restricted, for
// symmetry with the now-ratified membership/invite capability below.
func TestServiceAccount_TeamScopeCanMintAnotherAccountForItsOwnTeam(t *testing.T) {
s := newTS(t)
instanceKey := createServiceAccount(t, s, s.key, "terdut-operator", models.ServiceAccountScopeInstance, 0)
teamA := createTeamAs(t, s, instanceKey, "team-a")
keyA := createServiceAccount(t, s, instanceKey, "team-a-sa", models.ServiceAccountScopeTeam, teamA)
resp := s.reqAs(t, keyA, http.MethodPost, "/api/service-accounts",
map[string]any{"name": "team-a-sa-2", "scope": models.ServiceAccountScopeTeam, "team_id": teamA})
if resp.StatusCode != http.StatusCreated {
t.Errorf("team-scoped account minting another account for its own team: %d", resp.StatusCode)
}
resp.Body.Close()
}
// SERVICE-ACCOUNTS.md ratifies this explicitly: a team-scoped account is
// owner-equivalent for every requireTeamOwner endpoint, membership and
// invites included — terdut-operator's own invite-minting feature depends on
// exactly this. No test exercised handleCreateInvite from a service account
// before this change, and it would have 500'd (created_by written as a bare
// zero value against a NOT-validated-but-FK'd column) rather than succeeded;
// see the signup_test.go addition for that half.
func TestServiceAccount_TeamScopeManagesItsOwnInvites(t *testing.T) {
s := newTS(t)
instanceKey := createServiceAccount(t, s, s.key, "terdut-operator", models.ServiceAccountScopeInstance, 0)
teamA := createTeamAs(t, s, instanceKey, "team-a")
keyA := createServiceAccount(t, s, instanceKey, "team-a-sa", models.ServiceAccountScopeTeam, teamA)
var invite struct {
ID int64 `json:"id"`
}
resp := s.reqAs(t, keyA, http.MethodPost, "/api/teams/"+id64(teamA)+"/invites", map[string]any{})
if resp.StatusCode != http.StatusCreated {
t.Fatalf("team-scoped account creating an invite: %d", resp.StatusCode)
}
decode(t, resp, &invite)
var list []map[string]any
decode(t, s.reqAs(t, keyA, http.MethodGet, "/api/teams/"+id64(teamA)+"/invites", nil), &list)
if len(list) != 1 {
t.Errorf("expected the invite to list back, got %d", len(list))
}
if resp := s.reqAs(t, keyA, http.MethodDelete,
"/api/teams/"+id64(teamA)+"/invites/"+id64(invite.ID), nil); resp.StatusCode != http.StatusNoContent {
t.Errorf("team-scoped account revoking its own invite: %d", resp.StatusCode)
} else {
resp.Body.Close()
}
}
// ---------------------------------------------------------------------------
// Key rotation
// ---------------------------------------------------------------------------
// terdut-operator#3: an instance-scoped account is already trusted to CREATE
// a team-scoped account for any team (handleCreateServiceAccount's own
// isInstanceServiceAccount branch) — this pins that it is equally trusted to
// manage/rotate a key on one that already exists and that it did not just
// create in this call, which is the exact shape of terdut-operator's own
// crash-window recovery (mint succeeds, a later step is interrupted before
// persisting the credential locally, and the next reconcile retries into a
// 409 then needs to mint a fresh key on the now-existing account). Before
// this fix, the second POST .../keys below 403'd forever.
func TestServiceAccount_InstanceScopeAdoptsAnExistingTeamScopedAccountsKey(t *testing.T) {
s := newTS(t)
instanceKey := createServiceAccount(t, s, s.key, "terdut-operator", models.ServiceAccountScopeInstance, 0)
teamA := createTeamAs(t, s, instanceKey, "team-a")
resp := s.reqAs(t, instanceKey, http.MethodPost, "/api/service-accounts",
map[string]any{"name": "team-a-sa", "scope": models.ServiceAccountScopeTeam, "team_id": teamA})
if resp.StatusCode != http.StatusCreated {
t.Fatalf("create team-scoped account: %d", resp.StatusCode)
}
var created struct {
ServiceAccount struct {
ID int64 `json:"id"`
} `json:"service_account"`
}
decode(t, resp, &created)
// Simulates the adopt-on-409 recovery path: this instance-scoped caller
// did not just create this account in this call (a fresh *tdclient.Client
// request, same as a second, independent reconcile would issue), yet
// still needs to mint it a fresh key.
rotateResp := s.reqAs(t, instanceKey, http.MethodPost,
"/api/service-accounts/"+id64(created.ServiceAccount.ID)+"/keys", map[string]string{"name": "adopted"})
if rotateResp.StatusCode != http.StatusCreated {
t.Errorf("instance-scoped account adopting a team-scoped account's key: %d", rotateResp.StatusCode)
}
rotateResp.Body.Close()
}
// ---------------------------------------------------------------------------
// AdminOnly / requireSelfOrAdmin — unchanged after the Caller refactor
// ---------------------------------------------------------------------------
// The Caller abstraction must not have widened AdminOnly/requireSelfOrAdmin:
// user management and /api/admin/settings stay human-only, for every scope
// of service account, exactly as before.
func TestAdminOnly_RefusesEveryServiceAccountScope(t *testing.T) {
s := newTS(t)
instanceKey := createServiceAccount(t, s, s.key, "terdut-operator", models.ServiceAccountScopeInstance, 0)
teamA := createTeamAs(t, s, instanceKey, "team-a")
teamKey := createServiceAccount(t, s, instanceKey, "team-a-sa", models.ServiceAccountScopeTeam, teamA)
for _, key := range []string{instanceKey, teamKey} {
if resp := s.reqAs(t, key, http.MethodGet, "/api/admin/settings", nil); resp.StatusCode != http.StatusForbidden {
t.Errorf("expected 403 for a service account reading /api/admin/settings, got %d", resp.StatusCode)
} else {
resp.Body.Close()
}
if resp := s.reqAs(t, key, http.MethodPost, "/api/users",
map[string]string{"username": "nope", "email": "nope@example.com"}); resp.StatusCode != http.StatusForbidden {
t.Errorf("expected 403 for a service account creating a user, got %d", resp.StatusCode)
} else {
resp.Body.Close()
}
if resp := s.reqAs(t, key, http.MethodGet, "/api/me", nil); resp.StatusCode != http.StatusForbidden {
t.Errorf("expected 403 for a service account calling /api/me, got %d", resp.StatusCode)
} else {
resp.Body.Close()
}
}
}
func TestServiceAccount_SelfRotatesItsOwnKey(t *testing.T) {
s := newTS(t)
instanceKey := createServiceAccount(t, s, s.key, "terdut-operator", models.ServiceAccountScopeInstance, 0)
// Self-lookup by name, the pattern that turns /api/bootstrap's 403 into a
// normal flow instead of an unhandled error.
var accounts []map[string]any
decode(t, s.reqAs(t, instanceKey, http.MethodGet, "/api/service-accounts?name=terdut-operator", nil), &accounts)
if len(accounts) != 1 {
t.Fatalf("expected exactly one match for ?name=terdut-operator, got %d", len(accounts))
}
id := int64(accounts[0]["id"].(float64))
resp := s.reqAs(t, instanceKey, http.MethodPost, "/api/service-accounts/"+id64(id)+"/keys",
map[string]string{"name": "rotated"})
if resp.StatusCode != http.StatusCreated {
t.Fatalf("self-rotation: %d", resp.StatusCode)
}
var newKey struct {
Key string `json:"key"`
}
decode(t, resp, &newKey)
if resp := s.reqAs(t, newKey.Key, http.MethodPost, "/api/teams", map[string]string{"name": "after-rotation"}); resp.StatusCode != http.StatusCreated {
t.Errorf("expected the newly rotated key to work, got %d", resp.StatusCode)
} else {
resp.Body.Close()
}
// Rotation adds a key, it does not itself revoke the old one.
if resp := s.reqAs(t, instanceKey, http.MethodGet, "/api/service-accounts?name=terdut-operator", nil); resp.StatusCode != http.StatusOK {
t.Errorf("expected the original key to still work until explicitly revoked, got %d", resp.StatusCode)
} else {
resp.Body.Close()
}
}
// ---------------------------------------------------------------------------
// Operator mode
// ---------------------------------------------------------------------------
// Operator mode is exercised against a second router over an
// already-configured database, rather than turning it on for newTSWith's own
// setup: that setup creates the default integration with the admin's (human)
// key, which is precisely the write operator mode exists to refuse, and in
// the real deployment this flag targets that setup was never done by a human
// to begin with — the operator itself would have provisioned it.
func TestOperatorMode_BlocksHumanWritesButAllowsServiceAccounts(t *testing.T) {
s := newTS(t)
conf := testConfig()
conf.OperatorMode = true
opSrv := httptest.NewServer(api.NewRouter(s.db, s.notify, conf, "test"))
t.Cleanup(opSrv.Close)
do := func(key, method, path string, body any) *http.Response {
t.Helper()
var r io.Reader
if body != nil {
data, _ := json.Marshal(body)
r = bytes.NewReader(data)
}
req, _ := http.NewRequest(method, opSrv.URL+path, r)
req.Header.Set("Authorization", "Bearer "+key)
if body != nil {
req.Header.Set("Content-Type", "application/json")
}
resp, err := http.DefaultClient.Do(req)
if err != nil {
t.Fatalf("%s %s: %v", method, path, err)
}
return resp
}
// The bootstrap admin's own key is a human credential: refused.
resp := do(s.key, http.MethodPost, "/api/teams", map[string]string{"name": "human-team"})
if resp.StatusCode != http.StatusForbidden {
t.Fatalf("expected 403 for a human write under operator mode, got %d", resp.StatusCode)
}
var refusal map[string]string
decode(t, resp, &refusal)
if refusal["reason"] != "operator_managed" {
t.Errorf("expected reason=operator_managed, got %q", refusal["reason"])
}
// Creating the service account itself is not gated by operator mode —
// it is how an operator identifies itself, not one of the resources it
// manages.
resp2 := do(s.key, http.MethodPost, "/api/service-accounts",
map[string]any{"name": "terdut-operator", "scope": models.ServiceAccountScopeInstance})
if resp2.StatusCode != http.StatusCreated {
t.Fatalf("create service account under operator mode: %d", resp2.StatusCode)
}
var result struct {
Key struct {
Key string `json:"key"`
} `json:"key"`
}
decode(t, resp2, &result)
resp3 := do(result.Key.Key, http.MethodPost, "/api/teams", map[string]string{"name": "operator-team"})
if resp3.StatusCode != http.StatusCreated {
t.Fatalf("expected 201 for a service-account write under operator mode, got %d", resp3.StatusCode)
}
resp3.Body.Close()
// Reads are unaffected regardless of caller.
if resp := do(s.key, http.MethodGet, "/api/teams", nil); resp.StatusCode != http.StatusOK {
t.Errorf("expected reads to stay open under operator mode, got %d", resp.StatusCode)
} else {
resp.Body.Close()
}
}
func TestOperatorMode_OffLeavesHumanWritesAlone(t *testing.T) {
s := newTS(t) // testConfig(): OperatorMode false
resp := s.req(t, http.MethodPost, "/api/teams", map[string]string{"name": "still-fine"})
if resp.StatusCode != http.StatusCreated {
t.Errorf("expected a human write to succeed with operator mode off, got %d", resp.StatusCode)
}
resp.Body.Close()
}
// ---------------------------------------------------------------------------
// Version
// ---------------------------------------------------------------------------
func TestVersion(t *testing.T) {
s := newTS(t)
resp, err := http.Get(s.URL + "/api/version")
if err != nil {
t.Fatalf("get version: %v", err)
}
var v struct {
Version string `json:"version"`
}
decode(t, resp, &v)
if v.Version != "test" {
t.Errorf("expected version %q, got %q", "test", v.Version)
}
}
// ---------------------------------------------------------------------------
// Dead man's switch update-in-place
// ---------------------------------------------------------------------------
func TestDeadman_UpdateInPlacePreservesID(t *testing.T) {
s := newTS(t)
var created struct {
ID int64 `json:"id"`
}
decode(t, s.req(t, http.MethodPost, "/api/teams/"+defaultTeam+"/deadman/switches",
map[string]any{"matcher": "alertname=Watchdog", "timeout_seconds": 900, "severity": "critical"}), &created)
resp := s.req(t, http.MethodPut, "/api/teams/"+defaultTeam+"/deadman/switches/"+id64(created.ID),
map[string]any{"name": "renamed", "matcher": "alertname=Watchdog", "timeout_seconds": 1200, "severity": "warning"})
if resp.StatusCode != http.StatusOK {
t.Fatalf("update switch: %d", resp.StatusCode)
}
var updated struct {
ID int64 `json:"id"`
Name string `json:"name"`
TimeoutSeconds int64 `json:"timeout_seconds"`
Severity string `json:"severity"`
}
decode(t, resp, &updated)
if updated.ID != created.ID {
t.Errorf("expected id to stay %d, got %d", created.ID, updated.ID)
}
if updated.Name != "renamed" || updated.TimeoutSeconds != 1200 || updated.Severity != "warning" {
t.Errorf("expected the update to apply, got %+v", updated)
}
var list []map[string]any
decode(t, s.req(t, http.MethodGet, "/api/teams/"+defaultTeam+"/deadman/switches", nil), &list)
if len(list) != 1 {
t.Errorf("expected the update to replace in place, not add a row, got %d switches", len(list))
}
}
+24 -4
View File
@@ -359,7 +359,19 @@ func handleCreateInvite(db *sql.DB, publicURL string) http.HandlerFunc {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
caller, _ := userFromContext(r.Context())
// created_by is nullable (ON DELETE SET NULL) for exactly this
// reason: the caller minting an invite is not always a human with a
// real users row. A team-scoped service account is owner-equivalent
// here (requireTeamOwner above already let it through), and this
// must leave created_by NULL for one the same way
// handleCreateServiceAccount already does for the analogous case —
// an unchecked zero value would violate the users(id) foreign key
// instead of recording "nobody" cleanly.
var createdBy *int64
if u, ok := userFromContext(r.Context()); ok {
id := u.ID
createdBy = &id
}
expires := time.Now().Add(inviteTTL)
var out inviteJSON
@@ -368,7 +380,7 @@ func handleCreateInvite(db *sql.DB, publicURL string) http.HandlerFunc {
INSERT INTO invites (token_hash, team_id, role, created_by, expires_at, max_uses)
VALUES ($1, $2, $3, $4, $5, $6)
RETURNING id, team_id, role, created_at, expires_at, max_uses, uses`,
hash, teamID, req.Role, caller.ID, expires.Unix(), req.MaxUses).
hash, teamID, req.Role, createdBy, expires.Unix(), req.MaxUses).
Scan(&out.ID, &out.TeamID, &out.Role, &created, &expiresAt, &out.MaxUses, &out.Uses); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
@@ -429,7 +441,11 @@ func handleTestNotification(cfg NotifyConfig, db *sql.DB) http.HandlerFunc {
errResp("this server has no ntfy configured, so it can send nothing"))
return
}
caller, _ := userFromContext(r.Context())
caller, ok := userFromContext(r.Context())
if !ok {
respond(w, http.StatusForbidden, errResp("this endpoint is for human accounts only"))
return
}
var topic *string
if err := db.QueryRowContext(r.Context(),
@@ -470,7 +486,11 @@ func handleDismissOnboarding(db *sql.DB) http.HandlerFunc {
respond(w, http.StatusBadRequest, errResp("dismissed is required"))
return
}
caller, _ := userFromContext(r.Context())
caller, ok := userFromContext(r.Context())
if !ok {
respond(w, http.StatusForbidden, errResp("this endpoint is for human accounts only"))
return
}
var err error
if *req.Dismissed {
+56
View File
@@ -2,10 +2,14 @@ package api_test
import (
"bytes"
"database/sql"
"encoding/json"
"net/http"
"net/http/cookiejar"
"testing"
"git.ryuvia.com/niklas/terdut-server/internal/api"
"git.ryuvia.com/niklas/terdut-server/internal/models"
)
// signup posts to the unauthenticated sign-up endpoint, the way the form does,
@@ -268,3 +272,55 @@ func TestSignup_ValidatesLikeTheRestOfTheServer(t *testing.T) {
t.Errorf("an existing username: expected 409, got %d", taken.StatusCode)
}
}
// A team-scoped service account has no users row to attribute created_by to.
// Before this fix, handleCreateInvite wrote its zero-value caller.ID straight
// into that (nullable, ON DELETE SET NULL) foreign key instead of leaving it
// NULL the way handleCreateServiceAccount already does for the same
// situation — a 500, not the 201 TestServiceAccount_TeamScopeManagesItsOwnInvites
// now confirms. This pins the column itself ends up NULL, not just "some
// response came back".
func TestSignup_InviteCreatedByAServiceAccountLeavesCreatedByNull(t *testing.T) {
s := newTS(t)
instanceKey := createServiceAccount(t, s, s.key, "terdut-operator", models.ServiceAccountScopeInstance, 0)
teamA := createTeamAs(t, s, instanceKey, "team-a")
keyA := createServiceAccount(t, s, instanceKey, "team-a-sa", models.ServiceAccountScopeTeam, teamA)
var created struct {
ID int64 `json:"id"`
}
decode(t, s.reqAs(t, keyA, http.MethodPost, "/api/teams/"+id64(teamA)+"/invites", map[string]any{}), &created)
var createdBy sql.NullInt64
if err := s.db.QueryRow("SELECT created_by FROM invites WHERE id = $1", created.ID).Scan(&createdBy); err != nil {
t.Fatalf("read back invites.created_by: %v", err)
}
if createdBy.Valid {
t.Errorf("expected created_by to be NULL for a service-account-minted invite, got %d", createdBy.Int64)
}
}
// Neither of these has a real user_id to act on behalf of; both must 403 a
// service account explicitly rather than 500 (handleTestNotification, which
// used to query ntfy_topic for user id 0) or silently no-op (handleDismissOnboarding,
// which used to UPDATE ... WHERE id = 0, affecting nothing and still
// returning 204).
func TestServiceAccount_HumanOnlyEndpointsRefuseExplicitly(t *testing.T) {
// BaseURL set (even to a fake, unreachable address) so handleTestNotification
// reaches its AsHuman() check instead of short-circuiting on "ntfy not
// configured" first — this test is about the human-only check, not ntfy.
s := newTS(t, api.NotifyConfig{BaseURL: "http://ntfy.invalid"})
instanceKey := createServiceAccount(t, s, s.key, "terdut-operator", models.ServiceAccountScopeInstance, 0)
if resp := s.reqAs(t, instanceKey, http.MethodPost, "/api/me/notify/test", nil); resp.StatusCode != http.StatusForbidden {
t.Errorf("expected 403 for a service account testing notifications, got %d", resp.StatusCode)
} else {
resp.Body.Close()
}
if resp := s.reqAs(t, instanceKey, http.MethodPut, "/api/me/onboarding",
map[string]bool{"dismissed": true}); resp.StatusCode != http.StatusForbidden {
t.Errorf("expected 403 for a service account dismissing onboarding, got %d", resp.StatusCode)
} else {
resp.Body.Close()
}
}
+158 -11
View File
@@ -13,12 +13,19 @@ import (
"github.com/go-chi/chi/v5"
)
// handleListTeams lists the caller's own teams, each with their role in it. An
// administrator listing every team goes through the admin endpoint instead:
// this one answers "what am I part of", which is what the UI's team filter and
// the combined queue are built from.
// handleListTeams lists the caller's own teams, each with their role in it,
// or — with ?name= — looks up one team by exact name regardless of caller
// identity (TEAM-LOOKUP.md). An administrator listing every team goes
// through the admin endpoint instead: the no-name case here answers "what am
// I part of", which is what the UI's team filter and the combined queue are
// built from.
func handleListTeams(db *sql.DB) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
if name := strings.TrimSpace(r.URL.Query().Get("name")); name != "" {
handleListTeamsByName(db, w, r, name)
return
}
caller, _ := userFromContext(r.Context())
rows, err := db.QueryContext(r.Context(), `
SELECT t.id, t.name, t.created_at, m.role, m.source
@@ -51,6 +58,31 @@ func handleListTeams(db *sql.DB) http.HandlerFunc {
}
}
// handleListTeamsByName answers "is there a team named exactly this", open to
// any authenticated caller including a service account (TEAM-LOOKUP.md) —
// mirrors handleListServiceAccounts' own ?name= lookup: a one-or-zero-length
// array, never an error on no match, and no caller-identity filtering at
// all, since what it discloses (a name is taken, nothing about who's in it
// or any of its data) is the same low sensitivity that lookup already
// accepts for service-account names.
func handleListTeamsByName(db *sql.DB, w http.ResponseWriter, r *http.Request, name string) {
var t models.Team
var created int64
err := db.QueryRowContext(r.Context(),
"SELECT id, name, created_at FROM teams WHERE name = $1", name,
).Scan(&t.ID, &t.Name, &created)
if errors.Is(err, sql.ErrNoRows) {
respond(w, http.StatusOK, []models.Team{})
return
}
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
t.CreatedAt = time.Unix(created, 0).UTC()
respond(w, http.StatusOK, []models.Team{t})
}
// handleUserTeams lists one user's teams, for the admin page's per-user view:
// "what is this person in", which /api/teams cannot answer because it is always
// about the caller.
@@ -116,6 +148,15 @@ func handleUserTeams(db *sql.DB) http.HandlerFunc {
// handleCreateTeam creates a team and makes its creator the first owner. A team
// with no owner would need an administrator to repair before anybody could use
// it, so the two happen in one transaction.
//
// An instance-scoped service account may also create a team (SERVICE-ACCOUNTS.md:
// it acts with the same reach system administration has over teams), but it
// is not a users row and cannot become an owner the way a person does. The
// team it creates starts with no human owner at all — not a bug, the expected
// shape for one terdut-operator is about to provision: a system administrator
// can always act as owner to repair or hand it off (requireTeamOwner), and
// the account that created it mints itself a team-scoped credential for it
// next, via POST /api/service-accounts.
func handleCreateTeam(db *sql.DB) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
var req struct {
@@ -131,7 +172,14 @@ func handleCreateTeam(db *sql.DB) http.HandlerFunc {
return
}
caller, _ := userFromContext(r.Context())
caller, isUser := userFromContext(r.Context())
if !isUser && !isInstanceServiceAccount(r.Context()) {
// A team-scoped service account authenticates as owner of exactly
// one team already (see serveAsServiceAccount); letting it create
// another would reach outside that boundary.
respond(w, http.StatusForbidden, errResp("instance-scoped service account or user access required"))
return
}
tx, err := db.BeginTx(r.Context(), nil)
if err != nil {
@@ -152,11 +200,13 @@ func handleCreateTeam(db *sql.DB) http.HandlerFunc {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
if _, err := tx.ExecContext(r.Context(),
"INSERT INTO team_members (team_id, user_id, role) VALUES ($1, $2, $3)",
team.ID, caller.ID, models.RoleOwner); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
if isUser {
if _, err := tx.ExecContext(r.Context(),
"INSERT INTO team_members (team_id, user_id, role) VALUES ($1, $2, $3)",
team.ID, caller.ID, models.RoleOwner); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
}
if err := tx.Commit(); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
@@ -164,7 +214,9 @@ func handleCreateTeam(db *sql.DB) http.HandlerFunc {
}
team.CreatedAt = time.Unix(created, 0).UTC()
team.Role = models.RoleOwner
if isUser {
team.Role = models.RoleOwner
}
respond(w, http.StatusCreated, team)
}
}
@@ -861,6 +913,101 @@ func handleCreateTeamDeadman(db *sql.DB) http.HandlerFunc {
}
}
// handleUpdateTeamDeadman replaces one switch's configuration in place.
// Added alongside create/delete so an automated caller (terdut-operator) can
// reconcile a spec change without deleting and recreating the switch, which
// would otherwise be the only option and would needlessly rotate its id for
// no reason a reconciler's diff should ever manufacture.
func handleUpdateTeamDeadman(db *sql.DB) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
teamID, ok := teamParam(w, r)
if !ok {
return
}
if !requireTeamOwner(w, r, teamID) {
return
}
switchID, err := strconv.ParseInt(chi.URLParam(r, "switchID"), 10, 64)
if err != nil {
respond(w, http.StatusBadRequest, errResp("invalid switch id"))
return
}
var req deadmanSwitchRequest
if err := decodeJSON(r, &req); err != nil {
respond(w, http.StatusBadRequest, errResp("invalid request body"))
return
}
req.Matcher = strings.TrimSpace(req.Matcher)
req.Name = strings.TrimSpace(req.Name)
if req.Severity == "" {
req.Severity = "critical"
}
if !deadmanSeverities[req.Severity] {
respond(w, http.StatusBadRequest, errResp("severity must be critical, error, warning or info"))
return
}
if req.TimeoutSeconds <= 0 {
respond(w, http.StatusBadRequest, errResp("timeout_seconds must be positive"))
return
}
if strings.Contains(req.Matcher, ";") {
respond(w, http.StatusBadRequest, errResp("one matcher per switch: add another switch instead of separating with ;"))
return
}
m, err := parseDeadmanMatcher(req.Matcher)
if err != nil {
respond(w, http.StatusBadRequest, errResp(
"unusable matcher ("+err.Error()+"): each must name an alertname, as in alertname=Watchdog,cluster=prod"))
return
}
if req.Name == "" {
req.Name = m.config()
}
if len(req.Name) > 100 {
respond(w, http.StatusBadRequest, errResp("name is too long"))
return
}
res, err := db.ExecContext(r.Context(), `
UPDATE deadman_switches
SET name = $1, matcher = $2, timeout_seconds = $3, severity = $4
WHERE id = $5 AND team_id = $6`,
req.Name, m.config(), req.TimeoutSeconds, req.Severity, switchID, teamID)
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
if n, _ := res.RowsAffected(); n == 0 {
respond(w, http.StatusNotFound, errResp("switch not found"))
return
}
// The full status, not the bare request echoed back: an update can
// change whether the switch is dormant, alive or dead (a longer
// timeout can revive one that just went dead), and a caller
// reconciling against status deserves the same view
// handleListTeamDeadman would give it.
set, err := deadmanSetForTeam(r.Context(), db, teamID)
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
statuses, err := deadmanStatuses(r.Context(), db, teamID, set, time.Now())
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
for _, s := range statuses {
if s.ID == switchID {
respond(w, http.StatusOK, s)
return
}
}
respond(w, http.StatusInternalServerError, errResp("internal error"))
}
}
// handleDeleteTeamDeadman removes a switch. An incident it already opened stays
// open until somebody resolves it: deleting the switch says "stop watching", not
// "the problem is gone".
+74
View File
@@ -6,6 +6,8 @@ import (
"io"
"net/http"
"testing"
"git.ryuvia.com/niklas/terdut-server/internal/models"
)
// The whole point of #4: two teams sharing one server must not see each other's
@@ -454,3 +456,75 @@ func TestTeams_OutsiderSeesNothing(t *testing.T) {
t.Errorf("blue's team list: %v", teams)
}
}
// ---------------------------------------------------------------------------
// GET /api/teams?name= (TEAM-LOOKUP.md)
// ---------------------------------------------------------------------------
func TestListTeamsByName_FindsExactMatch(t *testing.T) {
s := newTS(t)
instanceKey := createServiceAccount(t, s, s.key, "terdut-operator", models.ServiceAccountScopeInstance, 0)
teamID := createTeamAs(t, s, instanceKey, "platform")
teams := list(t, s.reqAs(t, instanceKey, http.MethodGet, "/api/teams?name=platform", nil))
if len(teams) != 1 {
t.Fatalf("expected exactly one match for ?name=platform, got %d: %v", len(teams), teams)
}
if int64(teams[0]["id"].(float64)) != teamID {
t.Errorf("id = %v, want %d", teams[0]["id"], teamID)
}
// No membership, so no role to report (models.Team's own doc comment:
// "empty when nobody in particular is asking").
if _, has := teams[0]["role"]; has {
t.Errorf("expected no role on a name-lookup match, got %v", teams[0]["role"])
}
}
func TestListTeamsByName_NoMatchIsAnEmptyArrayNotAnError(t *testing.T) {
s := newTS(t)
instanceKey := createServiceAccount(t, s, s.key, "terdut-operator", models.ServiceAccountScopeInstance, 0)
resp := s.reqAs(t, instanceKey, http.MethodGet, "/api/teams?name=does-not-exist", nil)
if resp.StatusCode != http.StatusOK {
t.Fatalf("expected 200 on no match, got %d", resp.StatusCode)
}
teams := list(t, resp)
if len(teams) != 0 {
t.Errorf("expected an empty array, got %v", teams)
}
}
// The actual motivating scenario (TEAM-LOOKUP.md): a service account that
// already created a team, interrupted before it could remember the id,
// recovers it via ?name= on the same name its own POST 409s on.
func TestListTeamsByName_RecoversAfterCreateConflict(t *testing.T) {
s := newTS(t)
instanceKey := createServiceAccount(t, s, s.key, "terdut-operator", models.ServiceAccountScopeInstance, 0)
original := createTeamAs(t, s, instanceKey, "recovered")
conflict := s.reqAs(t, instanceKey, http.MethodPost, "/api/teams", map[string]string{"name": "recovered"})
if conflict.StatusCode != http.StatusConflict {
t.Fatalf("expected 409 recreating the same name, got %d", conflict.StatusCode)
}
conflict.Body.Close()
teams := list(t, s.reqAs(t, instanceKey, http.MethodGet, "/api/teams?name=recovered", nil))
if len(teams) != 1 || int64(teams[0]["id"].(float64)) != original {
t.Fatalf("expected to recover the original team %d via ?name=, got %v", original, teams)
}
}
// Not gated by isInstanceServiceAccount or AdminOnly (TEAM-LOOKUP.md): any
// authenticated caller may ask whether a name is taken, the same low
// sensitivity GET /api/service-accounts?name= already accepts.
func TestListTeamsByName_OpenToAnyAuthenticatedCaller(t *testing.T) {
s := newTS(t)
red := newTeam(t, s, "red")
_ = createTeamAs(t, s, s.key, "blue-target")
// red's own member, not a member of "blue-target", still gets a match.
teams := list(t, red.call(http.MethodGet, "/api/teams?name=blue-target", nil))
if len(teams) != 1 || teams[0]["name"] != "blue-target" {
t.Errorf("expected a non-member caller to still find the team by name, got %v", teams)
}
}
+10
View File
@@ -72,6 +72,14 @@ type Config struct {
// OIDC configures single sign-on. The zero value, with no Issuer, is off.
OIDC OIDC
// OperatorMode declares this install gitops-managed: writes to teams,
// escalation policies, dead man's switches and integrations from a human
// (a session or a user's own API key) are refused, while a service
// account's are not. Deploy-time and restart-required, like the rest of
// "where this server is plugged in" — it is a statement about who owns
// this install's configuration, not a per-request toggle.
OperatorMode bool
}
// OIDC is the single sign-on configuration. Groups from the provider decide
@@ -151,6 +159,8 @@ func Load() Config {
DisablePasswordLogin: !boolean("TERDUT_PASSWORD_LOGIN", true),
OIDC: loadOIDC(),
OperatorMode: boolean("TERDUT_OPERATOR_MODE", false),
}
}
+27
View File
@@ -1,6 +1,7 @@
package db
import (
"context"
"database/sql"
"embed"
"fmt"
@@ -62,6 +63,16 @@ func Open(dsn string) (*sql.DB, error) {
}
}
// migrationLockKey is the Postgres advisory lock Migrate holds for its whole
// run. Two replicas starting at once would otherwise race the check-then-apply
// loop below against schema_migrations: the loser could crash on a
// duplicate-key insert, or contend with the winner's uncommitted DDL. Blocking
// (pg_advisory_lock, not pg_try_advisory_lock as the archiver and notifier
// use): on boot there is no later tick to defer to, so the right behaviour is
// to wait for the other replica to finish migrating, not to skip ahead and
// start serving against an unmigrated schema.
const migrationLockKey int64 = 7265_0003
// Migrate applies every embedded migration that has not been applied yet, in
// filename order, recording each in schema_migrations.
//
@@ -69,6 +80,22 @@ func Open(dsn string) (*sql.DB, error) {
// migration that failed half way used to leave the schema in whatever state it
// had reached. Postgres has transactional DDL, so the rollback is real.
func Migrate(db *sql.DB) error {
ctx := context.Background()
conn, err := db.Conn(ctx)
if err != nil {
return fmt.Errorf("migrate: acquire connection: %w", err)
}
defer conn.Close()
if _, err := conn.ExecContext(ctx, "SELECT pg_advisory_lock($1)", migrationLockKey); err != nil {
return fmt.Errorf("migrate: acquire advisory lock: %w", err)
}
defer func() {
if _, err := conn.ExecContext(ctx, "SELECT pg_advisory_unlock($1)", migrationLockKey); err != nil {
log.Printf("migrate: release advisory lock: %v", err)
}
}()
if _, err := db.Exec(`CREATE TABLE IF NOT EXISTS schema_migrations (
version TEXT PRIMARY KEY,
applied_at BIGINT NOT NULL DEFAULT FLOOR(EXTRACT(EPOCH FROM now()))::bigint
+125
View File
@@ -0,0 +1,125 @@
package db_test
import (
"database/sql"
"fmt"
"net/url"
"os"
"strings"
"sync"
"testing"
"git.ryuvia.com/niklas/terdut-server/internal/db"
_ "github.com/jackc/pgx/v5/stdlib"
)
// TERDUT_TEST_DSN must point at a database the test role may create schemas
// in; see internal/api/testdb_test.go for the fuller rationale this mirrors.
// An unset DSN fails rather than skips, deliberately.
const testDSNEnv = "TERDUT_TEST_DSN"
// TestMigrate_ConcurrentCallersDoNotRace reproduces two replicas starting at
// once against a brand-new, unmigrated schema: both call db.Migrate at the
// same time. Before migrationLockKey, the loser could crash on a
// duplicate-key insert into schema_migrations, or contend with the winner's
// uncommitted DDL; with the advisory lock, one blocks until the other
// finishes and both return cleanly.
func TestMigrate_ConcurrentCallersDoNotRace(t *testing.T) {
dsn := os.Getenv(testDSNEnv)
if dsn == "" {
t.Fatalf("%s is not set: these tests need Postgres.\n"+
"Run `make test-db` for a local one, then\n"+
" export %s=postgres://terdut:terdut@localhost:5432/terdut_test?sslmode=disable",
testDSNEnv, testDSNEnv)
}
schema := fmt.Sprintf("migrate_race_%d", os.Getpid())
admin, err := sql.Open("pgx", dsn)
if err != nil {
t.Fatalf("connect to %s: %v", testDSNEnv, err)
}
defer admin.Close()
if _, err := admin.Exec("CREATE SCHEMA " + schema); err != nil {
t.Fatalf("create schema %s: %v", schema, err)
}
t.Cleanup(func() {
cleanup, err := sql.Open("pgx", dsn)
if err != nil {
return
}
defer cleanup.Close()
if _, err := cleanup.Exec("DROP SCHEMA " + schema + " CASCADE"); err != nil {
t.Logf("drop schema %s: %v", schema, err)
}
})
scoped := withSearchPath(dsn, schema)
const callers = 2
errs := make([]error, callers)
var wg sync.WaitGroup
for i := range callers {
wg.Add(1)
go func(i int) {
defer wg.Done()
database, err := db.Open(scoped)
if err != nil {
errs[i] = fmt.Errorf("open: %w", err)
return
}
defer database.Close()
errs[i] = db.Migrate(database)
}(i)
}
wg.Wait()
for i, err := range errs {
if err != nil {
t.Fatalf("Migrate #%d: %v", i, err)
}
}
entries, err := os.ReadDir("migrations")
if err != nil {
t.Fatalf("read migrations dir: %v", err)
}
var want int
for _, e := range entries {
if !e.IsDir() && strings.HasSuffix(e.Name(), ".sql") {
want++
}
}
check, err := sql.Open("pgx", scoped)
if err != nil {
t.Fatalf("connect for verification: %v", err)
}
defer check.Close()
var got int
if err := check.QueryRow("SELECT COUNT(*) FROM schema_migrations").Scan(&got); err != nil {
t.Fatalf("count schema_migrations: %v", err)
}
if got != want {
t.Fatalf("schema_migrations has %d row(s) after two concurrent Migrate calls, want %d (one per migration file, no duplicates)", got, want)
}
}
// withSearchPath pins a DSN to one schema. Copied from
// internal/api/testdb_test.go rather than shared: that helper lives in the
// api_test package, a separate compiled package this one cannot import.
func withSearchPath(dsn, schema string) string {
opt := "-csearch_path=" + schema
if strings.HasPrefix(dsn, "postgres://") || strings.HasPrefix(dsn, "postgresql://") {
u, err := url.Parse(dsn)
if err == nil {
q := u.Query()
q.Set("options", opt)
u.RawQuery = q.Encode()
return u.String()
}
}
return dsn + " options='" + opt + "'"
}
@@ -0,0 +1,43 @@
-- Service accounts: a scoped, non-human credential for automation (e.g.
-- terdut-operator) that needs to manage teams, escalation policies, dead
-- man's switches, integrations and OIDC group bindings without impersonating
-- a human user. See SERVICE-ACCOUNTS.md for the design this implements.
--
-- Deliberately not a users row: no password_hash, no is_admin, no
-- user_identities linkage, so a service account can never be pulled into
-- OIDC group sync or password login, and is never mistaken for a human in an
-- audit trail.
--
-- scope is 'instance' (acts with the same reach system administration has
-- over teams: create one, list them, mint a 'team'-scoped account against
-- any of them) or 'team' (acts as that one team's owner, and nothing else).
-- The CHECK ties team_id's presence to scope directly, rather than leaving it
-- to application code to keep the two consistent.
CREATE TABLE service_accounts (
id BIGINT GENERATED BY DEFAULT AS IDENTITY PRIMARY KEY,
name TEXT NOT NULL UNIQUE,
scope TEXT NOT NULL CHECK (scope IN ('instance', 'team')),
team_id BIGINT REFERENCES teams(id) ON DELETE CASCADE,
created_by BIGINT REFERENCES users(id) ON DELETE SET NULL,
created_at BIGINT NOT NULL DEFAULT FLOOR(EXTRACT(EPOCH FROM now()))::bigint,
CONSTRAINT service_accounts_scope_team_id_chk CHECK (
(scope = 'team' AND team_id IS NOT NULL) OR
(scope = 'instance' AND team_id IS NULL)
)
);
CREATE INDEX service_accounts_team_id_idx ON service_accounts(team_id);
-- One account, many keys: rotation is minting a new one and revoking the
-- old, the same shape api_keys already has, so an account's identity and
-- audit history survive a rotation instead of being recreated by it.
CREATE TABLE service_account_keys (
id BIGINT GENERATED BY DEFAULT AS IDENTITY PRIMARY KEY,
service_account_id BIGINT NOT NULL REFERENCES service_accounts(id) ON DELETE CASCADE,
key_hash TEXT NOT NULL UNIQUE,
name TEXT NOT NULL,
created_at BIGINT NOT NULL DEFAULT FLOOR(EXTRACT(EPOCH FROM now()))::bigint,
last_used_at BIGINT
);
CREATE INDEX service_account_keys_service_account_id_idx ON service_account_keys(service_account_id);
+37
View File
@@ -0,0 +1,37 @@
package models
import "time"
// Service account scopes. Instance acts with the same reach system
// administration has over teams: create one, list them, mint a team-scoped
// account against any of them. Team acts as that one team's owner, and
// nothing else.
const (
ServiceAccountScopeInstance = "instance"
ServiceAccountScopeTeam = "team"
)
// ServiceAccount is a non-human credential: not a users row, so it never
// touches OIDC group sync, login, or the is_admin flag, and is never mistaken
// for a human in an audit trail (see api_keys' user_id, which every service
// account key deliberately does not have).
type ServiceAccount struct {
ID int64 `json:"id"`
Name string `json:"name"`
Scope string `json:"scope"`
TeamID *int64 `json:"team_id,omitempty"`
CreatedBy *int64 `json:"created_by,omitempty"`
CreatedAt time.Time `json:"created_at"`
}
// ServiceAccountKey is one bearer credential on a ServiceAccount. Multiple
// keys per account, the same shape as APIKey, are what let rotation mint a
// new one and revoke the old without recreating the account.
type ServiceAccountKey struct {
ID int64 `json:"id"`
ServiceAccountID int64 `json:"service_account_id"`
Name string `json:"name"`
CreatedAt time.Time `json:"created_at"`
LastUsedAt *time.Time `json:"last_used_at,omitempty"`
Key string `json:"key,omitempty"` // populated only on creation, never stored
}
+122 -19
View File
@@ -47,12 +47,27 @@
--font: system-ui, -apple-system, "Segoe UI", Roboto, "Helvetica Neue", Arial, sans-serif;
--mono: ui-monospace, SFMono-Regular, Menlo, Consolas, monospace;
/* A ~4-size type scale, per issue #26's "use one type scale" ask. This
file still has a dozen one-off font-size values below; migrating the
low-risk, purely cosmetic ones (standalone titles with no dimensional
or functional constraint) onto these tokens is a start, not the whole
job — the rest (sizes tied to a fixed shape like the avatar circle, to
a deliberately prominent display like a stat tile or the on-call name,
or to a non-negotiable constraint like the 16px that stops iOS zooming
into an input) stay as either their own pixel value or a documented
exception, since guessing at those without seeing them render risks
trading one inconsistency for a worse one. */
--fs-xs: 12px;
--fs-sm: 13px;
--fs-base: 14px;
--fs-lg: 18px;
--fs-xl: 21px;
--topbar-h: 52px;
/* No bottom tab bar on any breakpoint any more — mobile uses the hamburger
menu in the topbar, desktop the sidebar — so this stays 0. Kept as a
variable rather than deleted since .view, .toast and .nav's own height
calc still read it. */
--tabbar-h: 0px;
/* The phone-width bottom tab bar's height. Unused above 900px: the
desktop block overrides .nav/.view/.toast directly rather than reading
this back down to 0. */
--tabbar-h: 58px;
--safe-top: env(safe-area-inset-top, 0px);
--safe-bottom: env(safe-area-inset-bottom, 0px);
}
@@ -212,9 +227,6 @@ input:focus, textarea:focus { outline: none; border-color: var(--accent); box-sh
font-size: 18px; font-weight: 700; letter-spacing: -0.01em;
overflow: hidden; text-overflow: ellipsis; white-space: nowrap; min-width: 0;
}
#menu-btn { position: relative; }
.nav-badge.menu-btn-badge { top: 2px; left: auto; right: 2px; }
.open-pill {
display: inline-flex; align-items: center; gap: 6px;
padding: 3px 10px; border-radius: 999px;
@@ -226,11 +238,11 @@ input:focus, textarea:focus { outline: none; border-color: var(--accent); box-sh
.open-pill.has-triggered::before { background: var(--crit); }
.open-pill.all-acked::before { background: var(--warn); }
/* Hidden on phones — mobile navigates through the hamburger menu in the
topbar instead (see #menu-btn / openNavMenu in app.js). Reappears as the
left sidebar from 900px, where the desktop block below redeclares display. */
/* Bottom tab bar on phones (Queue, On-call, Alerts, Team, More); becomes the
left sidebar from 900px, where the desktop block below redeclares display
and shows every section flat, .nav-link-secondary included. */
.nav {
display: none;
display: flex;
position: fixed; left: 0; right: 0; bottom: 0; z-index: 20;
height: calc(var(--tabbar-h) + var(--safe-bottom));
padding-bottom: var(--safe-bottom);
@@ -240,14 +252,24 @@ input:focus, textarea:focus { outline: none; border-color: var(--accent); box-sh
border-top: 1px solid var(--border);
}
.nav-brand { display: none; }
/* Hidden here (shown from 900px below): on the phone bar the team switcher
lives in the topbar instead, as #team-selector-mobile. */
.nav-team-selector { display: none; }
.nav-link {
position: relative;
/* flex: 1 spreads the tabs evenly across the bar's width; the desktop
block below cancels it back to a natural-width row item. */
flex: 1 1 0;
display: flex; flex-direction: column; align-items: center; justify-content: center; gap: 2px;
color: var(--faint); font-size: 11px; font-weight: 600;
/* min-width lets a column shrink below its label's natural width, which is
what stops six tabs widening the bar past the screen. */
min-width: 0; padding: 0 2px;
}
/* Must come after .nav-link above: same specificity (one class each), so
whichever is later in the file wins for an element wearing both classes,
and this needs to beat .nav-link's display:flex here on the phone bar. */
.nav-link-secondary { display: none; }
.nav-label {
max-width: 100%; overflow: hidden; text-overflow: ellipsis; white-space: nowrap;
}
@@ -261,6 +283,34 @@ input:focus, textarea:focus { outline: none; border-color: var(--accent); box-sh
font-size: 11px; font-weight: 700; line-height: 18px; text-align: center;
}
/* ---------- team selector ---------- */
/* The global control for which team the app is scoped to. Hidden (via the
`hidden` attribute, set from teamselector.js) for anybody in fewer than two
teams, the same rule every other team-aware control in this file follows. */
.nav-team-selector,
.team-selector-mobile {
display: inline-flex; align-items: center; gap: 8px;
border: 1px solid var(--border-strong); border-radius: 999px;
background: var(--surface); color: var(--text);
font-size: 13px; font-weight: 600; cursor: pointer;
padding: 4px 12px; max-width: 100%;
}
.team-selector-mobile { padding: 4px 10px; font-size: 12px; max-width: 120px; }
.team-selector-label { overflow: hidden; text-overflow: ellipsis; white-space: nowrap; }
.team-selector-chevron { width: 14px; height: 14px; flex: none; color: var(--faint); margin-left: -2px; }
/* A team's identity colour — not a status, so never the severity palette. Six
colours, then they repeat; teamColorClass() in format.js picks one by the
team's id, the same rcN convention the rota's per-person chips use. */
.team-dot { flex: none; width: 8px; height: 8px; border-radius: 50%; background: var(--border-strong); }
.team-dot.rc1 { background: var(--accent); }
.team-dot.rc2 { background: var(--ok); }
.team-dot.rc3 { background: var(--snooze); }
.team-dot.rc4 { background: var(--warn); }
.team-dot.rc5 { background: var(--teal); }
.team-dot.rc6 { background: var(--pink); }
.view { padding-bottom: calc(var(--tabbar-h) + var(--safe-bottom)); }
.view-page { padding-left: 16px; padding-right: 16px; }
.view-page > * { max-width: 760px; margin-left: auto; margin-right: auto; }
@@ -276,12 +326,14 @@ input:focus, textarea:focus { outline: none; border-color: var(--accent); box-sh
/* ---------- chips ---------- */
.chips {
position: relative;
display: flex; gap: 6px;
padding: 12px 16px 8px;
overflow-x: auto; scrollbar-width: none;
}
.chips::-webkit-scrollbar { display: none; }
.chip {
display: inline-flex; align-items: center; gap: 6px;
flex: none;
min-height: 34px; padding: 0 12px;
border: 1px solid var(--border-strong); border-radius: 999px;
@@ -290,6 +342,17 @@ input:focus, textarea:focus { outline: none; border-color: var(--accent); box-sh
}
.chip[aria-selected="true"] { background: var(--text); border-color: var(--text); color: var(--bg); }
.chip .count { margin-left: 4px; opacity: 0.7; }
/* An overlay, not a flex item: absolute against .chips' own (non-scrolling)
box stays flush with its real right edge regardless of scroll position,
which turned out not to be true of position:sticky here — as a flex
item, its sticky offset interacted with the row's gap and its own
negative margin, landing short of the edge by about one gap's width. */
.chips-fade {
position: absolute; top: 0; right: 0; bottom: 0;
width: 24px;
background: linear-gradient(to right, transparent, var(--bg));
pointer-events: none;
}
/* ---------- lists ---------- */
@@ -353,7 +416,7 @@ input:focus, textarea:focus { outline: none; border-color: var(--accent); box-sh
.empty {
padding: 48px 16px; text-align: center; color: var(--muted);
}
.empty strong { display: block; color: var(--text); font-size: 16px; margin-bottom: 4px; }
.empty strong { display: block; color: var(--text); font-size: var(--fs-lg); margin-bottom: 4px; }
.empty .icon { width: 36px; height: 36px; color: var(--ok); margin-bottom: 8px; }
.load-error {
@@ -406,8 +469,12 @@ input:focus, textarea:focus { outline: none; border-color: var(--accent); box-sh
.detail-head .copy { margin-left: auto; }
.clip-buffer { position: fixed; top: 0; left: 0; opacity: 0; pointer-events: none; }
.detail-head .crumb { font-weight: 600; color: var(--muted); font-size: 14px; }
.detail-title { font-size: 21px; font-weight: 750; letter-spacing: -0.01em; margin: 16px 0 8px; overflow-wrap: anywhere; }
.detail-title { font-size: var(--fs-xl); font-weight: 750; letter-spacing: -0.01em; margin: 16px 0 8px; overflow-wrap: anywhere; }
.detail-badges { display: flex; flex-wrap: wrap; gap: 6px; margin-bottom: 14px; }
/* A copy of the sticky actionbar's primary button, right under the status
it responds to — see quickActions() in incident.js. */
.detail-quick-actions { margin-bottom: 14px; }
.detail-quick-actions .btn-primary { font-size: 16px; min-height: 44px; }
.card {
background: var(--surface);
@@ -422,6 +489,9 @@ input:focus, textarea:focus { outline: none; border-color: var(--accent); box-sh
.facts dt { color: var(--muted); }
.facts dd { margin: 0; overflow-wrap: anywhere; }
.facts .sub { color: var(--faint); }
.facts dd.fact-summary { display: flex; flex-wrap: wrap; align-items: center; gap: 8px; }
.fact-chip { display: inline-flex; align-items: center; gap: 4px; }
.fact-icon { width: 15px; height: 15px; color: var(--faint); }
.section { margin-top: 22px; }
.section-title {
@@ -453,6 +523,17 @@ details > summary::before { content: "▸ "; }
details[open] > summary::before { content: "▾ "; }
details[open] > summary { margin-bottom: 8px; }
/* One .tl-phase per status the incident has been through (see
timelinePhases() in incident.js) — each with its own .timeline <ol>, so
the existing :first-child/:last-child rail-capping below gives each phase
its own self-contained connecting line rather than one running through
the headings. */
.tl-phase + .tl-phase { border-top: 1px solid var(--border); }
.tl-phase-title {
padding: 10px 14px 0;
font-size: 11px; font-weight: 700; text-transform: uppercase; letter-spacing: 0.06em;
color: var(--faint);
}
.timeline { list-style: none; margin: 0; padding: 4px 0; }
.tl-item {
position: relative;
@@ -485,7 +566,7 @@ details[open] > summary { margin-bottom: 8px; }
}
.note-fix { background: var(--ok-soft); border-left: 3px solid var(--ok); }
.similar { list-style: none; margin: 0; padding: 0; }
.similar-item { padding: 10px 0; }
.similar-item { padding: 10px 14px; }
.similar-item + .similar-item { border-top: 1px solid var(--border, var(--surface-2)); }
.check { display: flex; align-items: center; gap: 8px; font-size: 14px; color: var(--muted); }
.note-actions { display: flex; justify-content: flex-end; }
@@ -524,7 +605,7 @@ details[open] > summary { margin-bottom: 8px; }
@keyframes sheet-up { from { transform: translateY(24px); opacity: 0.6; } }
.sheet-inner { padding: 8px 16px calc(16px + var(--safe-bottom)); }
.sheet-grab { width: 40px; height: 4px; margin: 0 auto 12px; border-radius: 2px; background: var(--border-strong); }
.sheet-title { font-size: 17px; font-weight: 700; margin: 0 0 4px; }
.sheet-title { font-size: var(--fs-lg); font-weight: 700; margin: 0 0 4px; }
.sheet-text { color: var(--muted); margin: 0 0 14px; font-size: 14px; }
.sheet-form { display: grid; gap: 12px; }
.sheet-actions { display: flex; gap: 8px; margin-top: 16px; }
@@ -576,9 +657,18 @@ details[open] > summary { margin-bottom: 8px; }
.now-label { color: var(--muted); font-size: 13px; font-weight: 600; }
.now-name { font-size: 20px; font-weight: 750; }
.you { color: var(--accent); font-weight: 650; font-size: 13px; margin-left: 6px; }
/* Oncall's own "you" indicator only (see you() in oncall.js) — a pill badge
is easier to spot there than this plain accent-coloured text. */
.you-badge { margin-left: 6px; }
.week-nav { display: flex; align-items: center; gap: 4px; }
.week-nav .label { font-size: 14px; font-weight: 650; min-width: 9em; text-align: center; }
/* The week-nav button that shows the date range, "28 Sep – 4 Oct", with the
ISO week number as secondary text inside it — a separate class from
.label above (team.js's month-nav uses that one) so its <small> isn't
caught by the unrelated .label > span styling meant for label chips. */
.week-nav .week-label { font-size: 14px; font-weight: 650; min-width: 11.5em; text-align: center; white-space: nowrap; }
.week-label small { color: var(--faint); font-weight: 600; font-size: 11px; margin-left: 2px; }
.days { list-style: none; margin: 0; padding: 0; }
.day { display: grid; grid-template-columns: 3.2em 4.2em 1fr; align-items: center; gap: 8px; min-height: 50px; padding: 0 14px; }
.day + .day { border-top: 1px solid var(--border); }
@@ -586,11 +676,18 @@ details[open] > summary { margin-bottom: 8px; }
.day-date { color: var(--faint); font-size: 13px; }
.day-who { overflow: hidden; text-overflow: ellipsis; white-space: nowrap; }
.day-who.nobody { color: var(--faint); font-style: italic; }
/* A run of several days held by the same person (or left empty), replacing
what used to be one identical row per day — see weekRuns() in oncall.js. */
.day.range { grid-template-columns: 1fr auto; }
.day-range { font-weight: 650; }
.day.today { background: var(--accent-soft); }
.day.today:first-child { border-radius: var(--radius) var(--radius) 0 0; }
.day.today:last-child { border-radius: 0 0 var(--radius) var(--radius); }
.day.today .day-name { color: var(--accent); }
.day.today .day-name, .day.today .day-range { color: var(--accent); }
.day.past { opacity: 0.6; }
/* Highlights whichever row is yours, same soft tint as .today — they already
read fine layered (today's own row is almost always one of yours too). */
.day.mine { background: var(--accent-soft); }
.shift-list { list-style: none; margin: 0; padding: 0; }
.shift-list li { display: flex; justify-content: space-between; padding: 12px 14px; }
@@ -611,6 +708,8 @@ details[open] > summary { margin-bottom: 8px; }
.account-name { font-size: 18px; font-weight: 750; }
.account-email { color: var(--muted); font-size: 14px; overflow-wrap: anywhere; }
.pw-form { display: grid; gap: 12px; padding: 16px; }
.account-fold { border-top: 1px solid var(--border); padding-top: 12px; }
.account-fold .stacked-form { margin-top: 8px; }
.form-ok {
margin: 0; padding: 10px 12px;
background: var(--ok-soft); color: var(--ok);
@@ -645,8 +744,11 @@ kbd {
display: flex; align-items: center; gap: 10px;
padding: 4px 10px 18px; font-size: 18px; font-weight: 750; letter-spacing: -0.01em;
}
.nav-team-selector { display: inline-flex; margin: -8px 10px 14px; width: calc(100% - 20px); }
.nav-link-secondary { display: flex; }
.nav-more-btn { display: none; }
.nav-link {
flex-direction: row; justify-content: flex-start; gap: 12px;
flex: none; flex-direction: row; justify-content: flex-start; gap: 12px;
min-height: 40px; padding: 0 10px; border-radius: var(--radius-sm);
color: var(--muted); font-size: 14px;
}
@@ -670,6 +772,7 @@ kbd {
is hidden, so the chips wrap here instead: Archived stays reachable. */
.pane-list .chips { flex-wrap: wrap; overflow-x: visible; }
.pane-list .chip-sep { display: none; }
.chips-fade { display: none; }
.view-queue:not(.has-detail) .pane-detail { display: block; }
/* On desktop the list stays visible next to the detail. */
@@ -784,7 +887,6 @@ kbd {
.stacked-form label { display: flex; align-items: center; gap: 6px; flex-wrap: wrap; font-size: 14px; }
.stacked-form label.checkbox { gap: 8px; }
.stacked-form input.wide { min-width: min(420px, 100%); }
.team-picker { margin-top: 8px; max-width: 100%; }
/* The rota, a month at a time. A name is too wide to print thirty times and
too alike down a column to read, so a day carries an initial in that
@@ -936,6 +1038,7 @@ button.rota-week:hover { background: var(--surface-2); color: var(--text); }
.overview-head { display: flex; align-items: baseline; gap: 8px; }
.overview-count { margin-left: auto; color: var(--muted); font-size: 18px; font-weight: 700; }
.overview-item p { margin: 4px 0 0; }
.overview-note-icon { width: 13px; height: 13px; vertical-align: -2px; color: var(--ok); }
/* ---------- stats page ---------- */
+18 -7
View File
@@ -86,6 +86,9 @@
<img src="/icon.svg" alt="" width="28" height="28">
<span>terdut</span>
</a>
<!-- Which team the app is scoped to. Hidden unless the signed-in user is
in more than one; teamselector.js fills it in and wires the click. -->
<button class="nav-team-selector" id="team-selector" type="button" hidden></button>
<a class="nav-link" href="/" data-section="queue" aria-label="Queue">
<svg viewBox="0 0 24 24" aria-hidden="true"><path d="M4 6h16M4 12h16M4 18h10"/></svg>
<span class="nav-label">Queue</span>
@@ -99,7 +102,10 @@
<svg viewBox="0 0 24 24" aria-hidden="true"><path d="M6 16V11a6 6 0 0 1 12 0v5l1.5 2h-15z"/><path d="M10 20.5a2 2 0 0 0 4 0"/></svg>
<span class="nav-label">Alerts</span>
</a>
<a class="nav-link" href="/stats" data-section="stats" aria-label="Stats">
<!-- Secondary: full-width in the desktop sidebar, folded into the
"More" tab's sheet on the phone-width bottom bar instead (see
.nav-link-secondary in app.css and openNavMenu in app.js). -->
<a class="nav-link nav-link-secondary" href="/stats" data-section="stats" aria-label="Stats">
<svg viewBox="0 0 24 24" aria-hidden="true"><path d="M4 20h16M7 20v-7M12 20V6M17 20v-10"/></svg>
<span class="nav-label">Stats</span>
</a>
@@ -110,22 +116,27 @@
<!-- Hidden unless the signed-in user is a system administrator; app.js
unhides it once /api/me says so. The server refuses every admin
endpoint regardless, so this is a courtesy and not a gate. -->
<a class="nav-link" href="/admin" data-section="admin" aria-label="Admin" id="nav-admin" hidden>
<a class="nav-link nav-link-secondary" href="/admin" data-section="admin" aria-label="Admin" id="nav-admin" hidden>
<svg viewBox="0 0 24 24" aria-hidden="true"><path d="M12 3l7 3v6c0 4-3 7-7 9-4-2-7-5-7-9V6z"/></svg>
<span class="nav-label">Admin</span>
</a>
<a class="nav-link" href="/more" data-section="more" aria-label="Account">
<a class="nav-link nav-link-secondary" href="/more" data-section="more" aria-label="Account">
<svg viewBox="0 0 24 24" aria-hidden="true"><circle cx="12" cy="8" r="3.5"/><path d="M5 20a7 7 0 0 1 14 0"/></svg>
<span class="nav-label">Account</span>
</a>
<!-- Phone-width only (see .nav-more-btn in app.css): opens the same
sheet the old hamburger button did, for the sections the bottom
bar has no room for. Not shown on the desktop sidebar, which lists
every section already. -->
<button class="nav-link nav-more-btn" id="nav-more-btn" type="button" aria-label="More sections" aria-haspopup="menu">
<svg viewBox="0 0 24 24" aria-hidden="true"><path d="M5 12h.01M12 12h.01M19 12h.01"/></svg>
<span class="nav-label">More</span>
</button>
</nav>
<header class="topbar">
<div class="topbar-left">
<button class="btn btn-ghost btn-icon" id="menu-btn" type="button" aria-label="Menu" aria-haspopup="menu">
<svg viewBox="0 0 24 24" aria-hidden="true"><path d="M4 7h16M4 12h16M4 17h16"/></svg>
<span class="nav-badge menu-btn-badge" data-badge hidden></span>
</button>
<button class="team-selector-mobile" id="team-selector-mobile" type="button" hidden></button>
<h1 class="topbar-title" id="topbar-title">Queue</h1>
</div>
<span class="open-pill" id="open-pill" hidden></span>
+53 -41
View File
@@ -43,9 +43,12 @@ function render() {
//
// The topic is the whole address — the server it is published to is the
// install's one ntfy, set in the deployment and not something a user picks.
function notifyStatus(topic) {
return topic ? `Topic: ${topic}` : 'No topic set — pages go to the team’s fallback topic.';
}
function notifyForm(user) {
const err = h('p', { class: 'form-error', role: 'alert', hidden: true });
const ok = h('p', { class: 'form-ok', role: 'status', hidden: true });
const topic = h('input', {
name: 'ntfy_topic', type: 'text', autocomplete: 'off',
autocapitalize: 'none', spellcheck: false,
@@ -54,29 +57,7 @@ function notifyForm(user) {
});
const submit = h('button', { class: 'btn btn-primary', type: 'submit', text: 'Save topic' });
// Only offered once a topic is saved: the test publishes to whatever the
// server has stored, not to whatever is half-typed in the field.
const test = h('button', {
class: 'btn', type: 'button', text: 'Send a test push',
hidden: !user.ntfy_topic,
onclick: async () => {
err.hidden = true;
ok.hidden = true;
test.disabled = true;
try {
await api.testNotification();
ok.textContent = 'Sent. If nothing arrives, the topic is wrong or ntfy is not reachable.';
ok.hidden = false;
} catch (ex) {
err.textContent = ex.message;
err.hidden = false;
} finally {
test.disabled = false;
}
},
});
const form = h('form', { class: 'card pw-form' },
const form = h('form', { class: 'stacked-form' },
h('label', {},
h('span', { text: 'ntfy topic' }),
topic),
@@ -89,25 +70,46 @@ function notifyForm(user) {
h('p', { class: 'muted small' },
'Anyone who knows the topic can read your pages and publish to it, so ',
'pick something unguessable rather than your name.'),
err, ok,
h('div', { class: 'row-actions' }, submit, test),
err,
submit,
);
const status = h('p', { class: 'muted', text: notifyStatus(user.ntfy_topic) });
const summary = h('summary', { text: user.ntfy_topic ? 'Change topic' : 'Set a topic' });
const details = h('details', { class: 'account-fold' }, summary, form);
// Only offered once a topic is saved: the test publishes to whatever the
// server has stored, not to whatever is half-typed in the field.
const test = h('button', {
class: 'btn', type: 'button', text: 'Send a test push',
hidden: !user.ntfy_topic,
onclick: async () => {
test.disabled = true;
try {
await api.testNotification();
toast('Sent. If nothing arrives, the topic is wrong or ntfy is not reachable.');
} catch (ex) {
toast(ex.message, 'error');
} finally {
test.disabled = false;
}
},
});
form.addEventListener('submit', async (e) => {
e.preventDefault();
err.hidden = true;
ok.hidden = true;
submit.disabled = true;
try {
const updated = await api.setNotifyTarget(user.id, topic.value.trim());
// Keep the cached user in step, so the onboarding checklist stops
// asking for this and the test button appears without a reload.
state.me.user = updated;
ok.textContent = updated.ntfy_topic
? 'Topic saved.'
: 'Topic cleared. Your pages go to the team’s fallback topic.';
ok.hidden = false;
status.textContent = notifyStatus(updated.ntfy_topic);
summary.textContent = updated.ntfy_topic ? 'Change topic' : 'Set a topic';
test.hidden = !updated.ntfy_topic;
details.open = false;
toast(updated.ntfy_topic ? 'Topic saved' : 'Topic cleared. Your pages go to the team’s fallback topic.');
} catch (ex) {
err.textContent = ex.message;
err.hidden = false;
@@ -115,7 +117,11 @@ function notifyForm(user) {
submit.disabled = false;
}
});
return form;
return h('div', { class: 'card pw-form' },
status,
h('div', { class: 'row-actions' }, test),
details);
}
// With password login switched off a password opens nothing, so somebody who
@@ -131,14 +137,13 @@ function passwordSection(user, hasPassword) {
];
}
return [
h('div', { class: 'page-head' }, h('h2', { text: hasPassword ? 'Change password' : 'Set a password' })),
h('div', { class: 'page-head' }, h('h2', { text: 'Password' })),
passwordForm(user, hasPassword),
];
}
function passwordForm(user, hasPassword) {
const err = h('p', { class: 'form-error', role: 'alert', hidden: true });
const ok = h('p', { class: 'form-ok', role: 'status', hidden: true });
const current = hasPassword
? h('input', { name: 'current', type: 'password', autocomplete: 'current-password', required: true })
: null;
@@ -148,18 +153,24 @@ function passwordForm(user, hasPassword) {
// A hidden username field lets password managers file the new password
// under the right account.
const form = h('form', { class: 'card pw-form', autocomplete: 'on' },
const form = h('form', { class: 'stacked-form', autocomplete: 'on' },
h('input', { type: 'text', name: 'username', autocomplete: 'username', value: user.username, hidden: true, readonly: true }),
current && h('label', {}, h('span', { text: 'Current password' }), current),
h('label', {}, h('span', { text: 'New password' }), next),
h('label', {}, h('span', { text: 'Repeat new password' }), again),
err, ok, submit,
err, submit,
);
const status = h('p', {
class: 'muted',
text: hasPassword ? 'Password set.' : 'No password set — sign-in needs one of the other methods.',
});
const summary = h('summary', { text: hasPassword ? 'Change password' : 'Set a password' });
const details = h('details', { class: 'account-fold' }, summary, form);
form.addEventListener('submit', async (e) => {
e.preventDefault();
err.hidden = true;
ok.hidden = true;
if (next.value !== again.value) {
err.textContent = 'The new passwords do not match.';
err.hidden = false;
@@ -171,13 +182,14 @@ function passwordForm(user, hasPassword) {
state.me.has_password = true;
form.reset();
if (!current) {
// From now on the form needs the current-password field.
// From now on the form needs the current-password field, and a fresh
// render already comes up with the fold closed.
render();
toast('Password saved');
return;
}
ok.textContent = 'Password saved. Other devices have been signed out.';
ok.hidden = false;
details.open = false;
toast('Password saved. Other devices have been signed out.');
} catch (ex) {
err.textContent = ex.message;
err.hidden = false;
@@ -185,7 +197,7 @@ function passwordForm(user, hasPassword) {
submit.disabled = false;
}
});
return form;
return h('div', { class: 'card pw-form' }, status, details);
}
function shortcuts() {
+5 -1
View File
@@ -60,7 +60,11 @@ function render() {
let body;
if (error && !items) body = h('div', { class: 'load-error', text: error });
else if (!items) body = spinner();
else if (!items.length) body = emptyState(filter === 'firing' ? 'Nothing firing' : 'No alerts', '', filter === 'firing' ? 'checkCircle' : null);
else if (!items.length) {
body = filter === 'firing'
? emptyState('Nothing firing', 'No alerts are currently firing.', 'checkCircle')
: emptyState('No alerts', 'None match this filter.', 'bell');
}
else {
body = h('div', { class: 'list' },
error && h('div', { class: 'load-error', text: `Showing older data: ${error}` }),
+21 -15
View File
@@ -11,6 +11,7 @@ import * as alerts from './alerts.js';
import * as stats from './stats.js';
import * as account from './account.js';
import * as team from './team.js';
import * as teamselector from './teamselector.js';
import * as admin from './admin.js';
import * as adminuser from './adminuser.js';
import * as adminteam from './adminteam.js';
@@ -35,16 +36,17 @@ const SECTIONS = {
device: { title: 'Sign in a terminal', view: device },
};
// The mobile hamburger menu's contents — the same sections the desktop
// sidebar's .nav-link list carries in index.html, in the same order.
// The desktop sidebar's .nav-link list in index.html, in the same order.
// `secondary` marks the ones that fold into the phone bottom bar's "More"
// sheet (openNavMenu below) instead of getting a tab of their own there.
const NAV_ITEMS = [
{ path: '/', section: 'queue', label: 'Queue', icon: 'queueList' },
{ path: '/oncall', section: 'oncall', label: 'On-call', icon: 'calendar' },
{ path: '/alerts', section: 'alerts', label: 'Alerts', icon: 'bell' },
{ path: '/stats', section: 'stats', label: 'Stats', icon: 'chart' },
{ path: '/stats', section: 'stats', label: 'Stats', icon: 'chart', secondary: true },
{ path: '/team', section: 'team', label: 'Team', icon: 'team' },
{ path: '/admin', section: 'admin', label: 'Admin', icon: 'shield', adminOnly: true },
{ path: '/more', section: 'more', label: 'Account', icon: 'user' },
{ path: '/admin', section: 'admin', label: 'Admin', icon: 'shield', adminOnly: true, secondary: true },
{ path: '/more', section: 'more', label: 'Account', icon: 'user', secondary: true },
];
function parseRoute(pathname) {
@@ -151,15 +153,16 @@ function render() {
// ---------- nav menu ----------
// The mobile hamburger menu: same shape as the sheet-based action menus in
// incident.js (openSheet + a <ul class="menu"> of menu-item buttons), one
// item per NAV_ITEMS entry, resolving with a path for navigate() to use.
// The phone bottom bar's "More" sheet: same shape as the sheet-based action
// menus in incident.js (openSheet + a <ul class="menu"> of menu-item
// buttons), one item per secondary NAV_ITEMS entry — the ones the bar itself
// has no room for, since Queue/On-call/Alerts/Team already have their own
// tab and don't need to be reachable here too.
function openNavMenu() {
const current = SECTIONS[route.section].nav || route.section;
const triggered = state.open.filter((i) => i.status === 'triggered').length;
const items = NAV_ITEMS.filter((n) => !n.adminOnly || state.me?.user?.is_admin);
const items = NAV_ITEMS.filter((n) => n.secondary && (!n.adminOnly || state.me?.user?.is_admin));
ui.openSheet(() => [
ui.h('h2', { class: 'sheet-title', text: 'Sections' }),
ui.h('h2', { class: 'sheet-title', text: 'More' }),
ui.h('ul', { class: 'menu', role: 'menu' }, items.map((n) =>
ui.h('li', {}, ui.h('button', {
class: 'menu-item',
@@ -170,7 +173,6 @@ function openNavMenu() {
},
ui.icon(n.icon),
n.label,
n.section === 'queue' && triggered > 0 && ui.badge(String(triggered), 'st-triggered menu-sub'),
))),
),
]).then((path) => {
@@ -204,8 +206,8 @@ function updateBadges() {
pill.classList.toggle('has-triggered', triggered > 0);
pill.classList.toggle('all-acked', open > 0 && triggered === 0);
// Two badges carry this count: the sidebar's Queue tab (desktop) and the
// hamburger button (mobile) — only one of the two is ever visible at once.
// Just the Queue tab's own badge now — phone bottom bar and desktop
// sidebar both read the same [data-badge] span on that one nav-link.
for (const badge of document.querySelectorAll('[data-badge]')) {
badge.hidden = triggered === 0;
badge.textContent = String(triggered);
@@ -229,7 +231,8 @@ async function boot() {
document.addEventListener('keydown', onKey);
$('login-form').addEventListener('submit', onLogin);
$('signup-form').addEventListener('submit', onSignup);
$('menu-btn').addEventListener('click', openNavMenu);
$('nav-more-btn').addEventListener('click', openNavMenu);
teamselector.init();
ssoErrorCode = takeSSOError();
// /signup is the one route that works without a session.
@@ -245,6 +248,7 @@ async function boot() {
await loadAuthConfig();
state.me = await api.me();
await loadTeams();
teamselector.render();
// The Admin tab exists only for an administrator. Somebody who types /admin
// anyway gets the view's own "ask an administrator" card, not a blank page.
$('nav-admin').hidden = !state.me?.user?.is_admin;
@@ -328,6 +332,7 @@ async function onSignup(e) {
history.replaceState({ depth: 0 }, '', '/');
route = parseRoute('/');
await loadTeams();
teamselector.render();
$('nav-admin').hidden = !state.me?.user?.is_admin;
showApp();
} catch (ex) {
@@ -351,6 +356,7 @@ function ssoErrorText(code, name) {
no_email: `${sso} did not send an email address for you, which terdut needs.`,
email_conflict: 'An account with your email address already exists and could not be linked to this sign-in. Ask an administrator.',
disabled: 'Your account is disabled. Ask an administrator.',
not_bootstrapped: 'This install is still setting up. Try again in a moment.',
}[code] || `Signing in with ${sso} failed.`;
}
+8
View File
@@ -98,6 +98,14 @@ export function severityClass(sev) {
return '';
}
// A stable identity colour for a team, so the same team always reads the same
// colour without the server needing to store one. Teams have no colour field;
// this hashes the id into the six-colour rcN palette app.css already has for
// the rota's per-person chips (a team is not a status, so never severity).
export function teamColorClass(teamID) {
return `rc${(((teamID % 6) + 6) % 6) + 1}`;
}
// A one-line summary of the group labels, without the one the title already shows.
export function labelSummary(labels, skip = 'alertname') {
return Object.entries(labels || {})
+90 -20
View File
@@ -7,7 +7,7 @@ import {
h, clear, icon, badge, labelChip, openSheet, closeSheet, confirm, toast, spinner, emptyState,
} from './ui.js';
import {
ago, when, until, isFuture, severityClass, STATUS_LABEL,
ago, when, until, duration, isFuture, severityClass, STATUS_LABEL,
} from './format.js';
import { myID, users } from './state.js';
import { back } from './app.js';
@@ -86,6 +86,7 @@ function render() {
error && h('div', { class: 'load-error', text: `Showing older data: ${error}` }),
h('h1', { class: 'detail-title', text: inc.title }),
h('div', { class: 'detail-badges' }, statusBadges()),
quickActions(),
facts(),
groupLabels(),
alertsSection(),
@@ -125,6 +126,20 @@ function who(id, name) {
function facts() {
const rows = [];
const add = (k, ...v) => rows.push(h('dt', { text: k }), h('dd', {}, ...v));
// Duration, severity and who's on it, in one scannable row up top — the
// rest of this card has each of those too, but spread across rows that
// take reading top to bottom to piece together.
const elapsedTo = inc.resolved_at ? Date.parse(inc.resolved_at) : Date.now();
const responsible = inc.assigned_to_id != null ? who(inc.assigned_to_id, inc.assigned_to)
: inc.acknowledged_by_id != null ? who(inc.acknowledged_by_id, inc.acknowledged_by)
: 'Unassigned';
rows.push(h('dt', { text: 'At a glance' }), h('dd', { class: 'fact-summary' },
h('span', { class: 'fact-chip' }, icon('clock', 'icon fact-icon'), duration(elapsedTo - Date.parse(inc.triggered_at))),
inc.severity && badge(inc.severity, `plain ${severityClass(inc.severity)}`),
h('span', { class: 'fact-chip' }, icon('user', 'icon fact-icon'), responsible),
));
add('Triggered', when(inc.triggered_at), h('span', { class: 'sub', text: ` · ${ago(inc.triggered_at)}` }));
if (inc.acknowledged_at) {
add('Acknowledged', `${who(inc.acknowledged_by_id, inc.acknowledged_by)} · ${when(inc.acknowledged_at)}`);
@@ -148,9 +163,12 @@ function facts() {
function groupLabels() {
const entries = Object.entries(inc.group_labels || {});
if (!entries.length) return null;
// Collapsed by default, the same disclosure alertItem() below uses for an
// alert's own labels — this is background, not something to scan past.
return h('section', { class: 'section' },
h('h2', { class: 'section-title', text: 'Grouped by' }),
h('div', { class: 'labels-wrap' }, entries.map(([k, v]) => labelChip(k, v))),
h('details', {},
h('summary', { text: `Grouped by (${entries.length})` }),
h('div', { class: 'labels-wrap' }, entries.map(([k, v]) => labelChip(k, v)))),
);
}
@@ -206,12 +224,20 @@ function eventText(ev, named = false) {
case 'snoozed': return [strong(person), ` snoozed until ${ev.detail ? when(ev.detail) : '…'}`];
case 'unsnoozed': return [strong(person), ' ended the snooze'];
case 'resolved': return person ? [strong(person), ' resolved the incident'] : ['Resolved: every alert stopped firing'];
// Falls through to the generic `${ev.type}: ${ev.detail}` below otherwise
// — this just capitalises it and drops the redundant "escalated:" prefix
// from detail (already "level 2: alice, bob" or "escalation exhausted: …").
case 'escalated': return [`Escalated — ${ev.detail}`];
case 'note': return [strong(person), ' added a note'];
case 'resolution_note': return [strong(person), ' noted what fixed it'];
case 'notified': {
const to = person ? strong(person) : 'the fallback topic';
if (ev.detail === 'reminder') return ['Reminder sent to ', to];
if (ev.detail === 'resolved') return ['Resolution sent to ', to];
// 'escalated' is a second, later page — the next level firing, not the
// same page landing twice — so it reads as a bug unless told apart
// from the initial 'triggered' page below.
if (ev.detail === 'escalated') return ['Escalation paged ', to];
return ['Paged ', to];
}
case 'notify_failed': return ['Notification failed', ev.detail ? `: ${ev.detail}` : ''];
@@ -236,8 +262,34 @@ function similarSection() {
);
}
// Splits the already-sorted timeline on the status transitions that matter —
// first acknowledged, then resolved — so a long incident reads as "before
// anyone had it" / "while someone did" / "after it closed" instead of one
// undifferentiated list. A later re-acknowledge (after an unacknowledge)
// doesn't open a second "Acknowledged" phase; it's still the same spell of
// somebody owning it.
function timelinePhases(sorted) {
const phases = [{ label: 'Triggered', events: [] }];
let acked = false;
for (const ev of sorted) {
if (ev.type === 'acknowledged' && !acked) {
phases.push({ label: 'Acknowledged', events: [] });
acked = true;
} else if (ev.type === 'resolved') {
phases.push({ label: 'Resolved', events: [] });
}
phases[phases.length - 1].events.push(ev);
}
return phases.filter((p) => p.events.length);
}
function timelineSection() {
const sorted = [...events].sort((a, b) => Date.parse(a.created_at) - Date.parse(b.created_at) || a.id - b.id);
const phases = timelinePhases(sorted);
// A single phase (the common case: most incidents are acked once and
// resolved) names nothing extra — only a split timeline needs the
// headings to make sense of.
const named = phases.length > 1;
return h('section', { class: 'section' },
h('h2', { class: 'section-title' },
h('span', { text: 'Timeline' }),
@@ -245,8 +297,10 @@ function timelineSection() {
icon('note'), 'Add note')),
h('div', { class: 'card' },
sorted.length
? h('ol', { class: 'timeline' }, sorted.map(timelineItem))
: emptyState('No events yet', '')),
? phases.map((p) => h('div', { class: 'tl-phase' },
named && h('div', { class: 'tl-phase-title', text: p.label }),
h('ol', { class: 'timeline' }, p.events.map(timelineItem))))
: emptyState('No events yet', 'Nothing has happened on this incident yet.', 'clock')),
);
}
@@ -365,27 +419,43 @@ async function copyIncident() {
const isOpen = () => inc.status !== 'resolved';
const isSnoozed = () => isOpen() && isFuture(inc.snoozed_until);
function actionBar() {
let primary;
let secondary;
// primaryAction and secondaryAction are factories, not shared nodes — a
// button can only live in one place, and quickActions() below needs its own
// copy of the primary one rather than the actionbar's.
function primaryAction() {
if (inc.status === 'triggered') {
primary = h('button', { class: 'btn btn-primary', type: 'button', onclick: acknowledge }, icon('check'), 'Acknowledge');
} else if (inc.status === 'acknowledged') {
primary = h('button', { class: 'btn btn-primary', type: 'button', onclick: resolve }, icon('checkCircle'), 'Resolve');
} else {
primary = inc.archived_at
? h('button', { class: 'btn btn-primary', type: 'button', onclick: unarchive }, icon('undo'), 'Unarchive')
: h('button', { class: 'btn btn-primary', type: 'button', onclick: archive }, icon('archive'), 'Archive');
return h('button', { class: 'btn btn-primary', type: 'button', onclick: acknowledge }, icon('check'), 'Acknowledge');
}
if (inc.status === 'acknowledged') {
return h('button', { class: 'btn btn-primary', type: 'button', onclick: resolve }, icon('checkCircle'), 'Resolve');
}
return inc.archived_at
? h('button', { class: 'btn btn-primary', type: 'button', onclick: unarchive }, icon('undo'), 'Unarchive')
: h('button', { class: 'btn btn-primary', type: 'button', onclick: archive }, icon('archive'), 'Archive');
}
function secondaryAction() {
if (isOpen()) {
secondary = isSnoozed()
return isSnoozed()
? h('button', { class: 'btn', type: 'button', onclick: unsnooze }, icon('bell'), 'Unsnooze')
: h('button', { class: 'btn', type: 'button', onclick: snooze }, icon('clock'), 'Snooze');
} else {
secondary = h('button', { class: 'btn', type: 'button', onclick: addNote }, icon('note'), 'Note');
}
const more = h('button', { class: 'btn btn-icon', type: 'button', 'aria-label': 'More actions', onclick: moreMenu }, icon('more'));
const bar = h('div', { class: 'actionbar' }, primary, secondary, more);
return h('button', { class: 'btn', type: 'button', onclick: addNote }, icon('note'), 'Note');
}
// A copy of the primary action (Acknowledge/Resolve/…) up where it's seen
// right away, next to the status it responds to. The sticky actionbar below
// keeps carrying every action, primary included, for whenever the page has
// been scrolled past it.
function quickActions() {
const div = h('div', { class: 'detail-quick-actions' }, primaryAction());
if (busy) for (const b of div.querySelectorAll('button')) b.disabled = true;
return div;
}
function actionBar() {
const more = h('button', { class: 'btn', type: 'button', 'aria-label': 'More actions', onclick: moreMenu }, icon('more'), 'More');
const bar = h('div', { class: 'actionbar' }, primaryAction(), secondaryAction(), more);
if (busy) for (const b of bar.querySelectorAll('button')) b.disabled = true;
return bar;
}
+72 -24
View File
@@ -6,8 +6,8 @@
// every team the viewer is in, because somebody on two rotas wants both.
import * as api from './api.js';
import { h, clear, icon, spinner } from './ui.js';
import { isoDate, mondayOf, addDays, isoWeek, initial } from './format.js';
import { h, clear, badge, icon, spinner } from './ui.js';
import { isoDate, mondayOf, addDays, isoWeek, initial, duration } from './format.js';
import { myID, currentTeam } from './state.js';
const view = () => document.getElementById('view-oncall');
@@ -68,7 +68,7 @@ function render() {
}
function you(userID) {
return userID === myID() ? h('span', { class: 'you', text: 'you' }) : null;
return userID === myID() ? badge('you', 'plain st-oncall you-badge') : null;
}
// One card per team with somebody on call, and a single empty card when there
@@ -99,20 +99,48 @@ function nowCard() {
)));
}
// Groups the week's 7 days into runs held by the same person (or the same
// empty slot) — the week's own version of the consecutive-day grouping
// myShifts does for a single person's own dates, below. Seven identical rows
// for one person all week collapses to the one bar this way.
function weekRuns(byDate) {
const runs = [];
for (let i = 0; i < 7; i++) {
const date = isoDate(addDays(weekStart, i));
const e = byDate.get(date) || null;
const uid = e ? e.user_id : null;
const last = runs[runs.length - 1];
if (last && last.uid === uid) last.to = date;
else runs.push({ uid, entry: e, from: date, to: date });
}
return runs;
}
function weekCard() {
const byDate = new Map(data.week.map((e) => [e.date, e]));
const today = isoDate(new Date());
const days = [];
for (let i = 0; i < 7; i++) {
const d = addDays(weekStart, i);
const key = isoDate(d);
const e = byDate.get(key);
days.push(h('li', { class: `day ${key === today ? 'today' : ''} ${key < today ? 'past' : ''}` },
h('span', { class: 'day-name', text: dayName.format(d) }),
h('span', { class: 'day-date', text: dayDate.format(d) }),
h('span', { class: `day-who ${e ? '' : 'nobody'}` }, e ? e.username : 'nobody', e && you(e.user_id)),
));
}
const mine = myID();
const days = weekRuns(byDate).map((r) => {
const single = r.from === r.to;
const cls = [
'day',
single ? '' : 'range',
r.from <= today && today <= r.to ? 'today' : '',
r.to < today ? 'past' : '',
r.uid === mine ? 'mine' : '',
].filter(Boolean).join(' ');
const label = single
? [h('span', { class: 'day-name', text: dayName.format(parse(r.from)) }),
h('span', { class: 'day-date', text: dayDate.format(parse(r.from)) })]
: [h('span', {
class: 'day-range',
text: `${dayName.format(parse(r.from))} ${dayDate.format(parse(r.from))} – ${dayName.format(parse(r.to))} ${dayDate.format(parse(r.to))}`,
})];
return h('li', { class: cls },
...label,
h('span', { class: `day-who ${r.entry ? '' : 'nobody'}` }, r.entry ? r.entry.username : 'nobody', r.entry && you(r.entry.user_id)),
);
});
const thisWeek = isoDate(weekStart) === isoDate(mondayOf(new Date()));
return [
h('div', { class: 'page-head' },
@@ -121,12 +149,12 @@ function weekCard() {
h('button', { class: 'btn btn-ghost btn-icon', type: 'button', 'aria-label': 'Previous week', onclick: () => shiftWeek(-1) },
icon('chevronLeft')),
h('button', {
class: 'btn btn-ghost label',
class: 'btn btn-ghost week-label',
type: 'button',
title: 'Back to this week',
onclick: () => { weekStart = mondayOf(new Date()); refresh(); },
text: `Week ${isoWeek(weekStart)}`,
}),
text: `${dayDate.format(weekStart)} – ${dayDate.format(addDays(weekStart, 6))}`,
}, h('small', { text: ` Week ${isoWeek(weekStart)}` })),
h('button', { class: 'btn btn-ghost btn-icon', type: 'button', 'aria-label': 'Next week', onclick: () => shiftWeek(1) },
icon('chevronRight')),
),
@@ -135,7 +163,9 @@ function weekCard() {
];
}
// myShifts groups your upcoming dates into runs of consecutive days.
// myShifts groups your upcoming dates into runs of consecutive days, then
// splits off the one you're already in — listing it again under "Next
// shifts" told people they hadn't started a shift they were already on.
function myShifts() {
const mine = data.upcoming.filter((e) => e.user_id === myID()).map((e) => e.date).sort();
const runs = [];
@@ -145,16 +175,34 @@ function myShifts() {
else runs.push({ from: date, to: date });
}
const fmt = (s) => `${dayName.format(parse(s))} ${dayDate.format(parse(s))}`;
return [
h('div', { class: 'page-head' }, h('h2', { text: 'Your next shifts' })),
const label = (r) => (r.from === r.to ? fmt(r.from) : `${fmt(r.from)} – ${fmt(r.to)}`);
const today = isoDate(new Date());
const current = runs[0] && runs[0].from <= today ? runs[0] : null;
const next = current ? runs.slice(1) : runs;
const currentCard = current ? [
h('div', { class: 'page-head' }, h('h2', { text: 'Current shift' })),
h('div', { class: 'card' },
runs.length
? h('ul', { class: 'shift-list' }, runs.slice(0, 8).map((r) =>
h('ul', { class: 'shift-list' }, h('li', {},
h('span', { text: label(current) }),
h('span', { class: 'muted', text: `ends in ${duration(addDays(parse(current.to), 1) - Date.now())}` })))),
] : [];
return [
...currentCard,
h('div', { class: 'page-head' }, h('h2', { text: current ? 'Next shifts' : 'Your next shifts' })),
h('div', { class: 'card' },
next.length
? h('ul', { class: 'shift-list' }, next.slice(0, 8).map((r) =>
h('li', {},
h('span', { text: r.from === r.to ? fmt(r.from) : `${fmt(r.from)} – ${fmt(r.to)}` }),
h('span', { text: label(r) }),
h('span', { class: 'muted', text: days(r) })),
))
: h('div', { class: 'empty', text: 'Nothing scheduled in the next 60 days.' })),
: h('div', {
class: 'empty',
text: current ? 'Nothing else scheduled in the next 60 days.' : 'Nothing scheduled in the next 60 days.',
})),
];
}
+60 -45
View File
@@ -2,8 +2,8 @@
import * as api from './api.js';
import { h, clear, badge, emptyState, spinner } from './ui.js';
import { age, until, isFuture, severityClass, labelSummary } from './format.js';
import { state, myID } from './state.js';
import { ago, until, isFuture, severityClass, labelSummary, teamColorClass } from './format.js';
import { state, myID, setSelectedTeam, onTeamChange } from './state.js';
import * as onboarding from './onboarding.js';
import { navigate } from './app.js';
@@ -18,43 +18,30 @@ const FILTERS = [
];
const EMPTY = {
open: ['All clear', 'Nothing open right now.'],
triggered: ['Nothing triggered', 'Every open incident has been acknowledged.'],
acknowledged: ['Nothing acknowledged', 'No one is working an incident right now.'],
snoozed: ['Nothing snoozed', 'Snoozed incidents show up here until the snooze runs out.'],
resolved: ['Nothing resolved', 'Resolved incidents are archived after a while.'],
archived: ['Nothing archived', ''],
open: ['All clear', 'Nothing open right now.', 'checkCircle'],
triggered: ['Nothing triggered', 'Every open incident has been acknowledged.', 'checkCircle'],
acknowledged: ['Nothing acknowledged', 'No one is working an incident right now.', 'checkCircle'],
snoozed: ['Nothing snoozed', 'Snoozed incidents show up here until the snooze runs out.', 'clock'],
resolved: ['Nothing resolved', 'Resolved incidents are archived after a while.', null],
archived: ['Nothing archived', 'Resolved incidents land here once archived.', 'archive'],
};
onboarding.onRerender(() => renderList());
// The queue used to keep its own team filter (a per-tab sessionStorage value,
// out of step with team.js's own picker); both now defer to the global
// selector's shared state, so re-render whenever it changes.
onTeamChange(() => {
renderChips();
refresh({ fresh: true });
});
let filter = loadFilter();
let teamFilter = loadTeamFilter(); // '' for every team the viewer is in
let items = null; // null while loading
let error = null;
let selected = null;
let cursor = -1; // keyboard position in the list
let built = false;
function loadTeamFilter() {
try {
return sessionStorage.getItem('terdut.queue.team') || '';
} catch {
return '';
}
}
function setTeamFilter(id) {
teamFilter = id;
try {
sessionStorage.setItem('terdut.queue.team', id);
} catch {
/* storage unavailable */
}
renderChips();
refresh({ fresh: true });
}
function loadFilter() {
try {
const f = sessionStorage.getItem('terdut.queue.filter');
@@ -89,8 +76,8 @@ export async function refresh({ fresh = false } = {}) {
// The open list is already fetched for the badges; no need to ask twice.
// The cached open queue covers every team, so it can only be reused when
// no team filter is applied.
const query = teamFilter ? { ...f.query, team_id: teamFilter } : f.query;
const cached = filter === 'open' && !fresh && !teamFilter;
const query = state.selectedTeamID != null ? { ...f.query, team_id: state.selectedTeamID } : f.query;
const cached = filter === 'open' && !fresh && state.selectedTeamID == null;
const result = cached ? state.open : await api.incidents(query);
await onboarding.load();
if (requested !== filter) return;
@@ -100,6 +87,10 @@ export async function refresh({ fresh = false } = {}) {
if (requested !== filter) return;
error = err.message;
}
// state.open (what the chip counts read) has just been refreshed too, by
// whichever caller updated it before calling here — app.js's poll, or the
// `cached` branch above.
renderChips();
renderList();
}
@@ -114,18 +105,32 @@ function setFilter(id) {
refresh({ fresh: true });
}
// Counts for the three chips derivable from the open list already fetched
// for the badges — Snoozed/Resolved/Archived would need a request of their
// own, so those chips stay count-less for now.
function chipCount(id) {
const open = state.selectedTeamID == null
? state.open
: state.open.filter((i) => i.team_id === state.selectedTeamID);
if (id === 'open') return open.length;
if (id === 'triggered') return open.filter((i) => i.status === 'triggered').length;
if (id === 'acknowledged') return open.filter((i) => i.status === 'acknowledged').length;
return null;
}
function renderChips() {
const el = document.getElementById('queue-filters');
const chips = FILTERS.map((f) =>
h('button', {
const chips = FILTERS.map((f) => {
const count = chipCount(f.id);
return h('button', {
class: 'chip',
type: 'button',
role: 'tab',
'aria-selected': String(f.id === filter),
onclick: () => setFilter(f.id),
text: f.label,
}),
);
}, count != null && h('span', { class: 'count', text: String(count) }));
});
// Somebody in one team has nothing to choose between, so the row of team
// chips appears only when there is more than one. The default is all of
@@ -136,22 +141,29 @@ function renderChips() {
class: 'chip',
type: 'button',
role: 'tab',
'aria-selected': String(teamFilter === ''),
onclick: () => setTeamFilter(''),
text: 'All teams',
}));
'aria-selected': String(state.selectedTeamID == null),
onclick: () => setSelectedTeam(null),
}, h('span', { class: 'team-dot' }), ' All teams'));
for (const team of state.teams) {
chips.push(h('button', {
class: 'chip',
type: 'button',
role: 'tab',
'aria-selected': String(teamFilter === String(team.id)),
onclick: () => setTeamFilter(String(team.id)),
text: team.name,
}));
'aria-selected': String(team.id === state.selectedTeamID),
onclick: () => setSelectedTeam(team.id),
},
h('span', { class: `team-dot ${teamColorClass(team.id)}` }),
' ' + team.name,
));
}
}
// A scroll hint for the phone-width row, where the chips can run off the
// right edge with nothing to suggest there's more; the desktop sidebar
// wraps instead of scrolling (see .pane-list .chips), so this fades out
// there via CSS rather than being left out here.
chips.push(h('span', { class: 'chips-fade', 'aria-hidden': 'true' }));
clear(el, chips);
}
@@ -167,8 +179,8 @@ function renderList() {
return;
}
if (!items.length) {
const [title, text] = EMPTY[filter];
clear(el, checklist, emptyState(title, text, filter === 'open' ? 'checkCircle' : null));
const [title, text, iconName] = EMPTY[filter];
clear(el, checklist, emptyState(title, text, iconName));
return;
}
clear(el,
@@ -213,9 +225,12 @@ function row(inc, index) {
dataset: { index: String(index) },
},
h('div', { class: 'row-title', text: inc.title }),
h('div', { class: 'row-age', title: inc.triggered_at, text: age(inc.triggered_at) }),
h('div', { class: 'row-age', title: inc.triggered_at, text: `Triggered ${ago(inc.triggered_at)}` }),
h('div', { class: 'row-meta' },
status,
// The left-border colour alone doesn't say what it means; spell it out
// too, same badge the incident detail page uses for severity.
inc.severity && badge(inc.severity, `plain ${severityClass(inc.severity)}`),
assignee,
team,
labels && h('span', { class: 'labels', text: labels }),
+50 -3
View File
@@ -10,12 +10,53 @@ export const state = {
auth: { password_login: true, oidc: { enabled: false, name: '' } },
open: [], // the default queue: open, not snoozed
teams: [], // the teams the viewer belongs to, each with their role
// Which team the whole app is scoped to right now; null means "All teams".
// Set only through setSelectedTeam below, never assigned directly, so every
// view stays in sync and the choice is remembered across reloads.
selectedTeamID: loadSelectedTeam(),
};
// The team whose schedule and settings the views act on. A viewer in one team —
// which is everybody until somebody makes a second — never has to choose.
const SELECTED_TEAM_KEY = 'terdut.selectedTeam';
function loadSelectedTeam() {
try {
const raw = localStorage.getItem(SELECTED_TEAM_KEY);
return raw ? Number(raw) : null;
} catch {
return null; // storage unavailable, or nothing saved yet
}
}
// Callbacks to run whenever the selected team changes, so every view that
// cares — the queue's filter, the Team settings page, the selector's own
// trigger — stays in sync without a general event bus, following the one
// precedent for this in the codebase: onboarding.js's onRerender.
const teamListeners = [];
export function onTeamChange(cb) {
teamListeners.push(cb);
}
// setSelectedTeam changes which team the app is scoped to (id, or null for
// "All teams"), persists it — a durable preference, unlike the per-tab
// sessionStorage filter this replaces — and tells every registered listener.
export function setSelectedTeam(id) {
state.selectedTeamID = id;
try {
if (id == null) localStorage.removeItem(SELECTED_TEAM_KEY);
else localStorage.setItem(SELECTED_TEAM_KEY, String(id));
} catch {
/* storage unavailable */
}
for (const cb of teamListeners) cb();
}
// The team whose schedule and settings the views act on: the selected team,
// falling back to the first one the viewer belongs to — which is everybody's
// only team until somebody makes a second, or the stored selection naming a
// team the account has since left.
export function currentTeam() {
return state.teams[0] || null;
const teams = state.teams || [];
return teams.find((t) => t.id === state.selectedTeamID) || teams[0] || null;
}
export function myID() {
@@ -36,6 +77,12 @@ export async function users() {
export async function loadTeams() {
state.teams = await api.teams();
// A stored id that no longer names one of the account's teams — left it, or
// this is simply a different account signed in on the same browser — is as
// good as unset.
if (state.selectedTeamID != null && !state.teams.some((t) => t.id === state.selectedTeamID)) {
state.selectedTeamID = null;
}
return state.teams;
}
+1 -1
View File
@@ -80,7 +80,7 @@ function render() {
if (error && !data) body = h('div', { class: 'load-error', text: error });
else if (!data) body = spinner();
else if (!data.incidents.total && !data.byHour.some((x) => x.count)) {
body = emptyState('No data in this range', '', 'chart');
body = emptyState('No data in this range', 'Nothing happened in this window.', 'chart');
} else {
body = h('div', { class: 'stats' },
error && h('div', { class: 'load-error', text: `Showing older data: ${error}` }),
+21 -28
View File
@@ -19,7 +19,7 @@
import * as api from './api.js';
import { h, clear, spinner, confirm, icon, openSheet, closeSheet, menuCard, badge, labelChip, ssoBadge, SSO_MANAGED } from './ui.js';
import { state, currentTeam, users as allUsers, myID } from './state.js';
import { state, currentTeam, onTeamChange, users as allUsers, myID } from './state.js';
import { isoDate, addDays, mondayOf, isoWeek, initial, ago, when, duration } from './format.js';
const view = () => document.getElementById('view-team');
@@ -41,6 +41,9 @@ export const TABS = [
{ tab: 'deadman', path: '/team/deadman', label: 'Switches', title: 'Dead man’s switches' },
];
// Cached from currentTeam() on each refresh(), for the many actions below
// (assignSchedule, addTeamMember, ...) that need a plain id rather than a
// round trip through state.
let teamID = null;
// Which sub-section is open. Remembered rather than passed, because the poll
// loop calls refresh() with no route.
@@ -49,6 +52,13 @@ let data = null; // { team, ... }; which fields are present varies by tab
let error = null;
let freshKey = null; // an integration key, shown once, until the view is left
// The global team selector is what changes which team this page shows now;
// re-fetch under whichever sub-section is open when it fires.
onTeamChange(() => {
data = null;
refresh();
});
export function show(route) {
const next = route?.tab ?? null;
// A different sub-section wants different data, so the old answer goes
@@ -61,13 +71,8 @@ export function show(route) {
refresh();
}
function selectedTeam() {
const teams = state.teams || [];
return teams.find((t) => t.id === teamID) || currentTeam();
}
export async function refresh() {
const team = selectedTeam();
const team = currentTeam();
if (!team) {
data = null;
render();
@@ -172,24 +177,12 @@ function subnav() {
})));
}
// Only shown to somebody in more than one team, like the queue's filter chips.
// It is above the sections rather than inside one because it changes the
// subject of all six.
// Names which team's settings the six sections below belong to. It used to be
// a picker of its own for somebody in more than one team; that job now belongs
// to the global team selector in the nav, which is what onTeamChange above
// reacts to.
function teamPicker() {
if ((state.teams || []).length < 2) {
return h('div', { class: 'card' }, h('h2', { text: data.team.name }));
}
const select = h('select', { class: 'team-picker' },
...state.teams.map((t) => h('option', {
value: String(t.id), text: t.name, selected: t.id === teamID,
})));
select.addEventListener('change', () => {
teamID = Number(select.value);
data = null;
freshKey = null;
refresh();
});
return h('div', { class: 'card' }, h('h2', { text: 'Team' }), select);
return h('h1', { class: 'detail-title', text: data.team.name });
}
// --- overview --------------------------------------------------------------
@@ -211,18 +204,18 @@ function overview() {
menuCard('/team/rota', 'Rota', null,
onToday ? `${onToday.username} is on call today.` : 'Nobody is on call today.'),
menuCard('/team/members', 'Members', (data.members || []).length,
owners === 1 ? 'One owner.' : `${owners} owners.`),
owners === 1 ? '1 owner, full access.' : `${owners} owners, full access.`),
menuCard('/team/escalation', 'Escalation', levels || null,
levels
? `${levels === 1 ? 'One level' : `${levels} levels`}${data.escalation.fallback_topic ? ', then a fallback topic.' : '.'}`
? `${levels} escalation level${levels === 1 ? '' : 's'}${data.escalation.fallback_topic ? ', then a fallback topic.' : '.'}`
: 'No ladder — nobody but the first person is woken.'),
menuCard('/team/sources', 'Alert sources', keys || null,
keys
? (unused ? `${unused} of them never used.` : 'All in use.')
? (unused ? `${unused} of ${keys} never used.` : `${keys} key${keys === 1 ? '' : 's'}, all used recently.`)
: 'No key yet, so nothing can reach this team.'),
menuCard('/team/deadman', 'Dead man’s switches', switches || null,
switches
? (dead ? `${dead} of them silent.` : 'All quiet, as they should be.')
? (dead ? `${dead} of ${switches} silent.` : [icon('checkCircle', 'icon overview-note-icon'), ' Dead man’s switch: healthy.'])
: 'Nothing watched.'),
);
}
+66
View File
@@ -0,0 +1,66 @@
// The global team selector: a small control, once per layout (the desktop
// sidebar and the mobile topbar each have their own button in index.html),
// showing the current team's colour and name — or "All teams" — and opening a
// sheet to switch. Shown only once there is more than one team to choose
// between, the same rule every other team-aware control in this app follows;
// see state.js's currentTeam() for why nobody with just one ever has to.
import { h, icon, openSheet, closeSheet } from './ui.js';
import { state, currentTeam, setSelectedTeam, onTeamChange } from './state.js';
import { teamColorClass } from './format.js';
// Not a real team id (ids are positive), so it can never collide with one —
// the value closeSheet resolves with for "All teams", distinct from the null
// a dismissed sheet resolves with.
const ALL_TEAMS = '__all__';
const buttons = () => [
document.getElementById('team-selector'),
document.getElementById('team-selector-mobile'),
].filter(Boolean);
// init wires the buttons once, at boot. render (below) is what actually fills
// them in and is called again by state.js whenever the selection changes.
export function init() {
for (const btn of buttons()) btn.addEventListener('click', open);
onTeamChange(render);
}
export function render() {
const multiTeam = (state.teams || []).length > 1;
const team = currentTeam();
const label = team ? team.name : 'All teams';
const dotClass = team ? `team-dot ${teamColorClass(team.id)}` : 'team-dot';
for (const btn of buttons()) {
btn.hidden = !multiTeam;
btn.replaceChildren(
h('span', { class: dotClass }),
h('span', { class: 'team-selector-label', text: label }),
// Without this the pill reads as a tag rather than something you can
// open; the chevron is the only thing that says "dropdown" here.
icon('chevronDown', 'icon team-selector-chevron'),
);
}
}
function open() {
const teams = state.teams || [];
openSheet(() => [
h('h2', { class: 'sheet-title', text: 'Switch team' }),
h('ul', { class: 'menu', role: 'menu' },
h('li', {}, h('button', {
class: 'menu-item', type: 'button', role: 'menuitemradio',
'aria-checked': String(state.selectedTeamID == null),
onclick: () => closeSheet(ALL_TEAMS),
}, h('span', { class: 'team-dot' }), ' All teams')),
teams.map((t) => h('li', {}, h('button', {
class: 'menu-item', type: 'button', role: 'menuitemradio',
'aria-checked': String(t.id === state.selectedTeamID),
onclick: () => closeSheet(t.id),
}, h('span', { class: `team-dot ${teamColorClass(t.id)}` }), ' ' + t.name))),
),
]).then((choice) => {
if (choice == null) return; // dismissed: backdrop, escape, or cancel
setSelectedTeam(choice === ALL_TEAMS ? null : choice);
});
}
+6 -1
View File
@@ -52,6 +52,7 @@ const ICONS = {
trash: ['M4 7h16', 'M9 7V4h6v3', 'M6 7l1 13h10l1-13'],
chevronLeft: ['M15 18l-6-6 6-6'],
chevronRight: ['M9 6l6 6-6 6'],
chevronDown: ['M6 9l6 6 6-6'],
external: ['M14 4h6v6', 'M20 4l-9 9', 'M18 14v6H4V6h6'],
logout: ['M15 4h4v16h-4', 'M10 17l5-5-5-5', 'M15 12H4'],
};
@@ -186,12 +187,16 @@ export function labelChip(k, v) {
// One entry in a section's overview: a card that is a link, carrying the count
// only that section can state. Both the Admin tab and the Team tab open on one
// of these menus, and a menu item is a shape rather than a page's own idea.
// note is usually just a string, but can be an array of children instead —
// deadman's "healthy" note below pairs text with a status icon.
export function menuCard(href, label, count, note) {
return h('a', { class: 'card overview-item', href },
h('div', { class: 'overview-head' },
h('strong', { text: label }),
count != null && h('span', { class: 'overview-count', text: String(count) })),
h('p', { class: 'muted small', text: note }),
typeof note === 'string'
? h('p', { class: 'muted small', text: note })
: h('p', { class: 'muted small' }, note),
);
}