Compare commits

...

23 Commits

Author SHA1 Message Date
Niklas Ye 7efd1bbba7 Set the chart's placeholder version to 0.37.0
CI / chart (push) Successful in 1s
CI / test (push) Successful in 7s
CI / security (push) Successful in 12s
Release / test (push) Successful in 6s
Release / chart (push) Successful in 1s
Release / binaries (push) Successful in 21s
Release / image (push) Successful in 1m3s
Release / scan-image (push) Successful in 1s
2026-10-03 16:17:20 +02:00
Niklas Ye 3bf94a5d7f Default to 2 replicas and RollingUpdate now that the singleton jobs are locked
CI / chart (push) Successful in 0s
CI / test (push) Successful in 7s
CI / security (push) Successful in 12s
replicas and strategy: Recreate were the chart's only guard against the
archiver, notifier and migration races; v0.36.0 closed all three with
advisory locks and a conflict-resolving incident insert, which made
that guard redundant rather than load-bearing. Expose replicaCount
(new, no values.yaml key existed before) defaulting to 2, and switch
to strategy: RollingUpdate with no explicit maxUnavailable/maxSurge --
the 25%/25% default rounds to 0/1 at replicaCount: 2, which is already
zero-downtime.

The chart does not gate this on image.tag, so pointing it at a
pre-v0.36.0 image with the new default is a foot-gun by omission --
noted in both the values.yaml comment and the deployment.yaml comment,
not guarded in code, same as the chart does for every other
version-coupled assumption today.

Co-authored-by: Claude <noreply@anthropic.com>
2026-10-03 12:30:27 +02:00
Niklas Ye 5c4e0bdd0e Set the chart's placeholder version to 0.36.0
CI / chart (push) Successful in 1s
CI / security (push) Successful in 16s
CI / test (push) Successful in 4m50s
Release / test (push) Successful in 7s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 48s
Release / image (push) Successful in 1m12s
Release / scan-image (push) Successful in 2s
2026-10-03 12:03:28 +02:00
Niklas Ye 9e5b085d8b Resolve the new-incident insert conflict instead of dropping the payload
CI / chart (push) Successful in 1s
CI / security (push) Successful in 20s
CI / test (push) Has been cancelled
openIncident's INSERT had no ON CONFLICT clause, relying entirely on
incidentForGroup's earlier SELECT to avoid a duplicate. On more than
one replica, two webhook deliveries for the very first occurrence of
a brand-new groupKey can both pass that SELECT before either INSERTs;
the loser then hit incidents_open_group_key_idx's unique violation,
which rolled back its whole transaction — including that payload's
alert upserts, done earlier in the same transaction. ingest's error is
only logged and receiveWebhook answers 200 regardless, so nothing
retried it: the loser's alerts silently never existed.

Add ON CONFLICT (team_id, group_key) WHERE resolved_at IS NULL DO
NOTHING to the INSERT, matching the partial unique index. Postgres
only resolves that conflict after the winning transaction commits (or
rolls back), so by the time RETURNING comes back empty,
existingOpenIncident's follow-up SELECT is guaranteed to see the
winner's row. The loser attaches to it instead of failing outright,
and the rest of its payload commits normally. Covers both callers,
since the dead man's switch sweeper shares this same function.

New test (package api_test, fires N webhook deliveries for one
groupKey from a synchronized start with distinct fingerprints, so
they aren't accidentally serialized by upsertAlerts' own per-
fingerprint lock) confirmed meaningful: with the ON CONFLICT clause
reverted, it fails 10/10 on a missing alert fingerprint; restored, 0/10.
Note while building it: "exactly one incident" alone cannot
distinguish fixed from broken, since the DB's own unique index already
guarantees that either way — the real signal is the loser's payload
surviving.

Chart comment updated: all three of the chart's original reasons for
Recreate are now addressed in code, though replicas stays at 1 and the
strategy stays Recreate pending a deliberate decision to raise it.

Co-authored-by: Claude <noreply@anthropic.com>
2026-10-03 11:58:39 +02:00
Niklas Ye 0050738ca0 Advisory-lock DB migrations against concurrent replica startup
Migrate's check-then-apply loop against schema_migrations had no
locking: two replicas booting at once against a fresh or
partially-migrated database could both pass the "not yet applied"
check for the same file and race applying it, crashing whichever lost
the duplicate-key insert (confirmed: reverting the lock fails the new
test 10/10 on a duplicate-key violation, racing as early as the
CREATE TABLE IF NOT EXISTS schema_migrations statement itself).

Hold a Postgres advisory lock for Migrate's whole run, on a dedicated
connection reserved via db.Conn so lock and unlock happen on the same
session. Blocking (pg_advisory_lock), unlike the archiver/notifier's
pg_try_advisory_lock: on boot there's no later tick to defer to, so a
second replica should wait for the first to finish migrating rather
than skip ahead.

Adds internal/db's first test file, exercising two concurrent Migrate
calls against a fresh schema.

Still open: the new-incident-insert race on a webhook for a brand-new
groupKey, noted in the chart's updated comment. Login rate limiting
staying in-process, diluted across replicas, is an accepted tradeoff.

Co-authored-by: Claude <noreply@anthropic.com>
2026-10-03 11:43:09 +02:00
Niklas Ye 42180948d1 Guard the archiver and notifier passes with a Postgres advisory lock
Both background loops run unconditionally on every instance with no
coordination between them, which the chart's replicas: 1 + strategy:
Recreate exists specifically to paper over: with more than one replica,
every one of them would sweep and deliver notifications independently,
and two overlapping during a rollout would both page for the same
incident.

Add withAdvisoryLock, which takes a Postgres advisory lock on a
dedicated connection and runs a pass only if it gets the lock,
otherwise skipping until the next tick. Wire StartArchiver and
StartNotifier through it with their own lock keys, so Sweep and
NotifySweep themselves are untouched and every existing test calling
them directly keeps working unchanged.

This also closes the notifier's double-delivery race in passing: two
replicas can no longer both be inside deliverPending at once, since
only one can hold notifierLockKey at a time.

Deliberately not addressed here, and still blocking a replica count
above 1: the in-memory login rate limiter, the unlocked migration
runner, and the new-incident-insert race on a webhook for a brand-new
groupKey. Noted in the updated chart comment.

Co-authored-by: Claude <noreply@anthropic.com>
2026-10-03 11:36:48 +02:00
Niklas Ye c83c7c2a8b Set the chart's placeholder version to 0.35.1
CI / chart (push) Successful in 2s
CI / security (push) Successful in 20s
CI / test (push) Successful in 4m45s
Release / test (push) Successful in 3s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 50s
Release / image (push) Successful in 1m21s
Release / scan-image (push) Successful in 3s
Cosmetic, as before (497086c, 4358e84): make helm-package passes
--version and --app-version from the tag, so this decides nothing
about what gets published. Kept in step so the tree heading for
v0.35.1 doesn't say 0.35.0 to a reader who hasn't yet seen the tag.

🤖 Generated with [Claude Code](https://claude.ai/code)

Co-Authored-By: Claude <noreply@anthropic.com>
2026-10-03 09:39:01 +02:00
niklas 43beda9a30 Merge pull request 'Fix the queue chip fade falling short of the right edge' (#30) from fix-chips-fade-edge into main
CI / chart (push) Successful in 1s
CI / security (push) Successful in 21s
CI / test (push) Has been cancelled
2026-10-03 07:36:12 +00:00
Niklas Ye 710521a73c Fix the queue chip fade falling short of the right edge
CI / chart (pull_request) Successful in 1s
CI / security (pull_request) Successful in 16s
CI / test (pull_request) Successful in 5m13s
position: sticky, as a flex item of the row it's pinning itself
against, interacted with that row's gap and its own negative margin in
a way that landed it short of the true edge -- visibly, a sliver of
the next chip stayed poking out past where the fade should have
covered it, which is the "ends before the screen edge" bug reported
against the release.

Replaced with the simpler, better-supported pattern for this: an
absolutely positioned overlay against a position:relative, overflow:
auto parent. Unlike a sticky descendant, an absolutely positioned one
is resolved against the parent's own (non-scrolling) box, so it stays
flush with the real edge regardless of scroll position, without the
flex-gap/margin interaction that caused this.

Verified visually: a standalone reproduction of both versions,
screenshotted with Chromium's headless_shell (no browser automation
tool available in this environment, but the binary's right there) --
the old version shows the next chip's edge peeking past the fade, the
new one doesn't.

🤖 Generated with [Claude Code](https://claude.ai/code)

Co-Authored-By: Claude <noreply@anthropic.com>
2026-10-03 09:33:33 +02:00
niklas d9492913ed Merge pull request 'Fix Stats/Admin/Account showing in the bottom bar too' (#29) from fix-nav-secondary-cascade into main
CI / chart (push) Successful in 0s
CI / test (push) Successful in 8s
CI / security (push) Successful in 14s
Reviewed-on: #29
2026-10-03 07:30:04 +00:00
Niklas Ye 1770e5d945 Fix Stats/Admin/Account showing in the bottom bar too
CI / chart (pull_request) Successful in 1s
CI / security (pull_request) Successful in 16s
CI / test (pull_request) Successful in 5m2s
.nav-link-secondary's display:none sat before .nav-link's own
display:flex in the file. Both are single-class selectors, so they tie
on specificity, and a tie is broken by which one comes later in the
file -- not by which class the element happens to carry. .nav-link's
declaration, being later, won for every element wearing both classes,
so Stats/Admin/Account rendered as three extra tabs on the phone bar
instead of folding into "More" as intended.

Moved the rule below .nav-link instead of changing either declaration,
since nothing about the values was wrong -- only their order was.

🤖 Generated with [Claude Code](https://claude.ai/code)

Co-Authored-By: Claude <noreply@anthropic.com>
2026-10-03 09:22:48 +02:00
Niklas Ye 497086cb51 Set the chart's placeholder version to 0.35.0
CI / test (push) Successful in 4s
CI / chart (push) Successful in 0s
CI / security (push) Successful in 12s
Release / test (push) Successful in 4s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 51s
Release / image (push) Successful in 1m5s
Release / scan-image (push) Successful in 23s
Cosmetic: make helm-package passes --version and --app-version from the
tag, so these two fields decide nothing about what gets published. Kept
in step anyway, the same as 4358e84 and fd26fef before it, so the tree
heading for v0.35.0 does not say 0.34.0 to a reader who hasn't yet seen
the tag.

🤖 Generated with [Claude Code](https://claude.ai/code)

Co-Authored-By: Claude <noreply@anthropic.com>
2026-10-03 09:10:30 +02:00
niklas e3090d2779 Merge pull request 'Web: visual design pass (issue #26)' (#28) from design-polish-issue-26 into main
CI / test (push) Successful in 4s
CI / chart (push) Successful in 1s
CI / security (push) Successful in 13s
Reviewed-on: #28
2026-10-03 07:08:19 +00:00
Niklas Ye 91f03c21e8 Web: visual design pass (issue #26)
CI / chart (pull_request) Successful in 2s
CI / security (pull_request) Successful in 18s
CI / test (pull_request) Successful in 4m49s
Addresses the screenshot-review feedback in #26. No framework or build
step added — all of this stays within the existing plain HTML/CSS/
vanilla-JS + go:embed architecture.

- Nav: re-enable the bottom tab bar that was already built and
  switched off (Queue/On-call/Alerts/Team + a "More" sheet for
  Stats/Admin/Account), replacing the hamburger on phone width.
- Queue: chip counts, a scroll fade on the filter row, a
  "Triggered Xh ago" + severity label per row, a chevron on the team
  switcher so it reads as a dropdown.
- On-call: collapse repeated same-person days into shift bars (week
  view and "your shifts" both), show the week as a date range with
  the ISO week number as secondary text, split "Current shift" out
  from "Next shifts" with "ends in Nd", a pill badge + row highlight
  for "you".
- Incident detail: fix the actual bug behind the duplicate
  "acknowledged" timeline entries (acknowledgeIncident's UPDATE had no
  guard on the incident's current status, so acknowledging an
  already-acknowledged incident silently re-logged the event — now
  idempotent, with regression tests on both the authenticated route
  and the ntfy ack-button route). Relabel escalation re-pages so they
  don't look like the same page landing twice. Copy the primary action
  up near the top. Label the "···" button. Group the timeline by
  phase (triggered/acknowledged/resolved). Add an "at a glance"
  summary row (duration/severity/responsible) and collapse the group
  labels by default.
- Team overview: reword the vague copy ("One owner." etc.) into plain
  labels.
- Empty states: fill in missing icons/one-liners across queue,
  alerts, stats and the incident timeline.
- CSS: fix card padding bugs, verify link contrast already passes AA,
  introduce a --fs-* type-scale token set and migrate the few
  genuinely isolated cases onto it (left sizes tied to a fixed shape,
  a deliberately prominent display, or a non-negotiable constraint
  like the iOS-zoom-prevention input size as documented exceptions
  rather than guess at a render this change can't see).

Verified with the full fmt/lint/test/helm-lint gate, plus a live
instance against the test DB with seeded incidents and schedule data
to trace the on-call grouping and timeline phase-splitting logic
against real API responses.

🤖 Generated with [Claude Code](https://claude.ai/code)

Co-Authored-By: Claude <noreply@anthropic.com>
2026-10-03 09:01:54 +02:00
Niklas Ye 4358e84b24 Set the chart's placeholder version to 0.34.0
CI / test (push) Successful in 4s
CI / chart (push) Successful in 1s
CI / security (push) Successful in 12s
Release / test (push) Successful in 4s
Release / chart (push) Successful in 7s
Release / image (push) Successful in 2m27s
Release / binaries (push) Successful in 2m35s
Release / scan-image (push) Successful in 3s
2026-10-02 22:08:35 +02:00
niklas f45dc2f925 Merge pull request 'internal/api: unify human/service-account authz into one Caller type' (#24) from unify-caller-authz into main
CI / chart (push) Successful in 1s
CI / security (push) Successful in 22s
CI / test (push) Successful in 27s
2026-10-02 20:06:37 +00:00
Niklas Ye 774fdfcaa8 internal/api: unify human/service-account authz into one Caller type
CI / chart (pull_request) Successful in 1s
CI / security (pull_request) Successful in 15s
CI / test (pull_request) Successful in 5m21s
ctxUser/ctxTeams (human) and ctxServiceAccount (+ a synthetic ctxTeams
entry, service account) used to be two parallel, un-unified context
representations -- every authz predicate had to remember which one(s) it
needed, and every place that forgot either wrongly 403'd a service account
(terdut-server#23, terdut-operator#3), crashed on an unchecked zero-value
user id, or silently no-op'd. New internal/api/caller.go collapses both
into one Caller, stored under one ctxCaller key by serveAs/serveAsServiceAccount;
every existing predicate (userFromContext, callerTeamIDs, callerRole,
callerIsAdmin, isInstanceServiceAccount, AdminOnly, requireSelfOrAdmin,
requireTeamOwner, OperatorModeBlock) now reads through it, with identical
behavior for every untouched call site (alerts.go, incidents.go,
schedule.go, stats.go, etc.) -- confirmed by the full existing suite
passing unchanged.

Four real fixes land alongside the refactor, not just the restructuring:

1. callerMayManageServiceAccount gains the one load-bearing branch this
   exists for: an instance-scoped service account may now manage (mint or
   revoke a key on) any team-scoped account, not only a human admin, that
   team's human owner, or the account itself. handleCreateServiceAccount
   already let an instance-scoped caller *create* a team-scoped account for
   any team; adopting or rotating one it didn't just create in the same
   call -- terdut-operator's own documented crash-window recovery -- had no
   equivalent permission and 403'd forever. Closes terdut-operator#3.

2. handleCreateInvite wrote a service-account caller's zero-value user id
   straight into invites.created_by (nullable, but never passed as nil),
   which foreign-key-violates against users(id) -- a 500, not success, for
   any team-scoped service account minting an invite. Fixed the same way
   handleCreateServiceAccount already handles the analogous case. Found
   live while verifying this change, not filed separately since it's fixed
   in the same place it was found.

3. handleMe and handleTestNotification 500'd for a service-account caller
   (fetchUser/the ntfy_topic lookup against a zero-value user id that
   matches no row); handleDismissOnboarding silently no-op'd (UPDATE ...
   WHERE id = 0). All three now call Caller.AsHuman() and return an
   explicit 403 ("this endpoint is for human accounts only").

4. Ratifies, rather than further narrows, two capabilities a team-scoped
   service account already had by construction and this document's own
   text once called "a gap acknowledged rather than closed": owner-equivalent
   reach over membership/invites, and minting another service account for
   its own team. terdut-operator's new TerdutTeam invite-minting feature is
   about to depend on the first one, so this makes it documented, tested,
   intentional behavior instead of an accident nobody was supposed to rely
   on.

AdminOnly/requireSelfOrAdmin are unchanged in effect: still human-only,
forever, for every scope of service account -- confirmed by
TestAdminOnly_RefusesEveryServiceAccountScope. terdut-server#23's named
routes (POST /api/users, PUT /api/admin/settings) were never the right
thing to widen; its real fix is the terdut-operator invite feature,
recorded in SERVICE-ACCOUNTS.md's "What this unblocks" and closing that
issue once it ships.

SERVICE-ACCOUNTS.md amended in place (not a new file, its own established
convention) to describe the as-built Caller model, correct its own
aspirational claim about AdminOnly that TEAM-LOOKUP.md had already flagged
as not matching shipped code, and record all of the above.
2026-10-02 21:52:43 +02:00
Niklas Ye fd26fef1ba Set the chart's placeholder version to 0.33.2
CI / test (push) Successful in 4s
CI / chart (push) Successful in 1s
CI / security (push) Successful in 12s
Release / test (push) Successful in 3s
Release / chart (push) Successful in 3s
Release / binaries (push) Successful in 16s
Release / image (push) Successful in 1m8s
Release / scan-image (push) Successful in 1s
make helm-package passes --version and --app-version from the tag, so
these fields decide nothing about what is published -- but a tree heading
for v0.33.2 that still says 0.33.1 tells its reader something false.
Same as 4b15079 and 6f8499f before it.
2026-10-02 09:42:50 +02:00
Niklas Ye a9d788cc83 Wait for Postgres to accept connections before the main container starts
A Deployment created before Postgres has finished its very first boot --
initdb plus Patroni leader election, on a from-scratch postgres-operator
cluster -- crash-looped a few times. db.Open()'s own ping-retry budget
(pingAttempts/pingRetryDelay, internal/db/db.go) is sized for a much
shorter, different race -- NetworkPolicy propagation, a few seconds -- not
for genuine first-time cluster creation, which routinely takes longer, so
it exhausted and the process exited before ever binding its HTTP port. A
startupProbe cannot fix that: the crash happens before there is anything
to probe.

Added a wait-for-postgres init container instead: it loops pg_isready
against database.dsn until Postgres actually answers, before the main
container's own, unchanged retry budget gets a chance to run out.
pg_isready needs no credentials -- it reports PQPING_OK on anything that
amounts to a Postgres backend answering, including an auth challenge --
so no PGPASSWORD is wired into it.

Chart-only; no Go code changed. database.waitForPostgres.enabled defaults
to true and can be turned off if something else already guarantees
Postgres is reachable before this Deployment is created.
2026-10-02 09:42:42 +02:00
Niklas Ye 4b15079ac2 Set the chart's placeholder version to 0.33.1
CI / chart (push) Successful in 1s
CI / security (push) Successful in 19s
CI / test (push) Successful in 4m42s
Release / test (push) Successful in 3s
Release / chart (push) Successful in 2s
Release / image (push) Successful in 2m11s
Release / scan-image (push) Successful in 3s
Release / binaries (push) Successful in 2m47s
Cosmetic: `make helm-package` passes --version/--app-version from the
tag, so this field decides nothing about what gets published. Still done
so the tree doesn't say 0.33.0 while heading for a v0.33.1 release.
Cites fc9f47c, the fix this version actually is.
2026-10-02 09:18:28 +02:00
Niklas Ye fc9f47cc8d Refuse an OIDC sign-in from creating the very first user
Closes the race terdut-operator#1 found: /api/bootstrap and OIDC
auto-provisioning both key off the same signal (SELECT COUNT(*) FROM
users) with no coordination between them, so an otherwise-ordinary OIDC
sign-in against a freshly-created, not-yet-bootstrapped install could
create user #1 itself and take the one slot /api/bootstrap expects to win
uncontested (terdut-operator's own design, DESIGN.md §1/§6, assumes it is
the only caller). The operator has no way to recover from losing that
race -- it never gets a credential, and nothing it owns can clear the
occupying user row.

resolveSSOUser now checks the same gate handleBootstrap already does,
right where it's about to create a brand-new user (an identity nobody has
linked yet, that also matches no existing local account by email) -- not
anywhere else, since every other sign-in on an already-bootstrapped
install is unaffected. New sso_error code `not_bootstrapped`: the person
sees "this install is still setting up, try again in a moment" and a
second attempt once something has actually bootstrapped succeeds
normally, same as any other first sign-in.

Does not fix the other half of that issue (BootstrapStateLost's own
"delete and recreate" instructions still don't work once something has
occupied the slot some other way) -- this closes the specific race, not
every path to that state.
2026-10-02 09:12:32 +02:00
Niklas Ye bc9f793f1f Add GET /api/teams?name= (TEAM-LOOKUP.md)
CI / chart (push) Successful in 1s
CI / security (push) Successful in 1m15s
CI / test (push) Successful in 5m48s
Resolves the gap TEAM-LOOKUP.md raised: an instance-scoped service
account had no way to recover a team's id after a 409 on POST
/api/teams, unlike the already-solved equivalent for service accounts
themselves (GET /api/service-accounts?name=).

Same route, extended the same way handleListServiceAccounts already
branches on ?name=: unset behaves exactly as before (the caller's own
teams via team_members); set looks up one team by exact name, open to
any authenticated caller -- not gated by isInstanceServiceAccount or
AdminOnly, since what it discloses (a name is taken, nothing about who's
in it) is the same low sensitivity that lookup already accepts for
service-account names.

Tests cover the exact motivating scenario (create, 409 on a retry,
recover the id via ?name=), the empty-array-not-an-error case, that no
role/source is reported for a non-member match, and that a caller who
isn't a member of the matched team still gets it.
2026-10-01 10:49:11 +02:00
Niklas Ye 871274a3a0 Request: team lookup for service accounts (TEAM-LOOKUP.md)
Raised by terdut-operator's TerdutTeam controller (ROADMAP.md Stage 2):
an instance-scoped service account has no way to recover a team's id
after a 409 on POST /api/teams, unlike the equivalent, already-solved
case for service accounts themselves (GET /api/service-accounts?name=).
Also corrects a claim in SERVICE-ACCOUNTS.md's own text that doesn't
match AdminOnly's actual code -- confirmed against source, not assumed.
2026-10-01 10:43:09 +02:00
41 changed files with 1762 additions and 215 deletions
+1 -1
View File
@@ -793,7 +793,7 @@ attempted.
| `POST` | `/api/oidc/device/token` | `{"device_code"}` → `202 {"status":"pending"}`, then `200` with the session cookie once approved (once only). `410` with `{"error":"expired"}` or `{"error":"denied"}`; `429 {"error":"slow_down"}` if polled faster than `interval` |
| `POST` | `/api/oidc/device/approve` | **session** — `{"user_code"}`. Approves a pending device login as the caller. `403` for an API key; `404` for an unknown, expired or already decided code |
| `POST` | `/api/oidc/device/deny` | **session** — `{"user_code"}`. Refuses it |
| `GET` | `/api/oidc/callback` | Where the provider sends the browser back. Sets the session cookie and redirects to `/`, or to `/?sso_error=<code>` — one of `denied`, `expired`, `failed`, `unavailable`, `not_allowed`, `no_email`, `email_conflict`, `disabled` |
| `GET` | `/api/oidc/callback` | Where the provider sends the browser back. Sets the session cookie and redirects to `/`, or to `/?sso_error=<code>` — one of `denied`, `expired`, `failed`, `unavailable`, `not_allowed`, `no_email`, `email_conflict`, `disabled`, `not_bootstrapped` (no user exists on this install yet — sign in again once something has called `/api/bootstrap`) |
| `POST` | `/api/logout` | Ends the session and clears the cookie |
| `GET` | `/api/me` | The caller: `{user, has_password}` |
+108 -31
View File
@@ -119,38 +119,100 @@ new key, revoke the old one," not "recreate the account."
### Auth middleware
`internal/api/middleware.go`'s existing dual resolution (`Authorization: Bearer`
→ `apiKeyUser()`, or session cookie → `sessionUser()`, both landing on the same
`models.User` + team-membership context) gains a third path: a bearer token that
hashes to a `service_account_keys.key_hash` resolves to a distinct principal
type, not a synthesized `models.User`. `requireTeamMember`/`requireTeamOwner`
treat a matching team-scoped service account as owner-equivalent for that one
team (satisfies the same checks a real team owner would), and an instance-scoped
one as satisfying `AdminOnly` for team-creation/listing purposes **and** for
minting a `team`-scoped service account against any team (`POST
/api/service-accounts {"scope":"team","teamID":...}`) — this second permission
is what lets an operator-style caller create a team, then immediately mint that
team its own narrower credential, without a human in the loop for every team.
Neither permission extends to user-management endpoints (`POST /api/users`,
`PUT /api/users/{id}/admin`, etc.), which stay human-admin-only. Anywhere
identity is recorded for a human (incident timeline
`acknowledged_by`/`assigned_to`, audit-relevant fields), a service-account
principal is stored and displayed distinctly, e.g. `service-account:terdut-operator`,
never coerced into a `user_id` FK.
**Revised** (this section originally described an aspiration that didn't
match what shipped — `TEAM-LOOKUP.md` already caught one instance of that,
and a fuller audit found three more; this is the corrected, as-built
description, not the original proposal).
**Team scope, as implemented, is owner-equivalent for every `requireTeamOwner`
endpoint, membership and invites included — nothing server-side carves those
two out.** That's broader than what `terdut-operator`'s CRDs actually need
(escalation/deadman/integrations/OIDC-bindings only; membership is explicitly
never gitops-managed, see its DESIGN.md §4.2), a gap acknowledged rather than
closed here: narrowing this to exclude
`POST/DELETE /api/teams/{teamID}/members*` and
`.../invites*` specifically for a service-account caller is a small, isolated
follow-up (special-case those handlers rather than `requireTeamOwner` itself,
which every other owner-gated endpoint still wants shared). Until then, what
actually keeps membership out of automation's hands is that no operator built
against this scope should ever call those two endpoints — not a server-side
refusal.
`internal/api/middleware.go`'s dual resolution (`Authorization: Bearer` →
`apiKeyUser()`, or session cookie → `sessionUser()`) and the service-account
path (`serviceAccountFor()`) both resolve into one `Caller` type
(`internal/api/caller.go`), not two parallel, un-unified context
representations the way an earlier version of this server kept them. Every
authorization predicate reads `Caller`'s methods:
- `Caller.IsAdmin()` — true **only** for a human system administrator, never
for a service account of either scope, under any circumstance. `AdminOnly`
and `requireSelfOrAdmin` key on this alone — user management
(`POST /api/users`, `PUT /api/users/{id}/admin`, etc.) and
`GET/PUT /api/admin/settings` stay human-only, forever. The original text
here claimed an instance-scoped service account satisfies `AdminOnly` "for
team-creation/listing purposes" — that was never true of the shipped code
(`TEAM-LOOKUP.md` caught the listing half; the creation half was always a
separate, bespoke check in `handleCreateTeam`, not `AdminOnly` itself) and
is not being made true now. Don't widen `AdminOnly`: every time this has
come up, the fix has been a narrower, purpose-built capability instead
(`?name=` lookups for teams and service accounts; now
`terdut-operator`'s own invite-minting feature for the one real gap this
boundary left — how a human ever gets a first login on a no-OIDC,
operator-managed install. See the bottom of "What this unblocks.")
- `Caller.IsInstanceServiceAccount()` — true only for an instance-scoped
service account, never for a human (including a human admin).
`handleCreateTeam` uses exactly this: a human creates a team by being a
human (and becomes its owner); an instance-scoped service account creates
one with no human owner at all. The two paths are not interchangeable, so
this predicate deliberately does not also admit a human admin.
- `Caller.Role(teamID)`/`TeamIDs()` — a human's real `team_members` rows, or
a team-scoped service account's single synthetic owner membership
(`serveAsServiceAccount`). This is what makes `requireTeamMember`/
`requireTeamOwner` treat a team-scoped service account as owner-equivalent
for that one team, with no separate branch needed in either function.
- `Caller.ServiceAccountID()` — used by `OperatorModeBlock` ("any service
account passes") and by `callerMayManageServiceAccount`'s self-rotation
check.
- `Caller.AsHuman()` — the accessor every handler that needs a real
`user_id` to act on behalf of must call and check, instead of reading a
user off context unconditionally. Before the `Caller` type existed, four
handlers did the latter and silently misbehaved for a service-account
caller: `handleMe` and `handleTestNotification` 500'd (a zero-value user id
that matches no row), `handleDismissOnboarding` silently no-op'd (`UPDATE
... WHERE id = 0` affects nothing, still returns 204), and
`handleCreateInvite` wrote that same zero value into `invites.created_by`
— a real foreign-key violation, not just a wrong answer, since that column
is nullable but was never passed as `nil`. All four now call `AsHuman()`
and return an explicit 403 ("this endpoint is for human accounts only")
or, for the invite case, leave `created_by` `NULL` the same way
`handleCreateServiceAccount` already did for the analogous situation.
**Team scope is owner-equivalent for every `requireTeamOwner` endpoint,
membership and invites included — by design, not by an unclosed gap.** An
earlier version of this document flagged this as "acknowledged rather than
closed," kept in check only by the social convention that nobody *builds*
automation against those two routes. That convention is retired:
`terdut-operator`'s `TerdutTeam` controller now mints and revokes its own
team's invite link through exactly this capability (its existing
team-scoped credential, `POST`/`DELETE /api/teams/{teamID}/invites`), which
is the real fix for the human-onboarding gap below — not a narrower
carve-out of this capability. `service_accounts_test.go`'s
`TestServiceAccount_TeamScopeManagesItsOwnInvites` pins it.
**A team-scoped account can also mint another service account scoped to its
own team** (`handleCreateServiceAccount`'s `callerOwnsTeam` branch, which a
team-scoped caller already satisfies for its own team via the synthetic
membership above). Kept, not restricted, for the same reason: a team-scoped
credential is that team's owner's reach, full stop — carving this one
capability out while leaving membership/invites alone would be an arbitrary
asymmetry. Pinned by
`TestServiceAccount_TeamScopeCanMintAnotherAccountForItsOwnTeam`.
**`callerMayManageServiceAccount` gained the one load-bearing fix this
redesign exists for:** an instance-scoped service account may manage
(mint/revoke a key on) *any* team-scoped account, not only one admin, that
team's human owner, or the account itself. `handleCreateServiceAccount`
already let an instance-scoped caller *create* a team-scoped account for
any team; this closes the gap where adopting or rotating one it didn't just
create in the same call — exactly `terdut-operator`'s documented
adopt-on-409 crash-window recovery (its own `DESIGN.md` §5) — 403'd forever
instead of succeeding (`terdut-operator#3`). Pinned by
`TestServiceAccount_InstanceScopeAdoptsAnExistingTeamScopedAccountsKey`.
Anywhere identity is recorded for a human (incident timeline
`acknowledged_by`/`assigned_to`, audit-relevant fields), a service-account
caller is still coerced into a bare `user_id` of `0` today — `Caller`'s new
`Identity()` accessor exists for exactly this follow-up, but wiring it in
needs a schema migration (an actor-attribution column distinct from
`user_id`) and is deliberately out of scope here. Tracked separately, not by
this document.
## What this unblocks
@@ -177,6 +239,21 @@ Directly resolves `terdut-operator` DESIGN.md §6's two broken assumptions:
holding many credentials in one place (the operator's namespace) an
acceptable trade rather than reintroducing the mirrored design's
server-admin-equivalent-everywhere problem.
4. **A human can get a first login on a no-OIDC, operator-managed install —
without ever touching `AdminOnly` or `/api/admin/settings`.** This was
filed as `terdut-server#23` ("no API path to create a human login after
bootstrap") and diagnosed, at the time, as this server needing to let a
service account through `AdminOnly`. It doesn't: the fix lives entirely
in `terdut-operator`, because a team-scoped credential was *already*
owner-equivalent for `POST /api/teams/{teamID}/invites`, and invite
redemption (`POST /api/signup` with an `invite` token) bypasses
`signup_mode` entirely — `terdut-operator` just never grew a feature to
use either fact. Its `TerdutTeam` controller now mints and surfaces one
via its own existing team-scoped credential (`spec.invite`,
`status.inviteSecretRef`, see that repo's own docs), so a human joins a
CRD-managed team by a real invite link, the same way anyone else would.
`terdut-server#23` is closed with this note once that feature ships — its
named routes stay human-only, correctly, not a gap.
## Suggested sequencing
+73
View File
@@ -0,0 +1,73 @@
# Team lookup for service accounts: closing terdut-operator's create-path crash window
This is a design note for a feature, not an implementation plan — same posture as
`SERVICE-ACCOUNTS.md`, and raised for the same reason: `terdut-operator`'s `TerdutTeam`
controller (ROADMAP.md Stage 2, a separate repo, no shared code) hit a gap this server has
no answer for yet.
## The problem
`POST /api/teams` (`handleCreateTeam`, confirmed against `internal/api/teams.go`) lets an
instance-scoped service account create a team — it has its own explicit
`isInstanceServiceAccount(...)` branch alongside the human-user path, not gated by
`AdminOnly`. If that call succeeds server-side but the caller (`TerdutTeam`'s controller)
crashes before persisting the resulting team ID locally, a retry's `POST` 409s on the name's
unique constraint (confirmed: the `isUniqueViolation` branch in the same handler).
Recovering from that 409 means looking the team up by name, and nothing today permits that
for a service account:
- `GET /api/teams` (`handleListTeams`) answers "what teams does the *caller* belong to", via
a `team_members` join keyed on `userFromContext`'s `caller.ID` — confirmed against source.
A service account is never a member of anything, so this always returns empty for one,
regardless of what exists.
- `GET /api/admin/teams` (`handleAdminListTeams`) is gated by `AdminOnly`, and `AdminOnly`'s
actual code (`internal/api/middleware.go`) checks only `userFromContext(...).IsAdmin` — no
branch for a service account at all, confirmed against source. This contradicts
`SERVICE-ACCOUNTS.md`'s own text, which claims "an instance-scoped [service account
satisfies] `AdminOnly` for team-creation/listing purposes" — that claim doesn't match this
endpoint's actual, shipped code. (Team *creation* is fine: `handleCreateTeam` isn't behind
`AdminOnly` at all, it has its own check. Only the listing half of that sentence is wrong.)
This is exactly the shape of gap `SERVICE-ACCOUNTS.md`'s own `GET /api/service-accounts?name=`
closed for service accounts themselves (confirmed: that endpoint's own comment —
"the name lookup is open to any authenticated caller... what lets a service account find its
own account on the 403 that follows a second POST"). Teams never got the equivalent, because
nothing needed it until an operator started creating them unattended.
## Goals
- A service-account-accessible way to look up one team by exact name, mirroring
`GET /api/service-accounts?name=` as closely as possible — same shape, same reasoning,
same low sensitivity of what it discloses.
- No change to today's behavior for an empty/no-name request.
## Proposed shape
Extend `GET /api/teams` itself, the same way `handleListServiceAccounts` already branches on
a `?name=` query param, rather than adding a new route:
- `name` unset (today's behavior, unchanged): the caller's own teams, via `team_members`.
- `name=<value>` set: look up that one team by exact name — a one-or-zero-length array, not
an error on no match, mirroring `GET /api/service-accounts?name=`'s own response shape and
status codes exactly. Deliberately **not** gated by `isInstanceServiceAccount` or
`AdminOnly`: a human caller who's already a member sees this same information in their own
team list regardless, and a non-member learning only that a name is taken — not who's in
the team, not any of its data — is the same low-sensitivity disclosure
`GET /api/service-accounts?name=` already accepts for service-account names.
## What this unblocks
Directly resolves the crash-window gap in `terdut-operator`'s `TerdutTeam` controller: on a
409 from `POST /api/teams`, `GET /api/teams?name=<the same name>` — authenticated with the
same instance-scoped credential that just got the 409 — finds the id, and the controller
proceeds as if its own create had returned it directly. The same adopt-on-409 pattern already
proven for service accounts (that repo's `DESIGN.md` §6 point 1, §5's general rule), not a
new one.
## Suggested sequencing
Land this before `TerdutTeam`'s create path is implemented — the same reasoning
`SERVICE-ACCOUNTS.md` gave for its own sequencing: writing that code against today's gap as a
"known-temporary workaround" is wasted effort when the fix is this small and this
well-precedented.
+2 -2
View File
@@ -15,5 +15,5 @@ type: application
# appVersion and image.tag in values.yaml no longer agree, and that is not an oversight:
# image.tag stays "latest", which is what a local install actually pulls. appVersion is
# metadata and drives nothing.
version: 0.33.0
appVersion: "v0.33.0"
version: 0.37.0
appVersion: "v0.37.0"
+26 -5
View File
@@ -6,21 +6,42 @@ metadata:
labels:
{{- include "terdut-server.labels" . | nindent 4 }}
spec:
replicas: 1
replicas: {{ .Values.replicaCount }}
selector:
matchLabels:
{{- include "terdut-server.selectorLabels" . | nindent 6 }}
# Recreate, not RollingUpdate, even though the PVC that forced it is gone: the
# sweeper and the notifier are unsynchronised singletons, and two replicas
# overlapping during a rollout would both page for the same incident.
# RollingUpdate, not Recreate: the sweeper, notifier and migration runner
# each take a Postgres advisory lock around their own pass, and new-incident
# creation on the first webhook for a brand-new groupKey resolves its own
# insert conflict -- so two replicas overlapping during a rollout no longer
# double-page, race a migration, or drop a webhook payload (v0.36.0). No
# explicit maxUnavailable/maxSurge: the 25%/25% default rounds to 0/1 at
# replicaCount: 2, which is zero-downtime already.
strategy:
type: Recreate
type: RollingUpdate
template:
metadata:
labels:
{{- include "terdut-server.selectorLabels" . | nindent 8 }}
spec:
enableServiceLinks: false
{{- if .Values.database.waitForPostgres.enabled }}
initContainers:
- name: wait-for-postgres
image: "{{ .Values.database.waitForPostgres.image.repository }}:{{ .Values.database.waitForPostgres.image.tag }}"
imagePullPolicy: {{ .Values.database.waitForPostgres.image.pullPolicy }}
env:
- name: TERDUT_DB_DSN
value: {{ required "database.dsn is required" .Values.database.dsn | quote }}
command:
- sh
- -c
- |
until pg_isready -d "$TERDUT_DB_DSN"; do
echo "wait-for-postgres: not ready yet, retrying in 2s"
sleep 2
done
{{- end }}
containers:
- name: terdut-server
image: "{{ .Values.image.repository }}:{{ .Values.image.tag }}"
+27
View File
@@ -1,3 +1,10 @@
# Safe above 1 since v0.36.0: the sweeper, notifier and migration runner each
# take a Postgres advisory lock around their own pass, and a webhook that
# loses the race to open a brand-new incident attaches to the winner's row
# instead of dropping its payload. An image older than v0.36.0 does not have
# these guards -- do not raise this against one.
replicaCount: 2
networking:
hostname: "terdut.example.com"
servicePort: 8080
@@ -33,6 +40,26 @@ database:
passwordSecret:
name: ""
key: password
# Blocks the main container from starting until Postgres accepts
# connections. Without this, a Deployment created before Postgres has
# finished its very first boot -- initdb plus Patroni leader election, on a
# from-scratch postgres-operator cluster -- crash-loops a few times: the
# app's own ping-retry budget on startup (pingAttempts/pingRetryDelay in
# internal/db/db.go) is sized for a much shorter, different race --
# NetworkPolicy propagation, a few seconds -- not for genuine first-time
# cluster creation, which routinely takes longer, so it exhausts and the
# process exits before ever binding its HTTP port. A startupProbe cannot
# help here: the crash happens before there is anything to probe.
#
# pg_isready needs no credentials -- it reports PQPING_OK on anything that
# amounts to "a Postgres backend answered", including an auth challenge --
# so no PGPASSWORD is wired into this container.
waitForPostgres:
enabled: true
image:
repository: postgres
tag: "17-alpine"
pullPolicy: IfNotPresent
service:
type: ClusterIP
+122
View File
@@ -0,0 +1,122 @@
package api
// This file is internal (package api, not api_test) because withAdvisoryLock is
// unexported and these tests exercise its locking semantics directly rather than
// through the full StartArchiver/StartNotifier loop, which would make the "does
// not run while held" case timing-dependent instead of deterministic. It opens a
// plain connection to TERDUT_TEST_DSN rather than reusing testdb_test.go's
// newTestDB, since that helper lives in the separate, already-compiled
// api_test package and a Postgres advisory lock needs no schema or migration
// to exercise.
import (
"context"
"database/sql"
"os"
"testing"
_ "github.com/jackc/pgx/v5/stdlib"
)
// advisoryTestDB opens a plain, unmigrated connection to the test database. An
// unset DSN fails rather than skips, matching testdb_test.go's rationale: a
// suite that quietly tests nothing is worse than one that does not run.
func advisoryTestDB(t *testing.T) *sql.DB {
t.Helper()
dsn := os.Getenv("TERDUT_TEST_DSN")
if dsn == "" {
t.Fatalf("TERDUT_TEST_DSN is not set: these tests need Postgres.\n" +
"Run `make test-db` for a local one, then\n" +
" export TERDUT_TEST_DSN=postgres://terdut:terdut@localhost:5432/terdut_test?sslmode=disable")
}
db, err := sql.Open("pgx", dsn)
if err != nil {
t.Fatalf("connect to TERDUT_TEST_DSN: %v", err)
}
t.Cleanup(func() { db.Close() })
return db
}
func TestWithAdvisoryLock_RunsWhenFree(t *testing.T) {
db := advisoryTestDB(t)
ctx := context.Background()
ran := false
withAdvisoryLock(ctx, db, archiverLockKey, "test", func() { ran = true })
if !ran {
t.Fatal("fn did not run although the lock was free")
}
}
func TestWithAdvisoryLock_SkipsWhileHeldElsewhere(t *testing.T) {
db := advisoryTestDB(t)
ctx := context.Background()
// Hold the lock on a connection of our own, standing in for another
// replica mid-pass.
holder, err := db.Conn(ctx)
if err != nil {
t.Fatalf("acquire holder connection: %v", err)
}
defer holder.Close()
if _, err := holder.ExecContext(ctx, "SELECT pg_advisory_lock($1)", archiverLockKey); err != nil {
t.Fatalf("pre-acquire lock: %v", err)
}
ran := false
withAdvisoryLock(ctx, db, archiverLockKey, "test", func() { ran = true })
if ran {
t.Fatal("fn ran although another connection already held the lock")
}
if _, err := holder.ExecContext(ctx, "SELECT pg_advisory_unlock($1)", archiverLockKey); err != nil {
t.Fatalf("release held lock: %v", err)
}
// Now that the holder released it, the next caller should get it.
ran = false
withAdvisoryLock(ctx, db, archiverLockKey, "test", func() { ran = true })
if !ran {
t.Fatal("fn did not run after the other connection released the lock")
}
}
func TestWithAdvisoryLock_ReleasesAfterFnReturns(t *testing.T) {
db := advisoryTestDB(t)
ctx := context.Background()
withAdvisoryLock(ctx, db, notifierLockKey, "test", func() {})
// If the first call had leaked the lock, this one would see it held and
// skip, leaving ran false.
ran := false
withAdvisoryLock(ctx, db, notifierLockKey, "test", func() { ran = true })
if !ran {
t.Fatal("fn did not run on a later call: the earlier call leaked its lock")
}
}
func TestWithAdvisoryLock_KeysAreIndependent(t *testing.T) {
db := advisoryTestDB(t)
ctx := context.Background()
holder, err := db.Conn(ctx)
if err != nil {
t.Fatalf("acquire holder connection: %v", err)
}
defer holder.Close()
if _, err := holder.ExecContext(ctx, "SELECT pg_advisory_lock($1)", archiverLockKey); err != nil {
t.Fatalf("pre-acquire archiver lock: %v", err)
}
defer holder.ExecContext(ctx, "SELECT pg_advisory_unlock($1)", archiverLockKey)
// Holding archiverLockKey must not block notifierLockKey.
ran := false
withAdvisoryLock(ctx, db, notifierLockKey, "test", func() { ran = true })
if !ran {
t.Fatal("fn did not run under a different key although only archiverLockKey was held")
}
}
+35 -1
View File
@@ -376,6 +376,17 @@ func incidentForGroup(ctx context.Context, tx *sql.Tx, notify NotifyConfig, team
// its own. Hence the querier rather than a *sql.Tx. A nil severity leaves the
// column for refreshSeverity to fill from the member alerts; the sweeper passes
// one because its incidents have no members to derive it from.
//
// Both callers get here only after their own SELECT found no open incident for
// this group_key — but on more than one replica, two webhook deliveries for the
// very first occurrence of a brand-new group_key can both pass that SELECT
// before either INSERTs. ON CONFLICT DO NOTHING against
// incidents_open_group_key_idx is what makes the loser's INSERT a no-op instead
// of a unique-violation error that would otherwise roll back its entire
// payload; existingOpenIncident then hands it the winner's row. Postgres
// resolves that conflict only once the winner's transaction has committed (or
// rolled back), so by the time this RETURNING comes back empty, the winner's
// row is guaranteed visible to that follow-up SELECT.
func openIncident(ctx context.Context, q querier, notify NotifyConfig, teamID int64, groupKey, title string, groupLabels map[string]string, severity *string) (int64, error) {
onCall, err := currentOnCall(ctx, q, teamID)
if err != nil {
@@ -391,10 +402,18 @@ func openIncident(ctx context.Context, q querier, notify NotifyConfig, teamID in
err = q.QueryRowContext(ctx, `
INSERT INTO incidents (team_id, group_key, title, group_labels, signature, status, severity, triggered_at, assigned_to)
VALUES ($1, $2, $3, $4::jsonb, $5, 'triggered', $6, $7, $8)
ON CONFLICT (team_id, group_key) WHERE resolved_at IS NULL DO NOTHING
RETURNING id`,
teamID, groupKey, title, string(labelsJSON), incidentSignature(groupLabels, title), severity,
time.Now().Unix(), onCall).Scan(&id)
if err != nil {
switch {
case err == sql.ErrNoRows:
// Lost the race: someone else's incident for this group_key exists now.
// Everything below — the trigger event, assignment, page, escalation
// clock — already happened for that row when it was created; attach to
// it rather than fail this call (and the whole payload) outright.
return existingOpenIncident(ctx, q, teamID, groupKey)
case err != nil:
return 0, err
}
@@ -423,6 +442,21 @@ func openIncident(ctx context.Context, q querier, notify NotifyConfig, teamID in
return id, nil
}
// existingOpenIncident looks up the open incident openIncident's own INSERT just
// lost a conflict against — the same lookup incidentForGroup does before ever
// calling openIncident, repeated here for the caller that arrived second.
func existingOpenIncident(ctx context.Context, q querier, teamID int64, groupKey string) (int64, error) {
var id int64
err := q.QueryRowContext(ctx,
"SELECT id FROM incidents WHERE team_id = $1 AND group_key = $2 AND resolved_at IS NULL",
teamID, groupKey,
).Scan(&id)
if err != nil {
return 0, err
}
return id, nil
}
// linkAlert adds an alert to an incident, emitting a timeline entry only the
// first time. Re-sends of an already-linked alert are silent.
func linkAlert(ctx context.Context, tx *sql.Tx, incidentID, alertID int64) error {
+112
View File
@@ -0,0 +1,112 @@
package api_test
import (
"bytes"
"encoding/json"
"fmt"
"net/http"
"sync"
"testing"
)
// TestWebhook_ConcurrentFirstOccurrenceOpensOneIncident reproduces two
// replicas racing the very first webhook delivery for a brand-new group_key:
// both see no open incident yet (incidentForGroup's own SELECT finds
// nothing) and race openIncident's INSERT.
//
// The DB's own unique index already guarantees at most one incident either
// way, with or without this fix — so "exactly one incident" alone cannot
// tell a fixed run from a broken one. What ON CONFLICT handling actually
// changes is what happens to the *loser*: before it, the loser's INSERT hit
// incidents_open_group_key_idx's unique violation, which — since
// upsertAlerts ran earlier in that same transaction — rolled back its whole
// payload, alert insert included. ingest's error is only logged and
// receiveWebhook answers 200 regardless, so nothing ever retried it: the
// loser's alert silently never existed. That is the regression signal this
// test checks — every caller's fingerprint must show up in /api/alerts, not
// just the winner's.
func TestWebhook_ConcurrentFirstOccurrenceOpensOneIncident(t *testing.T) {
s := newTS(t)
const callers = 8
const groupKey = "race-group"
// Every caller needs its own fingerprint. A shared one would serialize all
// of them at upsertAlerts' own ON CONFLICT (team_id, fingerprint) row lock,
// long before any of them reached incidentForGroup — which would hide the
// very race this test exists to force.
bodies := make([][]byte, callers)
for i := range callers {
payload := map[string]any{
"version": "4",
"status": "firing",
"groupKey": groupKey,
"groupLabels": map[string]string{"alertname": "RaceAlert"},
"alerts": []map[string]any{amAlert(fmt.Sprintf("fp-race-%d", i), "RaceAlert", "firing",
"2026-05-20T10:00:00Z", "0001-01-01T00:00:00Z", nil)},
}
bodies[i], _ = json.Marshal(payload)
}
// A start line, so every request is fired as close to simultaneously as
// goroutine scheduling allows, rather than trickling out one dial at a
// time — the race window is the gap between incidentForGroup's SELECT and
// openIncident's INSERT, which a staggered start could easily miss.
var ready sync.WaitGroup
start := make(chan struct{})
statuses := make([]int, callers)
var wg sync.WaitGroup
for i := range callers {
ready.Add(1)
wg.Add(1)
go func(i int) {
defer wg.Done()
ready.Done()
<-start
resp, err := http.Post(s.URL+"/api/integrations/"+s.ingestKey+"/alertmanager",
"application/json", bytes.NewReader(bodies[i]))
if err != nil {
t.Errorf("post webhook #%d: %v", i, err)
return
}
defer resp.Body.Close()
statuses[i] = resp.StatusCode
}(i)
}
ready.Wait()
close(start)
wg.Wait()
for i, code := range statuses {
if code != http.StatusOK {
t.Errorf("webhook #%d returned %d, want 200", i, code)
}
}
var matched []any
for _, inc := range listIncidents(t, s, "") {
if inc["group_key"] == groupKey {
matched = append(matched, inc["id"])
}
}
if len(matched) != 1 {
t.Fatalf("expected exactly 1 incident for group_key %q after %d concurrent deliveries, got %d: %v",
groupKey, callers, len(matched), matched)
}
var alerts []map[string]any
decode(t, s.req(t, http.MethodGet, "/api/alerts", nil), &alerts)
seen := map[string]bool{}
for _, a := range alerts {
if fp, ok := a["fingerprint"].(string); ok {
seen[fp] = true
}
}
for i := range callers {
fp := fmt.Sprintf("fp-race-%d", i)
if !seen[fp] {
t.Errorf("alert %q is missing: its delivery's whole payload was silently rolled back "+
"when it lost the race for the incident", fp)
}
}
}
+18 -2
View File
@@ -15,6 +15,12 @@ const (
// expiryGrace absorbs clock skew and notification latency before an alert
// whose ends_at watermark has passed is treated as stale.
expiryGrace = 5 * time.Minute
// archiverLockKey is the Postgres advisory lock the sweeper takes for the
// duration of each pass, so that running more than one replica does not run
// the sweep concurrently on all of them. Its value has no meaning beyond
// being distinct from notifierLockKey.
archiverLockKey int64 = 7265_0001
)
// StartArchiver runs the alert sweeper until ctx is cancelled, starting with an
@@ -23,15 +29,25 @@ const (
// the fallback, not the setting: each pass reads the current value from the
// settings table, so an administrator's change takes effect on the next tick
// instead of at the next restart.
//
// Each pass runs under archiverLockKey (see withAdvisoryLock), so that on more
// than one replica only whichever instance's tick takes the lock first actually
// sweeps; the rest skip that tick rather than racing the same pass.
func StartArchiver(ctx context.Context, db *sql.DB, archiveAfter, staleAfter time.Duration, notify NotifyConfig) {
ticker := time.NewTicker(sweepInterval)
defer ticker.Stop()
Sweep(ctx, db, archiveAfter, staleAfter, notify)
sweep := func() {
withAdvisoryLock(ctx, db, archiverLockKey, "sweeper", func() {
Sweep(ctx, db, archiveAfter, staleAfter, notify)
})
}
sweep()
for {
select {
case <-ticker.C:
Sweep(ctx, db, archiveAfter, staleAfter, notify)
sweep()
case <-ctx.Done():
return
}
+5 -1
View File
@@ -279,7 +279,11 @@ type meResponse struct {
// between the login form and the app, since it cannot read its own cookie.
func handleMe(db *sql.DB) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
caller, _ := userFromContext(r.Context())
caller, ok := userFromContext(r.Context())
if !ok {
respond(w, http.StatusForbidden, errResp("this endpoint is for human accounts only"))
return
}
user, err := fetchUser(r.Context(), db, caller.ID)
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
+123
View File
@@ -0,0 +1,123 @@
package api
import (
"context"
"fmt"
"git.ryuvia.com/niklas/terdut-server/internal/models"
)
// Caller is the one principal type every authorization predicate in this
// package reads from. Before this, a human (ctxUser + ctxTeams) and a
// service account (ctxServiceAccount + a synthetic ctxTeams entry) were two
// parallel, un-unified context representations — every predicate had to
// remember which one(s) it needed to check, and the ones that forgot either
// 403'd a service account that should have been let through (terdut-server#23,
// terdut-operator#3), crashed on an unchecked zero-value user ID (handleMe,
// handleTestNotification), or silently no-op'd (handleDismissOnboarding).
// serveAs and serveAsServiceAccount now both build exactly one Caller and
// store it under one context key; everything else in this file is a read
// of one of its methods.
type Caller struct {
// user is set for a human caller (session cookie or a user's own API
// key), nil for a service account of either scope.
user *models.User
// sa is set for a service-account caller, nil for a human.
sa *serviceAccountPrincipal
// memberships is the caller's real team_members rows for a human, or —
// for a team-scoped service account — the single synthetic owner
// membership serveAsServiceAccount injects (see its own comment for
// why). Always nil for an instance-scoped service account: it acts on
// teams by id, not by belonging to one.
memberships []membership
}
// AsHuman returns the real user behind this caller, or false for a service
// account of either scope. Every handler that needs a real user_id to act
// on behalf of — not just "is this caller sufficiently privileged" — calls
// this and handles the false case explicitly, replacing the unchecked
// userFromContext(ctx) zero-value reads that used to silently misbehave for
// a service-account caller.
func (c Caller) AsHuman() (models.User, bool) {
if c.user == nil {
return models.User{}, false
}
return *c.user, true
}
// IsAdmin is true only for a human system administrator — never for a
// service account, of either scope, under any circumstance. AdminOnly and
// requireSelfOrAdmin key on this and nothing else: user management and
// /api/admin/settings stay human-only forever, by design (SERVICE-ACCOUNTS.md).
func (c Caller) IsAdmin() bool {
return c.user != nil && c.user.IsAdmin
}
// IsInstanceServiceAccount reports whether this caller is specifically an
// instance-scoped service account — never true for a human, including a
// human admin. handleCreateTeam needs exactly this: a human creates a team
// by being a human (and becomes its owner as a side effect), an
// instance-scoped service account creates one with no human owner at all;
// the two paths are not interchangeable, so this predicate must not also
// admit a human admin the way MayActAsInstanceAdmin deliberately does.
func (c Caller) IsInstanceServiceAccount() bool {
return c.sa != nil && c.sa.scope == models.ServiceAccountScopeInstance
}
// Role reports the caller's role in teamID, and whether they belong to it
// at all.
func (c Caller) Role(teamID int64) (string, bool) {
for _, m := range c.memberships {
if m.teamID == teamID {
return m.role, true
}
}
return "", false
}
// TeamIDs lists every team this caller belongs to: a human's real
// memberships, or a team-scoped service account's own single team. Always
// empty for an instance-scoped service account.
func (c Caller) TeamIDs() []int64 {
ids := make([]int64, 0, len(c.memberships))
for _, m := range c.memberships {
ids = append(ids, m.teamID)
}
return ids
}
// ServiceAccountID reports this caller's own service-account id, for the
// "may manage/rotate its own credential" self-check in
// callerMayManageServiceAccount, and for OperatorModeBlock's "any service
// account passes" rule.
func (c Caller) ServiceAccountID() (int64, bool) {
if c.sa == nil {
return 0, false
}
return c.sa.id, true
}
// Identity is a stable, log/audit-facing string distinguishing a human
// caller from a service account — "user:42" or "service-account:7". Not
// wired into any database column today (incidents.go's acknowledged_by/
// assigned_to/user_id are explicitly out of scope for this change — that
// needs its own schema migration, tracked separately), but this is the one
// place in the request path that already knows which kind of caller this
// is, and that follow-up will want exactly this accessor.
func (c Caller) Identity() string {
switch {
case c.user != nil:
return fmt.Sprintf("user:%d", c.user.ID)
case c.sa != nil:
return fmt.Sprintf("service-account:%d", c.sa.id)
default:
return "unknown"
}
}
func callerFromContext(ctx context.Context) (Caller, bool) {
c, ok := ctx.Value(ctxCaller).(Caller)
return c, ok
}
+39
View File
@@ -1,8 +1,11 @@
package api
import (
"context"
"database/sql"
"encoding/json"
"errors"
"log"
"net/http"
"strconv"
"strings"
@@ -77,3 +80,39 @@ func decodeJSON(r *http.Request, v any) error {
func errResp(msg string) map[string]string {
return map[string]string{"error": msg}
}
// withAdvisoryLock runs fn only if it can take the named Postgres advisory lock on a
// dedicated connection, and skips fn otherwise. This is what keeps the archiver and
// notifier safe to run on more than one replica: whichever instance's tick gets there
// first does the work; the rest see the lock held and simply wait for their next tick
// instead of running the same pass concurrently.
//
// pg_try_advisory_lock is session-scoped, so taking and releasing it must happen on the
// same connection, reserved via db.Conn rather than borrowed from the pool's shared
// connections fn itself may use — and released (unlocked, then closed) before returning,
// since a session lock otherwise outlives this call and leaks onto whatever reuses the
// pooled connection next.
func withAdvisoryLock(ctx context.Context, db *sql.DB, key int64, name string, fn func()) {
conn, err := db.Conn(ctx)
if err != nil {
log.Printf("%s: advisory lock: acquire connection: %v", name, err)
return
}
defer conn.Close()
var locked bool
if err := conn.QueryRowContext(ctx, "SELECT pg_try_advisory_lock($1)", key).Scan(&locked); err != nil {
log.Printf("%s: advisory lock: %v", name, err)
return
}
if !locked {
return // another replica is already running this pass
}
defer func() {
if _, err := conn.ExecContext(ctx, "SELECT pg_advisory_unlock($1)", key); err != nil {
log.Printf("%s: advisory unlock: %v", name, err)
}
}()
fn()
}
+7 -4
View File
@@ -261,14 +261,17 @@ func resolveIfSettled(ctx context.Context, q querier, incidentID int64) (bool, e
}
// acknowledgeIncident records that userID has picked an incident up, and reports
// whether it changed anything — an already-resolved incident is left alone.
// Shared by the authenticated handler and the Acknowledge button in a push
// notification, so both write the same state and the same timeline entry.
// whether it changed anything — an already-resolved or already-acknowledged
// incident is left alone, so a second acknowledge (a retried request, or a
// stale push notification tapped after the web UI already acked it) is a
// no-op rather than a second "acknowledged" timeline entry. Shared by the
// authenticated handler and the Acknowledge button in a push notification,
// so both write the same state and the same timeline entry.
func acknowledgeIncident(ctx context.Context, q querier, incidentID, userID int64) (bool, error) {
res, err := q.ExecContext(ctx, `
UPDATE incidents
SET status = 'acknowledged', acknowledged_by = $1, acknowledged_at = $2
WHERE id = $3 AND resolved_at IS NULL`,
WHERE id = $3 AND status = 'triggered'`,
userID, time.Now().Unix(), incidentID)
if err != nil {
return false, err
+12 -2
View File
@@ -197,10 +197,20 @@ func handleIncidentAcknowledge(db *sql.DB) http.HandlerFunc {
return
}
if !acked {
if !incidentExists(w, r, db, id) {
// incidentIDParam above already confirmed the incident exists, so this
// is either resolved, or already acknowledged — the latter is now a
// no-op rather than an error, since the caller's desired state
// (acknowledged) already holds.
inc, err := fetchIncident(r.Context(), db, id)
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
respond(w, http.StatusConflict, errResp("incident is resolved"))
if inc.Status == "resolved" {
respond(w, http.StatusConflict, errResp("incident is resolved"))
return
}
respond(w, http.StatusOK, inc)
return
}
respondIncident(w, r, db, id)
+42
View File
@@ -338,6 +338,48 @@ func TestIncident_Acknowledge(t *testing.T) {
}
}
// A second acknowledge — a retried request, or a stale push notification
// tapped after the web UI already acked it — must be a no-op: same state,
// no second "acknowledged" timeline entry. Regression test for the bug
// described in issue #26 ("two acknowledged entries look like a bug").
func TestIncident_AcknowledgeTwiceIsIdempotent(t *testing.T) {
s := newTS(t)
postWebhook(t, s, []map[string]any{
amAlert("fp-ack2", "Y", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
})
resp := s.req(t, http.MethodPost, "/api/incidents/1/acknowledge", nil)
if resp.StatusCode != http.StatusOK {
t.Fatalf("first acknowledge returned %d", resp.StatusCode)
}
var first map[string]any
decode(t, resp, &first)
resp = s.req(t, http.MethodPost, "/api/incidents/1/acknowledge", nil)
if resp.StatusCode != http.StatusOK {
t.Fatalf("second acknowledge returned %d, want 200 (idempotent)", resp.StatusCode)
}
var second map[string]any
decode(t, resp, &second)
if second["status"] != "acknowledged" {
t.Errorf("expected status still acknowledged, got %v", second["status"])
}
if second["acknowledged_by"] != first["acknowledged_by"] {
t.Errorf("expected the same acknowledged_by, got %v then %v", first["acknowledged_by"], second["acknowledged_by"])
}
types := eventTypes(timeline(t, s, 1))
n := 0
for _, ty := range types {
if ty == "acknowledged" {
n++
}
}
if n != 1 {
t.Errorf("expected exactly one acknowledged event, got %d in %v", n, types)
}
}
func TestIncident_ManualResolveIsTerminal(t *testing.T) {
s := newTS(t)
postWebhook(t, s, []map[string]any{
+29 -36
View File
@@ -16,10 +16,13 @@ import (
type contextKey string
const (
ctxUser contextKey = "user"
ctxSession contextKey = "session"
ctxTeams contextKey = "teams"
ctxServiceAccount contextKey = "service_account"
// ctxCaller holds the one Caller (see caller.go) every authorization
// predicate in this package reads from — a human and a service account
// used to be two parallel, un-unified context keys (ctxUser/ctxTeams vs.
// ctxServiceAccount); this is why that was a mistake, not a smaller
// version of the same idea.
ctxCaller contextKey = "caller"
ctxSession contextKey = "session"
)
// AuthMiddleware accepts either of the two credentials the server issues: an
@@ -179,8 +182,7 @@ func serveAs(w http.ResponseWriter, r *http.Request, next http.Handler, db *sql.
return
}
ctx := context.WithValue(r.Context(), ctxTeams, teams)
ctx = context.WithValue(ctx, ctxUser, u)
ctx := context.WithValue(r.Context(), ctxCaller, Caller{user: &u, memberships: teams})
if sessionID != 0 {
ctx = context.WithValue(ctx, ctxSession, sessionID)
}
@@ -192,9 +194,13 @@ func hashToken(token string) string {
return hex.EncodeToString(h[:])
}
// userFromContext is a thin compatibility wrapper over Caller.AsHuman(), so
// every call site written before the Caller abstraction (alerts.go,
// incidents.go, schedule.go, stats.go, and more) needs no change and keeps
// its exact existing behavior.
func userFromContext(ctx context.Context) (models.User, bool) {
u, ok := ctx.Value(ctxUser).(models.User)
return u, ok
c, _ := callerFromContext(ctx)
return c.AsHuman()
}
// serviceAccountPrincipal is a service account as resolved from its key:
@@ -244,26 +250,21 @@ func serviceAccountFor(ctx context.Context, db *sql.DB, token string) (serviceAc
// key is only ever set by the client that holds it, never attached by a
// browser to a request another site makes.
func serveAsServiceAccount(w http.ResponseWriter, r *http.Request, next http.Handler, sa serviceAccountPrincipal) {
ctx := r.Context()
var memberships []membership
if sa.scope == models.ServiceAccountScopeTeam {
ctx = context.WithValue(ctx, ctxTeams, []membership{{teamID: sa.teamID, role: models.RoleOwner}})
memberships = []membership{{teamID: sa.teamID, role: models.RoleOwner}}
}
ctx = context.WithValue(ctx, ctxServiceAccount, sa)
ctx := context.WithValue(r.Context(), ctxCaller, Caller{sa: &sa, memberships: memberships})
next.ServeHTTP(w, r.WithContext(ctx))
}
func serviceAccountFromContext(ctx context.Context) (serviceAccountPrincipal, bool) {
sa, ok := ctx.Value(ctxServiceAccount).(serviceAccountPrincipal)
return sa, ok
}
// isInstanceServiceAccount reports whether the caller is an instance-scoped
// service account — the one identity allowed to create a team and mint a
// team-scoped account against any of them, the two things system
// administration can already do that this extends to automation.
// isInstanceServiceAccount is a thin compatibility wrapper over
// Caller.IsInstanceServiceAccount(), for call sites outside this package's
// core predicates (handleCreateTeam, handleCreateServiceAccount) that
// needed this exact, narrow check before the Caller abstraction existed.
func isInstanceServiceAccount(ctx context.Context) bool {
sa, ok := serviceAccountFromContext(ctx)
return ok && sa.scope == models.ServiceAccountScopeInstance
c, _ := callerFromContext(ctx)
return c.IsInstanceServiceAccount()
}
// operatorReason marks a write that operator mode refused as such, distinct
@@ -291,7 +292,8 @@ func OperatorModeBlock(cfg config.Config) func(http.Handler) http.Handler {
return next
}
return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
if _, ok := serviceAccountFromContext(r.Context()); ok {
caller, _ := callerFromContext(r.Context())
if _, ok := caller.ServiceAccountID(); ok {
next.ServeHTTP(w, r)
return
}
@@ -333,24 +335,15 @@ func callerMemberships(ctx context.Context, db *sql.DB, userID int64) ([]members
// administration is about accounts, not about reading other people's incidents,
// and an admin who needs to see a team's queue can add themselves to it.
func callerTeamIDs(ctx context.Context) []int64 {
ms, _ := ctx.Value(ctxTeams).([]membership)
ids := make([]int64, 0, len(ms))
for _, m := range ms {
ids = append(ids, m.teamID)
}
return ids
c, _ := callerFromContext(ctx)
return c.TeamIDs()
}
// callerRole reports the caller's role in one team, and whether they are in it
// at all.
func callerRole(ctx context.Context, teamID int64) (string, bool) {
ms, _ := ctx.Value(ctxTeams).([]membership)
for _, m := range ms {
if m.teamID == teamID {
return m.role, true
}
}
return "", false
c, _ := callerFromContext(ctx)
return c.Role(teamID)
}
// requireTeamMember answers the request and reports false unless the caller
+19 -2
View File
@@ -38,6 +38,13 @@ const (
ackTokenTTL = 24 * time.Hour
)
// notifierLockKey is the Postgres advisory lock the notifier takes for the
// duration of each pass, so that running more than one replica does not
// deliver (or double-deliver) the same notification from more than one of
// them at once. Its value has no meaning beyond being distinct from
// archiverLockKey.
const notifierLockKey int64 = 7265_0002
// Notification kinds, recording why a push was sent.
const (
notifyTriggered = "triggered"
@@ -97,6 +104,10 @@ var notifyClient = &http.Client{Timeout: 10 * time.Second}
// StartNotifier delivers queued notifications until ctx is cancelled, starting
// with an immediate pass so a restart flushes whatever the last one left behind.
//
// Each pass runs under notifierLockKey (see withAdvisoryLock), so that on more
// than one replica only whichever instance's tick takes the lock first actually
// delivers; the rest skip that tick rather than racing the same pass.
func StartNotifier(ctx context.Context, db *sql.DB, cfg NotifyConfig) {
if !cfg.enabled() {
log.Print("notifier: disabled (no ntfy URL configured)")
@@ -107,11 +118,17 @@ func StartNotifier(ctx context.Context, db *sql.DB, cfg NotifyConfig) {
ticker := time.NewTicker(notifyInterval)
defer ticker.Stop()
NotifySweep(ctx, db, cfg)
sweep := func() {
withAdvisoryLock(ctx, db, notifierLockKey, "notifier", func() {
NotifySweep(ctx, db, cfg)
})
}
sweep()
for {
select {
case <-ticker.C:
NotifySweep(ctx, db, cfg)
sweep()
case <-ctx.Done():
return
}
+11 -4
View File
@@ -66,12 +66,19 @@ func handleNotifyAck(db *sql.DB) http.HandlerFunc {
return
}
if !acked {
// The incident closed between the page and the tap. Nothing to do,
// and nothing the responder did wrong — report the state, not an error,
// so ntfy shows a success toast rather than a failure.
// Either the incident closed between the page and the tap, or it was
// already acknowledged (e.g. from the web UI, or an earlier tap of
// the same button) — either way nothing the responder did wrong, so
// report the actual state rather than assuming "resolved", and let
// ntfy show a success toast rather than a failure.
inc, err := fetchIncident(r.Context(), db, incidentID)
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
respond(w, http.StatusOK, map[string]any{
"incident_id": incidentID,
"status": "resolved",
"status": inc.Status,
})
return
}
+41
View File
@@ -408,6 +408,47 @@ func TestNotify_AckButtonAcknowledgesIncident(t *testing.T) {
}
}
// The ack token isn't single-use (it stays valid for a day, in case the
// first tap never reaches the server), so tapping the same notification's
// Acknowledge button twice is a real scenario, not just a retried request.
// It must report the incident's actual state, not assume "resolved" —
// see handleNotifyAck's !acked branch — and must not log a second
// "acknowledged" event.
func TestNotify_AckButtonTwiceIsIdempotent(t *testing.T) {
s, f := notifyTS(t, api.NotifyConfig{PublicURL: "https://terdut.example.com"})
fireCritical(t, s)
s.sweepNotify(t)
ackURL := f.messages()[0].Actions[0].URL
path := ackURL[strings.Index(ackURL, "/api/notify/ack/"):]
for i := range 2 {
resp, err := http.Post(s.URL+path, "application/json", nil)
if err != nil {
t.Fatalf("ack %d: %v", i+1, err)
}
var body map[string]any
decode(t, resp, &body)
if resp.StatusCode != http.StatusOK {
t.Fatalf("ack %d returned %d", i+1, resp.StatusCode)
}
if body["status"] != "acknowledged" {
t.Errorf("ack %d: expected status acknowledged, got %v", i+1, body["status"])
}
}
n := 0
for _, ty := range eventTypes(timeline(t, s, 1)) {
if ty == "acknowledged" {
n++
}
}
if n != 1 {
t.Errorf("expected exactly one acknowledged event after two taps, got %d", n)
}
}
func TestNotify_AckRejectsUnknownToken(t *testing.T) {
s, _ := notifyTS(t, api.NotifyConfig{PublicURL: "https://terdut.example.com"})
fireCritical(t, s)
+25
View File
@@ -48,6 +48,17 @@ const (
ssoNoEmail ssoError = "no_email" // the provider sent no email address
ssoEmailConflict ssoError = "email_conflict" // a local account has this email and cannot be linked
ssoDisabled ssoError = "disabled" // the linked account is disabled
// ssoNotBootstrapped: this identity has no existing account, and no user
// exists on this install yet either -- creating one here would race
// POST /api/bootstrap for the one gitops-managed installs expect to win
// it (terdut-operator's own DESIGN.md §1, §6), which has no way to
// recover if it loses. The person sees this for at most as long as it
// takes whatever is bootstrapping this install to finish; signing in
// again afterward hits the ordinary first-sign-in path. Found by
// terdut-operator#1: nothing stopped an otherwise-ordinary OIDC sign-in
// from quietly winning this race against an operator that assumed it
// was the only caller.
ssoNotBootstrapped ssoError = "not_bootstrapped"
)
// handleAuthConfig says how this server can be signed in to, so the login form
@@ -366,6 +377,20 @@ func resolveSSOUser(ctx context.Context, tx *sql.Tx, cfg config.OIDC, id *oidc.I
return 0, ssoEmailConflict
}
case errors.Is(err, sql.ErrNoRows):
// Creating the very first user is /api/bootstrap's own job (same
// gate, same table: SELECT COUNT(*) FROM users in handleBootstrap).
// An identity nobody has linked yet, on an install with no users at
// all, is exactly the race terdut-operator#1 found: whoever gets
// here first wins a slot the other side has no way to recover from
// losing. Refusing it here costs an otherwise-ordinary sign-in
// nothing but a retry once bootstrap has actually run.
var userCount int
if err := tx.QueryRowContext(ctx, "SELECT COUNT(*) FROM users").Scan(&userCount); err != nil {
return 0, err
}
if userCount == 0 {
return 0, ssoNotBootstrapped
}
userID, err = createSSOUser(ctx, tx, id)
if err != nil {
return 0, err
+39
View File
@@ -543,6 +543,45 @@ func TestSSO_DisabledUserIsRefused(t *testing.T) {
}
}
// TestSSO_FirstUserIsRefusedUntilBootstrap is terdut-operator#1: an
// otherwise-ordinary OIDC sign-in against a brand-new, not-yet-bootstrapped
// install must not be allowed to create the first user and win the race
// POST /api/bootstrap expects to win uncontested. Built directly over
// api.NewRouter rather than newSSOTS/newTS, both of which bootstrap before
// a test body ever runs -- exactly the state this test needs to not have yet.
func TestSSO_FirstUserIsRefusedUntilBootstrap(t *testing.T) {
idp := newFakeIdP(t)
database := newTestDB(t)
srv := httptest.NewServer(api.NewRouter(database, api.NotifyConfig{PublicURL: "http://terdut.test"}, ssoConfig(idp), "test"))
t.Cleanup(srv.Close)
first := newBrowser(t, srv.URL)
first.CheckRedirect = func(*http.Request, []*http.Request) error { return http.ErrUseLastResponse }
if loc := signInSSO(t, idp, first, alice); loc != "/?sso_error=not_bootstrapped" {
t.Fatalf("sent to %q, want not_bootstrapped", loc)
}
// Bootstrap the install for real, the way terdut-operator's own
// reconcileBootstrap does.
resp, err := http.Post(srv.URL+"/api/bootstrap", "application/json",
strings.NewReader(`{"username":"admin","email":"admin@test.com"}`))
if err != nil {
t.Fatal(err)
}
resp.Body.Close()
if resp.StatusCode != http.StatusCreated {
t.Fatalf("bootstrap: %d", resp.StatusCode)
}
// The same identity, signing in again, is this install's ordinary first
// SSO user now -- no longer refused.
second := newBrowser(t, srv.URL)
second.CheckRedirect = func(*http.Request, []*http.Request) error { return http.ErrUseLastResponse }
if loc := signInSSO(t, idp, second, alice); loc != "/" {
t.Errorf("sent to %q after bootstrap, want success", loc)
}
}
func TestSSO_NoEmailIsRefused(t *testing.T) {
idp := newFakeIdP(t)
s := newSSOTS(t, idp)
+31 -8
View File
@@ -41,10 +41,14 @@ func callerIsAdmin(ctx context.Context) bool {
return ok && u.IsAdmin
}
// callerOwnsTeam reports whether the caller is a human owner of teamID. Built
// on callerRole/ctxTeams like requireTeamOwner, but without writing a
// response: callers here need to combine it with other ways of being
// allowed, not stop at the first no.
// callerOwnsTeam reports whether the caller is owner-equivalent for teamID:
// a human owner, or that team's own team-scoped service account (its single
// synthetic membership, serveAsServiceAccount — ratified in
// SERVICE-ACCOUNTS.md as intentional, not an accident: a team-scoped
// credential is that team's owner's reach, full stop, membership and
// invites included). Built on callerRole like requireTeamOwner, but without
// writing a response: callers here need to combine it with other ways of
// being allowed, not stop at the first no.
func callerOwnsTeam(ctx context.Context, teamID int64) bool {
role, ok := callerRole(ctx, teamID)
return ok && role == models.RoleOwner
@@ -173,9 +177,24 @@ func fetchServiceAccount(ctx context.Context, db *sql.DB, id int64) (models.Serv
// callerMayManageServiceAccount reports whether the caller may mint or revoke
// a key on sa: a system administrator, that team-scoped account's own human
// owner, or the account rotating its own credential — which is not a
// privilege escalation, the same reasoning requireSelfOrAdmin already rests
// on for a user's own API keys.
// owner, the account rotating its own credential (not a privilege
// escalation, the same reasoning requireSelfOrAdmin already rests on for a
// user's own API keys) — or, new, an instance-scoped service account
// managing any team-scoped account.
//
// That last branch closes terdut-operator#3: handleCreateServiceAccount
// already lets an instance-scoped caller *create* a team-scoped account for
// any team (the branch below it, isInstanceServiceAccount(ctx)) — this
// account didn't have an equivalent reach to *adopt or rotate* one it
// didn't just create in the same call, which is exactly the recovery path
// terdut-operator's own documented crash-window handling depends on
// (DESIGN.md §5's general adopt-on-conflict rule): a reconcile that creates
// the account successfully but crashes before persisting its credential
// locally retries into a 409, and without this branch the only available
// recovery — minting a fresh key on the now-existing account — 403'd
// forever, with no way out. Granting it here is not a new power: it
// mirrors the create-time reach this scope already has, just extended to
// the retry path DESIGN.md's own crash-window reasoning requires.
func callerMayManageServiceAccount(ctx context.Context, sa models.ServiceAccount) bool {
if callerIsAdmin(ctx) {
return true
@@ -183,7 +202,11 @@ func callerMayManageServiceAccount(ctx context.Context, sa models.ServiceAccount
if sa.TeamID != nil && callerOwnsTeam(ctx, *sa.TeamID) {
return true
}
if self, ok := serviceAccountFromContext(ctx); ok && self.id == sa.ID {
caller, _ := callerFromContext(ctx)
if id, ok := caller.ServiceAccountID(); ok && id == sa.ID {
return true
}
if sa.TeamID != nil && caller.IsInstanceServiceAccount() {
return true
}
return false
+126
View File
@@ -181,10 +181,136 @@ func TestServiceAccount_TeamScopeManagesItsResources(t *testing.T) {
resp2.Body.Close()
}
// Documents the capability already granted at create time (handleCreateServiceAccount's
// own callerOwnsTeam branch) also applies here: a team-scoped account is that
// team's owner's reach, membership and further accounts included, not just
// the handful of endpoints exercised above. Kept, not restricted, for
// symmetry with the now-ratified membership/invite capability below.
func TestServiceAccount_TeamScopeCanMintAnotherAccountForItsOwnTeam(t *testing.T) {
s := newTS(t)
instanceKey := createServiceAccount(t, s, s.key, "terdut-operator", models.ServiceAccountScopeInstance, 0)
teamA := createTeamAs(t, s, instanceKey, "team-a")
keyA := createServiceAccount(t, s, instanceKey, "team-a-sa", models.ServiceAccountScopeTeam, teamA)
resp := s.reqAs(t, keyA, http.MethodPost, "/api/service-accounts",
map[string]any{"name": "team-a-sa-2", "scope": models.ServiceAccountScopeTeam, "team_id": teamA})
if resp.StatusCode != http.StatusCreated {
t.Errorf("team-scoped account minting another account for its own team: %d", resp.StatusCode)
}
resp.Body.Close()
}
// SERVICE-ACCOUNTS.md ratifies this explicitly: a team-scoped account is
// owner-equivalent for every requireTeamOwner endpoint, membership and
// invites included — terdut-operator's own invite-minting feature depends on
// exactly this. No test exercised handleCreateInvite from a service account
// before this change, and it would have 500'd (created_by written as a bare
// zero value against a NOT-validated-but-FK'd column) rather than succeeded;
// see the signup_test.go addition for that half.
func TestServiceAccount_TeamScopeManagesItsOwnInvites(t *testing.T) {
s := newTS(t)
instanceKey := createServiceAccount(t, s, s.key, "terdut-operator", models.ServiceAccountScopeInstance, 0)
teamA := createTeamAs(t, s, instanceKey, "team-a")
keyA := createServiceAccount(t, s, instanceKey, "team-a-sa", models.ServiceAccountScopeTeam, teamA)
var invite struct {
ID int64 `json:"id"`
}
resp := s.reqAs(t, keyA, http.MethodPost, "/api/teams/"+id64(teamA)+"/invites", map[string]any{})
if resp.StatusCode != http.StatusCreated {
t.Fatalf("team-scoped account creating an invite: %d", resp.StatusCode)
}
decode(t, resp, &invite)
var list []map[string]any
decode(t, s.reqAs(t, keyA, http.MethodGet, "/api/teams/"+id64(teamA)+"/invites", nil), &list)
if len(list) != 1 {
t.Errorf("expected the invite to list back, got %d", len(list))
}
if resp := s.reqAs(t, keyA, http.MethodDelete,
"/api/teams/"+id64(teamA)+"/invites/"+id64(invite.ID), nil); resp.StatusCode != http.StatusNoContent {
t.Errorf("team-scoped account revoking its own invite: %d", resp.StatusCode)
} else {
resp.Body.Close()
}
}
// ---------------------------------------------------------------------------
// Key rotation
// ---------------------------------------------------------------------------
// terdut-operator#3: an instance-scoped account is already trusted to CREATE
// a team-scoped account for any team (handleCreateServiceAccount's own
// isInstanceServiceAccount branch) — this pins that it is equally trusted to
// manage/rotate a key on one that already exists and that it did not just
// create in this call, which is the exact shape of terdut-operator's own
// crash-window recovery (mint succeeds, a later step is interrupted before
// persisting the credential locally, and the next reconcile retries into a
// 409 then needs to mint a fresh key on the now-existing account). Before
// this fix, the second POST .../keys below 403'd forever.
func TestServiceAccount_InstanceScopeAdoptsAnExistingTeamScopedAccountsKey(t *testing.T) {
s := newTS(t)
instanceKey := createServiceAccount(t, s, s.key, "terdut-operator", models.ServiceAccountScopeInstance, 0)
teamA := createTeamAs(t, s, instanceKey, "team-a")
resp := s.reqAs(t, instanceKey, http.MethodPost, "/api/service-accounts",
map[string]any{"name": "team-a-sa", "scope": models.ServiceAccountScopeTeam, "team_id": teamA})
if resp.StatusCode != http.StatusCreated {
t.Fatalf("create team-scoped account: %d", resp.StatusCode)
}
var created struct {
ServiceAccount struct {
ID int64 `json:"id"`
} `json:"service_account"`
}
decode(t, resp, &created)
// Simulates the adopt-on-409 recovery path: this instance-scoped caller
// did not just create this account in this call (a fresh *tdclient.Client
// request, same as a second, independent reconcile would issue), yet
// still needs to mint it a fresh key.
rotateResp := s.reqAs(t, instanceKey, http.MethodPost,
"/api/service-accounts/"+id64(created.ServiceAccount.ID)+"/keys", map[string]string{"name": "adopted"})
if rotateResp.StatusCode != http.StatusCreated {
t.Errorf("instance-scoped account adopting a team-scoped account's key: %d", rotateResp.StatusCode)
}
rotateResp.Body.Close()
}
// ---------------------------------------------------------------------------
// AdminOnly / requireSelfOrAdmin — unchanged after the Caller refactor
// ---------------------------------------------------------------------------
// The Caller abstraction must not have widened AdminOnly/requireSelfOrAdmin:
// user management and /api/admin/settings stay human-only, for every scope
// of service account, exactly as before.
func TestAdminOnly_RefusesEveryServiceAccountScope(t *testing.T) {
s := newTS(t)
instanceKey := createServiceAccount(t, s, s.key, "terdut-operator", models.ServiceAccountScopeInstance, 0)
teamA := createTeamAs(t, s, instanceKey, "team-a")
teamKey := createServiceAccount(t, s, instanceKey, "team-a-sa", models.ServiceAccountScopeTeam, teamA)
for _, key := range []string{instanceKey, teamKey} {
if resp := s.reqAs(t, key, http.MethodGet, "/api/admin/settings", nil); resp.StatusCode != http.StatusForbidden {
t.Errorf("expected 403 for a service account reading /api/admin/settings, got %d", resp.StatusCode)
} else {
resp.Body.Close()
}
if resp := s.reqAs(t, key, http.MethodPost, "/api/users",
map[string]string{"username": "nope", "email": "nope@example.com"}); resp.StatusCode != http.StatusForbidden {
t.Errorf("expected 403 for a service account creating a user, got %d", resp.StatusCode)
} else {
resp.Body.Close()
}
if resp := s.reqAs(t, key, http.MethodGet, "/api/me", nil); resp.StatusCode != http.StatusForbidden {
t.Errorf("expected 403 for a service account calling /api/me, got %d", resp.StatusCode)
} else {
resp.Body.Close()
}
}
}
func TestServiceAccount_SelfRotatesItsOwnKey(t *testing.T) {
s := newTS(t)
instanceKey := createServiceAccount(t, s, s.key, "terdut-operator", models.ServiceAccountScopeInstance, 0)
+24 -4
View File
@@ -359,7 +359,19 @@ func handleCreateInvite(db *sql.DB, publicURL string) http.HandlerFunc {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
caller, _ := userFromContext(r.Context())
// created_by is nullable (ON DELETE SET NULL) for exactly this
// reason: the caller minting an invite is not always a human with a
// real users row. A team-scoped service account is owner-equivalent
// here (requireTeamOwner above already let it through), and this
// must leave created_by NULL for one the same way
// handleCreateServiceAccount already does for the analogous case —
// an unchecked zero value would violate the users(id) foreign key
// instead of recording "nobody" cleanly.
var createdBy *int64
if u, ok := userFromContext(r.Context()); ok {
id := u.ID
createdBy = &id
}
expires := time.Now().Add(inviteTTL)
var out inviteJSON
@@ -368,7 +380,7 @@ func handleCreateInvite(db *sql.DB, publicURL string) http.HandlerFunc {
INSERT INTO invites (token_hash, team_id, role, created_by, expires_at, max_uses)
VALUES ($1, $2, $3, $4, $5, $6)
RETURNING id, team_id, role, created_at, expires_at, max_uses, uses`,
hash, teamID, req.Role, caller.ID, expires.Unix(), req.MaxUses).
hash, teamID, req.Role, createdBy, expires.Unix(), req.MaxUses).
Scan(&out.ID, &out.TeamID, &out.Role, &created, &expiresAt, &out.MaxUses, &out.Uses); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
@@ -429,7 +441,11 @@ func handleTestNotification(cfg NotifyConfig, db *sql.DB) http.HandlerFunc {
errResp("this server has no ntfy configured, so it can send nothing"))
return
}
caller, _ := userFromContext(r.Context())
caller, ok := userFromContext(r.Context())
if !ok {
respond(w, http.StatusForbidden, errResp("this endpoint is for human accounts only"))
return
}
var topic *string
if err := db.QueryRowContext(r.Context(),
@@ -470,7 +486,11 @@ func handleDismissOnboarding(db *sql.DB) http.HandlerFunc {
respond(w, http.StatusBadRequest, errResp("dismissed is required"))
return
}
caller, _ := userFromContext(r.Context())
caller, ok := userFromContext(r.Context())
if !ok {
respond(w, http.StatusForbidden, errResp("this endpoint is for human accounts only"))
return
}
var err error
if *req.Dismissed {
+56
View File
@@ -2,10 +2,14 @@ package api_test
import (
"bytes"
"database/sql"
"encoding/json"
"net/http"
"net/http/cookiejar"
"testing"
"git.ryuvia.com/niklas/terdut-server/internal/api"
"git.ryuvia.com/niklas/terdut-server/internal/models"
)
// signup posts to the unauthenticated sign-up endpoint, the way the form does,
@@ -268,3 +272,55 @@ func TestSignup_ValidatesLikeTheRestOfTheServer(t *testing.T) {
t.Errorf("an existing username: expected 409, got %d", taken.StatusCode)
}
}
// A team-scoped service account has no users row to attribute created_by to.
// Before this fix, handleCreateInvite wrote its zero-value caller.ID straight
// into that (nullable, ON DELETE SET NULL) foreign key instead of leaving it
// NULL the way handleCreateServiceAccount already does for the same
// situation — a 500, not the 201 TestServiceAccount_TeamScopeManagesItsOwnInvites
// now confirms. This pins the column itself ends up NULL, not just "some
// response came back".
func TestSignup_InviteCreatedByAServiceAccountLeavesCreatedByNull(t *testing.T) {
s := newTS(t)
instanceKey := createServiceAccount(t, s, s.key, "terdut-operator", models.ServiceAccountScopeInstance, 0)
teamA := createTeamAs(t, s, instanceKey, "team-a")
keyA := createServiceAccount(t, s, instanceKey, "team-a-sa", models.ServiceAccountScopeTeam, teamA)
var created struct {
ID int64 `json:"id"`
}
decode(t, s.reqAs(t, keyA, http.MethodPost, "/api/teams/"+id64(teamA)+"/invites", map[string]any{}), &created)
var createdBy sql.NullInt64
if err := s.db.QueryRow("SELECT created_by FROM invites WHERE id = $1", created.ID).Scan(&createdBy); err != nil {
t.Fatalf("read back invites.created_by: %v", err)
}
if createdBy.Valid {
t.Errorf("expected created_by to be NULL for a service-account-minted invite, got %d", createdBy.Int64)
}
}
// Neither of these has a real user_id to act on behalf of; both must 403 a
// service account explicitly rather than 500 (handleTestNotification, which
// used to query ntfy_topic for user id 0) or silently no-op (handleDismissOnboarding,
// which used to UPDATE ... WHERE id = 0, affecting nothing and still
// returning 204).
func TestServiceAccount_HumanOnlyEndpointsRefuseExplicitly(t *testing.T) {
// BaseURL set (even to a fake, unreachable address) so handleTestNotification
// reaches its AsHuman() check instead of short-circuiting on "ntfy not
// configured" first — this test is about the human-only check, not ntfy.
s := newTS(t, api.NotifyConfig{BaseURL: "http://ntfy.invalid"})
instanceKey := createServiceAccount(t, s, s.key, "terdut-operator", models.ServiceAccountScopeInstance, 0)
if resp := s.reqAs(t, instanceKey, http.MethodPost, "/api/me/notify/test", nil); resp.StatusCode != http.StatusForbidden {
t.Errorf("expected 403 for a service account testing notifications, got %d", resp.StatusCode)
} else {
resp.Body.Close()
}
if resp := s.reqAs(t, instanceKey, http.MethodPut, "/api/me/onboarding",
map[string]bool{"dismissed": true}); resp.StatusCode != http.StatusForbidden {
t.Errorf("expected 403 for a service account dismissing onboarding, got %d", resp.StatusCode)
} else {
resp.Body.Close()
}
}
+36 -4
View File
@@ -13,12 +13,19 @@ import (
"github.com/go-chi/chi/v5"
)
// handleListTeams lists the caller's own teams, each with their role in it. An
// administrator listing every team goes through the admin endpoint instead:
// this one answers "what am I part of", which is what the UI's team filter and
// the combined queue are built from.
// handleListTeams lists the caller's own teams, each with their role in it,
// or — with ?name= — looks up one team by exact name regardless of caller
// identity (TEAM-LOOKUP.md). An administrator listing every team goes
// through the admin endpoint instead: the no-name case here answers "what am
// I part of", which is what the UI's team filter and the combined queue are
// built from.
func handleListTeams(db *sql.DB) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
if name := strings.TrimSpace(r.URL.Query().Get("name")); name != "" {
handleListTeamsByName(db, w, r, name)
return
}
caller, _ := userFromContext(r.Context())
rows, err := db.QueryContext(r.Context(), `
SELECT t.id, t.name, t.created_at, m.role, m.source
@@ -51,6 +58,31 @@ func handleListTeams(db *sql.DB) http.HandlerFunc {
}
}
// handleListTeamsByName answers "is there a team named exactly this", open to
// any authenticated caller including a service account (TEAM-LOOKUP.md) —
// mirrors handleListServiceAccounts' own ?name= lookup: a one-or-zero-length
// array, never an error on no match, and no caller-identity filtering at
// all, since what it discloses (a name is taken, nothing about who's in it
// or any of its data) is the same low sensitivity that lookup already
// accepts for service-account names.
func handleListTeamsByName(db *sql.DB, w http.ResponseWriter, r *http.Request, name string) {
var t models.Team
var created int64
err := db.QueryRowContext(r.Context(),
"SELECT id, name, created_at FROM teams WHERE name = $1", name,
).Scan(&t.ID, &t.Name, &created)
if errors.Is(err, sql.ErrNoRows) {
respond(w, http.StatusOK, []models.Team{})
return
}
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
t.CreatedAt = time.Unix(created, 0).UTC()
respond(w, http.StatusOK, []models.Team{t})
}
// handleUserTeams lists one user's teams, for the admin page's per-user view:
// "what is this person in", which /api/teams cannot answer because it is always
// about the caller.
+74
View File
@@ -6,6 +6,8 @@ import (
"io"
"net/http"
"testing"
"git.ryuvia.com/niklas/terdut-server/internal/models"
)
// The whole point of #4: two teams sharing one server must not see each other's
@@ -454,3 +456,75 @@ func TestTeams_OutsiderSeesNothing(t *testing.T) {
t.Errorf("blue's team list: %v", teams)
}
}
// ---------------------------------------------------------------------------
// GET /api/teams?name= (TEAM-LOOKUP.md)
// ---------------------------------------------------------------------------
func TestListTeamsByName_FindsExactMatch(t *testing.T) {
s := newTS(t)
instanceKey := createServiceAccount(t, s, s.key, "terdut-operator", models.ServiceAccountScopeInstance, 0)
teamID := createTeamAs(t, s, instanceKey, "platform")
teams := list(t, s.reqAs(t, instanceKey, http.MethodGet, "/api/teams?name=platform", nil))
if len(teams) != 1 {
t.Fatalf("expected exactly one match for ?name=platform, got %d: %v", len(teams), teams)
}
if int64(teams[0]["id"].(float64)) != teamID {
t.Errorf("id = %v, want %d", teams[0]["id"], teamID)
}
// No membership, so no role to report (models.Team's own doc comment:
// "empty when nobody in particular is asking").
if _, has := teams[0]["role"]; has {
t.Errorf("expected no role on a name-lookup match, got %v", teams[0]["role"])
}
}
func TestListTeamsByName_NoMatchIsAnEmptyArrayNotAnError(t *testing.T) {
s := newTS(t)
instanceKey := createServiceAccount(t, s, s.key, "terdut-operator", models.ServiceAccountScopeInstance, 0)
resp := s.reqAs(t, instanceKey, http.MethodGet, "/api/teams?name=does-not-exist", nil)
if resp.StatusCode != http.StatusOK {
t.Fatalf("expected 200 on no match, got %d", resp.StatusCode)
}
teams := list(t, resp)
if len(teams) != 0 {
t.Errorf("expected an empty array, got %v", teams)
}
}
// The actual motivating scenario (TEAM-LOOKUP.md): a service account that
// already created a team, interrupted before it could remember the id,
// recovers it via ?name= on the same name its own POST 409s on.
func TestListTeamsByName_RecoversAfterCreateConflict(t *testing.T) {
s := newTS(t)
instanceKey := createServiceAccount(t, s, s.key, "terdut-operator", models.ServiceAccountScopeInstance, 0)
original := createTeamAs(t, s, instanceKey, "recovered")
conflict := s.reqAs(t, instanceKey, http.MethodPost, "/api/teams", map[string]string{"name": "recovered"})
if conflict.StatusCode != http.StatusConflict {
t.Fatalf("expected 409 recreating the same name, got %d", conflict.StatusCode)
}
conflict.Body.Close()
teams := list(t, s.reqAs(t, instanceKey, http.MethodGet, "/api/teams?name=recovered", nil))
if len(teams) != 1 || int64(teams[0]["id"].(float64)) != original {
t.Fatalf("expected to recover the original team %d via ?name=, got %v", original, teams)
}
}
// Not gated by isInstanceServiceAccount or AdminOnly (TEAM-LOOKUP.md): any
// authenticated caller may ask whether a name is taken, the same low
// sensitivity GET /api/service-accounts?name= already accepts.
func TestListTeamsByName_OpenToAnyAuthenticatedCaller(t *testing.T) {
s := newTS(t)
red := newTeam(t, s, "red")
_ = createTeamAs(t, s, s.key, "blue-target")
// red's own member, not a member of "blue-target", still gets a match.
teams := list(t, red.call(http.MethodGet, "/api/teams?name=blue-target", nil))
if len(teams) != 1 || teams[0]["name"] != "blue-target" {
t.Errorf("expected a non-member caller to still find the team by name, got %v", teams)
}
}
+27
View File
@@ -1,6 +1,7 @@
package db
import (
"context"
"database/sql"
"embed"
"fmt"
@@ -62,6 +63,16 @@ func Open(dsn string) (*sql.DB, error) {
}
}
// migrationLockKey is the Postgres advisory lock Migrate holds for its whole
// run. Two replicas starting at once would otherwise race the check-then-apply
// loop below against schema_migrations: the loser could crash on a
// duplicate-key insert, or contend with the winner's uncommitted DDL. Blocking
// (pg_advisory_lock, not pg_try_advisory_lock as the archiver and notifier
// use): on boot there is no later tick to defer to, so the right behaviour is
// to wait for the other replica to finish migrating, not to skip ahead and
// start serving against an unmigrated schema.
const migrationLockKey int64 = 7265_0003
// Migrate applies every embedded migration that has not been applied yet, in
// filename order, recording each in schema_migrations.
//
@@ -69,6 +80,22 @@ func Open(dsn string) (*sql.DB, error) {
// migration that failed half way used to leave the schema in whatever state it
// had reached. Postgres has transactional DDL, so the rollback is real.
func Migrate(db *sql.DB) error {
ctx := context.Background()
conn, err := db.Conn(ctx)
if err != nil {
return fmt.Errorf("migrate: acquire connection: %w", err)
}
defer conn.Close()
if _, err := conn.ExecContext(ctx, "SELECT pg_advisory_lock($1)", migrationLockKey); err != nil {
return fmt.Errorf("migrate: acquire advisory lock: %w", err)
}
defer func() {
if _, err := conn.ExecContext(ctx, "SELECT pg_advisory_unlock($1)", migrationLockKey); err != nil {
log.Printf("migrate: release advisory lock: %v", err)
}
}()
if _, err := db.Exec(`CREATE TABLE IF NOT EXISTS schema_migrations (
version TEXT PRIMARY KEY,
applied_at BIGINT NOT NULL DEFAULT FLOOR(EXTRACT(EPOCH FROM now()))::bigint
+125
View File
@@ -0,0 +1,125 @@
package db_test
import (
"database/sql"
"fmt"
"net/url"
"os"
"strings"
"sync"
"testing"
"git.ryuvia.com/niklas/terdut-server/internal/db"
_ "github.com/jackc/pgx/v5/stdlib"
)
// TERDUT_TEST_DSN must point at a database the test role may create schemas
// in; see internal/api/testdb_test.go for the fuller rationale this mirrors.
// An unset DSN fails rather than skips, deliberately.
const testDSNEnv = "TERDUT_TEST_DSN"
// TestMigrate_ConcurrentCallersDoNotRace reproduces two replicas starting at
// once against a brand-new, unmigrated schema: both call db.Migrate at the
// same time. Before migrationLockKey, the loser could crash on a
// duplicate-key insert into schema_migrations, or contend with the winner's
// uncommitted DDL; with the advisory lock, one blocks until the other
// finishes and both return cleanly.
func TestMigrate_ConcurrentCallersDoNotRace(t *testing.T) {
dsn := os.Getenv(testDSNEnv)
if dsn == "" {
t.Fatalf("%s is not set: these tests need Postgres.\n"+
"Run `make test-db` for a local one, then\n"+
" export %s=postgres://terdut:terdut@localhost:5432/terdut_test?sslmode=disable",
testDSNEnv, testDSNEnv)
}
schema := fmt.Sprintf("migrate_race_%d", os.Getpid())
admin, err := sql.Open("pgx", dsn)
if err != nil {
t.Fatalf("connect to %s: %v", testDSNEnv, err)
}
defer admin.Close()
if _, err := admin.Exec("CREATE SCHEMA " + schema); err != nil {
t.Fatalf("create schema %s: %v", schema, err)
}
t.Cleanup(func() {
cleanup, err := sql.Open("pgx", dsn)
if err != nil {
return
}
defer cleanup.Close()
if _, err := cleanup.Exec("DROP SCHEMA " + schema + " CASCADE"); err != nil {
t.Logf("drop schema %s: %v", schema, err)
}
})
scoped := withSearchPath(dsn, schema)
const callers = 2
errs := make([]error, callers)
var wg sync.WaitGroup
for i := range callers {
wg.Add(1)
go func(i int) {
defer wg.Done()
database, err := db.Open(scoped)
if err != nil {
errs[i] = fmt.Errorf("open: %w", err)
return
}
defer database.Close()
errs[i] = db.Migrate(database)
}(i)
}
wg.Wait()
for i, err := range errs {
if err != nil {
t.Fatalf("Migrate #%d: %v", i, err)
}
}
entries, err := os.ReadDir("migrations")
if err != nil {
t.Fatalf("read migrations dir: %v", err)
}
var want int
for _, e := range entries {
if !e.IsDir() && strings.HasSuffix(e.Name(), ".sql") {
want++
}
}
check, err := sql.Open("pgx", scoped)
if err != nil {
t.Fatalf("connect for verification: %v", err)
}
defer check.Close()
var got int
if err := check.QueryRow("SELECT COUNT(*) FROM schema_migrations").Scan(&got); err != nil {
t.Fatalf("count schema_migrations: %v", err)
}
if got != want {
t.Fatalf("schema_migrations has %d row(s) after two concurrent Migrate calls, want %d (one per migration file, no duplicates)", got, want)
}
}
// withSearchPath pins a DSN to one schema. Copied from
// internal/api/testdb_test.go rather than shared: that helper lives in the
// api_test package, a separate compiled package this one cannot import.
func withSearchPath(dsn, schema string) string {
opt := "-csearch_path=" + schema
if strings.HasPrefix(dsn, "postgres://") || strings.HasPrefix(dsn, "postgresql://") {
u, err := url.Parse(dsn)
if err == nil {
q := u.Query()
q.Set("options", opt)
u.RawQuery = q.Encode()
return u.String()
}
}
return dsn + " options='" + opt + "'"
}
+92 -19
View File
@@ -47,12 +47,27 @@
--font: system-ui, -apple-system, "Segoe UI", Roboto, "Helvetica Neue", Arial, sans-serif;
--mono: ui-monospace, SFMono-Regular, Menlo, Consolas, monospace;
/* A ~4-size type scale, per issue #26's "use one type scale" ask. This
file still has a dozen one-off font-size values below; migrating the
low-risk, purely cosmetic ones (standalone titles with no dimensional
or functional constraint) onto these tokens is a start, not the whole
job — the rest (sizes tied to a fixed shape like the avatar circle, to
a deliberately prominent display like a stat tile or the on-call name,
or to a non-negotiable constraint like the 16px that stops iOS zooming
into an input) stay as either their own pixel value or a documented
exception, since guessing at those without seeing them render risks
trading one inconsistency for a worse one. */
--fs-xs: 12px;
--fs-sm: 13px;
--fs-base: 14px;
--fs-lg: 18px;
--fs-xl: 21px;
--topbar-h: 52px;
/* No bottom tab bar on any breakpoint any more — mobile uses the hamburger
menu in the topbar, desktop the sidebar — so this stays 0. Kept as a
variable rather than deleted since .view, .toast and .nav's own height
calc still read it. */
--tabbar-h: 0px;
/* The phone-width bottom tab bar's height. Unused above 900px: the
desktop block overrides .nav/.view/.toast directly rather than reading
this back down to 0. */
--tabbar-h: 58px;
--safe-top: env(safe-area-inset-top, 0px);
--safe-bottom: env(safe-area-inset-bottom, 0px);
}
@@ -212,9 +227,6 @@ input:focus, textarea:focus { outline: none; border-color: var(--accent); box-sh
font-size: 18px; font-weight: 700; letter-spacing: -0.01em;
overflow: hidden; text-overflow: ellipsis; white-space: nowrap; min-width: 0;
}
#menu-btn { position: relative; }
.nav-badge.menu-btn-badge { top: 2px; left: auto; right: 2px; }
.open-pill {
display: inline-flex; align-items: center; gap: 6px;
padding: 3px 10px; border-radius: 999px;
@@ -226,11 +238,11 @@ input:focus, textarea:focus { outline: none; border-color: var(--accent); box-sh
.open-pill.has-triggered::before { background: var(--crit); }
.open-pill.all-acked::before { background: var(--warn); }
/* Hidden on phones — mobile navigates through the hamburger menu in the
topbar instead (see #menu-btn / openNavMenu in app.js). Reappears as the
left sidebar from 900px, where the desktop block below redeclares display. */
/* Bottom tab bar on phones (Queue, On-call, Alerts, Team, More); becomes the
left sidebar from 900px, where the desktop block below redeclares display
and shows every section flat, .nav-link-secondary included. */
.nav {
display: none;
display: flex;
position: fixed; left: 0; right: 0; bottom: 0; z-index: 20;
height: calc(var(--tabbar-h) + var(--safe-bottom));
padding-bottom: var(--safe-bottom);
@@ -240,14 +252,24 @@ input:focus, textarea:focus { outline: none; border-color: var(--accent); box-sh
border-top: 1px solid var(--border);
}
.nav-brand { display: none; }
/* Hidden here (shown from 900px below): on the phone bar the team switcher
lives in the topbar instead, as #team-selector-mobile. */
.nav-team-selector { display: none; }
.nav-link {
position: relative;
/* flex: 1 spreads the tabs evenly across the bar's width; the desktop
block below cancels it back to a natural-width row item. */
flex: 1 1 0;
display: flex; flex-direction: column; align-items: center; justify-content: center; gap: 2px;
color: var(--faint); font-size: 11px; font-weight: 600;
/* min-width lets a column shrink below its label's natural width, which is
what stops six tabs widening the bar past the screen. */
min-width: 0; padding: 0 2px;
}
/* Must come after .nav-link above: same specificity (one class each), so
whichever is later in the file wins for an element wearing both classes,
and this needs to beat .nav-link's display:flex here on the phone bar. */
.nav-link-secondary { display: none; }
.nav-label {
max-width: 100%; overflow: hidden; text-overflow: ellipsis; white-space: nowrap;
}
@@ -276,6 +298,7 @@ input:focus, textarea:focus { outline: none; border-color: var(--accent); box-sh
}
.team-selector-mobile { padding: 4px 10px; font-size: 12px; max-width: 120px; }
.team-selector-label { overflow: hidden; text-overflow: ellipsis; white-space: nowrap; }
.team-selector-chevron { width: 14px; height: 14px; flex: none; color: var(--faint); margin-left: -2px; }
/* A team's identity colour — not a status, so never the severity palette. Six
colours, then they repeat; teamColorClass() in format.js picks one by the
@@ -303,6 +326,7 @@ input:focus, textarea:focus { outline: none; border-color: var(--accent); box-sh
/* ---------- chips ---------- */
.chips {
position: relative;
display: flex; gap: 6px;
padding: 12px 16px 8px;
overflow-x: auto; scrollbar-width: none;
@@ -318,6 +342,17 @@ input:focus, textarea:focus { outline: none; border-color: var(--accent); box-sh
}
.chip[aria-selected="true"] { background: var(--text); border-color: var(--text); color: var(--bg); }
.chip .count { margin-left: 4px; opacity: 0.7; }
/* An overlay, not a flex item: absolute against .chips' own (non-scrolling)
box stays flush with its real right edge regardless of scroll position,
which turned out not to be true of position:sticky here — as a flex
item, its sticky offset interacted with the row's gap and its own
negative margin, landing short of the edge by about one gap's width. */
.chips-fade {
position: absolute; top: 0; right: 0; bottom: 0;
width: 24px;
background: linear-gradient(to right, transparent, var(--bg));
pointer-events: none;
}
/* ---------- lists ---------- */
@@ -381,7 +416,7 @@ input:focus, textarea:focus { outline: none; border-color: var(--accent); box-sh
.empty {
padding: 48px 16px; text-align: center; color: var(--muted);
}
.empty strong { display: block; color: var(--text); font-size: 16px; margin-bottom: 4px; }
.empty strong { display: block; color: var(--text); font-size: var(--fs-lg); margin-bottom: 4px; }
.empty .icon { width: 36px; height: 36px; color: var(--ok); margin-bottom: 8px; }
.load-error {
@@ -434,8 +469,12 @@ input:focus, textarea:focus { outline: none; border-color: var(--accent); box-sh
.detail-head .copy { margin-left: auto; }
.clip-buffer { position: fixed; top: 0; left: 0; opacity: 0; pointer-events: none; }
.detail-head .crumb { font-weight: 600; color: var(--muted); font-size: 14px; }
.detail-title { font-size: 21px; font-weight: 750; letter-spacing: -0.01em; margin: 16px 0 8px; overflow-wrap: anywhere; }
.detail-title { font-size: var(--fs-xl); font-weight: 750; letter-spacing: -0.01em; margin: 16px 0 8px; overflow-wrap: anywhere; }
.detail-badges { display: flex; flex-wrap: wrap; gap: 6px; margin-bottom: 14px; }
/* A copy of the sticky actionbar's primary button, right under the status
it responds to — see quickActions() in incident.js. */
.detail-quick-actions { margin-bottom: 14px; }
.detail-quick-actions .btn-primary { font-size: 16px; min-height: 44px; }
.card {
background: var(--surface);
@@ -450,6 +489,9 @@ input:focus, textarea:focus { outline: none; border-color: var(--accent); box-sh
.facts dt { color: var(--muted); }
.facts dd { margin: 0; overflow-wrap: anywhere; }
.facts .sub { color: var(--faint); }
.facts dd.fact-summary { display: flex; flex-wrap: wrap; align-items: center; gap: 8px; }
.fact-chip { display: inline-flex; align-items: center; gap: 4px; }
.fact-icon { width: 15px; height: 15px; color: var(--faint); }
.section { margin-top: 22px; }
.section-title {
@@ -481,6 +523,17 @@ details > summary::before { content: "▸ "; }
details[open] > summary::before { content: "▾ "; }
details[open] > summary { margin-bottom: 8px; }
/* One .tl-phase per status the incident has been through (see
timelinePhases() in incident.js) — each with its own .timeline <ol>, so
the existing :first-child/:last-child rail-capping below gives each phase
its own self-contained connecting line rather than one running through
the headings. */
.tl-phase + .tl-phase { border-top: 1px solid var(--border); }
.tl-phase-title {
padding: 10px 14px 0;
font-size: 11px; font-weight: 700; text-transform: uppercase; letter-spacing: 0.06em;
color: var(--faint);
}
.timeline { list-style: none; margin: 0; padding: 4px 0; }
.tl-item {
position: relative;
@@ -513,7 +566,7 @@ details[open] > summary { margin-bottom: 8px; }
}
.note-fix { background: var(--ok-soft); border-left: 3px solid var(--ok); }
.similar { list-style: none; margin: 0; padding: 0; }
.similar-item { padding: 10px 0; }
.similar-item { padding: 10px 14px; }
.similar-item + .similar-item { border-top: 1px solid var(--border, var(--surface-2)); }
.check { display: flex; align-items: center; gap: 8px; font-size: 14px; color: var(--muted); }
.note-actions { display: flex; justify-content: flex-end; }
@@ -552,7 +605,7 @@ details[open] > summary { margin-bottom: 8px; }
@keyframes sheet-up { from { transform: translateY(24px); opacity: 0.6; } }
.sheet-inner { padding: 8px 16px calc(16px + var(--safe-bottom)); }
.sheet-grab { width: 40px; height: 4px; margin: 0 auto 12px; border-radius: 2px; background: var(--border-strong); }
.sheet-title { font-size: 17px; font-weight: 700; margin: 0 0 4px; }
.sheet-title { font-size: var(--fs-lg); font-weight: 700; margin: 0 0 4px; }
.sheet-text { color: var(--muted); margin: 0 0 14px; font-size: 14px; }
.sheet-form { display: grid; gap: 12px; }
.sheet-actions { display: flex; gap: 8px; margin-top: 16px; }
@@ -604,9 +657,18 @@ details[open] > summary { margin-bottom: 8px; }
.now-label { color: var(--muted); font-size: 13px; font-weight: 600; }
.now-name { font-size: 20px; font-weight: 750; }
.you { color: var(--accent); font-weight: 650; font-size: 13px; margin-left: 6px; }
/* Oncall's own "you" indicator only (see you() in oncall.js) — a pill badge
is easier to spot there than this plain accent-coloured text. */
.you-badge { margin-left: 6px; }
.week-nav { display: flex; align-items: center; gap: 4px; }
.week-nav .label { font-size: 14px; font-weight: 650; min-width: 9em; text-align: center; }
/* The week-nav button that shows the date range, "28 Sep – 4 Oct", with the
ISO week number as secondary text inside it — a separate class from
.label above (team.js's month-nav uses that one) so its <small> isn't
caught by the unrelated .label > span styling meant for label chips. */
.week-nav .week-label { font-size: 14px; font-weight: 650; min-width: 11.5em; text-align: center; white-space: nowrap; }
.week-label small { color: var(--faint); font-weight: 600; font-size: 11px; margin-left: 2px; }
.days { list-style: none; margin: 0; padding: 0; }
.day { display: grid; grid-template-columns: 3.2em 4.2em 1fr; align-items: center; gap: 8px; min-height: 50px; padding: 0 14px; }
.day + .day { border-top: 1px solid var(--border); }
@@ -614,11 +676,18 @@ details[open] > summary { margin-bottom: 8px; }
.day-date { color: var(--faint); font-size: 13px; }
.day-who { overflow: hidden; text-overflow: ellipsis; white-space: nowrap; }
.day-who.nobody { color: var(--faint); font-style: italic; }
/* A run of several days held by the same person (or left empty), replacing
what used to be one identical row per day — see weekRuns() in oncall.js. */
.day.range { grid-template-columns: 1fr auto; }
.day-range { font-weight: 650; }
.day.today { background: var(--accent-soft); }
.day.today:first-child { border-radius: var(--radius) var(--radius) 0 0; }
.day.today:last-child { border-radius: 0 0 var(--radius) var(--radius); }
.day.today .day-name { color: var(--accent); }
.day.today .day-name, .day.today .day-range { color: var(--accent); }
.day.past { opacity: 0.6; }
/* Highlights whichever row is yours, same soft tint as .today — they already
read fine layered (today's own row is almost always one of yours too). */
.day.mine { background: var(--accent-soft); }
.shift-list { list-style: none; margin: 0; padding: 0; }
.shift-list li { display: flex; justify-content: space-between; padding: 12px 14px; }
@@ -675,9 +744,11 @@ kbd {
display: flex; align-items: center; gap: 10px;
padding: 4px 10px 18px; font-size: 18px; font-weight: 750; letter-spacing: -0.01em;
}
.nav-team-selector { margin: -8px 10px 14px; width: calc(100% - 20px); }
.nav-team-selector { display: inline-flex; margin: -8px 10px 14px; width: calc(100% - 20px); }
.nav-link-secondary { display: flex; }
.nav-more-btn { display: none; }
.nav-link {
flex-direction: row; justify-content: flex-start; gap: 12px;
flex: none; flex-direction: row; justify-content: flex-start; gap: 12px;
min-height: 40px; padding: 0 10px; border-radius: var(--radius-sm);
color: var(--muted); font-size: 14px;
}
@@ -701,6 +772,7 @@ kbd {
is hidden, so the chips wrap here instead: Archived stays reachable. */
.pane-list .chips { flex-wrap: wrap; overflow-x: visible; }
.pane-list .chip-sep { display: none; }
.chips-fade { display: none; }
.view-queue:not(.has-detail) .pane-detail { display: block; }
/* On desktop the list stays visible next to the detail. */
@@ -966,6 +1038,7 @@ button.rota-week:hover { background: var(--surface-2); color: var(--text); }
.overview-head { display: flex; align-items: baseline; gap: 8px; }
.overview-count { margin-left: auto; color: var(--muted); font-size: 18px; font-weight: 700; }
.overview-item p { margin: 4px 0 0; }
.overview-note-icon { width: 13px; height: 13px; vertical-align: -2px; color: var(--ok); }
/* ---------- stats page ---------- */
+14 -7
View File
@@ -102,7 +102,10 @@
<svg viewBox="0 0 24 24" aria-hidden="true"><path d="M6 16V11a6 6 0 0 1 12 0v5l1.5 2h-15z"/><path d="M10 20.5a2 2 0 0 0 4 0"/></svg>
<span class="nav-label">Alerts</span>
</a>
<a class="nav-link" href="/stats" data-section="stats" aria-label="Stats">
<!-- Secondary: full-width in the desktop sidebar, folded into the
"More" tab's sheet on the phone-width bottom bar instead (see
.nav-link-secondary in app.css and openNavMenu in app.js). -->
<a class="nav-link nav-link-secondary" href="/stats" data-section="stats" aria-label="Stats">
<svg viewBox="0 0 24 24" aria-hidden="true"><path d="M4 20h16M7 20v-7M12 20V6M17 20v-10"/></svg>
<span class="nav-label">Stats</span>
</a>
@@ -113,22 +116,26 @@
<!-- Hidden unless the signed-in user is a system administrator; app.js
unhides it once /api/me says so. The server refuses every admin
endpoint regardless, so this is a courtesy and not a gate. -->
<a class="nav-link" href="/admin" data-section="admin" aria-label="Admin" id="nav-admin" hidden>
<a class="nav-link nav-link-secondary" href="/admin" data-section="admin" aria-label="Admin" id="nav-admin" hidden>
<svg viewBox="0 0 24 24" aria-hidden="true"><path d="M12 3l7 3v6c0 4-3 7-7 9-4-2-7-5-7-9V6z"/></svg>
<span class="nav-label">Admin</span>
</a>
<a class="nav-link" href="/more" data-section="more" aria-label="Account">
<a class="nav-link nav-link-secondary" href="/more" data-section="more" aria-label="Account">
<svg viewBox="0 0 24 24" aria-hidden="true"><circle cx="12" cy="8" r="3.5"/><path d="M5 20a7 7 0 0 1 14 0"/></svg>
<span class="nav-label">Account</span>
</a>
<!-- Phone-width only (see .nav-more-btn in app.css): opens the same
sheet the old hamburger button did, for the sections the bottom
bar has no room for. Not shown on the desktop sidebar, which lists
every section already. -->
<button class="nav-link nav-more-btn" id="nav-more-btn" type="button" aria-label="More sections" aria-haspopup="menu">
<svg viewBox="0 0 24 24" aria-hidden="true"><path d="M5 12h.01M12 12h.01M19 12h.01"/></svg>
<span class="nav-label">More</span>
</button>
</nav>
<header class="topbar">
<div class="topbar-left">
<button class="btn btn-ghost btn-icon" id="menu-btn" type="button" aria-label="Menu" aria-haspopup="menu">
<svg viewBox="0 0 24 24" aria-hidden="true"><path d="M4 7h16M4 12h16M4 17h16"/></svg>
<span class="nav-badge menu-btn-badge" data-badge hidden></span>
</button>
<button class="team-selector-mobile" id="team-selector-mobile" type="button" hidden></button>
<h1 class="topbar-title" id="topbar-title">Queue</h1>
</div>
+5 -1
View File
@@ -60,7 +60,11 @@ function render() {
let body;
if (error && !items) body = h('div', { class: 'load-error', text: error });
else if (!items) body = spinner();
else if (!items.length) body = emptyState(filter === 'firing' ? 'Nothing firing' : 'No alerts', '', filter === 'firing' ? 'checkCircle' : null);
else if (!items.length) {
body = filter === 'firing'
? emptyState('Nothing firing', 'No alerts are currently firing.', 'checkCircle')
: emptyState('No alerts', 'None match this filter.', 'bell');
}
else {
body = h('div', { class: 'list' },
error && h('div', { class: 'load-error', text: `Showing older data: ${error}` }),
+17 -15
View File
@@ -36,16 +36,17 @@ const SECTIONS = {
device: { title: 'Sign in a terminal', view: device },
};
// The mobile hamburger menu's contents — the same sections the desktop
// sidebar's .nav-link list carries in index.html, in the same order.
// The desktop sidebar's .nav-link list in index.html, in the same order.
// `secondary` marks the ones that fold into the phone bottom bar's "More"
// sheet (openNavMenu below) instead of getting a tab of their own there.
const NAV_ITEMS = [
{ path: '/', section: 'queue', label: 'Queue', icon: 'queueList' },
{ path: '/oncall', section: 'oncall', label: 'On-call', icon: 'calendar' },
{ path: '/alerts', section: 'alerts', label: 'Alerts', icon: 'bell' },
{ path: '/stats', section: 'stats', label: 'Stats', icon: 'chart' },
{ path: '/stats', section: 'stats', label: 'Stats', icon: 'chart', secondary: true },
{ path: '/team', section: 'team', label: 'Team', icon: 'team' },
{ path: '/admin', section: 'admin', label: 'Admin', icon: 'shield', adminOnly: true },
{ path: '/more', section: 'more', label: 'Account', icon: 'user' },
{ path: '/admin', section: 'admin', label: 'Admin', icon: 'shield', adminOnly: true, secondary: true },
{ path: '/more', section: 'more', label: 'Account', icon: 'user', secondary: true },
];
function parseRoute(pathname) {
@@ -152,15 +153,16 @@ function render() {
// ---------- nav menu ----------
// The mobile hamburger menu: same shape as the sheet-based action menus in
// incident.js (openSheet + a <ul class="menu"> of menu-item buttons), one
// item per NAV_ITEMS entry, resolving with a path for navigate() to use.
// The phone bottom bar's "More" sheet: same shape as the sheet-based action
// menus in incident.js (openSheet + a <ul class="menu"> of menu-item
// buttons), one item per secondary NAV_ITEMS entry — the ones the bar itself
// has no room for, since Queue/On-call/Alerts/Team already have their own
// tab and don't need to be reachable here too.
function openNavMenu() {
const current = SECTIONS[route.section].nav || route.section;
const triggered = state.open.filter((i) => i.status === 'triggered').length;
const items = NAV_ITEMS.filter((n) => !n.adminOnly || state.me?.user?.is_admin);
const items = NAV_ITEMS.filter((n) => n.secondary && (!n.adminOnly || state.me?.user?.is_admin));
ui.openSheet(() => [
ui.h('h2', { class: 'sheet-title', text: 'Sections' }),
ui.h('h2', { class: 'sheet-title', text: 'More' }),
ui.h('ul', { class: 'menu', role: 'menu' }, items.map((n) =>
ui.h('li', {}, ui.h('button', {
class: 'menu-item',
@@ -171,7 +173,6 @@ function openNavMenu() {
},
ui.icon(n.icon),
n.label,
n.section === 'queue' && triggered > 0 && ui.badge(String(triggered), 'st-triggered menu-sub'),
))),
),
]).then((path) => {
@@ -205,8 +206,8 @@ function updateBadges() {
pill.classList.toggle('has-triggered', triggered > 0);
pill.classList.toggle('all-acked', open > 0 && triggered === 0);
// Two badges carry this count: the sidebar's Queue tab (desktop) and the
// hamburger button (mobile) — only one of the two is ever visible at once.
// Just the Queue tab's own badge now — phone bottom bar and desktop
// sidebar both read the same [data-badge] span on that one nav-link.
for (const badge of document.querySelectorAll('[data-badge]')) {
badge.hidden = triggered === 0;
badge.textContent = String(triggered);
@@ -230,7 +231,7 @@ async function boot() {
document.addEventListener('keydown', onKey);
$('login-form').addEventListener('submit', onLogin);
$('signup-form').addEventListener('submit', onSignup);
$('menu-btn').addEventListener('click', openNavMenu);
$('nav-more-btn').addEventListener('click', openNavMenu);
teamselector.init();
ssoErrorCode = takeSSOError();
@@ -355,6 +356,7 @@ function ssoErrorText(code, name) {
no_email: `${sso} did not send an email address for you, which terdut needs.`,
email_conflict: 'An account with your email address already exists and could not be linked to this sign-in. Ask an administrator.',
disabled: 'Your account is disabled. Ask an administrator.',
not_bootstrapped: 'This install is still setting up. Try again in a moment.',
}[code] || `Signing in with ${sso} failed.`;
}
+90 -20
View File
@@ -7,7 +7,7 @@ import {
h, clear, icon, badge, labelChip, openSheet, closeSheet, confirm, toast, spinner, emptyState,
} from './ui.js';
import {
ago, when, until, isFuture, severityClass, STATUS_LABEL,
ago, when, until, duration, isFuture, severityClass, STATUS_LABEL,
} from './format.js';
import { myID, users } from './state.js';
import { back } from './app.js';
@@ -86,6 +86,7 @@ function render() {
error && h('div', { class: 'load-error', text: `Showing older data: ${error}` }),
h('h1', { class: 'detail-title', text: inc.title }),
h('div', { class: 'detail-badges' }, statusBadges()),
quickActions(),
facts(),
groupLabels(),
alertsSection(),
@@ -125,6 +126,20 @@ function who(id, name) {
function facts() {
const rows = [];
const add = (k, ...v) => rows.push(h('dt', { text: k }), h('dd', {}, ...v));
// Duration, severity and who's on it, in one scannable row up top — the
// rest of this card has each of those too, but spread across rows that
// take reading top to bottom to piece together.
const elapsedTo = inc.resolved_at ? Date.parse(inc.resolved_at) : Date.now();
const responsible = inc.assigned_to_id != null ? who(inc.assigned_to_id, inc.assigned_to)
: inc.acknowledged_by_id != null ? who(inc.acknowledged_by_id, inc.acknowledged_by)
: 'Unassigned';
rows.push(h('dt', { text: 'At a glance' }), h('dd', { class: 'fact-summary' },
h('span', { class: 'fact-chip' }, icon('clock', 'icon fact-icon'), duration(elapsedTo - Date.parse(inc.triggered_at))),
inc.severity && badge(inc.severity, `plain ${severityClass(inc.severity)}`),
h('span', { class: 'fact-chip' }, icon('user', 'icon fact-icon'), responsible),
));
add('Triggered', when(inc.triggered_at), h('span', { class: 'sub', text: ` · ${ago(inc.triggered_at)}` }));
if (inc.acknowledged_at) {
add('Acknowledged', `${who(inc.acknowledged_by_id, inc.acknowledged_by)} · ${when(inc.acknowledged_at)}`);
@@ -148,9 +163,12 @@ function facts() {
function groupLabels() {
const entries = Object.entries(inc.group_labels || {});
if (!entries.length) return null;
// Collapsed by default, the same disclosure alertItem() below uses for an
// alert's own labels — this is background, not something to scan past.
return h('section', { class: 'section' },
h('h2', { class: 'section-title', text: 'Grouped by' }),
h('div', { class: 'labels-wrap' }, entries.map(([k, v]) => labelChip(k, v))),
h('details', {},
h('summary', { text: `Grouped by (${entries.length})` }),
h('div', { class: 'labels-wrap' }, entries.map(([k, v]) => labelChip(k, v)))),
);
}
@@ -206,12 +224,20 @@ function eventText(ev, named = false) {
case 'snoozed': return [strong(person), ` snoozed until ${ev.detail ? when(ev.detail) : '…'}`];
case 'unsnoozed': return [strong(person), ' ended the snooze'];
case 'resolved': return person ? [strong(person), ' resolved the incident'] : ['Resolved: every alert stopped firing'];
// Falls through to the generic `${ev.type}: ${ev.detail}` below otherwise
// — this just capitalises it and drops the redundant "escalated:" prefix
// from detail (already "level 2: alice, bob" or "escalation exhausted: …").
case 'escalated': return [`Escalated — ${ev.detail}`];
case 'note': return [strong(person), ' added a note'];
case 'resolution_note': return [strong(person), ' noted what fixed it'];
case 'notified': {
const to = person ? strong(person) : 'the fallback topic';
if (ev.detail === 'reminder') return ['Reminder sent to ', to];
if (ev.detail === 'resolved') return ['Resolution sent to ', to];
// 'escalated' is a second, later page — the next level firing, not the
// same page landing twice — so it reads as a bug unless told apart
// from the initial 'triggered' page below.
if (ev.detail === 'escalated') return ['Escalation paged ', to];
return ['Paged ', to];
}
case 'notify_failed': return ['Notification failed', ev.detail ? `: ${ev.detail}` : ''];
@@ -236,8 +262,34 @@ function similarSection() {
);
}
// Splits the already-sorted timeline on the status transitions that matter —
// first acknowledged, then resolved — so a long incident reads as "before
// anyone had it" / "while someone did" / "after it closed" instead of one
// undifferentiated list. A later re-acknowledge (after an unacknowledge)
// doesn't open a second "Acknowledged" phase; it's still the same spell of
// somebody owning it.
function timelinePhases(sorted) {
const phases = [{ label: 'Triggered', events: [] }];
let acked = false;
for (const ev of sorted) {
if (ev.type === 'acknowledged' && !acked) {
phases.push({ label: 'Acknowledged', events: [] });
acked = true;
} else if (ev.type === 'resolved') {
phases.push({ label: 'Resolved', events: [] });
}
phases[phases.length - 1].events.push(ev);
}
return phases.filter((p) => p.events.length);
}
function timelineSection() {
const sorted = [...events].sort((a, b) => Date.parse(a.created_at) - Date.parse(b.created_at) || a.id - b.id);
const phases = timelinePhases(sorted);
// A single phase (the common case: most incidents are acked once and
// resolved) names nothing extra — only a split timeline needs the
// headings to make sense of.
const named = phases.length > 1;
return h('section', { class: 'section' },
h('h2', { class: 'section-title' },
h('span', { text: 'Timeline' }),
@@ -245,8 +297,10 @@ function timelineSection() {
icon('note'), 'Add note')),
h('div', { class: 'card' },
sorted.length
? h('ol', { class: 'timeline' }, sorted.map(timelineItem))
: emptyState('No events yet', '')),
? phases.map((p) => h('div', { class: 'tl-phase' },
named && h('div', { class: 'tl-phase-title', text: p.label }),
h('ol', { class: 'timeline' }, p.events.map(timelineItem))))
: emptyState('No events yet', 'Nothing has happened on this incident yet.', 'clock')),
);
}
@@ -365,27 +419,43 @@ async function copyIncident() {
const isOpen = () => inc.status !== 'resolved';
const isSnoozed = () => isOpen() && isFuture(inc.snoozed_until);
function actionBar() {
let primary;
let secondary;
// primaryAction and secondaryAction are factories, not shared nodes — a
// button can only live in one place, and quickActions() below needs its own
// copy of the primary one rather than the actionbar's.
function primaryAction() {
if (inc.status === 'triggered') {
primary = h('button', { class: 'btn btn-primary', type: 'button', onclick: acknowledge }, icon('check'), 'Acknowledge');
} else if (inc.status === 'acknowledged') {
primary = h('button', { class: 'btn btn-primary', type: 'button', onclick: resolve }, icon('checkCircle'), 'Resolve');
} else {
primary = inc.archived_at
? h('button', { class: 'btn btn-primary', type: 'button', onclick: unarchive }, icon('undo'), 'Unarchive')
: h('button', { class: 'btn btn-primary', type: 'button', onclick: archive }, icon('archive'), 'Archive');
return h('button', { class: 'btn btn-primary', type: 'button', onclick: acknowledge }, icon('check'), 'Acknowledge');
}
if (inc.status === 'acknowledged') {
return h('button', { class: 'btn btn-primary', type: 'button', onclick: resolve }, icon('checkCircle'), 'Resolve');
}
return inc.archived_at
? h('button', { class: 'btn btn-primary', type: 'button', onclick: unarchive }, icon('undo'), 'Unarchive')
: h('button', { class: 'btn btn-primary', type: 'button', onclick: archive }, icon('archive'), 'Archive');
}
function secondaryAction() {
if (isOpen()) {
secondary = isSnoozed()
return isSnoozed()
? h('button', { class: 'btn', type: 'button', onclick: unsnooze }, icon('bell'), 'Unsnooze')
: h('button', { class: 'btn', type: 'button', onclick: snooze }, icon('clock'), 'Snooze');
} else {
secondary = h('button', { class: 'btn', type: 'button', onclick: addNote }, icon('note'), 'Note');
}
const more = h('button', { class: 'btn btn-icon', type: 'button', 'aria-label': 'More actions', onclick: moreMenu }, icon('more'));
const bar = h('div', { class: 'actionbar' }, primary, secondary, more);
return h('button', { class: 'btn', type: 'button', onclick: addNote }, icon('note'), 'Note');
}
// A copy of the primary action (Acknowledge/Resolve/…) up where it's seen
// right away, next to the status it responds to. The sticky actionbar below
// keeps carrying every action, primary included, for whenever the page has
// been scrolled past it.
function quickActions() {
const div = h('div', { class: 'detail-quick-actions' }, primaryAction());
if (busy) for (const b of div.querySelectorAll('button')) b.disabled = true;
return div;
}
function actionBar() {
const more = h('button', { class: 'btn', type: 'button', 'aria-label': 'More actions', onclick: moreMenu }, icon('more'), 'More');
const bar = h('div', { class: 'actionbar' }, primaryAction(), secondaryAction(), more);
if (busy) for (const b of bar.querySelectorAll('button')) b.disabled = true;
return bar;
}
+72 -24
View File
@@ -6,8 +6,8 @@
// every team the viewer is in, because somebody on two rotas wants both.
import * as api from './api.js';
import { h, clear, icon, spinner } from './ui.js';
import { isoDate, mondayOf, addDays, isoWeek, initial } from './format.js';
import { h, clear, badge, icon, spinner } from './ui.js';
import { isoDate, mondayOf, addDays, isoWeek, initial, duration } from './format.js';
import { myID, currentTeam } from './state.js';
const view = () => document.getElementById('view-oncall');
@@ -68,7 +68,7 @@ function render() {
}
function you(userID) {
return userID === myID() ? h('span', { class: 'you', text: 'you' }) : null;
return userID === myID() ? badge('you', 'plain st-oncall you-badge') : null;
}
// One card per team with somebody on call, and a single empty card when there
@@ -99,20 +99,48 @@ function nowCard() {
)));
}
// Groups the week's 7 days into runs held by the same person (or the same
// empty slot) — the week's own version of the consecutive-day grouping
// myShifts does for a single person's own dates, below. Seven identical rows
// for one person all week collapses to the one bar this way.
function weekRuns(byDate) {
const runs = [];
for (let i = 0; i < 7; i++) {
const date = isoDate(addDays(weekStart, i));
const e = byDate.get(date) || null;
const uid = e ? e.user_id : null;
const last = runs[runs.length - 1];
if (last && last.uid === uid) last.to = date;
else runs.push({ uid, entry: e, from: date, to: date });
}
return runs;
}
function weekCard() {
const byDate = new Map(data.week.map((e) => [e.date, e]));
const today = isoDate(new Date());
const days = [];
for (let i = 0; i < 7; i++) {
const d = addDays(weekStart, i);
const key = isoDate(d);
const e = byDate.get(key);
days.push(h('li', { class: `day ${key === today ? 'today' : ''} ${key < today ? 'past' : ''}` },
h('span', { class: 'day-name', text: dayName.format(d) }),
h('span', { class: 'day-date', text: dayDate.format(d) }),
h('span', { class: `day-who ${e ? '' : 'nobody'}` }, e ? e.username : 'nobody', e && you(e.user_id)),
));
}
const mine = myID();
const days = weekRuns(byDate).map((r) => {
const single = r.from === r.to;
const cls = [
'day',
single ? '' : 'range',
r.from <= today && today <= r.to ? 'today' : '',
r.to < today ? 'past' : '',
r.uid === mine ? 'mine' : '',
].filter(Boolean).join(' ');
const label = single
? [h('span', { class: 'day-name', text: dayName.format(parse(r.from)) }),
h('span', { class: 'day-date', text: dayDate.format(parse(r.from)) })]
: [h('span', {
class: 'day-range',
text: `${dayName.format(parse(r.from))} ${dayDate.format(parse(r.from))} – ${dayName.format(parse(r.to))} ${dayDate.format(parse(r.to))}`,
})];
return h('li', { class: cls },
...label,
h('span', { class: `day-who ${r.entry ? '' : 'nobody'}` }, r.entry ? r.entry.username : 'nobody', r.entry && you(r.entry.user_id)),
);
});
const thisWeek = isoDate(weekStart) === isoDate(mondayOf(new Date()));
return [
h('div', { class: 'page-head' },
@@ -121,12 +149,12 @@ function weekCard() {
h('button', { class: 'btn btn-ghost btn-icon', type: 'button', 'aria-label': 'Previous week', onclick: () => shiftWeek(-1) },
icon('chevronLeft')),
h('button', {
class: 'btn btn-ghost label',
class: 'btn btn-ghost week-label',
type: 'button',
title: 'Back to this week',
onclick: () => { weekStart = mondayOf(new Date()); refresh(); },
text: `Week ${isoWeek(weekStart)}`,
}),
text: `${dayDate.format(weekStart)} – ${dayDate.format(addDays(weekStart, 6))}`,
}, h('small', { text: ` Week ${isoWeek(weekStart)}` })),
h('button', { class: 'btn btn-ghost btn-icon', type: 'button', 'aria-label': 'Next week', onclick: () => shiftWeek(1) },
icon('chevronRight')),
),
@@ -135,7 +163,9 @@ function weekCard() {
];
}
// myShifts groups your upcoming dates into runs of consecutive days.
// myShifts groups your upcoming dates into runs of consecutive days, then
// splits off the one you're already in — listing it again under "Next
// shifts" told people they hadn't started a shift they were already on.
function myShifts() {
const mine = data.upcoming.filter((e) => e.user_id === myID()).map((e) => e.date).sort();
const runs = [];
@@ -145,16 +175,34 @@ function myShifts() {
else runs.push({ from: date, to: date });
}
const fmt = (s) => `${dayName.format(parse(s))} ${dayDate.format(parse(s))}`;
return [
h('div', { class: 'page-head' }, h('h2', { text: 'Your next shifts' })),
const label = (r) => (r.from === r.to ? fmt(r.from) : `${fmt(r.from)} – ${fmt(r.to)}`);
const today = isoDate(new Date());
const current = runs[0] && runs[0].from <= today ? runs[0] : null;
const next = current ? runs.slice(1) : runs;
const currentCard = current ? [
h('div', { class: 'page-head' }, h('h2', { text: 'Current shift' })),
h('div', { class: 'card' },
runs.length
? h('ul', { class: 'shift-list' }, runs.slice(0, 8).map((r) =>
h('ul', { class: 'shift-list' }, h('li', {},
h('span', { text: label(current) }),
h('span', { class: 'muted', text: `ends in ${duration(addDays(parse(current.to), 1) - Date.now())}` })))),
] : [];
return [
...currentCard,
h('div', { class: 'page-head' }, h('h2', { text: current ? 'Next shifts' : 'Your next shifts' })),
h('div', { class: 'card' },
next.length
? h('ul', { class: 'shift-list' }, next.slice(0, 8).map((r) =>
h('li', {},
h('span', { text: r.from === r.to ? fmt(r.from) : `${fmt(r.from)} – ${fmt(r.to)}` }),
h('span', { text: label(r) }),
h('span', { class: 'muted', text: days(r) })),
))
: h('div', { class: 'empty', text: 'Nothing scheduled in the next 60 days.' })),
: h('div', {
class: 'empty',
text: current ? 'Nothing else scheduled in the next 60 days.' : 'Nothing scheduled in the next 60 days.',
})),
];
}
+41 -14
View File
@@ -2,7 +2,7 @@
import * as api from './api.js';
import { h, clear, badge, emptyState, spinner } from './ui.js';
import { age, until, isFuture, severityClass, labelSummary, teamColorClass } from './format.js';
import { ago, until, isFuture, severityClass, labelSummary, teamColorClass } from './format.js';
import { state, myID, setSelectedTeam, onTeamChange } from './state.js';
import * as onboarding from './onboarding.js';
import { navigate } from './app.js';
@@ -18,12 +18,12 @@ const FILTERS = [
];
const EMPTY = {
open: ['All clear', 'Nothing open right now.'],
triggered: ['Nothing triggered', 'Every open incident has been acknowledged.'],
acknowledged: ['Nothing acknowledged', 'No one is working an incident right now.'],
snoozed: ['Nothing snoozed', 'Snoozed incidents show up here until the snooze runs out.'],
resolved: ['Nothing resolved', 'Resolved incidents are archived after a while.'],
archived: ['Nothing archived', ''],
open: ['All clear', 'Nothing open right now.', 'checkCircle'],
triggered: ['Nothing triggered', 'Every open incident has been acknowledged.', 'checkCircle'],
acknowledged: ['Nothing acknowledged', 'No one is working an incident right now.', 'checkCircle'],
snoozed: ['Nothing snoozed', 'Snoozed incidents show up here until the snooze runs out.', 'clock'],
resolved: ['Nothing resolved', 'Resolved incidents are archived after a while.', null],
archived: ['Nothing archived', 'Resolved incidents land here once archived.', 'archive'],
};
onboarding.onRerender(() => renderList());
@@ -87,6 +87,10 @@ export async function refresh({ fresh = false } = {}) {
if (requested !== filter) return;
error = err.message;
}
// state.open (what the chip counts read) has just been refreshed too, by
// whichever caller updated it before calling here — app.js's poll, or the
// `cached` branch above.
renderChips();
renderList();
}
@@ -101,18 +105,32 @@ function setFilter(id) {
refresh({ fresh: true });
}
// Counts for the three chips derivable from the open list already fetched
// for the badges — Snoozed/Resolved/Archived would need a request of their
// own, so those chips stay count-less for now.
function chipCount(id) {
const open = state.selectedTeamID == null
? state.open
: state.open.filter((i) => i.team_id === state.selectedTeamID);
if (id === 'open') return open.length;
if (id === 'triggered') return open.filter((i) => i.status === 'triggered').length;
if (id === 'acknowledged') return open.filter((i) => i.status === 'acknowledged').length;
return null;
}
function renderChips() {
const el = document.getElementById('queue-filters');
const chips = FILTERS.map((f) =>
h('button', {
const chips = FILTERS.map((f) => {
const count = chipCount(f.id);
return h('button', {
class: 'chip',
type: 'button',
role: 'tab',
'aria-selected': String(f.id === filter),
onclick: () => setFilter(f.id),
text: f.label,
}),
);
}, count != null && h('span', { class: 'count', text: String(count) }));
});
// Somebody in one team has nothing to choose between, so the row of team
// chips appears only when there is more than one. The default is all of
@@ -140,6 +158,12 @@ function renderChips() {
}
}
// A scroll hint for the phone-width row, where the chips can run off the
// right edge with nothing to suggest there's more; the desktop sidebar
// wraps instead of scrolling (see .pane-list .chips), so this fades out
// there via CSS rather than being left out here.
chips.push(h('span', { class: 'chips-fade', 'aria-hidden': 'true' }));
clear(el, chips);
}
@@ -155,8 +179,8 @@ function renderList() {
return;
}
if (!items.length) {
const [title, text] = EMPTY[filter];
clear(el, checklist, emptyState(title, text, filter === 'open' ? 'checkCircle' : null));
const [title, text, iconName] = EMPTY[filter];
clear(el, checklist, emptyState(title, text, iconName));
return;
}
clear(el,
@@ -201,9 +225,12 @@ function row(inc, index) {
dataset: { index: String(index) },
},
h('div', { class: 'row-title', text: inc.title }),
h('div', { class: 'row-age', title: inc.triggered_at, text: age(inc.triggered_at) }),
h('div', { class: 'row-age', title: inc.triggered_at, text: `Triggered ${ago(inc.triggered_at)}` }),
h('div', { class: 'row-meta' },
status,
// The left-border colour alone doesn't say what it means; spell it out
// too, same badge the incident detail page uses for severity.
inc.severity && badge(inc.severity, `plain ${severityClass(inc.severity)}`),
assignee,
team,
labels && h('span', { class: 'labels', text: labels }),
+1 -1
View File
@@ -80,7 +80,7 @@ function render() {
if (error && !data) body = h('div', { class: 'load-error', text: error });
else if (!data) body = spinner();
else if (!data.incidents.total && !data.byHour.some((x) => x.count)) {
body = emptyState('No data in this range', '', 'chart');
body = emptyState('No data in this range', 'Nothing happened in this window.', 'chart');
} else {
body = h('div', { class: 'stats' },
error && h('div', { class: 'load-error', text: `Showing older data: ${error}` }),
+5 -5
View File
@@ -182,7 +182,7 @@ function subnav() {
// to the global team selector in the nav, which is what onTeamChange above
// reacts to.
function teamPicker() {
return h('div', { class: 'card' }, h('h2', { text: data.team.name }));
return h('h1', { class: 'detail-title', text: data.team.name });
}
// --- overview --------------------------------------------------------------
@@ -204,18 +204,18 @@ function overview() {
menuCard('/team/rota', 'Rota', null,
onToday ? `${onToday.username} is on call today.` : 'Nobody is on call today.'),
menuCard('/team/members', 'Members', (data.members || []).length,
owners === 1 ? 'One owner.' : `${owners} owners.`),
owners === 1 ? '1 owner, full access.' : `${owners} owners, full access.`),
menuCard('/team/escalation', 'Escalation', levels || null,
levels
? `${levels === 1 ? 'One level' : `${levels} levels`}${data.escalation.fallback_topic ? ', then a fallback topic.' : '.'}`
? `${levels} escalation level${levels === 1 ? '' : 's'}${data.escalation.fallback_topic ? ', then a fallback topic.' : '.'}`
: 'No ladder — nobody but the first person is woken.'),
menuCard('/team/sources', 'Alert sources', keys || null,
keys
? (unused ? `${unused} of them never used.` : 'All in use.')
? (unused ? `${unused} of ${keys} never used.` : `${keys} key${keys === 1 ? '' : 's'}, all used recently.`)
: 'No key yet, so nothing can reach this team.'),
menuCard('/team/deadman', 'Dead man’s switches', switches || null,
switches
? (dead ? `${dead} of them silent.` : 'All quiet, as they should be.')
? (dead ? `${dead} of ${switches} silent.` : [icon('checkCircle', 'icon overview-note-icon'), ' Dead man’s switch: healthy.'])
: 'Nothing watched.'),
);
}
+4 -1
View File
@@ -5,7 +5,7 @@
// between, the same rule every other team-aware control in this app follows;
// see state.js's currentTeam() for why nobody with just one ever has to.
import { h, openSheet, closeSheet } from './ui.js';
import { h, icon, openSheet, closeSheet } from './ui.js';
import { state, currentTeam, setSelectedTeam, onTeamChange } from './state.js';
import { teamColorClass } from './format.js';
@@ -36,6 +36,9 @@ export function render() {
btn.replaceChildren(
h('span', { class: dotClass }),
h('span', { class: 'team-selector-label', text: label }),
// Without this the pill reads as a tag rather than something you can
// open; the chevron is the only thing that says "dropdown" here.
icon('chevronDown', 'icon team-selector-chevron'),
);
}
}
+6 -1
View File
@@ -52,6 +52,7 @@ const ICONS = {
trash: ['M4 7h16', 'M9 7V4h6v3', 'M6 7l1 13h10l1-13'],
chevronLeft: ['M15 18l-6-6 6-6'],
chevronRight: ['M9 6l6 6-6 6'],
chevronDown: ['M6 9l6 6 6-6'],
external: ['M14 4h6v6', 'M20 4l-9 9', 'M18 14v6H4V6h6'],
logout: ['M15 4h4v16h-4', 'M10 17l5-5-5-5', 'M15 12H4'],
};
@@ -186,12 +187,16 @@ export function labelChip(k, v) {
// One entry in a section's overview: a card that is a link, carrying the count
// only that section can state. Both the Admin tab and the Team tab open on one
// of these menus, and a menu item is a shape rather than a page's own idea.
// note is usually just a string, but can be an array of children instead —
// deadman's "healthy" note below pairs text with a status icon.
export function menuCard(href, label, count, note) {
return h('a', { class: 'card overview-item', href },
h('div', { class: 'overview-head' },
h('strong', { text: label }),
count != null && h('span', { class: 'overview-count', text: String(count) })),
h('p', { class: 'muted small', text: note }),
typeof note === 'string'
? h('p', { class: 'muted small', text: note })
: h('p', { class: 'muted small' }, note),
);
}