Commit Graph

171 Commits

Author SHA1 Message Date
Niklas Ye 0f88574a41 Set the chart's placeholder version to 0.39.0
CI / chart (push) Successful in 4s
CI / security (push) Successful in 45s
CI / test (push) Successful in 6m1s
Release / test (push) Successful in 8s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 25s
Release / image (push) Successful in 1m12s
Release / scan-image (push) Successful in 6s
Cosmetic: make helm-package passes --version and --app-version from the
tag, so these two fields decide nothing about what gets published. Still
done, as with 584d344 (0.38.0) and df83adf (0.37.2) before it, because a
tree heading for v0.39.0 that still says 0.38.0 tells its reader
something false.
v0.39.0
2026-10-08 09:02:07 +02:00
Niklas Ye 3613fd5732 Check rows.Err() in the remaining Next() loops (#27)
Incident list, the three stats breakdowns and the user list could return a
truncated result as if complete when the scan failed partway.
2026-10-08 09:00:54 +02:00
Niklas Ye dc62278788 Check rows.Err() after the timeline scan loop (#27) 2026-10-08 08:58:19 +02:00
Niklas Ye 0aaea8efb5 Record the actor on assign, archive and unarchive (#35)
Assign logged only the assignee (user_id), archive/unarchive logged nothing.
Migration 018 adds actor_user_id/actor_service_account_id to incident_events
for 'assigned'; archive/unarchive now log archived/unarchived events with the
caller via callerActorIDs. Timeline JSON gains actor_* fields; web timeline
renders them. Service accounts are still not assignable.
2026-10-08 08:53:04 +02:00
Niklas Ye 584d3441fc Set the chart's placeholder version to 0.38.0
CI / chart (push) Successful in 1s
CI / security (push) Successful in 2m9s
CI / test (push) Successful in 6m7s
Release / test (push) Successful in 8s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 35s
Release / image (push) Successful in 1m10s
Release / scan-image (push) Successful in 2s
Cosmetic: make helm-package passes --version and --app-version from the
tag, so these two fields decide nothing about what gets published. Still
done, as with df83adf (0.37.2) and 2b2609e (0.37.1) before it, because a
tree heading for v0.38.0 that still says 0.37.2 tells its reader
something false.
v0.38.0
2026-10-07 22:39:00 +02:00
Niklas Ye 926aa2d3ec Add optional API key expiry and a missing way to list them
Part of the same security-hardening pass as the last five commits, and
the last item in its backlog. User API keys had no expiry at all --
unlike service-account keys, visibly distinct only by their "tdsa_"
prefix -- and, it turns out while implementing this, no way to list
them either: only create (returns the raw key once) and delete-by-id
existed, so a key's owner had no way to even discover what keys they
had short of remembering IDs from creation time.

handleCreateAPIKey takes an optional expires_in_days (0, the default,
keeps today's behavior: never expires, so no existing integration is
affected). apiKeyUser's lookup now carries `expires_at IS NULL OR
expires_at > now` as part of the query itself, the same way serveAs's
disabled_at check already works -- an expired key simply fails to
resolve, like a wrong one, rather than resolving and being caught
after the fact. New GET /api/users/{id}/api-keys (requireSelfOrAdmin,
same as create/delete) lists id/name/created_at/last_used_at/expires_at,
never the raw key.

Scoped down from the original plan on request: no web UI change, since
there turned out to be no existing API-keys UI at all to extend --
building one from scratch would have been a real feature addition, not
a hardening tweak.

Mirrored the additive expires_at field in terdut-tui's APIKey struct
(separate commit, separate repo) per this workspace's version-coupling
rule; the TUI does not create or list expiring keys itself yet.

New tests (api_keys_test.go): default never-expires, expires_in_days
sets expires_at, out-of-range values rejected, an expired key fails
auth after a fresh one worked, the listing never includes the raw key.
Also added the new GET route to authz_scope_test.go's self-or-admin
table from the previous commit.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2026-10-07 22:31:44 +02:00
Niklas Ye a2ca9c25d0 Add a regression test for the team/admin/self authorization pattern
Part of the same security-hardening pass as the last four commits.
terdut-server's authz is already solid -- centralized predicates
(requireTeamMember, requireTeamOwner, requireSelfOrAdmin, AdminOnly)
rather than ad hoc per-handler checks, confirmed by spot-checking several
handlers while writing this. But it is enforced by convention, not the
type system: a future handler that forgets its guard would compile and
read fine on review, exactly like one that remembers it.

authz_scope_test.go builds two teams and, for every team-scoped route
(members, OIDC groups, invites, escalation, dead man's switches,
integrations, schedule, plus every /api/incidents/{id}/... route, scoped
by the incident's own team_id through incidentIDParam's single
chokepoint), calls it as one team's owner against the other team's
resources and asserts 404 -- requireTeamMember and requireTeamOwner both
answer a non-member that way. Separate tests cover AdminOnly's routes
(403 for a non-admin) and requireSelfOrAdmin's (403 for a non-admin
acting on someone else's account).

Verified the test actually catches a regression, not just that it
passes today: temporarily removed handleListTeamMembers' requireTeamMember
call, confirmed exactly that one subtest failed and nothing else did,
then put it back.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2026-10-07 22:21:46 +02:00
Niklas Ye 92959cac38 Run the pod as non-root with a read-only filesystem
Part of the same security-hardening pass as the last three commits.
Neither the Dockerfile nor the chart's Deployment set any securityContext
at all, so the container ran as root by default — scratch has no
/etc/passwd for a USER directive to resolve against, so nobody had set one.

Dockerfile now ends with USER 65532:65532 (numeric, since scratch has no
user database; 65532 is the common "nonroot" convention, distroless's own
uid). The chart's Deployment adds a matching pod-level securityContext
(runAsNonRoot, runAsUser/runAsGroup: 65532, seccompProfile: RuntimeDefault)
plus per-container hardening (allowPrivilegeEscalation: false,
capabilities dropped, readOnlyRootFilesystem: true) on both the app
container and the wait-for-postgres init container — neither writes
anything to disk, so the root filesystem can stay read-only.

Verified with helm-lint and a manual `helm template` render of both the
terdut-server and terdut-demo charts. Not yet verified: an actual pod
starting with these in place — readOnlyRootFilesystem is exactly where a
non-obvious write (a temp file, a cache dir) would surface as a crash
rather than a lint error, so that needs a real rollout to confirm.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2026-10-07 22:16:06 +02:00
Niklas Ye f15db0e20a Add gosec to CI; fix the one real finding it surfaced
Part of the same security-hardening pass as the last two commits. make
lint was go vet only; govulncheck and gitleaks already scanned deps and
secrets on every push, but nothing read this repo's own source for
risky patterns (weak crypto, injection shapes, insecure cookies, ...).

New `make security-code` runs gosec, wired into ci.yaml's security job
alongside the other two. G104 (unchecked error) is excluded at the
Makefile level: every one of its 41 initial hits was this codebase's
existing, deliberate idiom for a best-effort write or an already-
reviewed json.Unmarshal of its own JSONB, predating gosec, and the rule
cannot tell that apart from a mistake -- seventeen individual #nosec
comments would hide a future real G104 regression in the suppression
noise rather than surface it. Reasoning is on the Makefile target.

Of the 12 remaining hits:
  - Genuinely real: oidc.go's callback logged error_description (and,
    two call sites down, identity.Subject) via %s before the request's
    state was even checked against its cookie -- an attacker-reachable
    value going into the log unquoted. Switched to %q, matching
    identity.Username's existing treatment, so a value holding a
    newline can't forge a second log line.
  - False positives, annotated inline rather than globally suppressed:
    4x G124 on cookies that already set Secure via cookieSecure(...)
    (a function call, not the literal `true` the rule wants), 3x G202
    on sqlArgs-built queries that only ever splice in a "$N"
    placeholder, never a value, and the remaining 5x G706 on log lines
    that were already %q-quoted -- gosec's taint analysis doesn't
    model format verbs, so it flags the tainted argument regardless.

Also fixed handleMe's swallowed Scan error (gosec's catch, pre-fix):
a transient DB error left hash/dismissed at their zero values and the
response claimed no password and no onboarding dismissal regardless
of the truth, rather than surfacing a 500.

Checked both workflow files for the injection class letsvisit found
there (a `${{ }}` expression spliced straight into a `run:` block):
every one here already goes through `env:` as a quoted shell variable,
documented in ci.yaml's own header comment. Nothing to fix.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2026-10-07 22:12:31 +02:00
Niklas Ye b82c10acf4 Back the login/signup/OIDC/device rate limiters with Postgres
Part of the same security-hardening pass as the body-size/header commit.
loginLimiter was an in-memory, per-process sync.Mutex+map -- fine for one
replica, but charts/terdut-server/values.yaml has set replicaCount: 2 in
production since v0.37.0. Each pod counted only its own traffic, so every
limit it guarded (failed logins, sign-ups, OIDC/device-login starts) was
effectively twice as generous as the constants say, not just in theory.

loginLimiter now stores its counters in a new rate_limit_counters table
(migration 016) instead of a map; blocked/fail/clear take a context and
query/upsert/delete a row keyed by the same strings callers already used
(username, client address, "signup:"+address, ...). Semantics are
unchanged -- a fixed window that resets rather than slides -- so no call
site's behavior changes, only where the count lives. Added purgeRateLimits
to the sweeper, alongside purgeSessions/purgeAckTokens, so expired windows
don't accumulate.

New internal (package api) tests in rate_limiter_test.go cover the basic
behavior plus the regression this exists to fix: two loginLimiter values
sharing one database, standing in for two replicas, now see one combined
count instead of each keeping their own.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2026-10-07 22:03:51 +02:00
Niklas Ye 7cd6fbf571 Cap request body size and add baseline security headers
Part of a security-hardening pass (see wiki for the full backlog).
decodeJSON had no size limit at all, so every JSON endpoint -- including
the two unauthenticated ones (bootstrap, the Alertmanager webhook) --
would buffer an attacker-supplied body of unbounded size before it was
even validated. decodeJSON now wraps the body in http.MaxBytesReader at
a 1 MiB default; the webhook gets its own 8 MiB cap via decodeJSONLimit,
since a real Alertmanager batch can be bigger than an ordinary API body.

Also adds a securityHeaders middleware, applied globally: nosniff on
every response (previously only the static site got it), and HSTS
(180-day max-age, conservative on purpose) whenever cookieSecure's
signal says the browser is on HTTPS. Checked the chart/gateway config
first -- neither sets HSTS anywhere, so this was a real gap.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2026-10-07 21:57:00 +02:00
Niklas Ye df83adfe47 Set the chart's placeholder version to 0.37.2
CI / chart (push) Successful in 1s
CI / security (push) Successful in 20s
CI / test (push) Successful in 5m28s
Release / test (push) Successful in 8s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 49s
Release / image (push) Successful in 1m17s
Release / scan-image (push) Successful in 26s
v0.37.2
2026-10-07 21:12:51 +02:00
Niklas Ye 9da913080f Stop 500ing when a service account acts on an incident
Every incident-mutation handler read userFromContext(ctx) and wrote the
result's .ID into acknowledged_by/incident_events.user_id without checking
the ok bool. For a team-scoped service-account caller this returned a
zero-value user id, which violated the users(id) FK and 500'd on
acknowledge, unacknowledge, resolve, snooze, unsnooze and create-note.
handleDeleteNote didn't crash but silently matched zero rows instead
(WHERE user_id = 0), so a service account could never delete its own note.

Add acknowledged_by_service_account_id (incidents) and service_account_id
(incident_events) as nullable FKs to service_accounts(id), parallel to and
mutually exclusive with the existing human columns (migration 015, with a
CHECK enforcing the exclusion). Route every one of the six handlers plus
delete-note through a new callerActorIDs() helper that branches on
Caller.AsHuman()/ServiceAccountID() instead of assuming a human, and thread
a serviceAccountID parameter through logEvent and the new
acknowledgeIncidentAs (acknowledgeIncident itself is untouched: its only
other caller, the push-notification Acknowledge button, is always human).
Render the new actor distinctly from both a human and "the server acted"
in the web UI's incident timeline and facts card.

handleIncidentAssign, handleIncidentArchive and handleIncidentUnarchive are
deliberately not touched here — they track no actor at all today, for
anyone, which is a separate pre-existing gap (follow-up issue to come).

Fixes #25
2026-10-07 21:09:31 +02:00
Niklas Ye 2b2609e98f Set the chart's placeholder version to 0.37.1
CI / chart (push) Successful in 1s
CI / security (push) Successful in 1m18s
CI / test (push) Successful in 6m25s
Release / test (push) Successful in 7s
Release / chart (push) Successful in 7s
Release / image (push) Successful in 2m22s
Release / scan-image (push) Successful in 26s
Release / binaries (push) Successful in 2m53s
Cosmetic: `make helm-package` passes --version and --app-version from
the tag, so these fields decide nothing about what gets published. But
a tree heading for v0.37.1 that still says 0.37.0 tells a reader
something false. Same as 7efd1bb, which cites 5c4e0bd.
v0.37.1
2026-10-04 14:26:31 +02:00
Niklas Ye fa6d82d6e5 Fix desktop layout bugs and show more incident actions directly
On-call, the incident queue and the account page all had latent CSS bugs
that only show up once the browser is wide enough to hit the desktop
breakpoint (900px+):

- On-call: .days reset its own margin to 0, which canceled the page-wide
  auto-centering on just that element, leaving the day list pinned to the
  left edge while every other card on the page centered normally.
- Queue: the list pane stayed capped at 340-420px even with nothing
  selected, leaving the rest of the screen empty. It now fills the width
  until an incident is picked, then goes back to list+detail.
- Account, admin user and admin team: .btn and .back-link are inline-flex,
  and margin:auto only centers a block box, so the Sign-out button and the
  two admin back-links sat left of their sibling cards instead of matching
  their width. Wrapped each in a block div.

The incident detail action bar also folded Assign, Add note, Resolve and
Clear acknowledgement into a "More" sheet sized for a phone's width.
Desktop has the room, so it now shows them as direct buttons and hides
More instead; Copy incident stays out of the bar since the header already
has its own button for it.

Filed as niklas/terdut-server#31, #32, #33, #34, each with a screenshot.
2026-10-04 14:26:16 +02:00
Niklas Ye 7efd1bbba7 Set the chart's placeholder version to 0.37.0
CI / chart (push) Successful in 1s
CI / test (push) Successful in 7s
CI / security (push) Successful in 12s
Release / test (push) Successful in 6s
Release / chart (push) Successful in 1s
Release / binaries (push) Successful in 21s
Release / image (push) Successful in 1m3s
Release / scan-image (push) Successful in 1s
v0.37.0
2026-10-03 16:17:20 +02:00
Niklas Ye 3bf94a5d7f Default to 2 replicas and RollingUpdate now that the singleton jobs are locked
CI / chart (push) Successful in 0s
CI / test (push) Successful in 7s
CI / security (push) Successful in 12s
replicas and strategy: Recreate were the chart's only guard against the
archiver, notifier and migration races; v0.36.0 closed all three with
advisory locks and a conflict-resolving incident insert, which made
that guard redundant rather than load-bearing. Expose replicaCount
(new, no values.yaml key existed before) defaulting to 2, and switch
to strategy: RollingUpdate with no explicit maxUnavailable/maxSurge --
the 25%/25% default rounds to 0/1 at replicaCount: 2, which is already
zero-downtime.

The chart does not gate this on image.tag, so pointing it at a
pre-v0.36.0 image with the new default is a foot-gun by omission --
noted in both the values.yaml comment and the deployment.yaml comment,
not guarded in code, same as the chart does for every other
version-coupled assumption today.

Co-authored-by: Claude <noreply@anthropic.com>
2026-10-03 12:30:27 +02:00
Niklas Ye 5c4e0bdd0e Set the chart's placeholder version to 0.36.0
CI / chart (push) Successful in 1s
CI / security (push) Successful in 16s
CI / test (push) Successful in 4m50s
Release / test (push) Successful in 7s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 48s
Release / image (push) Successful in 1m12s
Release / scan-image (push) Successful in 2s
v0.36.0
2026-10-03 12:03:28 +02:00
Niklas Ye 9e5b085d8b Resolve the new-incident insert conflict instead of dropping the payload
CI / chart (push) Successful in 1s
CI / security (push) Successful in 20s
CI / test (push) Has been cancelled
openIncident's INSERT had no ON CONFLICT clause, relying entirely on
incidentForGroup's earlier SELECT to avoid a duplicate. On more than
one replica, two webhook deliveries for the very first occurrence of
a brand-new groupKey can both pass that SELECT before either INSERTs;
the loser then hit incidents_open_group_key_idx's unique violation,
which rolled back its whole transaction — including that payload's
alert upserts, done earlier in the same transaction. ingest's error is
only logged and receiveWebhook answers 200 regardless, so nothing
retried it: the loser's alerts silently never existed.

Add ON CONFLICT (team_id, group_key) WHERE resolved_at IS NULL DO
NOTHING to the INSERT, matching the partial unique index. Postgres
only resolves that conflict after the winning transaction commits (or
rolls back), so by the time RETURNING comes back empty,
existingOpenIncident's follow-up SELECT is guaranteed to see the
winner's row. The loser attaches to it instead of failing outright,
and the rest of its payload commits normally. Covers both callers,
since the dead man's switch sweeper shares this same function.

New test (package api_test, fires N webhook deliveries for one
groupKey from a synchronized start with distinct fingerprints, so
they aren't accidentally serialized by upsertAlerts' own per-
fingerprint lock) confirmed meaningful: with the ON CONFLICT clause
reverted, it fails 10/10 on a missing alert fingerprint; restored, 0/10.
Note while building it: "exactly one incident" alone cannot
distinguish fixed from broken, since the DB's own unique index already
guarantees that either way — the real signal is the loser's payload
surviving.

Chart comment updated: all three of the chart's original reasons for
Recreate are now addressed in code, though replicas stays at 1 and the
strategy stays Recreate pending a deliberate decision to raise it.

Co-authored-by: Claude <noreply@anthropic.com>
2026-10-03 11:58:39 +02:00
Niklas Ye 0050738ca0 Advisory-lock DB migrations against concurrent replica startup
Migrate's check-then-apply loop against schema_migrations had no
locking: two replicas booting at once against a fresh or
partially-migrated database could both pass the "not yet applied"
check for the same file and race applying it, crashing whichever lost
the duplicate-key insert (confirmed: reverting the lock fails the new
test 10/10 on a duplicate-key violation, racing as early as the
CREATE TABLE IF NOT EXISTS schema_migrations statement itself).

Hold a Postgres advisory lock for Migrate's whole run, on a dedicated
connection reserved via db.Conn so lock and unlock happen on the same
session. Blocking (pg_advisory_lock), unlike the archiver/notifier's
pg_try_advisory_lock: on boot there's no later tick to defer to, so a
second replica should wait for the first to finish migrating rather
than skip ahead.

Adds internal/db's first test file, exercising two concurrent Migrate
calls against a fresh schema.

Still open: the new-incident-insert race on a webhook for a brand-new
groupKey, noted in the chart's updated comment. Login rate limiting
staying in-process, diluted across replicas, is an accepted tradeoff.

Co-authored-by: Claude <noreply@anthropic.com>
2026-10-03 11:43:09 +02:00
Niklas Ye 42180948d1 Guard the archiver and notifier passes with a Postgres advisory lock
Both background loops run unconditionally on every instance with no
coordination between them, which the chart's replicas: 1 + strategy:
Recreate exists specifically to paper over: with more than one replica,
every one of them would sweep and deliver notifications independently,
and two overlapping during a rollout would both page for the same
incident.

Add withAdvisoryLock, which takes a Postgres advisory lock on a
dedicated connection and runs a pass only if it gets the lock,
otherwise skipping until the next tick. Wire StartArchiver and
StartNotifier through it with their own lock keys, so Sweep and
NotifySweep themselves are untouched and every existing test calling
them directly keeps working unchanged.

This also closes the notifier's double-delivery race in passing: two
replicas can no longer both be inside deliverPending at once, since
only one can hold notifierLockKey at a time.

Deliberately not addressed here, and still blocking a replica count
above 1: the in-memory login rate limiter, the unlocked migration
runner, and the new-incident-insert race on a webhook for a brand-new
groupKey. Noted in the updated chart comment.

Co-authored-by: Claude <noreply@anthropic.com>
2026-10-03 11:36:48 +02:00
Niklas Ye c83c7c2a8b Set the chart's placeholder version to 0.35.1
CI / chart (push) Successful in 2s
CI / security (push) Successful in 20s
CI / test (push) Successful in 4m45s
Release / test (push) Successful in 3s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 50s
Release / image (push) Successful in 1m21s
Release / scan-image (push) Successful in 3s
Cosmetic, as before (497086c, 4358e84): make helm-package passes
--version and --app-version from the tag, so this decides nothing
about what gets published. Kept in step so the tree heading for
v0.35.1 doesn't say 0.35.0 to a reader who hasn't yet seen the tag.

🤖 Generated with [Claude Code](https://claude.ai/code)

Co-Authored-By: Claude <noreply@anthropic.com>
v0.35.1
2026-10-03 09:39:01 +02:00
niklas 43beda9a30 Merge pull request 'Fix the queue chip fade falling short of the right edge' (#30) from fix-chips-fade-edge into main
CI / chart (push) Successful in 1s
CI / security (push) Successful in 21s
CI / test (push) Has been cancelled
2026-10-03 07:36:12 +00:00
Niklas Ye 710521a73c Fix the queue chip fade falling short of the right edge
CI / chart (pull_request) Successful in 1s
CI / security (pull_request) Successful in 16s
CI / test (pull_request) Successful in 5m13s
position: sticky, as a flex item of the row it's pinning itself
against, interacted with that row's gap and its own negative margin in
a way that landed it short of the true edge -- visibly, a sliver of
the next chip stayed poking out past where the fade should have
covered it, which is the "ends before the screen edge" bug reported
against the release.

Replaced with the simpler, better-supported pattern for this: an
absolutely positioned overlay against a position:relative, overflow:
auto parent. Unlike a sticky descendant, an absolutely positioned one
is resolved against the parent's own (non-scrolling) box, so it stays
flush with the real edge regardless of scroll position, without the
flex-gap/margin interaction that caused this.

Verified visually: a standalone reproduction of both versions,
screenshotted with Chromium's headless_shell (no browser automation
tool available in this environment, but the binary's right there) --
the old version shows the next chip's edge peeking past the fade, the
new one doesn't.

🤖 Generated with [Claude Code](https://claude.ai/code)

Co-Authored-By: Claude <noreply@anthropic.com>
2026-10-03 09:33:33 +02:00
niklas d9492913ed Merge pull request 'Fix Stats/Admin/Account showing in the bottom bar too' (#29) from fix-nav-secondary-cascade into main
CI / chart (push) Successful in 0s
CI / test (push) Successful in 8s
CI / security (push) Successful in 14s
Reviewed-on: #29
2026-10-03 07:30:04 +00:00
Niklas Ye 1770e5d945 Fix Stats/Admin/Account showing in the bottom bar too
CI / chart (pull_request) Successful in 1s
CI / security (pull_request) Successful in 16s
CI / test (pull_request) Successful in 5m2s
.nav-link-secondary's display:none sat before .nav-link's own
display:flex in the file. Both are single-class selectors, so they tie
on specificity, and a tie is broken by which one comes later in the
file -- not by which class the element happens to carry. .nav-link's
declaration, being later, won for every element wearing both classes,
so Stats/Admin/Account rendered as three extra tabs on the phone bar
instead of folding into "More" as intended.

Moved the rule below .nav-link instead of changing either declaration,
since nothing about the values was wrong -- only their order was.

🤖 Generated with [Claude Code](https://claude.ai/code)

Co-Authored-By: Claude <noreply@anthropic.com>
2026-10-03 09:22:48 +02:00
Niklas Ye 497086cb51 Set the chart's placeholder version to 0.35.0
CI / test (push) Successful in 4s
CI / chart (push) Successful in 0s
CI / security (push) Successful in 12s
Release / test (push) Successful in 4s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 51s
Release / image (push) Successful in 1m5s
Release / scan-image (push) Successful in 23s
Cosmetic: make helm-package passes --version and --app-version from the
tag, so these two fields decide nothing about what gets published. Kept
in step anyway, the same as 4358e84 and fd26fef before it, so the tree
heading for v0.35.0 does not say 0.34.0 to a reader who hasn't yet seen
the tag.

🤖 Generated with [Claude Code](https://claude.ai/code)

Co-Authored-By: Claude <noreply@anthropic.com>
v0.35.0
2026-10-03 09:10:30 +02:00
niklas e3090d2779 Merge pull request 'Web: visual design pass (issue #26)' (#28) from design-polish-issue-26 into main
CI / test (push) Successful in 4s
CI / chart (push) Successful in 1s
CI / security (push) Successful in 13s
Reviewed-on: #28
2026-10-03 07:08:19 +00:00
Niklas Ye 91f03c21e8 Web: visual design pass (issue #26)
CI / chart (pull_request) Successful in 2s
CI / security (pull_request) Successful in 18s
CI / test (pull_request) Successful in 4m49s
Addresses the screenshot-review feedback in #26. No framework or build
step added — all of this stays within the existing plain HTML/CSS/
vanilla-JS + go:embed architecture.

- Nav: re-enable the bottom tab bar that was already built and
  switched off (Queue/On-call/Alerts/Team + a "More" sheet for
  Stats/Admin/Account), replacing the hamburger on phone width.
- Queue: chip counts, a scroll fade on the filter row, a
  "Triggered Xh ago" + severity label per row, a chevron on the team
  switcher so it reads as a dropdown.
- On-call: collapse repeated same-person days into shift bars (week
  view and "your shifts" both), show the week as a date range with
  the ISO week number as secondary text, split "Current shift" out
  from "Next shifts" with "ends in Nd", a pill badge + row highlight
  for "you".
- Incident detail: fix the actual bug behind the duplicate
  "acknowledged" timeline entries (acknowledgeIncident's UPDATE had no
  guard on the incident's current status, so acknowledging an
  already-acknowledged incident silently re-logged the event — now
  idempotent, with regression tests on both the authenticated route
  and the ntfy ack-button route). Relabel escalation re-pages so they
  don't look like the same page landing twice. Copy the primary action
  up near the top. Label the "···" button. Group the timeline by
  phase (triggered/acknowledged/resolved). Add an "at a glance"
  summary row (duration/severity/responsible) and collapse the group
  labels by default.
- Team overview: reword the vague copy ("One owner." etc.) into plain
  labels.
- Empty states: fill in missing icons/one-liners across queue,
  alerts, stats and the incident timeline.
- CSS: fix card padding bugs, verify link contrast already passes AA,
  introduce a --fs-* type-scale token set and migrate the few
  genuinely isolated cases onto it (left sizes tied to a fixed shape,
  a deliberately prominent display, or a non-negotiable constraint
  like the iOS-zoom-prevention input size as documented exceptions
  rather than guess at a render this change can't see).

Verified with the full fmt/lint/test/helm-lint gate, plus a live
instance against the test DB with seeded incidents and schedule data
to trace the on-call grouping and timeline phase-splitting logic
against real API responses.

🤖 Generated with [Claude Code](https://claude.ai/code)

Co-Authored-By: Claude <noreply@anthropic.com>
2026-10-03 09:01:54 +02:00
Niklas Ye 4358e84b24 Set the chart's placeholder version to 0.34.0
CI / test (push) Successful in 4s
CI / chart (push) Successful in 1s
CI / security (push) Successful in 12s
Release / test (push) Successful in 4s
Release / chart (push) Successful in 7s
Release / image (push) Successful in 2m27s
Release / binaries (push) Successful in 2m35s
Release / scan-image (push) Successful in 3s
v0.34.0
2026-10-02 22:08:35 +02:00
niklas f45dc2f925 Merge pull request 'internal/api: unify human/service-account authz into one Caller type' (#24) from unify-caller-authz into main
CI / chart (push) Successful in 1s
CI / security (push) Successful in 22s
CI / test (push) Successful in 27s
2026-10-02 20:06:37 +00:00
Niklas Ye 774fdfcaa8 internal/api: unify human/service-account authz into one Caller type
CI / chart (pull_request) Successful in 1s
CI / security (pull_request) Successful in 15s
CI / test (pull_request) Successful in 5m21s
ctxUser/ctxTeams (human) and ctxServiceAccount (+ a synthetic ctxTeams
entry, service account) used to be two parallel, un-unified context
representations -- every authz predicate had to remember which one(s) it
needed, and every place that forgot either wrongly 403'd a service account
(terdut-server#23, terdut-operator#3), crashed on an unchecked zero-value
user id, or silently no-op'd. New internal/api/caller.go collapses both
into one Caller, stored under one ctxCaller key by serveAs/serveAsServiceAccount;
every existing predicate (userFromContext, callerTeamIDs, callerRole,
callerIsAdmin, isInstanceServiceAccount, AdminOnly, requireSelfOrAdmin,
requireTeamOwner, OperatorModeBlock) now reads through it, with identical
behavior for every untouched call site (alerts.go, incidents.go,
schedule.go, stats.go, etc.) -- confirmed by the full existing suite
passing unchanged.

Four real fixes land alongside the refactor, not just the restructuring:

1. callerMayManageServiceAccount gains the one load-bearing branch this
   exists for: an instance-scoped service account may now manage (mint or
   revoke a key on) any team-scoped account, not only a human admin, that
   team's human owner, or the account itself. handleCreateServiceAccount
   already let an instance-scoped caller *create* a team-scoped account for
   any team; adopting or rotating one it didn't just create in the same
   call -- terdut-operator's own documented crash-window recovery -- had no
   equivalent permission and 403'd forever. Closes terdut-operator#3.

2. handleCreateInvite wrote a service-account caller's zero-value user id
   straight into invites.created_by (nullable, but never passed as nil),
   which foreign-key-violates against users(id) -- a 500, not success, for
   any team-scoped service account minting an invite. Fixed the same way
   handleCreateServiceAccount already handles the analogous case. Found
   live while verifying this change, not filed separately since it's fixed
   in the same place it was found.

3. handleMe and handleTestNotification 500'd for a service-account caller
   (fetchUser/the ntfy_topic lookup against a zero-value user id that
   matches no row); handleDismissOnboarding silently no-op'd (UPDATE ...
   WHERE id = 0). All three now call Caller.AsHuman() and return an
   explicit 403 ("this endpoint is for human accounts only").

4. Ratifies, rather than further narrows, two capabilities a team-scoped
   service account already had by construction and this document's own
   text once called "a gap acknowledged rather than closed": owner-equivalent
   reach over membership/invites, and minting another service account for
   its own team. terdut-operator's new TerdutTeam invite-minting feature is
   about to depend on the first one, so this makes it documented, tested,
   intentional behavior instead of an accident nobody was supposed to rely
   on.

AdminOnly/requireSelfOrAdmin are unchanged in effect: still human-only,
forever, for every scope of service account -- confirmed by
TestAdminOnly_RefusesEveryServiceAccountScope. terdut-server#23's named
routes (POST /api/users, PUT /api/admin/settings) were never the right
thing to widen; its real fix is the terdut-operator invite feature,
recorded in SERVICE-ACCOUNTS.md's "What this unblocks" and closing that
issue once it ships.

SERVICE-ACCOUNTS.md amended in place (not a new file, its own established
convention) to describe the as-built Caller model, correct its own
aspirational claim about AdminOnly that TEAM-LOOKUP.md had already flagged
as not matching shipped code, and record all of the above.
2026-10-02 21:52:43 +02:00
Niklas Ye fd26fef1ba Set the chart's placeholder version to 0.33.2
CI / test (push) Successful in 4s
CI / chart (push) Successful in 1s
CI / security (push) Successful in 12s
Release / test (push) Successful in 3s
Release / chart (push) Successful in 3s
Release / binaries (push) Successful in 16s
Release / image (push) Successful in 1m8s
Release / scan-image (push) Successful in 1s
make helm-package passes --version and --app-version from the tag, so
these fields decide nothing about what is published -- but a tree heading
for v0.33.2 that still says 0.33.1 tells its reader something false.
Same as 4b15079 and 6f8499f before it.
v0.33.2
2026-10-02 09:42:50 +02:00
Niklas Ye a9d788cc83 Wait for Postgres to accept connections before the main container starts
A Deployment created before Postgres has finished its very first boot --
initdb plus Patroni leader election, on a from-scratch postgres-operator
cluster -- crash-looped a few times. db.Open()'s own ping-retry budget
(pingAttempts/pingRetryDelay, internal/db/db.go) is sized for a much
shorter, different race -- NetworkPolicy propagation, a few seconds -- not
for genuine first-time cluster creation, which routinely takes longer, so
it exhausted and the process exited before ever binding its HTTP port. A
startupProbe cannot fix that: the crash happens before there is anything
to probe.

Added a wait-for-postgres init container instead: it loops pg_isready
against database.dsn until Postgres actually answers, before the main
container's own, unchanged retry budget gets a chance to run out.
pg_isready needs no credentials -- it reports PQPING_OK on anything that
amounts to a Postgres backend answering, including an auth challenge --
so no PGPASSWORD is wired into it.

Chart-only; no Go code changed. database.waitForPostgres.enabled defaults
to true and can be turned off if something else already guarantees
Postgres is reachable before this Deployment is created.
2026-10-02 09:42:42 +02:00
Niklas Ye 4b15079ac2 Set the chart's placeholder version to 0.33.1
CI / chart (push) Successful in 1s
CI / security (push) Successful in 19s
CI / test (push) Successful in 4m42s
Release / test (push) Successful in 3s
Release / chart (push) Successful in 2s
Release / image (push) Successful in 2m11s
Release / scan-image (push) Successful in 3s
Release / binaries (push) Successful in 2m47s
Cosmetic: `make helm-package` passes --version/--app-version from the
tag, so this field decides nothing about what gets published. Still done
so the tree doesn't say 0.33.0 while heading for a v0.33.1 release.
Cites fc9f47c, the fix this version actually is.
v0.33.1
2026-10-02 09:18:28 +02:00
Niklas Ye fc9f47cc8d Refuse an OIDC sign-in from creating the very first user
Closes the race terdut-operator#1 found: /api/bootstrap and OIDC
auto-provisioning both key off the same signal (SELECT COUNT(*) FROM
users) with no coordination between them, so an otherwise-ordinary OIDC
sign-in against a freshly-created, not-yet-bootstrapped install could
create user #1 itself and take the one slot /api/bootstrap expects to win
uncontested (terdut-operator's own design, DESIGN.md §1/§6, assumes it is
the only caller). The operator has no way to recover from losing that
race -- it never gets a credential, and nothing it owns can clear the
occupying user row.

resolveSSOUser now checks the same gate handleBootstrap already does,
right where it's about to create a brand-new user (an identity nobody has
linked yet, that also matches no existing local account by email) -- not
anywhere else, since every other sign-in on an already-bootstrapped
install is unaffected. New sso_error code `not_bootstrapped`: the person
sees "this install is still setting up, try again in a moment" and a
second attempt once something has actually bootstrapped succeeds
normally, same as any other first sign-in.

Does not fix the other half of that issue (BootstrapStateLost's own
"delete and recreate" instructions still don't work once something has
occupied the slot some other way) -- this closes the specific race, not
every path to that state.
2026-10-02 09:12:32 +02:00
Niklas Ye bc9f793f1f Add GET /api/teams?name= (TEAM-LOOKUP.md)
CI / chart (push) Successful in 1s
CI / security (push) Successful in 1m15s
CI / test (push) Successful in 5m48s
Resolves the gap TEAM-LOOKUP.md raised: an instance-scoped service
account had no way to recover a team's id after a 409 on POST
/api/teams, unlike the already-solved equivalent for service accounts
themselves (GET /api/service-accounts?name=).

Same route, extended the same way handleListServiceAccounts already
branches on ?name=: unset behaves exactly as before (the caller's own
teams via team_members); set looks up one team by exact name, open to
any authenticated caller -- not gated by isInstanceServiceAccount or
AdminOnly, since what it discloses (a name is taken, nothing about who's
in it) is the same low sensitivity that lookup already accepts for
service-account names.

Tests cover the exact motivating scenario (create, 409 on a retry,
recover the id via ?name=), the empty-array-not-an-error case, that no
role/source is reported for a non-member match, and that a caller who
isn't a member of the matched team still gets it.
2026-10-01 10:49:11 +02:00
Niklas Ye 871274a3a0 Request: team lookup for service accounts (TEAM-LOOKUP.md)
Raised by terdut-operator's TerdutTeam controller (ROADMAP.md Stage 2):
an instance-scoped service account has no way to recover a team's id
after a 409 on POST /api/teams, unlike the equivalent, already-solved
case for service accounts themselves (GET /api/service-accounts?name=).
Also corrects a claim in SERVICE-ACCOUNTS.md's own text that doesn't
match AdminOnly's actual code -- confirmed against source, not assumed.
2026-10-01 10:43:09 +02:00
Niklas Ye 6f8499fa42 Set the chart's placeholder version to 0.33.0
CI / chart (push) Successful in 1s
CI / security (push) Successful in 18s
CI / test (push) Successful in 4m26s
Release / test (push) Successful in 9s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 29s
Release / image (push) Successful in 1m10s
Release / scan-image (push) Successful in 31s
Cosmetic: make helm-package passes --version and --app-version from the
tag, so these two fields decide nothing about what gets published. Kept
in step anyway, the same as 0ee576f (0.32.0) and 949d659 (0.31.0) before
it, so a tree heading for v0.33.0 does not read as still being on 0.32.0.
v0.33.0
2026-09-29 21:33:15 +02:00
Niklas Ye ef731e85c5 Document service accounts and operator mode in the README
The last commit (a4dd60f) shipped the code with no README update, against
this repo's own convention of documenting the whole API surface there.
Adds the Service accounts section and table, the Authentication and Teams
prose covering the new principal and TERDUT_OPERATOR_MODE, and the
/api/version row.

Corrects one thing along the way: a first draft claimed team-scoped
accounts are refused on team membership/invite endpoints. Checked against
teams.go and that is false — nothing server-side carves those two out,
only convention (no sane operator would call them) keeps them out of
automation's hands. README and SERVICE-ACCOUNTS.md now say that plainly
instead of the stronger, incorrect claim.
2026-09-29 21:33:04 +02:00
Niklas Ye a4dd60f6b8 Add service accounts and operator mode
Service accounts (SERVICE-ACCOUNTS.md) are a scoped, non-human credential:
not a users row, so they never touch OIDC sync, login or the is_admin flag.
Instance scope can create a team and mint a team-scoped account for it;
team scope is owner-equivalent for that one team and nothing else. This is
what unblocks terdut-operator's DESIGN.md §6 — no more impersonating a
human admin, and a real rotation story instead of the unworkable
delete-and-re-bootstrap /api/bootstrap can't actually do.

- migration 014: service_accounts + service_account_keys
- POST /api/service-accounts, POST/DELETE .../keys, GET ?name= self-lookup
- AuthMiddleware resolves a tdsa_-prefixed key to a distinct principal;
  a team-scoped account gets a synthetic single membership so
  requireTeamMember/requireTeamOwner work on it unmodified
- handleCreateTeam accepts an instance-scoped caller; the team it creates
  has no human owner, which is the expected shape for one an operator is
  about to hand a team-scoped credential to

Operator mode (TERDUT_OPERATOR_MODE / values.operatorMode) declares an
install gitops-managed: session and user-API-key writes to teams,
escalation policies, dead man's switches and integrations get 403
reason=operator_managed, while a service account's writes still go
through. Team membership/invites and the schedule are deliberately left
out — never gitops-managed by design, and still human day-to-day work.
/api/auth/config reports operator_mode so the web UI can grey these
sections out from the start rather than only after a write fails.

Also: GET /api/version (both terdut-tui and terdut-operator currently
detect server capability by route-probing; this gives them a real answer),
and a PUT for dead man's switches so a reconciler can update one in place
instead of deleting and recreating it.
2026-09-29 21:25:17 +02:00
Niklas Ye b5573fbca2 Add design note: scoped service-account/token type
Proposes a non-human credential type — service_accounts +
service_account_keys, instance- or team-scoped, distinct from both
user API keys (always tied to a human's full rights) and integration
keys (narrow, one-way webhook auth only). Directly unblocks
terdut-operator's DESIGN.md §6, whose bootstrap/rotation plan doesn't
work against /api/bootstrap's actual single-shot-per-install behavior.

Design note only, no implementation yet.
2026-09-29 20:35:28 +02:00
Niklas Ye 0ee576f793 Set the chart's placeholder version to 0.32.0
CI / chart (push) Successful in 1s
CI / security (push) Successful in 16s
CI / test (push) Successful in 4m2s
Release / test (push) Successful in 4s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 31s
Release / image (push) Successful in 1m10s
Release / scan-image (push) Successful in 3s
Cosmetic: make helm-package sets the published version and appVersion
from the tag, so these two fields decide nothing (see the comment
above them). Kept in step anyway, same as 949d659 and e5b4df7, so a
tree heading for v0.32.0 doesn't say 0.31.0.
v0.32.0
2026-09-27 22:05:26 +02:00
Niklas Ye b610b1817a Fold the account page's ntfy and password forms behind disclosures
Both sat open by default, competing with the rest of the page for
attention on every visit even though most visits need neither. Team's
rota already has the same problem for its bulk-assign form and solves
it with a native <details>/<summary> disclosure, styled generically in
app.css; this reuses that idiom rather than inventing a JS toggle.

Each section now shows a one-line status (the topic, or whether a
password is set) with the actual form folded under a summary naming
the action ("Set a topic" / "Change topic", "Set a password" /
"Change password"). A successful save closes the fold and confirms
with a toast, since the point of folding is that a saved form goes
back to being just a status line; a validation or API error keeps the
fold open and shows inline, next to the field it's about.

The password section's heading no longer says "Change password" or
"Set a password" itself, since that verb now lives on the summary; it
just says "Password", matching the existing SSO-off case.
2026-09-27 22:04:59 +02:00
Niklas Ye 949d6595ba Set the chart's placeholder version to 0.31.0
CI / chart (push) Successful in 1s
CI / security (push) Successful in 17s
CI / test (push) Successful in 3m48s
Release / test (push) Successful in 5s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 41s
Release / image (push) Successful in 1m3s
Release / scan-image (push) Successful in 5s
Cosmetic: make helm-package sets the published version and appVersion
from the tag, so these two fields decide nothing (see the comment
above them). Kept in step anyway, same as e5b4df7 and 97a4814, so a
tree heading for v0.31.0 doesn't say 0.30.0.
v0.31.0
2026-09-27 18:23:43 +02:00
Niklas Ye 33356ca978 Add a global, colour-coded team selector to the nav
state.js's currentTeam() was hard-coded to teams[0] and never really meant
"the team currently selected" — team.js's settings page and queue.js's
filter chips each kept their own separate, unsynchronized notion of "which
team" instead, so picking one on one page had no effect on the other.

Replaces both with a single state.selectedTeamID, set only through the new
setSelectedTeam (persisted in localStorage, unlike the queue's old per-tab
sessionStorage filter) and broadcast to listeners via onTeamChange. A new
teamselector.js control — a coloured dot plus the team's name, or "All
teams" — sits at the top of both the desktop sidebar and the mobile topbar,
opening the existing bottom-sheet menu to switch. Shown only once someone is
in more than one team, matching every other team-aware control in this app.

Colours come from a new teamColorClass() in format.js, hashing a team's id
into the six-colour rc1..rc6 palette already used for the rota's per-person
chips, so no schema or API change is needed. The queue's team filter chips
pick up the same colours.
2026-09-27 18:14:54 +02:00
Niklas Ye e5b4df7c03 Set the chart's placeholder version to 0.30.0
CI / chart (push) Successful in 1s
CI / security (push) Successful in 25s
CI / test (push) Successful in 4m17s
Release / test (push) Successful in 4s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 28s
Release / image (push) Successful in 1m9s
Release / scan-image (push) Successful in 27s
Cosmetic: make helm-package sets the published version and appVersion
from the tag, so these two fields decide nothing (see the comment
above them). Kept in step anyway, same as 97a4814 and 155f27c, so a
tree heading for v0.30.0 doesn't say 0.29.1.
v0.30.0
2026-09-27 11:44:16 +02:00
Niklas Ye 5b4683febf Let each team name its own OIDC group, not a global mapping
Team membership from single sign-on used to come from one env var,
TERDUT_OIDC_GROUP_MAPPINGS, matched against a team by name and creating
the team if none existed. That put the decision in the server's
environment rather than the team's own hands, needed a restart to
change, and let a typo in a team name silently create a stray team.

Each team now carries its own oidc_member_group and oidc_owner_group,
set by its owner (or an administrator) from the Members tab, or PUT
/api/teams/{teamID}/oidc-groups. The "highest role wins" rule
TERDUT_OIDC_GROUP_MAPPINGS used to apply across mappings now applies
across one team's own two fields: being in both makes somebody an
owner. The sync no longer creates a team by name; a group only ever
grants into a team that already exists.

This is a breaking change for anyone already using
TERDUT_OIDC_GROUP_MAPPINGS, deliberately not auto-migrated: an
OIDC-sourced membership is dropped at a user's next sign-in until its
team's owner re-sets the group. The README's OIDC section spells out
the migration and the risk of a visible access gap during it.

TERDUT_OIDC_ADMIN_GROUP and TERDUT_OIDC_ALLOWED_GROUPS are untouched --
only team membership moved. terdut-tui needs no change: it only reads
GET /api/teams and GET /api/teams/{id}/members, and neither response
shape moved.
2026-09-27 11:43:57 +02:00
Niklas Ye 97a4814c04 Set the chart's placeholder version to 0.29.1
CI / chart (push) Successful in 2s
CI / test (push) Successful in 12s
CI / security (push) Successful in 15s
Release / test (push) Successful in 7s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 22s
Release / image (push) Successful in 1m3s
Release / scan-image (push) Successful in 5s
Cosmetic: make helm-package sets the published version and appVersion
from the tag, so these two fields decide nothing (see the comment
above them). Kept in step anyway, same as 155f27c and c5be55d, so a
tree heading for v0.29.1 doesn't say 0.29.0.
v0.29.1
2026-09-26 22:05:33 +02:00
Niklas Ye a2dc9e3b03 Ship a CA bundle in the image so single sign-on can reach the provider
The image is built FROM scratch and carried only the binary, so it had no
trust store, and every HTTPS call failed with "x509: certificate signed by
unknown authority". Nothing needed one until v0.29.0: OIDC discovery and the
token exchange are HTTPS calls to the identity provider, and the first
sign-in against Authentik died in discovery. The tests could not see it,
because they run on the host, whose trust store is fine.

The builder's ca-certificates.crt is copied in by name, so a missing file
fails the build instead of shipping an image that cannot sign anybody in.
Verified by fetching the provider's discovery URL from a scratch image with
and without the bundle: the same x509 error, then 200.

Password login and everything that talks only to Postgres were unaffected.
2026-09-26 22:05:33 +02:00