terdut-demo in Ryuvia/charts, a TerdutServer CR reconciled by terdut-operator,
and the kind demo in terdut-operator both pin this image, and chart-bump only
moves the terdut-server wrapper. terdut-demo therefore sat at v0.37.0 through
v0.41.0-v0.43.0 before it was synced on 2026-10-08, and nothing said it was
behind. CLAUDE.md now lists the step after the release one: bump the demo's
tag to the same digest and its chart version, as its own PR, and read the
range's migrations first, since the demo's database migrates forward at start.
Documentation only: nothing is built from this file, so it needs no release.
gosec's G202 flagged the string concatenation in handleListClusters, and the
security job gates CI. The pieces joined were only placeholders and fixed
clauses, never request data, so this was not an injection; but every other
handler here builds its SQL with fmt.Sprintf over placeholders, and this one
now reads the same way. The query and its result are unchanged.
Cosmetic: make helm-package passes --version and --app-version from the
tag, so these two fields decide nothing about what gets published. Still
done, as with db474ca (0.42.1) and def0f68 (0.42.0) before it, because a
tree heading for v0.43.0 that still says 0.42.1 tells its reader
something false.
A team with a cluster per alert source can now narrow the queue to one
cluster. A dropdown under the status chips lists the clusters, and appears
once there are two or more to choose between, the rule the team selector
follows. The choice is kept per browser, like the selected team, and dropped
if the server no longer knows the cluster rather than leaving an empty list
with no explanation. The status chip counts follow it, and a row leaves out
its cluster chip once the queue is narrowed to one.
The filter is on the server. The list is capped at 50 rows, so filtering what
is on screen would quietly miss older incidents in the Resolved and Archived
lists. GET /api/incidents takes ?cluster=<value>, matched against the
incident's `cluster` group label, and GET /api/incidents/clusters lists the
distinct values from the last 90 days (optionally for one team), so the
dropdown is not limited to what the current page happens to show. Both are
additive: no existing parameter or JSON shape changed, so terdut-tui keeps
working unchanged and has nothing it must mirror.
An incident carries the label only when `cluster` is in Alertmanager's
group_by, so the filter only sees those; the README says so.
Cosmetic: make helm-package passes --version and --app-version from the
tag, so these two fields decide nothing about what gets published. Still
done, as with def0f68 (0.42.0) and b848143 (0.41.3) before it, because a
tree heading for v0.42.1 that still says 0.42.0 tells its reader
something false.
A phone's lock screen cuts a long title off at the end, and an incident title
keeps its grouping labels there: "PodRestarting (cluster=prod-eu,
namespace=shop)". For a team with a Kubernetes cluster per alert source, the
cluster is the first thing a person wants and the first thing lost.
When the incident has a `cluster` group label, the title of the page now
leads with it, "[prod-eu] PodRestarting (namespace=shop)", and the label is
dropped from the parenthesis so it is not said twice, along with the
parenthesis itself if nothing else is left. The reminder and the resolution
use the same title, and the label is the one the web UI's chip reads. The
message body is unchanged.
An incident without a cluster label gets the same title as before, and
incidents that are already open keep theirs; only pages sent from now on
change. The constant originLabel names the label, as ORIGIN_LABEL does in the
web UI.
No endpoint or JSON shape changed, so nothing to mirror in terdut-tui.
Cosmetic: make helm-package passes --version and --app-version from the
tag, so these two fields decide nothing about what gets published. Still
done, as with b848143 (0.41.3) and 7665e5e (0.41.2) before it, because a
tree heading for v0.42.0 that still says 0.41.3 tells its reader
something false.
A team with one Alertmanager per Kubernetes cluster could not tell at a
glance where an incident started: the cluster was only a word inside the
title. When an alert carries a `cluster` label, the queue rows, the incident
page and the alerts list now show it as a chip in that cluster's colour, a
stable pick from the existing six-colour palette. A queue row drops
`cluster=...` from its title, since the chip says it, and the incident page
keeps the full title.
An incident has the label only when it is in Alertmanager's group_by, which
is also what keeps two clusters' identical alerts apart: incidents are matched
on the team and the groupKey, and the groupKey does not include external
labels. The README has a section on the two settings (Prometheus
externalLabels and group_by). The alerts list reads the label from the alert
itself, so it shows the chip with only the external label.
Web UI and docs only: no endpoint or JSON shape changed, so nothing to mirror
in terdut-tui. Nothing changes for a team whose alerts have no cluster label.
Cosmetic: make helm-package passes --version and --app-version from the
tag, so these two fields decide nothing about what gets published. Still
done, as with 7665e5e (0.41.2) and 3a96b20 (0.41.1) before it, because a
tree heading for v0.41.3 that still says 0.41.2 tells its reader
something false.
The fade that hints at more filters to the right was an element inside the
scrolling row, so it scrolled away with the chips instead of staying at the
edge; the comment on it claimed the opposite. It is now a mask on the strip
itself, applied only while there is more to scroll to (fadeOnOverflow in
ui.js sets data-more), so it stays put and disappears at the end instead of
dimming the last chip.
The Team and Admin tab strips had an always-on mask from the earlier
"Sources is cut off" fix, which dimmed their last tab even when fully
scrolled; they use the same mechanism now.
Web UI only: no endpoint or JSON shape changed, so nothing to mirror in
terdut-tui.
Cosmetic: make helm-package passes --version and --app-version from the
tag, so these two fields decide nothing about what gets published. Still
done, as with 3a96b20 (0.41.1) and 065b557 (0.41.0) before it, because a
tree heading for v0.41.2 that still says 0.41.1 tells its reader
something false.
The desktop sidebar is now grouped: Queue, On-call, Alerts and Stats; Team
and Admin; then Account at the foot, shown as the signed-in person with an
avatar and their name. The active item has a bar as well as a tint. The
queue's row of team chips is gone: the team selector is the one place the
team is chosen, and the chips offered the same choice a second time.
Tables get 16px between columns, so a right-aligned count no longer touches
the text beside it. An empty SSO group reads "Not configured" instead of a
dash, and the main action on each of those pages (New source, New switch,
Add member, Assign, and the admin Add, Create invite and Create) is a filled
button, with Edit and Cancel staying secondary. The bare triangles on
"Grouped by" and the label lists are real disclosure buttons: a chevron
that turns and a 44px target on a touch screen.
On a phone the More tab was a <button> that kept the browser's grey box, so
it looked highlighted next to four plain links. It now matches them, the
labels are 12px and the open tab's icon is filled.
Web UI only: no endpoint or JSON shape changed, so nothing to mirror in
terdut-tui.
Cosmetic: make helm-package passes --version and --app-version from the
tag, so these two fields decide nothing about what gets published. Still
done, as with 065b557 (0.41.0) and ead5df1 (0.40.0) before it, because a
tree heading for v0.41.1 that still says 0.41.0 tells its reader
something false.
The rota put an initial on every day, so one person covering a week was
seven identical chips and a legend to decode them. A shift is now one bar
across its days with the person's name on it; a flat end with a chevron
means it carries on across the row break. Days with nobody on call are a
hatched amber bar while they can still be fixed and a quiet grey one once
they are history, and the footer is a green banner when the month is
covered or an amber one when it is not. Tapping a day still opens it: the
bars ignore the pointer, so the tap reaches the cell underneath. The month
title was monospace because its button borrowed .label, which belongs to
label chips; it has its own class now, and a Today button sits beside the
arrows.
The escalation card was a five-column table. It is now the ladder it
describes: a numbered step per level with its status and who it pages, the
wait before the next level on its own line, and what happens after the last
one at the bottom. With no fallback topic that last step is an amber
callout, with an Add fallback button for owners that opens the editor on the
field. That gap is what incident #36 hit.
Web UI only: no endpoint or JSON shape changed, so nothing to mirror in
terdut-tui.
The Team settings cards (Escalation, Sources, Members, Switches, Single
sign-on) printed their text flush against the border with the button in the
corner: .card has never had padding and these never added any. A card with
a header row now pads itself, with the title left, the button right and a
divider before the content.
The incident page no longer carries the primary action twice. The copy up
by the status is gone and the sticky bar keeps it; Note is the timeline's
link, with Copy in the bar for a resolved incident, and the phone's More
sheet drops what the bar already shows. In the queue, a row's title wraps
to two lines so the namespace that tells rows apart is no longer cut off,
zero counts on the filter chips are dimmed, and a row omits the status the
filter already states and the team once the queue is narrowed to one.
Status colours failed 4.5:1 against their own fill in the light theme
(warning 3.6, info 4.1, critical 4.4, ok 4.45, snooze 4.49), so the light
tokens are darker; the dark theme already passed and is unchanged. Severity
badges now carry a shape as well as a colour, and each timeline event has an
icon. An escalation that ran out of levels, a failed notification and a
silent heartbeat stand out in amber or red.
Web UI only: no endpoint or JSON shape changed, so nothing to mirror in
terdut-tui.
Cosmetic: make helm-package passes --version and --app-version from the
tag, so these two fields decide nothing about what gets published. Still
done, as with ead5df1 (0.40.0) and 0f88574 (0.39.0) before it, because a
tree heading for v0.41.0 that still says 0.40.0 tells its reader
something false.
Secondary text was too dim to read: --faint sat at about 4.0:1 in the dark
theme and 3.3:1 in the light one, and it carries row ages, hints and labels.
It now clears 4.5:1 in both. The dark surfaces and borders are a step
further apart so cards stand out from the page, and the light borders a
touch stronger.
Account has an Appearance section with System, Light and Dark. System is
the old behaviour. The choice is per browser, kept in localStorage, and is
applied by a small js/theme.js loaded from <head> so there is no flash of
the other theme; the CSP allows no inline script, hence a file of its own.
On-call is redesigned: a hero card for who is on call now, with when the
shift ends for the selected team, and the week as seven day cells instead of
grouped rows. Today is marked with a neutral tint and a bar rather than the
accent colour, which now means only things you can act on; your own days are
marked by the "you" badge, not a fill. On a wide screen the hero and week sit
beside your shifts under a page title. The page stays read-only: shifts are
still edited in the TUI.
Smaller fixes: the team switcher no longer appears in the phone's bottom bar
as well as the top bar (a shared rule overrode the one that hides it); the
Team and Admin tab strips fade at the edge and scroll the open section into
view, so "Sources" is no longer cut off; and "All clear" is a green status
badge with a check instead of a grey pill that looked like a button.
Web UI only: no endpoint or JSON shape changed, so nothing to mirror in
terdut-tui.
Cosmetic: make helm-package passes --version and --app-version from the
tag, so these two fields decide nothing about what gets published. Still
done, as with 0f88574 (0.39.0) and 584d344 (0.38.0) before it, because a
tree heading for v0.40.0 that still says 0.39.0 tells its reader
something false.
The three team lists carried bare buttons on every row (Remove, Rename,
Revoke, Edit), which crowds a row that is meant for scanning and puts a
destructive action one stray click from every entry. A row now opens a
sheet with the facts the row has no room for, and the actions live
there: Edit and Delete for a switch, Rename and Revoke for a source,
Edit and Remove for a member. The guards are unchanged: owners only, the
last owner and SSO-managed members still cannot be removed, and every
destructive action still asks first.
Switches can finally be edited in place. The PUT endpoint has existed
since the operator needed it; the UI simply never called it, so changing
a timeout meant deleting the switch and losing its history. The edit
form is the add form, prefilled.
The first column of all three lists is now headed Name. Web UI only: no
change to any endpoint or JSON shape, so nothing to mirror in the TUI.
Cosmetic: make helm-package passes --version and --app-version from the
tag, so these two fields decide nothing about what gets published. Still
done, as with 584d344 (0.38.0) and df83adf (0.37.2) before it, because a
tree heading for v0.39.0 that still says 0.38.0 tells its reader
something false.
Assign logged only the assignee (user_id), archive/unarchive logged nothing.
Migration 018 adds actor_user_id/actor_service_account_id to incident_events
for 'assigned'; archive/unarchive now log archived/unarchived events with the
caller via callerActorIDs. Timeline JSON gains actor_* fields; web timeline
renders them. Service accounts are still not assignable.
Cosmetic: make helm-package passes --version and --app-version from the
tag, so these two fields decide nothing about what gets published. Still
done, as with df83adf (0.37.2) and 2b2609e (0.37.1) before it, because a
tree heading for v0.38.0 that still says 0.37.2 tells its reader
something false.
Part of the same security-hardening pass as the last five commits, and
the last item in its backlog. User API keys had no expiry at all --
unlike service-account keys, visibly distinct only by their "tdsa_"
prefix -- and, it turns out while implementing this, no way to list
them either: only create (returns the raw key once) and delete-by-id
existed, so a key's owner had no way to even discover what keys they
had short of remembering IDs from creation time.
handleCreateAPIKey takes an optional expires_in_days (0, the default,
keeps today's behavior: never expires, so no existing integration is
affected). apiKeyUser's lookup now carries `expires_at IS NULL OR
expires_at > now` as part of the query itself, the same way serveAs's
disabled_at check already works -- an expired key simply fails to
resolve, like a wrong one, rather than resolving and being caught
after the fact. New GET /api/users/{id}/api-keys (requireSelfOrAdmin,
same as create/delete) lists id/name/created_at/last_used_at/expires_at,
never the raw key.
Scoped down from the original plan on request: no web UI change, since
there turned out to be no existing API-keys UI at all to extend --
building one from scratch would have been a real feature addition, not
a hardening tweak.
Mirrored the additive expires_at field in terdut-tui's APIKey struct
(separate commit, separate repo) per this workspace's version-coupling
rule; the TUI does not create or list expiring keys itself yet.
New tests (api_keys_test.go): default never-expires, expires_in_days
sets expires_at, out-of-range values rejected, an expired key fails
auth after a fresh one worked, the listing never includes the raw key.
Also added the new GET route to authz_scope_test.go's self-or-admin
table from the previous commit.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude <noreply@anthropic.com>
Part of the same security-hardening pass as the last four commits.
terdut-server's authz is already solid -- centralized predicates
(requireTeamMember, requireTeamOwner, requireSelfOrAdmin, AdminOnly)
rather than ad hoc per-handler checks, confirmed by spot-checking several
handlers while writing this. But it is enforced by convention, not the
type system: a future handler that forgets its guard would compile and
read fine on review, exactly like one that remembers it.
authz_scope_test.go builds two teams and, for every team-scoped route
(members, OIDC groups, invites, escalation, dead man's switches,
integrations, schedule, plus every /api/incidents/{id}/... route, scoped
by the incident's own team_id through incidentIDParam's single
chokepoint), calls it as one team's owner against the other team's
resources and asserts 404 -- requireTeamMember and requireTeamOwner both
answer a non-member that way. Separate tests cover AdminOnly's routes
(403 for a non-admin) and requireSelfOrAdmin's (403 for a non-admin
acting on someone else's account).
Verified the test actually catches a regression, not just that it
passes today: temporarily removed handleListTeamMembers' requireTeamMember
call, confirmed exactly that one subtest failed and nothing else did,
then put it back.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude <noreply@anthropic.com>
Part of the same security-hardening pass as the last three commits.
Neither the Dockerfile nor the chart's Deployment set any securityContext
at all, so the container ran as root by default — scratch has no
/etc/passwd for a USER directive to resolve against, so nobody had set one.
Dockerfile now ends with USER 65532:65532 (numeric, since scratch has no
user database; 65532 is the common "nonroot" convention, distroless's own
uid). The chart's Deployment adds a matching pod-level securityContext
(runAsNonRoot, runAsUser/runAsGroup: 65532, seccompProfile: RuntimeDefault)
plus per-container hardening (allowPrivilegeEscalation: false,
capabilities dropped, readOnlyRootFilesystem: true) on both the app
container and the wait-for-postgres init container — neither writes
anything to disk, so the root filesystem can stay read-only.
Verified with helm-lint and a manual `helm template` render of both the
terdut-server and terdut-demo charts. Not yet verified: an actual pod
starting with these in place — readOnlyRootFilesystem is exactly where a
non-obvious write (a temp file, a cache dir) would surface as a crash
rather than a lint error, so that needs a real rollout to confirm.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude <noreply@anthropic.com>
Part of the same security-hardening pass as the last two commits. make
lint was go vet only; govulncheck and gitleaks already scanned deps and
secrets on every push, but nothing read this repo's own source for
risky patterns (weak crypto, injection shapes, insecure cookies, ...).
New `make security-code` runs gosec, wired into ci.yaml's security job
alongside the other two. G104 (unchecked error) is excluded at the
Makefile level: every one of its 41 initial hits was this codebase's
existing, deliberate idiom for a best-effort write or an already-
reviewed json.Unmarshal of its own JSONB, predating gosec, and the rule
cannot tell that apart from a mistake -- seventeen individual #nosec
comments would hide a future real G104 regression in the suppression
noise rather than surface it. Reasoning is on the Makefile target.
Of the 12 remaining hits:
- Genuinely real: oidc.go's callback logged error_description (and,
two call sites down, identity.Subject) via %s before the request's
state was even checked against its cookie -- an attacker-reachable
value going into the log unquoted. Switched to %q, matching
identity.Username's existing treatment, so a value holding a
newline can't forge a second log line.
- False positives, annotated inline rather than globally suppressed:
4x G124 on cookies that already set Secure via cookieSecure(...)
(a function call, not the literal `true` the rule wants), 3x G202
on sqlArgs-built queries that only ever splice in a "$N"
placeholder, never a value, and the remaining 5x G706 on log lines
that were already %q-quoted -- gosec's taint analysis doesn't
model format verbs, so it flags the tainted argument regardless.
Also fixed handleMe's swallowed Scan error (gosec's catch, pre-fix):
a transient DB error left hash/dismissed at their zero values and the
response claimed no password and no onboarding dismissal regardless
of the truth, rather than surfacing a 500.
Checked both workflow files for the injection class letsvisit found
there (a `${{ }}` expression spliced straight into a `run:` block):
every one here already goes through `env:` as a quoted shell variable,
documented in ci.yaml's own header comment. Nothing to fix.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude <noreply@anthropic.com>
Part of the same security-hardening pass as the body-size/header commit.
loginLimiter was an in-memory, per-process sync.Mutex+map -- fine for one
replica, but charts/terdut-server/values.yaml has set replicaCount: 2 in
production since v0.37.0. Each pod counted only its own traffic, so every
limit it guarded (failed logins, sign-ups, OIDC/device-login starts) was
effectively twice as generous as the constants say, not just in theory.
loginLimiter now stores its counters in a new rate_limit_counters table
(migration 016) instead of a map; blocked/fail/clear take a context and
query/upsert/delete a row keyed by the same strings callers already used
(username, client address, "signup:"+address, ...). Semantics are
unchanged -- a fixed window that resets rather than slides -- so no call
site's behavior changes, only where the count lives. Added purgeRateLimits
to the sweeper, alongside purgeSessions/purgeAckTokens, so expired windows
don't accumulate.
New internal (package api) tests in rate_limiter_test.go cover the basic
behavior plus the regression this exists to fix: two loginLimiter values
sharing one database, standing in for two replicas, now see one combined
count instead of each keeping their own.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude <noreply@anthropic.com>
Part of a security-hardening pass (see wiki for the full backlog).
decodeJSON had no size limit at all, so every JSON endpoint -- including
the two unauthenticated ones (bootstrap, the Alertmanager webhook) --
would buffer an attacker-supplied body of unbounded size before it was
even validated. decodeJSON now wraps the body in http.MaxBytesReader at
a 1 MiB default; the webhook gets its own 8 MiB cap via decodeJSONLimit,
since a real Alertmanager batch can be bigger than an ordinary API body.
Also adds a securityHeaders middleware, applied globally: nosniff on
every response (previously only the static site got it), and HSTS
(180-day max-age, conservative on purpose) whenever cookieSecure's
signal says the browser is on HTTPS. Checked the chart/gateway config
first -- neither sets HSTS anywhere, so this was a real gap.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude <noreply@anthropic.com>
Every incident-mutation handler read userFromContext(ctx) and wrote the
result's .ID into acknowledged_by/incident_events.user_id without checking
the ok bool. For a team-scoped service-account caller this returned a
zero-value user id, which violated the users(id) FK and 500'd on
acknowledge, unacknowledge, resolve, snooze, unsnooze and create-note.
handleDeleteNote didn't crash but silently matched zero rows instead
(WHERE user_id = 0), so a service account could never delete its own note.
Add acknowledged_by_service_account_id (incidents) and service_account_id
(incident_events) as nullable FKs to service_accounts(id), parallel to and
mutually exclusive with the existing human columns (migration 015, with a
CHECK enforcing the exclusion). Route every one of the six handlers plus
delete-note through a new callerActorIDs() helper that branches on
Caller.AsHuman()/ServiceAccountID() instead of assuming a human, and thread
a serviceAccountID parameter through logEvent and the new
acknowledgeIncidentAs (acknowledgeIncident itself is untouched: its only
other caller, the push-notification Acknowledge button, is always human).
Render the new actor distinctly from both a human and "the server acted"
in the web UI's incident timeline and facts card.
handleIncidentAssign, handleIncidentArchive and handleIncidentUnarchive are
deliberately not touched here — they track no actor at all today, for
anyone, which is a separate pre-existing gap (follow-up issue to come).
Fixes#25
Cosmetic: `make helm-package` passes --version and --app-version from
the tag, so these fields decide nothing about what gets published. But
a tree heading for v0.37.1 that still says 0.37.0 tells a reader
something false. Same as 7efd1bb, which cites 5c4e0bd.
On-call, the incident queue and the account page all had latent CSS bugs
that only show up once the browser is wide enough to hit the desktop
breakpoint (900px+):
- On-call: .days reset its own margin to 0, which canceled the page-wide
auto-centering on just that element, leaving the day list pinned to the
left edge while every other card on the page centered normally.
- Queue: the list pane stayed capped at 340-420px even with nothing
selected, leaving the rest of the screen empty. It now fills the width
until an incident is picked, then goes back to list+detail.
- Account, admin user and admin team: .btn and .back-link are inline-flex,
and margin:auto only centers a block box, so the Sign-out button and the
two admin back-links sat left of their sibling cards instead of matching
their width. Wrapped each in a block div.
The incident detail action bar also folded Assign, Add note, Resolve and
Clear acknowledgement into a "More" sheet sized for a phone's width.
Desktop has the room, so it now shows them as direct buttons and hides
More instead; Copy incident stays out of the bar since the header already
has its own button for it.
Filed as niklas/terdut-server#31, #32, #33, #34, each with a screenshot.
replicas and strategy: Recreate were the chart's only guard against the
archiver, notifier and migration races; v0.36.0 closed all three with
advisory locks and a conflict-resolving incident insert, which made
that guard redundant rather than load-bearing. Expose replicaCount
(new, no values.yaml key existed before) defaulting to 2, and switch
to strategy: RollingUpdate with no explicit maxUnavailable/maxSurge --
the 25%/25% default rounds to 0/1 at replicaCount: 2, which is already
zero-downtime.
The chart does not gate this on image.tag, so pointing it at a
pre-v0.36.0 image with the new default is a foot-gun by omission --
noted in both the values.yaml comment and the deployment.yaml comment,
not guarded in code, same as the chart does for every other
version-coupled assumption today.
Co-authored-by: Claude <noreply@anthropic.com>
openIncident's INSERT had no ON CONFLICT clause, relying entirely on
incidentForGroup's earlier SELECT to avoid a duplicate. On more than
one replica, two webhook deliveries for the very first occurrence of
a brand-new groupKey can both pass that SELECT before either INSERTs;
the loser then hit incidents_open_group_key_idx's unique violation,
which rolled back its whole transaction — including that payload's
alert upserts, done earlier in the same transaction. ingest's error is
only logged and receiveWebhook answers 200 regardless, so nothing
retried it: the loser's alerts silently never existed.
Add ON CONFLICT (team_id, group_key) WHERE resolved_at IS NULL DO
NOTHING to the INSERT, matching the partial unique index. Postgres
only resolves that conflict after the winning transaction commits (or
rolls back), so by the time RETURNING comes back empty,
existingOpenIncident's follow-up SELECT is guaranteed to see the
winner's row. The loser attaches to it instead of failing outright,
and the rest of its payload commits normally. Covers both callers,
since the dead man's switch sweeper shares this same function.
New test (package api_test, fires N webhook deliveries for one
groupKey from a synchronized start with distinct fingerprints, so
they aren't accidentally serialized by upsertAlerts' own per-
fingerprint lock) confirmed meaningful: with the ON CONFLICT clause
reverted, it fails 10/10 on a missing alert fingerprint; restored, 0/10.
Note while building it: "exactly one incident" alone cannot
distinguish fixed from broken, since the DB's own unique index already
guarantees that either way — the real signal is the loser's payload
surviving.
Chart comment updated: all three of the chart's original reasons for
Recreate are now addressed in code, though replicas stays at 1 and the
strategy stays Recreate pending a deliberate decision to raise it.
Co-authored-by: Claude <noreply@anthropic.com>
Migrate's check-then-apply loop against schema_migrations had no
locking: two replicas booting at once against a fresh or
partially-migrated database could both pass the "not yet applied"
check for the same file and race applying it, crashing whichever lost
the duplicate-key insert (confirmed: reverting the lock fails the new
test 10/10 on a duplicate-key violation, racing as early as the
CREATE TABLE IF NOT EXISTS schema_migrations statement itself).
Hold a Postgres advisory lock for Migrate's whole run, on a dedicated
connection reserved via db.Conn so lock and unlock happen on the same
session. Blocking (pg_advisory_lock), unlike the archiver/notifier's
pg_try_advisory_lock: on boot there's no later tick to defer to, so a
second replica should wait for the first to finish migrating rather
than skip ahead.
Adds internal/db's first test file, exercising two concurrent Migrate
calls against a fresh schema.
Still open: the new-incident-insert race on a webhook for a brand-new
groupKey, noted in the chart's updated comment. Login rate limiting
staying in-process, diluted across replicas, is an accepted tradeoff.
Co-authored-by: Claude <noreply@anthropic.com>
Both background loops run unconditionally on every instance with no
coordination between them, which the chart's replicas: 1 + strategy:
Recreate exists specifically to paper over: with more than one replica,
every one of them would sweep and deliver notifications independently,
and two overlapping during a rollout would both page for the same
incident.
Add withAdvisoryLock, which takes a Postgres advisory lock on a
dedicated connection and runs a pass only if it gets the lock,
otherwise skipping until the next tick. Wire StartArchiver and
StartNotifier through it with their own lock keys, so Sweep and
NotifySweep themselves are untouched and every existing test calling
them directly keeps working unchanged.
This also closes the notifier's double-delivery race in passing: two
replicas can no longer both be inside deliverPending at once, since
only one can hold notifierLockKey at a time.
Deliberately not addressed here, and still blocking a replica count
above 1: the in-memory login rate limiter, the unlocked migration
runner, and the new-incident-insert race on a webhook for a brand-new
groupKey. Noted in the updated chart comment.
Co-authored-by: Claude <noreply@anthropic.com>
Cosmetic, as before (497086c, 4358e84): make helm-package passes
--version and --app-version from the tag, so this decides nothing
about what gets published. Kept in step so the tree heading for
v0.35.1 doesn't say 0.35.0 to a reader who hasn't yet seen the tag.
🤖 Generated with [Claude Code](https://claude.ai/code)
Co-Authored-By: Claude <noreply@anthropic.com>
position: sticky, as a flex item of the row it's pinning itself
against, interacted with that row's gap and its own negative margin in
a way that landed it short of the true edge -- visibly, a sliver of
the next chip stayed poking out past where the fade should have
covered it, which is the "ends before the screen edge" bug reported
against the release.
Replaced with the simpler, better-supported pattern for this: an
absolutely positioned overlay against a position:relative, overflow:
auto parent. Unlike a sticky descendant, an absolutely positioned one
is resolved against the parent's own (non-scrolling) box, so it stays
flush with the real edge regardless of scroll position, without the
flex-gap/margin interaction that caused this.
Verified visually: a standalone reproduction of both versions,
screenshotted with Chromium's headless_shell (no browser automation
tool available in this environment, but the binary's right there) --
the old version shows the next chip's edge peeking past the fade, the
new one doesn't.
🤖 Generated with [Claude Code](https://claude.ai/code)
Co-Authored-By: Claude <noreply@anthropic.com>
.nav-link-secondary's display:none sat before .nav-link's own
display:flex in the file. Both are single-class selectors, so they tie
on specificity, and a tie is broken by which one comes later in the
file -- not by which class the element happens to carry. .nav-link's
declaration, being later, won for every element wearing both classes,
so Stats/Admin/Account rendered as three extra tabs on the phone bar
instead of folding into "More" as intended.
Moved the rule below .nav-link instead of changing either declaration,
since nothing about the values was wrong -- only their order was.
🤖 Generated with [Claude Code](https://claude.ai/code)
Co-Authored-By: Claude <noreply@anthropic.com>
Cosmetic: make helm-package passes --version and --app-version from the
tag, so these two fields decide nothing about what gets published. Kept
in step anyway, the same as 4358e84 and fd26fef before it, so the tree
heading for v0.35.0 does not say 0.34.0 to a reader who hasn't yet seen
the tag.
🤖 Generated with [Claude Code](https://claude.ai/code)
Co-Authored-By: Claude <noreply@anthropic.com>
Addresses the screenshot-review feedback in #26. No framework or build
step added — all of this stays within the existing plain HTML/CSS/
vanilla-JS + go:embed architecture.
- Nav: re-enable the bottom tab bar that was already built and
switched off (Queue/On-call/Alerts/Team + a "More" sheet for
Stats/Admin/Account), replacing the hamburger on phone width.
- Queue: chip counts, a scroll fade on the filter row, a
"Triggered Xh ago" + severity label per row, a chevron on the team
switcher so it reads as a dropdown.
- On-call: collapse repeated same-person days into shift bars (week
view and "your shifts" both), show the week as a date range with
the ISO week number as secondary text, split "Current shift" out
from "Next shifts" with "ends in Nd", a pill badge + row highlight
for "you".
- Incident detail: fix the actual bug behind the duplicate
"acknowledged" timeline entries (acknowledgeIncident's UPDATE had no
guard on the incident's current status, so acknowledging an
already-acknowledged incident silently re-logged the event — now
idempotent, with regression tests on both the authenticated route
and the ntfy ack-button route). Relabel escalation re-pages so they
don't look like the same page landing twice. Copy the primary action
up near the top. Label the "···" button. Group the timeline by
phase (triggered/acknowledged/resolved). Add an "at a glance"
summary row (duration/severity/responsible) and collapse the group
labels by default.
- Team overview: reword the vague copy ("One owner." etc.) into plain
labels.
- Empty states: fill in missing icons/one-liners across queue,
alerts, stats and the incident timeline.
- CSS: fix card padding bugs, verify link contrast already passes AA,
introduce a --fs-* type-scale token set and migrate the few
genuinely isolated cases onto it (left sizes tied to a fixed shape,
a deliberately prominent display, or a non-negotiable constraint
like the iOS-zoom-prevention input size as documented exceptions
rather than guess at a render this change can't see).
Verified with the full fmt/lint/test/helm-lint gate, plus a live
instance against the test DB with seeded incidents and schedule data
to trace the on-call grouping and timeline phase-splitting logic
against real API responses.
🤖 Generated with [Claude Code](https://claude.ai/code)
Co-Authored-By: Claude <noreply@anthropic.com>