1cb09525e39e11d6c2c77225e41c79307c8b3f4d
84 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
c4833067f0 |
Build the clusters query with Sprintf, as the other handlers do
CI / chart (push) Successful in 1s
CI / security (push) Successful in 20s
CI / test (push) Successful in 5m15s
Release / test (push) Successful in 7s
Release / chart (push) Successful in 4s
Release / binaries (push) Successful in 24s
Release / image (push) Successful in 1m12s
Release / scan-image (push) Successful in 24s
gosec's G202 flagged the string concatenation in handleListClusters, and the security job gates CI. The pieces joined were only placeholders and fixed clauses, never request data, so this was not an injection; but every other handler here builds its SQL with fmt.Sprintf over placeholders, and this one now reads the same way. The query and its result are unchanged. |
||
|
|
a955356821 |
Filter the queue by cluster
A team with a cluster per alert source can now narrow the queue to one cluster. A dropdown under the status chips lists the clusters, and appears once there are two or more to choose between, the rule the team selector follows. The choice is kept per browser, like the selected team, and dropped if the server no longer knows the cluster rather than leaving an empty list with no explanation. The status chip counts follow it, and a row leaves out its cluster chip once the queue is narrowed to one. The filter is on the server. The list is capped at 50 rows, so filtering what is on screen would quietly miss older incidents in the Resolved and Archived lists. GET /api/incidents takes ?cluster=<value>, matched against the incident's `cluster` group label, and GET /api/incidents/clusters lists the distinct values from the last 90 days (optionally for one team), so the dropdown is not limited to what the current page happens to show. Both are additive: no existing parameter or JSON shape changed, so terdut-tui keeps working unchanged and has nothing it must mirror. An incident carries the label only when `cluster` is in Alertmanager's group_by, so the filter only sees those; the README says so. |
||
|
|
a23e88c16d |
Put the cluster first in an ntfy page title
A phone's lock screen cuts a long title off at the end, and an incident title keeps its grouping labels there: "PodRestarting (cluster=prod-eu, namespace=shop)". For a team with a Kubernetes cluster per alert source, the cluster is the first thing a person wants and the first thing lost. When the incident has a `cluster` group label, the title of the page now leads with it, "[prod-eu] PodRestarting (namespace=shop)", and the label is dropped from the parenthesis so it is not said twice, along with the parenthesis itself if nothing else is left. The reminder and the resolution use the same title, and the label is the one the web UI's chip reads. The message body is unchanged. An incident without a cluster label gets the same title as before, and incidents that are already open keep theirs; only pages sent from now on change. The constant originLabel names the label, as ORIGIN_LABEL does in the web UI. No endpoint or JSON shape changed, so nothing to mirror in terdut-tui. |
||
|
|
fb86a18988 |
Show which cluster an incident came from, as a chip
A team with one Alertmanager per Kubernetes cluster could not tell at a glance where an incident started: the cluster was only a word inside the title. When an alert carries a `cluster` label, the queue rows, the incident page and the alerts list now show it as a chip in that cluster's colour, a stable pick from the existing six-colour palette. A queue row drops `cluster=...` from its title, since the chip says it, and the incident page keeps the full title. An incident has the label only when it is in Alertmanager's group_by, which is also what keeps two clusters' identical alerts apart: incidents are matched on the team and the groupKey, and the groupKey does not include external labels. The README has a section on the two settings (Prometheus externalLabels and group_by). The alerts list reads the label from the alert itself, so it shows the chip with only the external label. Web UI and docs only: no endpoint or JSON shape changed, so nothing to mirror in terdut-tui. Nothing changes for a team whose alerts have no cluster label. |
||
|
|
942517c7a8 |
Keep the scroll fade at the screen edge on the queue's filter row
The fade that hints at more filters to the right was an element inside the scrolling row, so it scrolled away with the chips instead of staying at the edge; the comment on it claimed the opposite. It is now a mask on the strip itself, applied only while there is more to scroll to (fadeOnOverflow in ui.js sets data-more), so it stays put and disappears at the end instead of dimming the last chip. The Team and Admin tab strips had an always-on mask from the earlier "Sources is cut off" fix, which dimmed their last tab even when fully scrolled; they use the same mechanism now. Web UI only: no endpoint or JSON shape changed, so nothing to mirror in terdut-tui. |
||
|
|
bcf1a3e99b |
Give the sidebar groups, and tidy the tables, disclosures and tab bar
The desktop sidebar is now grouped: Queue, On-call, Alerts and Stats; Team and Admin; then Account at the foot, shown as the signed-in person with an avatar and their name. The active item has a bar as well as a tint. The queue's row of team chips is gone: the team selector is the one place the team is chosen, and the chips offered the same choice a second time. Tables get 16px between columns, so a right-aligned count no longer touches the text beside it. An empty SSO group reads "Not configured" instead of a dash, and the main action on each of those pages (New source, New switch, Add member, Assign, and the admin Add, Create invite and Create) is a filled button, with Edit and Cancel staying secondary. The bare triangles on "Grouped by" and the label lists are real disclosure buttons: a chevron that turns and a 44px target on a touch screen. On a phone the More tab was a <button> that kept the browser's grey box, so it looked highlighted next to four plain links. It now matches them, the labels are 12px and the open tab's icon is filled. Web UI only: no endpoint or JSON shape changed, so nothing to mirror in terdut-tui. |
||
|
|
7b9d309d13 |
Draw the rota as shift bars, and the escalation chain as a ladder
The rota put an initial on every day, so one person covering a week was seven identical chips and a legend to decode them. A shift is now one bar across its days with the person's name on it; a flat end with a chevron means it carries on across the row break. Days with nobody on call are a hatched amber bar while they can still be fixed and a quiet grey one once they are history, and the footer is a green banner when the month is covered or an amber one when it is not. Tapping a day still opens it: the bars ignore the pointer, so the tap reaches the cell underneath. The month title was monospace because its button borrowed .label, which belongs to label chips; it has its own class now, and a Today button sits beside the arrows. The escalation card was a five-column table. It is now the ladder it describes: a numbered step per level with its status and who it pages, the wait before the next level on its own line, and what happens after the last one at the bottom. With no fallback topic that last step is an amber callout, with an Add fallback button for owners that opens the editor on the field. That gap is what incident #36 hit. Web UI only: no endpoint or JSON shape changed, so nothing to mirror in terdut-tui. |
||
|
|
7412456c5a |
Pad the settings cards, and make status and timeline readable at a glance
The Team settings cards (Escalation, Sources, Members, Switches, Single sign-on) printed their text flush against the border with the button in the corner: .card has never had padding and these never added any. A card with a header row now pads itself, with the title left, the button right and a divider before the content. The incident page no longer carries the primary action twice. The copy up by the status is gone and the sticky bar keeps it; Note is the timeline's link, with Copy in the bar for a resolved incident, and the phone's More sheet drops what the bar already shows. In the queue, a row's title wraps to two lines so the namespace that tells rows apart is no longer cut off, zero counts on the filter chips are dimmed, and a row omits the status the filter already states and the team once the queue is narrowed to one. Status colours failed 4.5:1 against their own fill in the light theme (warning 3.6, info 4.1, critical 4.4, ok 4.45, snooze 4.49), so the light tokens are darker; the dark theme already passed and is unchanged. Severity badges now carry a shape as well as a colour, and each timeline event has an icon. An escalation that ran out of levels, a failed notification and a silent heartbeat stand out in amber or red. Web UI only: no endpoint or JSON shape changed, so nothing to mirror in terdut-tui. |
||
|
|
a8dc89e23d |
Make the web UI easier to read, and add a theme toggle
Secondary text was too dim to read: --faint sat at about 4.0:1 in the dark theme and 3.3:1 in the light one, and it carries row ages, hints and labels. It now clears 4.5:1 in both. The dark surfaces and borders are a step further apart so cards stand out from the page, and the light borders a touch stronger. Account has an Appearance section with System, Light and Dark. System is the old behaviour. The choice is per browser, kept in localStorage, and is applied by a small js/theme.js loaded from <head> so there is no flash of the other theme; the CSP allows no inline script, hence a file of its own. On-call is redesigned: a hero card for who is on call now, with when the shift ends for the selected team, and the week as seven day cells instead of grouped rows. Today is marked with a neutral tint and a bar rather than the accent colour, which now means only things you can act on; your own days are marked by the "you" badge, not a fill. On a wide screen the hero and week sit beside your shifts under a page title. The page stays read-only: shifts are still edited in the TUI. Smaller fixes: the team switcher no longer appears in the phone's bottom bar as well as the top bar (a shared rule overrode the one that hides it); the Team and Admin tab strips fade at the edge and scroll the open section into view, so "Sources" is no longer cut off; and "All clear" is a green status badge with a check instead of a grey pill that looked like a button. Web UI only: no endpoint or JSON shape changed, so nothing to mirror in terdut-tui. |
||
|
|
3ced069134 |
Open switches, sources and members in a details sheet
The three team lists carried bare buttons on every row (Remove, Rename, Revoke, Edit), which crowds a row that is meant for scanning and puts a destructive action one stray click from every entry. A row now opens a sheet with the facts the row has no room for, and the actions live there: Edit and Delete for a switch, Rename and Revoke for a source, Edit and Remove for a member. The guards are unchanged: owners only, the last owner and SSO-managed members still cannot be removed, and every destructive action still asks first. Switches can finally be edited in place. The PUT endpoint has existed since the operator needed it; the UI simply never called it, so changing a timeout meant deleting the switch and losing its history. The edit form is the add form, prefilled. The first column of all three lists is now headed Name. Web UI only: no change to any endpoint or JSON shape, so nothing to mirror in the TUI. |
||
|
|
3613fd5732 |
Check rows.Err() in the remaining Next() loops (#27)
Incident list, the three stats breakdowns and the user list could return a truncated result as if complete when the scan failed partway. |
||
|
|
dc62278788 | Check rows.Err() after the timeline scan loop (#27) | ||
|
|
0aaea8efb5 |
Record the actor on assign, archive and unarchive (#35)
Assign logged only the assignee (user_id), archive/unarchive logged nothing. Migration 018 adds actor_user_id/actor_service_account_id to incident_events for 'assigned'; archive/unarchive now log archived/unarchived events with the caller via callerActorIDs. Timeline JSON gains actor_* fields; web timeline renders them. Service accounts are still not assignable. |
||
|
|
926aa2d3ec |
Add optional API key expiry and a missing way to list them
Part of the same security-hardening pass as the last five commits, and
the last item in its backlog. User API keys had no expiry at all --
unlike service-account keys, visibly distinct only by their "tdsa_"
prefix -- and, it turns out while implementing this, no way to list
them either: only create (returns the raw key once) and delete-by-id
existed, so a key's owner had no way to even discover what keys they
had short of remembering IDs from creation time.
handleCreateAPIKey takes an optional expires_in_days (0, the default,
keeps today's behavior: never expires, so no existing integration is
affected). apiKeyUser's lookup now carries `expires_at IS NULL OR
expires_at > now` as part of the query itself, the same way serveAs's
disabled_at check already works -- an expired key simply fails to
resolve, like a wrong one, rather than resolving and being caught
after the fact. New GET /api/users/{id}/api-keys (requireSelfOrAdmin,
same as create/delete) lists id/name/created_at/last_used_at/expires_at,
never the raw key.
Scoped down from the original plan on request: no web UI change, since
there turned out to be no existing API-keys UI at all to extend --
building one from scratch would have been a real feature addition, not
a hardening tweak.
Mirrored the additive expires_at field in terdut-tui's APIKey struct
(separate commit, separate repo) per this workspace's version-coupling
rule; the TUI does not create or list expiring keys itself yet.
New tests (api_keys_test.go): default never-expires, expires_in_days
sets expires_at, out-of-range values rejected, an expired key fails
auth after a fresh one worked, the listing never includes the raw key.
Also added the new GET route to authz_scope_test.go's self-or-admin
table from the previous commit.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude <noreply@anthropic.com>
|
||
|
|
a2ca9c25d0 |
Add a regression test for the team/admin/self authorization pattern
Part of the same security-hardening pass as the last four commits.
terdut-server's authz is already solid -- centralized predicates
(requireTeamMember, requireTeamOwner, requireSelfOrAdmin, AdminOnly)
rather than ad hoc per-handler checks, confirmed by spot-checking several
handlers while writing this. But it is enforced by convention, not the
type system: a future handler that forgets its guard would compile and
read fine on review, exactly like one that remembers it.
authz_scope_test.go builds two teams and, for every team-scoped route
(members, OIDC groups, invites, escalation, dead man's switches,
integrations, schedule, plus every /api/incidents/{id}/... route, scoped
by the incident's own team_id through incidentIDParam's single
chokepoint), calls it as one team's owner against the other team's
resources and asserts 404 -- requireTeamMember and requireTeamOwner both
answer a non-member that way. Separate tests cover AdminOnly's routes
(403 for a non-admin) and requireSelfOrAdmin's (403 for a non-admin
acting on someone else's account).
Verified the test actually catches a regression, not just that it
passes today: temporarily removed handleListTeamMembers' requireTeamMember
call, confirmed exactly that one subtest failed and nothing else did,
then put it back.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude <noreply@anthropic.com>
|
||
|
|
f15db0e20a |
Add gosec to CI; fix the one real finding it surfaced
Part of the same security-hardening pass as the last two commits. make
lint was go vet only; govulncheck and gitleaks already scanned deps and
secrets on every push, but nothing read this repo's own source for
risky patterns (weak crypto, injection shapes, insecure cookies, ...).
New `make security-code` runs gosec, wired into ci.yaml's security job
alongside the other two. G104 (unchecked error) is excluded at the
Makefile level: every one of its 41 initial hits was this codebase's
existing, deliberate idiom for a best-effort write or an already-
reviewed json.Unmarshal of its own JSONB, predating gosec, and the rule
cannot tell that apart from a mistake -- seventeen individual #nosec
comments would hide a future real G104 regression in the suppression
noise rather than surface it. Reasoning is on the Makefile target.
Of the 12 remaining hits:
- Genuinely real: oidc.go's callback logged error_description (and,
two call sites down, identity.Subject) via %s before the request's
state was even checked against its cookie -- an attacker-reachable
value going into the log unquoted. Switched to %q, matching
identity.Username's existing treatment, so a value holding a
newline can't forge a second log line.
- False positives, annotated inline rather than globally suppressed:
4x G124 on cookies that already set Secure via cookieSecure(...)
(a function call, not the literal `true` the rule wants), 3x G202
on sqlArgs-built queries that only ever splice in a "$N"
placeholder, never a value, and the remaining 5x G706 on log lines
that were already %q-quoted -- gosec's taint analysis doesn't
model format verbs, so it flags the tainted argument regardless.
Also fixed handleMe's swallowed Scan error (gosec's catch, pre-fix):
a transient DB error left hash/dismissed at their zero values and the
response claimed no password and no onboarding dismissal regardless
of the truth, rather than surfacing a 500.
Checked both workflow files for the injection class letsvisit found
there (a `${{ }}` expression spliced straight into a `run:` block):
every one here already goes through `env:` as a quoted shell variable,
documented in ci.yaml's own header comment. Nothing to fix.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude <noreply@anthropic.com>
|
||
|
|
b82c10acf4 |
Back the login/signup/OIDC/device rate limiters with Postgres
Part of the same security-hardening pass as the body-size/header commit. loginLimiter was an in-memory, per-process sync.Mutex+map -- fine for one replica, but charts/terdut-server/values.yaml has set replicaCount: 2 in production since v0.37.0. Each pod counted only its own traffic, so every limit it guarded (failed logins, sign-ups, OIDC/device-login starts) was effectively twice as generous as the constants say, not just in theory. loginLimiter now stores its counters in a new rate_limit_counters table (migration 016) instead of a map; blocked/fail/clear take a context and query/upsert/delete a row keyed by the same strings callers already used (username, client address, "signup:"+address, ...). Semantics are unchanged -- a fixed window that resets rather than slides -- so no call site's behavior changes, only where the count lives. Added purgeRateLimits to the sweeper, alongside purgeSessions/purgeAckTokens, so expired windows don't accumulate. New internal (package api) tests in rate_limiter_test.go cover the basic behavior plus the regression this exists to fix: two loginLimiter values sharing one database, standing in for two replicas, now see one combined count instead of each keeping their own. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com> |
||
|
|
7cd6fbf571 |
Cap request body size and add baseline security headers
Part of a security-hardening pass (see wiki for the full backlog). decodeJSON had no size limit at all, so every JSON endpoint -- including the two unauthenticated ones (bootstrap, the Alertmanager webhook) -- would buffer an attacker-supplied body of unbounded size before it was even validated. decodeJSON now wraps the body in http.MaxBytesReader at a 1 MiB default; the webhook gets its own 8 MiB cap via decodeJSONLimit, since a real Alertmanager batch can be bigger than an ordinary API body. Also adds a securityHeaders middleware, applied globally: nosniff on every response (previously only the static site got it), and HSTS (180-day max-age, conservative on purpose) whenever cookieSecure's signal says the browser is on HTTPS. Checked the chart/gateway config first -- neither sets HSTS anywhere, so this was a real gap. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com> |
||
|
|
9da913080f |
Stop 500ing when a service account acts on an incident
Every incident-mutation handler read userFromContext(ctx) and wrote the result's .ID into acknowledged_by/incident_events.user_id without checking the ok bool. For a team-scoped service-account caller this returned a zero-value user id, which violated the users(id) FK and 500'd on acknowledge, unacknowledge, resolve, snooze, unsnooze and create-note. handleDeleteNote didn't crash but silently matched zero rows instead (WHERE user_id = 0), so a service account could never delete its own note. Add acknowledged_by_service_account_id (incidents) and service_account_id (incident_events) as nullable FKs to service_accounts(id), parallel to and mutually exclusive with the existing human columns (migration 015, with a CHECK enforcing the exclusion). Route every one of the six handlers plus delete-note through a new callerActorIDs() helper that branches on Caller.AsHuman()/ServiceAccountID() instead of assuming a human, and thread a serviceAccountID parameter through logEvent and the new acknowledgeIncidentAs (acknowledgeIncident itself is untouched: its only other caller, the push-notification Acknowledge button, is always human). Render the new actor distinctly from both a human and "the server acted" in the web UI's incident timeline and facts card. handleIncidentAssign, handleIncidentArchive and handleIncidentUnarchive are deliberately not touched here — they track no actor at all today, for anyone, which is a separate pre-existing gap (follow-up issue to come). Fixes #25 |
||
|
|
fa6d82d6e5 |
Fix desktop layout bugs and show more incident actions directly
On-call, the incident queue and the account page all had latent CSS bugs that only show up once the browser is wide enough to hit the desktop breakpoint (900px+): - On-call: .days reset its own margin to 0, which canceled the page-wide auto-centering on just that element, leaving the day list pinned to the left edge while every other card on the page centered normally. - Queue: the list pane stayed capped at 340-420px even with nothing selected, leaving the rest of the screen empty. It now fills the width until an incident is picked, then goes back to list+detail. - Account, admin user and admin team: .btn and .back-link are inline-flex, and margin:auto only centers a block box, so the Sign-out button and the two admin back-links sat left of their sibling cards instead of matching their width. Wrapped each in a block div. The incident detail action bar also folded Assign, Add note, Resolve and Clear acknowledgement into a "More" sheet sized for a phone's width. Desktop has the room, so it now shows them as direct buttons and hides More instead; Copy incident stays out of the bar since the header already has its own button for it. Filed as niklas/terdut-server#31, #32, #33, #34, each with a screenshot. |
||
|
|
9e5b085d8b |
Resolve the new-incident insert conflict instead of dropping the payload
openIncident's INSERT had no ON CONFLICT clause, relying entirely on incidentForGroup's earlier SELECT to avoid a duplicate. On more than one replica, two webhook deliveries for the very first occurrence of a brand-new groupKey can both pass that SELECT before either INSERTs; the loser then hit incidents_open_group_key_idx's unique violation, which rolled back its whole transaction — including that payload's alert upserts, done earlier in the same transaction. ingest's error is only logged and receiveWebhook answers 200 regardless, so nothing retried it: the loser's alerts silently never existed. Add ON CONFLICT (team_id, group_key) WHERE resolved_at IS NULL DO NOTHING to the INSERT, matching the partial unique index. Postgres only resolves that conflict after the winning transaction commits (or rolls back), so by the time RETURNING comes back empty, existingOpenIncident's follow-up SELECT is guaranteed to see the winner's row. The loser attaches to it instead of failing outright, and the rest of its payload commits normally. Covers both callers, since the dead man's switch sweeper shares this same function. New test (package api_test, fires N webhook deliveries for one groupKey from a synchronized start with distinct fingerprints, so they aren't accidentally serialized by upsertAlerts' own per- fingerprint lock) confirmed meaningful: with the ON CONFLICT clause reverted, it fails 10/10 on a missing alert fingerprint; restored, 0/10. Note while building it: "exactly one incident" alone cannot distinguish fixed from broken, since the DB's own unique index already guarantees that either way — the real signal is the loser's payload surviving. Chart comment updated: all three of the chart's original reasons for Recreate are now addressed in code, though replicas stays at 1 and the strategy stays Recreate pending a deliberate decision to raise it. Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
0050738ca0 |
Advisory-lock DB migrations against concurrent replica startup
Migrate's check-then-apply loop against schema_migrations had no locking: two replicas booting at once against a fresh or partially-migrated database could both pass the "not yet applied" check for the same file and race applying it, crashing whichever lost the duplicate-key insert (confirmed: reverting the lock fails the new test 10/10 on a duplicate-key violation, racing as early as the CREATE TABLE IF NOT EXISTS schema_migrations statement itself). Hold a Postgres advisory lock for Migrate's whole run, on a dedicated connection reserved via db.Conn so lock and unlock happen on the same session. Blocking (pg_advisory_lock), unlike the archiver/notifier's pg_try_advisory_lock: on boot there's no later tick to defer to, so a second replica should wait for the first to finish migrating rather than skip ahead. Adds internal/db's first test file, exercising two concurrent Migrate calls against a fresh schema. Still open: the new-incident-insert race on a webhook for a brand-new groupKey, noted in the chart's updated comment. Login rate limiting staying in-process, diluted across replicas, is an accepted tradeoff. Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
42180948d1 |
Guard the archiver and notifier passes with a Postgres advisory lock
Both background loops run unconditionally on every instance with no coordination between them, which the chart's replicas: 1 + strategy: Recreate exists specifically to paper over: with more than one replica, every one of them would sweep and deliver notifications independently, and two overlapping during a rollout would both page for the same incident. Add withAdvisoryLock, which takes a Postgres advisory lock on a dedicated connection and runs a pass only if it gets the lock, otherwise skipping until the next tick. Wire StartArchiver and StartNotifier through it with their own lock keys, so Sweep and NotifySweep themselves are untouched and every existing test calling them directly keeps working unchanged. This also closes the notifier's double-delivery race in passing: two replicas can no longer both be inside deliverPending at once, since only one can hold notifierLockKey at a time. Deliberately not addressed here, and still blocking a replica count above 1: the in-memory login rate limiter, the unlocked migration runner, and the new-incident-insert race on a webhook for a brand-new groupKey. Noted in the updated chart comment. Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
43beda9a30 | Merge pull request 'Fix the queue chip fade falling short of the right edge' (#30) from fix-chips-fade-edge into main | ||
|
|
710521a73c |
Fix the queue chip fade falling short of the right edge
position: sticky, as a flex item of the row it's pinning itself against, interacted with that row's gap and its own negative margin in a way that landed it short of the true edge -- visibly, a sliver of the next chip stayed poking out past where the fade should have covered it, which is the "ends before the screen edge" bug reported against the release. Replaced with the simpler, better-supported pattern for this: an absolutely positioned overlay against a position:relative, overflow: auto parent. Unlike a sticky descendant, an absolutely positioned one is resolved against the parent's own (non-scrolling) box, so it stays flush with the real edge regardless of scroll position, without the flex-gap/margin interaction that caused this. Verified visually: a standalone reproduction of both versions, screenshotted with Chromium's headless_shell (no browser automation tool available in this environment, but the binary's right there) -- the old version shows the next chip's edge peeking past the fade, the new one doesn't. 🤖 Generated with [Claude Code](https://claude.ai/code) Co-Authored-By: Claude <noreply@anthropic.com> |
||
|
|
1770e5d945 |
Fix Stats/Admin/Account showing in the bottom bar too
.nav-link-secondary's display:none sat before .nav-link's own display:flex in the file. Both are single-class selectors, so they tie on specificity, and a tie is broken by which one comes later in the file -- not by which class the element happens to carry. .nav-link's declaration, being later, won for every element wearing both classes, so Stats/Admin/Account rendered as three extra tabs on the phone bar instead of folding into "More" as intended. Moved the rule below .nav-link instead of changing either declaration, since nothing about the values was wrong -- only their order was. 🤖 Generated with [Claude Code](https://claude.ai/code) Co-Authored-By: Claude <noreply@anthropic.com> |
||
|
|
91f03c21e8 |
Web: visual design pass (issue #26)
Addresses the screenshot-review feedback in #26. No framework or build step added — all of this stays within the existing plain HTML/CSS/ vanilla-JS + go:embed architecture. - Nav: re-enable the bottom tab bar that was already built and switched off (Queue/On-call/Alerts/Team + a "More" sheet for Stats/Admin/Account), replacing the hamburger on phone width. - Queue: chip counts, a scroll fade on the filter row, a "Triggered Xh ago" + severity label per row, a chevron on the team switcher so it reads as a dropdown. - On-call: collapse repeated same-person days into shift bars (week view and "your shifts" both), show the week as a date range with the ISO week number as secondary text, split "Current shift" out from "Next shifts" with "ends in Nd", a pill badge + row highlight for "you". - Incident detail: fix the actual bug behind the duplicate "acknowledged" timeline entries (acknowledgeIncident's UPDATE had no guard on the incident's current status, so acknowledging an already-acknowledged incident silently re-logged the event — now idempotent, with regression tests on both the authenticated route and the ntfy ack-button route). Relabel escalation re-pages so they don't look like the same page landing twice. Copy the primary action up near the top. Label the "···" button. Group the timeline by phase (triggered/acknowledged/resolved). Add an "at a glance" summary row (duration/severity/responsible) and collapse the group labels by default. - Team overview: reword the vague copy ("One owner." etc.) into plain labels. - Empty states: fill in missing icons/one-liners across queue, alerts, stats and the incident timeline. - CSS: fix card padding bugs, verify link contrast already passes AA, introduce a --fs-* type-scale token set and migrate the few genuinely isolated cases onto it (left sizes tied to a fixed shape, a deliberately prominent display, or a non-negotiable constraint like the iOS-zoom-prevention input size as documented exceptions rather than guess at a render this change can't see). Verified with the full fmt/lint/test/helm-lint gate, plus a live instance against the test DB with seeded incidents and schedule data to trace the on-call grouping and timeline phase-splitting logic against real API responses. 🤖 Generated with [Claude Code](https://claude.ai/code) Co-Authored-By: Claude <noreply@anthropic.com> |
||
|
|
774fdfcaa8 |
internal/api: unify human/service-account authz into one Caller type
ctxUser/ctxTeams (human) and ctxServiceAccount (+ a synthetic ctxTeams
entry, service account) used to be two parallel, un-unified context
representations -- every authz predicate had to remember which one(s) it
needed, and every place that forgot either wrongly 403'd a service account
(terdut-server#23, terdut-operator#3), crashed on an unchecked zero-value
user id, or silently no-op'd. New internal/api/caller.go collapses both
into one Caller, stored under one ctxCaller key by serveAs/serveAsServiceAccount;
every existing predicate (userFromContext, callerTeamIDs, callerRole,
callerIsAdmin, isInstanceServiceAccount, AdminOnly, requireSelfOrAdmin,
requireTeamOwner, OperatorModeBlock) now reads through it, with identical
behavior for every untouched call site (alerts.go, incidents.go,
schedule.go, stats.go, etc.) -- confirmed by the full existing suite
passing unchanged.
Four real fixes land alongside the refactor, not just the restructuring:
1. callerMayManageServiceAccount gains the one load-bearing branch this
exists for: an instance-scoped service account may now manage (mint or
revoke a key on) any team-scoped account, not only a human admin, that
team's human owner, or the account itself. handleCreateServiceAccount
already let an instance-scoped caller *create* a team-scoped account for
any team; adopting or rotating one it didn't just create in the same
call -- terdut-operator's own documented crash-window recovery -- had no
equivalent permission and 403'd forever. Closes terdut-operator#3.
2. handleCreateInvite wrote a service-account caller's zero-value user id
straight into invites.created_by (nullable, but never passed as nil),
which foreign-key-violates against users(id) -- a 500, not success, for
any team-scoped service account minting an invite. Fixed the same way
handleCreateServiceAccount already handles the analogous case. Found
live while verifying this change, not filed separately since it's fixed
in the same place it was found.
3. handleMe and handleTestNotification 500'd for a service-account caller
(fetchUser/the ntfy_topic lookup against a zero-value user id that
matches no row); handleDismissOnboarding silently no-op'd (UPDATE ...
WHERE id = 0). All three now call Caller.AsHuman() and return an
explicit 403 ("this endpoint is for human accounts only").
4. Ratifies, rather than further narrows, two capabilities a team-scoped
service account already had by construction and this document's own
text once called "a gap acknowledged rather than closed": owner-equivalent
reach over membership/invites, and minting another service account for
its own team. terdut-operator's new TerdutTeam invite-minting feature is
about to depend on the first one, so this makes it documented, tested,
intentional behavior instead of an accident nobody was supposed to rely
on.
AdminOnly/requireSelfOrAdmin are unchanged in effect: still human-only,
forever, for every scope of service account -- confirmed by
TestAdminOnly_RefusesEveryServiceAccountScope. terdut-server#23's named
routes (POST /api/users, PUT /api/admin/settings) were never the right
thing to widen; its real fix is the terdut-operator invite feature,
recorded in SERVICE-ACCOUNTS.md's "What this unblocks" and closing that
issue once it ships.
SERVICE-ACCOUNTS.md amended in place (not a new file, its own established
convention) to describe the as-built Caller model, correct its own
aspirational claim about AdminOnly that TEAM-LOOKUP.md had already flagged
as not matching shipped code, and record all of the above.
|
||
|
|
fc9f47cc8d |
Refuse an OIDC sign-in from creating the very first user
Closes the race terdut-operator#1 found: /api/bootstrap and OIDC auto-provisioning both key off the same signal (SELECT COUNT(*) FROM users) with no coordination between them, so an otherwise-ordinary OIDC sign-in against a freshly-created, not-yet-bootstrapped install could create user #1 itself and take the one slot /api/bootstrap expects to win uncontested (terdut-operator's own design, DESIGN.md §1/§6, assumes it is the only caller). The operator has no way to recover from losing that race -- it never gets a credential, and nothing it owns can clear the occupying user row. resolveSSOUser now checks the same gate handleBootstrap already does, right where it's about to create a brand-new user (an identity nobody has linked yet, that also matches no existing local account by email) -- not anywhere else, since every other sign-in on an already-bootstrapped install is unaffected. New sso_error code `not_bootstrapped`: the person sees "this install is still setting up, try again in a moment" and a second attempt once something has actually bootstrapped succeeds normally, same as any other first sign-in. Does not fix the other half of that issue (BootstrapStateLost's own "delete and recreate" instructions still don't work once something has occupied the slot some other way) -- this closes the specific race, not every path to that state. |
||
|
|
bc9f793f1f |
Add GET /api/teams?name= (TEAM-LOOKUP.md)
Resolves the gap TEAM-LOOKUP.md raised: an instance-scoped service account had no way to recover a team's id after a 409 on POST /api/teams, unlike the already-solved equivalent for service accounts themselves (GET /api/service-accounts?name=). Same route, extended the same way handleListServiceAccounts already branches on ?name=: unset behaves exactly as before (the caller's own teams via team_members); set looks up one team by exact name, open to any authenticated caller -- not gated by isInstanceServiceAccount or AdminOnly, since what it discloses (a name is taken, nothing about who's in it) is the same low sensitivity that lookup already accepts for service-account names. Tests cover the exact motivating scenario (create, 409 on a retry, recover the id via ?name=), the empty-array-not-an-error case, that no role/source is reported for a non-member match, and that a caller who isn't a member of the matched team still gets it. |
||
|
|
a4dd60f6b8 |
Add service accounts and operator mode
Service accounts (SERVICE-ACCOUNTS.md) are a scoped, non-human credential: not a users row, so they never touch OIDC sync, login or the is_admin flag. Instance scope can create a team and mint a team-scoped account for it; team scope is owner-equivalent for that one team and nothing else. This is what unblocks terdut-operator's DESIGN.md §6 — no more impersonating a human admin, and a real rotation story instead of the unworkable delete-and-re-bootstrap /api/bootstrap can't actually do. - migration 014: service_accounts + service_account_keys - POST /api/service-accounts, POST/DELETE .../keys, GET ?name= self-lookup - AuthMiddleware resolves a tdsa_-prefixed key to a distinct principal; a team-scoped account gets a synthetic single membership so requireTeamMember/requireTeamOwner work on it unmodified - handleCreateTeam accepts an instance-scoped caller; the team it creates has no human owner, which is the expected shape for one an operator is about to hand a team-scoped credential to Operator mode (TERDUT_OPERATOR_MODE / values.operatorMode) declares an install gitops-managed: session and user-API-key writes to teams, escalation policies, dead man's switches and integrations get 403 reason=operator_managed, while a service account's writes still go through. Team membership/invites and the schedule are deliberately left out — never gitops-managed by design, and still human day-to-day work. /api/auth/config reports operator_mode so the web UI can grey these sections out from the start rather than only after a write fails. Also: GET /api/version (both terdut-tui and terdut-operator currently detect server capability by route-probing; this gives them a real answer), and a PUT for dead man's switches so a reconciler can update one in place instead of deleting and recreating it. |
||
|
|
b610b1817a |
Fold the account page's ntfy and password forms behind disclosures
Both sat open by default, competing with the rest of the page for
attention on every visit even though most visits need neither. Team's
rota already has the same problem for its bulk-assign form and solves
it with a native <details>/<summary> disclosure, styled generically in
app.css; this reuses that idiom rather than inventing a JS toggle.
Each section now shows a one-line status (the topic, or whether a
password is set) with the actual form folded under a summary naming
the action ("Set a topic" / "Change topic", "Set a password" /
"Change password"). A successful save closes the fold and confirms
with a toast, since the point of folding is that a saved form goes
back to being just a status line; a validation or API error keeps the
fold open and shows inline, next to the field it's about.
The password section's heading no longer says "Change password" or
"Set a password" itself, since that verb now lives on the summary; it
just says "Password", matching the existing SSO-off case.
|
||
|
|
33356ca978 |
Add a global, colour-coded team selector to the nav
state.js's currentTeam() was hard-coded to teams[0] and never really meant "the team currently selected" — team.js's settings page and queue.js's filter chips each kept their own separate, unsynchronized notion of "which team" instead, so picking one on one page had no effect on the other. Replaces both with a single state.selectedTeamID, set only through the new setSelectedTeam (persisted in localStorage, unlike the queue's old per-tab sessionStorage filter) and broadcast to listeners via onTeamChange. A new teamselector.js control — a coloured dot plus the team's name, or "All teams" — sits at the top of both the desktop sidebar and the mobile topbar, opening the existing bottom-sheet menu to switch. Shown only once someone is in more than one team, matching every other team-aware control in this app. Colours come from a new teamColorClass() in format.js, hashing a team's id into the six-colour rc1..rc6 palette already used for the rota's per-person chips, so no schema or API change is needed. The queue's team filter chips pick up the same colours. |
||
|
|
5b4683febf |
Let each team name its own OIDC group, not a global mapping
Team membership from single sign-on used to come from one env var,
TERDUT_OIDC_GROUP_MAPPINGS, matched against a team by name and creating
the team if none existed. That put the decision in the server's
environment rather than the team's own hands, needed a restart to
change, and let a typo in a team name silently create a stray team.
Each team now carries its own oidc_member_group and oidc_owner_group,
set by its owner (or an administrator) from the Members tab, or PUT
/api/teams/{teamID}/oidc-groups. The "highest role wins" rule
TERDUT_OIDC_GROUP_MAPPINGS used to apply across mappings now applies
across one team's own two fields: being in both makes somebody an
owner. The sync no longer creates a team by name; a group only ever
grants into a team that already exists.
This is a breaking change for anyone already using
TERDUT_OIDC_GROUP_MAPPINGS, deliberately not auto-migrated: an
OIDC-sourced membership is dropped at a user's next sign-in until its
team's owner re-sets the group. The README's OIDC section spells out
the migration and the risk of a visible access gap during it.
TERDUT_OIDC_ADMIN_GROUP and TERDUT_OIDC_ALLOWED_GROUPS are untouched --
only team membership moved. terdut-tui needs no change: it only reads
GET /api/teams and GET /api/teams/{id}/members, and neither response
shape moved.
|
||
|
|
a27ff49171 |
Sign in through an OpenID Connect provider, and from a terminal
terdut can now sign people in through any OIDC provider (written against Authentik), and let groups at the provider decide who may sign in, which teams they belong to and whether they administer the install. Password login keeps working alongside it; TERDUT_PASSWORD_LOGIN=false turns it off, and is refused at startup unless SSO is configured. With no TERDUT_OIDC_* setting nothing changes, so every existing install behaves as before. Identity is (issuer, subject), never email or username: those are mutable at the provider and a recycled address must not inherit an account. An existing user is linked by email only when the provider marks it verified, or TERDUT_OIDC_TRUST_EMAIL is set, which Authentik needs. Group grants are marked source='oidc' on team_members and users, and the sync changes only those rows. Hand-made memberships and administrators are left alone, and the sync bypasses the last-owner and last-admin guards because the provider is the source of truth for what it grants. Editing managed access by hand is refused with 409, since the next sign-in would undo it. The web UI badges it as SSO and disables the controls. Groups are read only at sign-in, so an SSO session carries a hard ceiling (sessions.max_expires_at, 12h by default) that sliding never extends. There is no refresh token, which means API keys of somebody removed at the provider stay valid until an administrator disables the user. That is accepted and documented, not fixed. A client with no browser, the TUI over SSH, signs in with a device code run by terdut itself (POST /api/oidc/device and /device/token), so the terminal never talks to the provider and ends up with the ordinary terdut_session cookie. Only a browser session can approve a code; an API key cannot. /device?code= sends a signed-out visitor through sign-in and back, which is what oidc_logins.next is for. oauth2 is pinned to v0.36.0: v0.37 needs Go 1.26 and the Dockerfile builds on 1.25. Migrations 011 and 012 add tables and defaulted columns only. |
||
|
|
b2c3868619 |
Fix the web UI stuck on its loading spinner
format.js defined isoWeek twice.
|
||
|
|
9d1df2b611 |
Show week numbers on the rota, and assign a whole week from them
The rota grid starts each row with its ISO week number, and for an owner the number is a button: one tap opens a sheet for that week, showing who holds each of its seven days, and puts one person on all of them. A rota is usually handed out by the week, and seven taps on seven days was the only way to do it short of the range form. ISO 8601 numbering, because the grid already runs Monday to Sunday: week 1 is the one holding the year's first Thursday, which is taken from the Thursday of the row so the year boundaries come out right (2025-12-29 is week 1 of 2026, 2020-12-31 is week 53). Days already past are left alone. Who was on call last Tuesday is a fact, and "the whole week" should not rewrite it, so a half-elapsed week covers the days still to come and the sheet says so; a week that is entirely over has nothing to assign. By default the assignment replaces whoever holds those days, as the day sheet does and the sheet states, and a checkbox limits it to the days nobody has yet. The overhang into the neighbouring month is part of the same week and is included. Web UI only: the existing schedule endpoint already takes a list of dates and a replace flag, so nothing changed on the server and there is nothing to mirror in terdut-tui. |
||
|
|
e616c82646 |
Show team members as a list, with who is on call and who cannot be paged
Team -> Members was a two-column table and an inline add form. It is now a table in the style of Switches, Sources and Escalation: a status badge, the member, their role, their next rota day, when they were last active and when they joined. Adding a member and changing a role moved into sheets, and removing one asks first. The badge is the one that matters at 03:00: On call if the rota has them today, Reachable if they have an ntfy topic, and Can't be paged when a page to them would go nowhere -- no topic, or a disabled account -- with the reason under their name. Not being pageable wins over being on call, since an on-call person nobody can reach is the case worth seeing before an incident finds it. The rules are the notifier's own. The topic itself is never in the response, only whether one is set. Last active is the newer of a member's newest session and API-key use, and is shown to every member of the team like the rest of the list. Rota days are UTC dates, and the page formats them as such so a day cannot show up as the one before. Removing a member leaves the rota days already assigned to them alone, which the confirm says, so they are reassigned from the Rota tab rather than silently dropped. Demoting the last owner is now refused with 409, as removing them already was: it was the same outcome by another route, a team with nobody who can edit it. API: GET /members gains status, on_call, next_shift, pageable, problem and last_active_at; additive, no migration, and terdut-tui needs nothing. POST /members answers 409 for the last-owner demotion. |
||
|
|
1f1faa437c |
Show the escalation ladder as a list, with who it would page and where it is
Team -> Escalation was the draft form on the page, which showed the ladder only as inputs. It is now a table in the style of Switches and Sources: a row per level with a status badge, who it pages, the wait before the next level, and the open incidents currently waiting on it. Below it, the repeat count, the fallback topic and when the ladder last escalated (linking the incident). The editor moved into an "Edit ladder" sheet, so a poll of the page underneath can no longer throw away half an edit, and the page-level draft state went with it. Targets are resolved to who they mean today, and the badge says what would actually happen: Ready, Escalating (an unanswered incident has climbed to level 2 or higher), or Pages nobody. The last is the one worth seeing before an incident finds it: an empty rota, a person with no ntfy topic or a disabled account each make a rung a silence with a number on it, and the target says which. The rules are pageLevel's own, so the page cannot promise a page the notifier would skip. "Last escalated" comes from the escalated timeline events that already exist, so there is no migration. Acknowledging or resolving takes an incident off the ladder, so Escalating clears then while the history stays. API: GET /escalation gains status and waiting per level, username, reachable and problem per target, and last_escalated_at and last_escalated_incident_id. Output only and additive; PUT is unchanged and terdut-tui needs nothing. |
||
|
|
d675f8ec9b |
List alert sources with their status and last arrival on Team -> Sources
Like Team -> Switches, the page is now a table: a status badge (Active
if the key posted within a day, Quiet if it has but not lately, Never
used), when it last posted a webhook, when an alert last arrived on it,
how many distinct alerts it refreshed in the last 24 hours, and when it
was created. Adding a source moved into a "New source" sheet, and owners
can rename one from its row.
"Last alert" and the count needed alerts to remember which source they
came in on, which they never did, so migration 010 adds
alerts.integration_id and every accepted payload stamps it. Last sender
wins when two sources post the same fingerprint. It is not backfilled: a
NULL says "before this was recorded" rather than guessing, and it heals
by itself as Alertmanager re-sends each alert every repeat_interval.
Revoking a source keeps its alerts, unattributed.
Last webhook and last alert are separate on purpose: a payload with
nothing usable in it stamps the first and not the second. The Quiet
threshold is a fixed day, a colour and not an alarm, since silence that
should page is what dead man's switches are for.
The counts are indexed subqueries (alerts_integration_idx) rather than a
join, which would read every alert a source ever delivered.
API: the integrations list gains status, last_alert_at and alerts_24h,
and PATCH /api/teams/{id}/integrations/{id} renames. Both are additive;
terdut-tui needs nothing.
|
||
|
|
f3918b863c |
List dead man's switches with their status on Team -> Switches
The page was a bare form: it did not say which switches existed or
whether they were alive. It now lists them, each with a Healthy, Dead or
Dormant badge, when its heartbeat was last heard and when it last opened
an incident (linked while that incident is open). A matcher that several
clusters satisfy is broken down per cluster, since a live cluster must
not hide a dead one. The form moved into a "New switch" sheet, and each
row has a Remove with a confirm.
That needed a switch to be a thing, so switches are rows now
(migration 009) with their own name, matcher, timeout and severity,
instead of one string with one team-wide timeout in deadman_configs.
Existing configuration is split into one row per matcher; a team whose
timeout was zero simply has none. The sweeper and the status endpoint
share one death rule (deadmanAlert.dead), so the page cannot disagree
with the pager. Incident group keys are unchanged, so incidents that
are open across the upgrade keep working.
The environment defaults (TERDUT_DEADMAN_*) are seeded into teams once
per install, recorded in settings, so a team that deletes its last
switch does not get it back on the next restart. Installs that already
had per-team rows are marked as seeded by the migration.
Removing a switch stops the watching but leaves an incident it already
opened open until someone resolves it.
API: GET/PUT /api/teams/{id}/deadman are replaced by
GET/POST /deadman/switches and DELETE /deadman/switches/{switchID}.
terdut-tui does not call them, so nothing to mirror there.
|
||
|
|
591d5b8df0 |
Copy an incident to the clipboard as Markdown
A button in the incident header (also `y`, and "Copy incident" in the more menu) puts everything the page knows on the clipboard, for pasting into a chat or an agent prompt with no integration involved. The text carries the facts, every alert with all its labels and annotations (the page only shows summary or description), the timeline with notes in full, and the "Seen before" resolution notes. Times are ISO 8601 and users are named rather than "you", since relative and first-person wording is ambiguous once pasted elsewhere. The async clipboard API needs a secure context and this server is often reached over plain HTTP, so it falls back to execCommand. Web UI only: no endpoint or JSON shape changed, so nothing to mirror in terdut-tui. |
||
|
|
8b2789b9b2 |
Let the filter chips wrap in the desktop incident list
The list pane is 340-420px wide and its chip row scrolled sideways with the scrollbar hidden. That works by swipe on a phone, but a mouse has nothing to grab, so Archived (the last chip) could not be reached on a wide screen. In the desktop layout the row now wraps instead, and the divider between the status and team chips is hidden there, since it would sit mid-line. Phones keep the sideways scroll: the rule is inside the min-width: 900px block. Claude-Session: https://claude.ai/code/session_01MMados3BD1oSjevHxbmVqU |
||
|
|
60ebb75cd2 |
Show notes from similar earlier incidents
Each incident gets a signature: the alert name plus the group labels that
say what is broken, minus the ones that only say where it ran (instance,
pod, container, ...). GET /api/incidents/{id}/similar returns resolved
incidents in the same team with the same signature that have notes.
Notes can be marked as the resolution note, "what fixed it", either with a
resolution field on resolve or pinned on a note. Those lead the similar
list, show on the incident page as "Seen before", and the triggered
notification carries the latest one.
Claude-Session: https://claude.ai/code/session_01MMados3BD1oSjevHxbmVqU
|
||
|
|
423ed9b3a3 |
Add a Stats page to the web UI
Statistics used to live only in terdut-tui; the account page said so. The page shows the same figures as the TUI's Stats tab -- incident counts, MTTA and MTTR, top alerts, and alert frequency by hour (UTC) and by day of week -- and adds a range picker (Today, 7d, 30d, 90d, All) that the TUI does not have. The ranges are day-granular because the server reads from/to as whole UTC dates, so there is no 24h chip. No server change: the page uses the existing /api/stats/* endpoints, which already scope to the caller's teams. Charts are inline SVG and plain elements sized from script, because the CSP forbids inline styles, inline scripts and CDN libraries. Removes the "statistics are in terdut-tui" notes from the account page and the README. |
||
|
|
559be6de6e |
Retry the first database ping instead of dying on it
Every start crashed once or twice before going healthy: kube-router enforces this namespace's NetworkPolicy per-node, reacting to the new pod's creation event, and the app's first connection attempt can reach Postgres's node before that node's allow-set has been updated with the new pod's IP. The result is "connection refused" -- an active reject, not a timeout, which is how it was told apart from Postgres itself not being ready (it had been up for two days in the run that was diagnosed). That race resolves within several seconds in practice, so Open now retries the ping up to five times, two seconds apart, logging each failure, before giving up with the same wrapped error as before. Nothing else about Open's behaviour changed: a genuinely absent database still fails, just after ~8s instead of immediately. Claude-Session: https://claude.ai/code/session_01RHPj4ggeFdEjKKfm4SHbD7 |
||
|
|
e536fdd2c0 |
Replace the mobile tab bar with a hamburger menu
Six tabs (Queue, On-call, Alerts, Team, Admin, Account) had already
outgrown the bottom bar once:
|
||
|
|
3cdd5aee1f |
Give the Team tab sub-sections of its own
The Team tab was five cards stacked on one page: the rota, the escalation ladder, the alert sources, the dead man's switches and the membership. |
||
|
|
67d68ce058 |
Show the rota as a month rather than a list of dates
The Team tab printed the next thirty days as thirty rows of date, name and
a Clear button. That is a rota spelled out one day at a time, and it is the
one shape the question cannot be read in: what anybody wants from a rota is
who holds which stretch, and thirty names down a column hides a handover
between two rows that look the same. It was also the longest thing on the
page by a wide margin, so the escalation ladder and the alert sources sat
below a screen of dates.
It is a month now, Monday to Sunday, one coloured initial per day. A shift
becomes a run of one colour, which is the shape the answer actually has; a
gap becomes a hole you can see. The legend underneath says whose colour is
whose, and one line says how many days are left uncovered, counting only
from today -- an empty Tuesday last week is history, not a hole somebody
still has to fill.
Laid out like the on-call page's week, deliberately: heading and arrows
outside the card, days inside it. It is the same rota, and two pages
showing it two ways would be two things to learn.
Colours come from a person's place in the member list, so they hold still
as you page between months, and six of them repeat -- the initial inside
still tells two people apart, and a legend that has to explain nine hues is
not a legend. They are not the severity palette: nothing on a rota is
critical, and a red Thursday would read as one. --teal and --pink are new
in both themes for the two the palette was short.
The per-row Clear button had nowhere left to live, so a day opens the sheet
the app already uses for confirmations: who holds it, a picker, Assign and
Clear. That assign sends replace=true where the range form still asks
first, and the difference is the point -- the sheet has just named whoever
holds the day, so taking it from them is the thing that was asked for
rather than something to warn about. The range form is unchanged and folded
into a details, since filling a whole shift is what it is for; it opens on
the month above it rather than on today, so paging to March to fill March
does not hand you September.
The server is untouched. The month drawn is the month fetched -- the grid's
Monday overhang and its trailing days are real days and are fetched with
it -- so paging is one GET /api/teams/{id}/schedule per month with from and
to, where it used to be one fixed thirty-day window. No new endpoint, no
change to what the API returns, and terdut-tui is unaffected.
Nobody has looked at this in a browser either. What is checked is the
rendering: team.js's own refresh() was run against a stub fetch and a
pocket DOM for September 2026, and it produces 35 cells for a month whose
1st is a Tuesday, the right from/to on the schedule call, today marked on
the 22nd, three people in the legend with "you" on the viewer, the gap
count over a five-day hole, and -- as a member rather than an owner -- the
same grid as plain divs with no sheet and no range form. How it looks at
phone width, and whether the six colours hold up in dark mode, are not
checked.
Claude-Session: https://claude.ai/code/session_01RHPj4ggeFdEjKKfm4SHbD7
|
||
|
|
a6fa673e08 |
Give every team a page of its own
The Admin tab's team list was growing controls the way the user list did before |