1cb09525e39e11d6c2c77225e41c79307c8b3f4d
10 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
a955356821 |
Filter the queue by cluster
A team with a cluster per alert source can now narrow the queue to one cluster. A dropdown under the status chips lists the clusters, and appears once there are two or more to choose between, the rule the team selector follows. The choice is kept per browser, like the selected team, and dropped if the server no longer knows the cluster rather than leaving an empty list with no explanation. The status chip counts follow it, and a row leaves out its cluster chip once the queue is narrowed to one. The filter is on the server. The list is capped at 50 rows, so filtering what is on screen would quietly miss older incidents in the Resolved and Archived lists. GET /api/incidents takes ?cluster=<value>, matched against the incident's `cluster` group label, and GET /api/incidents/clusters lists the distinct values from the last 90 days (optionally for one team), so the dropdown is not limited to what the current page happens to show. Both are additive: no existing parameter or JSON shape changed, so terdut-tui keeps working unchanged and has nothing it must mirror. An incident carries the label only when `cluster` is in Alertmanager's group_by, so the filter only sees those; the README says so. |
||
|
|
fb86a18988 |
Show which cluster an incident came from, as a chip
A team with one Alertmanager per Kubernetes cluster could not tell at a glance where an incident started: the cluster was only a word inside the title. When an alert carries a `cluster` label, the queue rows, the incident page and the alerts list now show it as a chip in that cluster's colour, a stable pick from the existing six-colour palette. A queue row drops `cluster=...` from its title, since the chip says it, and the incident page keeps the full title. An incident has the label only when it is in Alertmanager's group_by, which is also what keeps two clusters' identical alerts apart: incidents are matched on the team and the groupKey, and the groupKey does not include external labels. The README has a section on the two settings (Prometheus externalLabels and group_by). The alerts list reads the label from the alert itself, so it shows the chip with only the external label. Web UI and docs only: no endpoint or JSON shape changed, so nothing to mirror in terdut-tui. Nothing changes for a team whose alerts have no cluster label. |
||
|
|
942517c7a8 |
Keep the scroll fade at the screen edge on the queue's filter row
The fade that hints at more filters to the right was an element inside the scrolling row, so it scrolled away with the chips instead of staying at the edge; the comment on it claimed the opposite. It is now a mask on the strip itself, applied only while there is more to scroll to (fadeOnOverflow in ui.js sets data-more), so it stays put and disappears at the end instead of dimming the last chip. The Team and Admin tab strips had an always-on mask from the earlier "Sources is cut off" fix, which dimmed their last tab even when fully scrolled; they use the same mechanism now. Web UI only: no endpoint or JSON shape changed, so nothing to mirror in terdut-tui. |
||
|
|
bcf1a3e99b |
Give the sidebar groups, and tidy the tables, disclosures and tab bar
The desktop sidebar is now grouped: Queue, On-call, Alerts and Stats; Team and Admin; then Account at the foot, shown as the signed-in person with an avatar and their name. The active item has a bar as well as a tint. The queue's row of team chips is gone: the team selector is the one place the team is chosen, and the chips offered the same choice a second time. Tables get 16px between columns, so a right-aligned count no longer touches the text beside it. An empty SSO group reads "Not configured" instead of a dash, and the main action on each of those pages (New source, New switch, Add member, Assign, and the admin Add, Create invite and Create) is a filled button, with Edit and Cancel staying secondary. The bare triangles on "Grouped by" and the label lists are real disclosure buttons: a chevron that turns and a 44px target on a touch screen. On a phone the More tab was a <button> that kept the browser's grey box, so it looked highlighted next to four plain links. It now matches them, the labels are 12px and the open tab's icon is filled. Web UI only: no endpoint or JSON shape changed, so nothing to mirror in terdut-tui. |
||
|
|
7412456c5a |
Pad the settings cards, and make status and timeline readable at a glance
The Team settings cards (Escalation, Sources, Members, Switches, Single sign-on) printed their text flush against the border with the button in the corner: .card has never had padding and these never added any. A card with a header row now pads itself, with the title left, the button right and a divider before the content. The incident page no longer carries the primary action twice. The copy up by the status is gone and the sticky bar keeps it; Note is the timeline's link, with Copy in the bar for a resolved incident, and the phone's More sheet drops what the bar already shows. In the queue, a row's title wraps to two lines so the namespace that tells rows apart is no longer cut off, zero counts on the filter chips are dimmed, and a row omits the status the filter already states and the team once the queue is narrowed to one. Status colours failed 4.5:1 against their own fill in the light theme (warning 3.6, info 4.1, critical 4.4, ok 4.45, snooze 4.49), so the light tokens are darker; the dark theme already passed and is unchanged. Severity badges now carry a shape as well as a colour, and each timeline event has an icon. An escalation that ran out of levels, a failed notification and a silent heartbeat stand out in amber or red. Web UI only: no endpoint or JSON shape changed, so nothing to mirror in terdut-tui. |
||
|
|
91f03c21e8 |
Web: visual design pass (issue #26)
Addresses the screenshot-review feedback in #26. No framework or build step added — all of this stays within the existing plain HTML/CSS/ vanilla-JS + go:embed architecture. - Nav: re-enable the bottom tab bar that was already built and switched off (Queue/On-call/Alerts/Team + a "More" sheet for Stats/Admin/Account), replacing the hamburger on phone width. - Queue: chip counts, a scroll fade on the filter row, a "Triggered Xh ago" + severity label per row, a chevron on the team switcher so it reads as a dropdown. - On-call: collapse repeated same-person days into shift bars (week view and "your shifts" both), show the week as a date range with the ISO week number as secondary text, split "Current shift" out from "Next shifts" with "ends in Nd", a pill badge + row highlight for "you". - Incident detail: fix the actual bug behind the duplicate "acknowledged" timeline entries (acknowledgeIncident's UPDATE had no guard on the incident's current status, so acknowledging an already-acknowledged incident silently re-logged the event — now idempotent, with regression tests on both the authenticated route and the ntfy ack-button route). Relabel escalation re-pages so they don't look like the same page landing twice. Copy the primary action up near the top. Label the "···" button. Group the timeline by phase (triggered/acknowledged/resolved). Add an "at a glance" summary row (duration/severity/responsible) and collapse the group labels by default. - Team overview: reword the vague copy ("One owner." etc.) into plain labels. - Empty states: fill in missing icons/one-liners across queue, alerts, stats and the incident timeline. - CSS: fix card padding bugs, verify link contrast already passes AA, introduce a --fs-* type-scale token set and migrate the few genuinely isolated cases onto it (left sizes tied to a fixed shape, a deliberately prominent display, or a non-negotiable constraint like the iOS-zoom-prevention input size as documented exceptions rather than guess at a render this change can't see). Verified with the full fmt/lint/test/helm-lint gate, plus a live instance against the test DB with seeded incidents and schedule data to trace the on-call grouping and timeline phase-splitting logic against real API responses. 🤖 Generated with [Claude Code](https://claude.ai/code) Co-Authored-By: Claude <noreply@anthropic.com> |
||
|
|
33356ca978 |
Add a global, colour-coded team selector to the nav
state.js's currentTeam() was hard-coded to teams[0] and never really meant "the team currently selected" — team.js's settings page and queue.js's filter chips each kept their own separate, unsynchronized notion of "which team" instead, so picking one on one page had no effect on the other. Replaces both with a single state.selectedTeamID, set only through the new setSelectedTeam (persisted in localStorage, unlike the queue's old per-tab sessionStorage filter) and broadcast to listeners via onTeamChange. A new teamselector.js control — a coloured dot plus the team's name, or "All teams" — sits at the top of both the desktop sidebar and the mobile topbar, opening the existing bottom-sheet menu to switch. Shown only once someone is in more than one team, matching every other team-aware control in this app. Colours come from a new teamColorClass() in format.js, hashing a team's id into the six-colour rc1..rc6 palette already used for the rota's per-person chips, so no schema or API change is needed. The queue's team filter chips pick up the same colours. |
||
|
|
b39aac36b7 |
Add the sign-up page and the first-run checklist
Second half of #7. The API could create accounts from invite links since the last change; this is the part somebody can actually use. /signup is the one route that works without a session. It asks the server what it may offer before showing anything: an invite link that is good names the team it leads to, a link that is not says so before somebody picks a password rather than after, and an invite-only server with no link says that instead of presenting a form it will refuse. The login card only offers "create one" when sign-up is open, so the door nobody can walk through is not advertised. Signing up signs you in and lands on the queue, because the alternative is a form saying "now go and log in" about the credential just chosen. The checklist is the other half. Four things have to be true before an alert reaches a phone -- a notification topic, somebody on the rota, an alert source, and an alert that has actually arrived -- and on a fresh install none of them are. It sits above the queue until they are. It is computed from the data rather than from stored progress: a topic is set or it is not, an integration exists or it does not. That means it cannot claim a step is done when it is not, and it comes back by itself if somebody deletes their integration a month later. The only stored state is the dismissal, which is per user and not per browser -- finishing on a laptop should not leave the phone nagging. The topic step is the only one the checklist can finish itself, and the only proof that counts is a phone buzzing, so there is a test push. POST /api/me/notify/test publishes directly rather than through the outbox, which requires an incident this deliberately does not have. Its failure is the useful part: a wrong topic, a rejected token and an ntfy that is down all look identical from the phone, which is silence, so the error comes back to the browser instead. Verified against a live server with a real ntfy stand-in, the whole path: an owner mints an invite, the sign-up page reports it valid and names the team, the invitee signs up and is signed in as a member of that team, the checklist's four questions answer correctly on a fresh install, a test push is refused with no topic and delivered with one -- "PAGED terdut-owner | terdut test" -- and the dismissal survives a reload. Claude-Session: https://claude.ai/code/session_01RHPj4ggeFdEjKKfm4SHbD7 |
||
|
|
74359c72ab |
Give each team its own dead man's switches, and the UI a team to show
The rest of #4. Two halves that belong together because they are the same sentence from opposite ends: a team decides which of its alerts are heartbeats, and the UI has to be able to say which team it is talking about. Switches were three environment variables, which made them one setting for the whole install. That was the last piece of the alerting path a team could not control: it could take its own alerts on its own key and still not say which of them were heartbeats, or how long a silence had to last. They are a row per team now, edited by an owner through PUT /api/teams/{teamID}/deadman, and the sweeper runs each team against its own matchers, timeout and severity. The environment variables become the starting point rather than the setting. Every team without a configuration is seeded from them at startup, so an upgrade keeps watching exactly what it was watching, and SeedDeadmanConfigs never overwrites -- a redeploy must not put the environment's value back over an owner's edit. A team created later watches nothing until somebody says otherwise: inheriting an install-wide heartbeat would page a new team about a source it has never heard of, and a switch nobody chose is the kind that gets muted rather than fixed. A matcher string with no alertname in it is refused at the door instead of stored. Storing it would produce a switch that watches nothing silently, which is the exact failure the feature exists to prevent. NewRouter and Sweep lose their DeadmanConfig parameter -- there is no longer one answer to hand them. The type stays, because parsing a matcher string is still parsing a matcher string. The UI side: rows in the queue carry a team badge, the filter row gains a team chip per team, and "on call now" shows one card per team. All three appear only when the viewer is in more than one team -- otherwise they are the same word repeated down a list, which is noise rather than information, and the single-team install reads exactly as it did before teams existed. Verified against a live two-team server as well as in tests: the combined queue labelled by team, the team_id filter, a heartbeat that is a heartbeat in one team and an ordinary alert in another, and a new team's switches starting empty while the upgraded team keeps the environment's. Claude-Session: https://claude.ai/code/session_01RHPj4ggeFdEjKKfm4SHbD7 |
||
|
|
dc3879eca6 |
Serve a web UI for the incident queue, built for phones
Whoever is on call gets paged on a phone, and until now the only ways to
act on a page were the notification's Acknowledge button or a terminal.
Tapping the notification itself opened /api/incidents/{id}, which a
browser can only answer with a 401 in JSON. The server now serves a web
UI at / covering the incident queue, each incident's alerts and timeline
with every action on it, who is on call, the alert feed, and changing
your own password. The notification link now points at /incidents/{id}
in that UI.
It is embedded in the binary and has no build step: plain HTML, CSS and
ES modules under internal/web/static, served with an ETag per file and a
CSP that allows nothing from any other origin. That is how rd-web is
built. It avoids adding a node toolchain to the Dockerfile and the
pipeline for a page this size, and it keeps the page on the same origin
as the API, so no CORS is needed and nothing else has to be deployed.
Paths without a file extension fall back to index.html, so a deep link
survives a reload. An unknown path under /api/ still gets a JSON 404
rather than the page.
Signing in uses a username and password, because pasting a 64-character
API key into a phone at 3am is not a sign-in flow. Users have no
password until one is set through PUT /api/users/{id}/password, or
optionally at bootstrap. A user without a password is exactly where they
were before this commit and can only use API keys. A login sets an
HttpOnly, SameSite=Lax session cookie. It lasts 30 days and slides
forward while in use, so an on-call phone does not sign itself out.
Only the token's hash is stored, as for API keys.
The cookie needs a CSRF guard where a bearer header does not, because
browsers attach cookies to requests other sites make. So cookie-
authenticated requests go through Go 1.25's http.CrossOriginProtection,
and bearer requests do not. A request carrying an Authorization header
is judged on that header alone and never falls back to the cookie.
Changing a password ends every other session of that user. Changing
your own requires the current password, so a phone left signed in
cannot be used to take the account over.
Failed logins are counted per username and per client address. Ten
failures for one username in 15 minutes refuse that username for the
rest of the window, even with the right password. That makes locking
somebody out possible for anyone who knows their username. It was
accepted because the alternative is unlimited guessing, and during a
lockout the notification's Acknowledge button and API keys keep
working. The address limit reads the first X-Forwarded-For hop, since
behind the gateway RemoteAddr is Envoy. It is looser, because a whole
office behind one NAT shares it.
The Secure flag follows TERDUT_PUBLIC_URL, since TLS terminates at the
gateway and the server itself only ever sees plain HTTP. The chart
already defaults that variable to https://<hostname>.
Schedule editing, statistics and user management stay in terdut-tui for
now. The API they use is unchanged, and bearer authentication behaves
exactly as before.
|