1770e5d94569e9b74c8879cbc39726ad123fc460
5 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
91f03c21e8 |
Web: visual design pass (issue #26)
Addresses the screenshot-review feedback in #26. No framework or build step added — all of this stays within the existing plain HTML/CSS/ vanilla-JS + go:embed architecture. - Nav: re-enable the bottom tab bar that was already built and switched off (Queue/On-call/Alerts/Team + a "More" sheet for Stats/Admin/Account), replacing the hamburger on phone width. - Queue: chip counts, a scroll fade on the filter row, a "Triggered Xh ago" + severity label per row, a chevron on the team switcher so it reads as a dropdown. - On-call: collapse repeated same-person days into shift bars (week view and "your shifts" both), show the week as a date range with the ISO week number as secondary text, split "Current shift" out from "Next shifts" with "ends in Nd", a pill badge + row highlight for "you". - Incident detail: fix the actual bug behind the duplicate "acknowledged" timeline entries (acknowledgeIncident's UPDATE had no guard on the incident's current status, so acknowledging an already-acknowledged incident silently re-logged the event — now idempotent, with regression tests on both the authenticated route and the ntfy ack-button route). Relabel escalation re-pages so they don't look like the same page landing twice. Copy the primary action up near the top. Label the "···" button. Group the timeline by phase (triggered/acknowledged/resolved). Add an "at a glance" summary row (duration/severity/responsible) and collapse the group labels by default. - Team overview: reword the vague copy ("One owner." etc.) into plain labels. - Empty states: fill in missing icons/one-liners across queue, alerts, stats and the incident timeline. - CSS: fix card padding bugs, verify link contrast already passes AA, introduce a --fs-* type-scale token set and migrate the few genuinely isolated cases onto it (left sizes tied to a fixed shape, a deliberately prominent display, or a non-negotiable constraint like the iOS-zoom-prevention input size as documented exceptions rather than guess at a render this change can't see). Verified with the full fmt/lint/test/helm-lint gate, plus a live instance against the test DB with seeded incidents and schedule data to trace the on-call grouping and timeline phase-splitting logic against real API responses. 🤖 Generated with [Claude Code](https://claude.ai/code) Co-Authored-By: Claude <noreply@anthropic.com> |
||
|
|
591d5b8df0 |
Copy an incident to the clipboard as Markdown
A button in the incident header (also `y`, and "Copy incident" in the more menu) puts everything the page knows on the clipboard, for pasting into a chat or an agent prompt with no integration involved. The text carries the facts, every alert with all its labels and annotations (the page only shows summary or description), the timeline with notes in full, and the "Seen before" resolution notes. Times are ISO 8601 and users are named rather than "you", since relative and first-person wording is ambiguous once pasted elsewhere. The async clipboard API needs a secure context and this server is often reached over plain HTTP, so it falls back to execCommand. Web UI only: no endpoint or JSON shape changed, so nothing to mirror in terdut-tui. |
||
|
|
60ebb75cd2 |
Show notes from similar earlier incidents
Each incident gets a signature: the alert name plus the group labels that
say what is broken, minus the ones that only say where it ran (instance,
pod, container, ...). GET /api/incidents/{id}/similar returns resolved
incidents in the same team with the same signature that have notes.
Notes can be marked as the resolution note, "what fixed it", either with a
resolution field on resolve or pinned on a note. Those lead the similar
list, show on the incident page as "Seen before", and the triggered
notification carries the latest one.
Claude-Session: https://claude.ai/code/session_01MMados3BD1oSjevHxbmVqU
|
||
|
|
d728af53b1 |
Put a team's own settings in the web UI
Closes #17. Everything a team owner configures was API-only: escalation, integrations, dead man's switches, membership, and the rota -- which the on-call view still described as the TUI's job, and the TUI has been broken against this server since teams landed. Setting up the feature this whole line of work exists for meant using curl. A Team tab now holds all of it, one team at a time, with a picker for somebody in more than one. An owner edits; a member sees the same page without the controls, because the server refuses their writes anyway -- hiding a button is a courtesy to the reader, not the thing enforcing anything. The escalation editor holds a draft and sends the whole ladder, because the API replaces it wholesale: the levels are an order, and patching one rung leaves the numbering of the others undecided. Adding a level defaults to five minutes and the rota, which is the shape almost every ladder starts as. An integration key is returned exactly once, so creating one opens a panel that says so, shows the URL large with a copy button, and renders the Alertmanager receiver snippet with the URL already in it -- the next thing anybody does with that key is paste it into a config. The panel stays until it is dismissed rather than disappearing on the next re-render. The incident view gains where an incident is on the ladder and when the next page is due, which is the question somebody looking at an unacknowledged incident actually has. The API carries it: the incident payload now includes escalation_level and escalation_due_at, the latter computed in the incident SELECT by joining the level's timeout, so a list costs no extra queries. Verified against a live server by making every call the page makes, including the writes: the six reads the Team tab issues, a two-level ladder saved and read back, an integration created and its key returned once, three days of rota assigned, switches set, and an incident showing level 1 with a due time five minutes out. Claude-Session: https://claude.ai/code/session_01RHPj4ggeFdEjKKfm4SHbD7 |
||
|
|
dc3879eca6 |
Serve a web UI for the incident queue, built for phones
Whoever is on call gets paged on a phone, and until now the only ways to
act on a page were the notification's Acknowledge button or a terminal.
Tapping the notification itself opened /api/incidents/{id}, which a
browser can only answer with a 401 in JSON. The server now serves a web
UI at / covering the incident queue, each incident's alerts and timeline
with every action on it, who is on call, the alert feed, and changing
your own password. The notification link now points at /incidents/{id}
in that UI.
It is embedded in the binary and has no build step: plain HTML, CSS and
ES modules under internal/web/static, served with an ETag per file and a
CSP that allows nothing from any other origin. That is how rd-web is
built. It avoids adding a node toolchain to the Dockerfile and the
pipeline for a page this size, and it keeps the page on the same origin
as the API, so no CORS is needed and nothing else has to be deployed.
Paths without a file extension fall back to index.html, so a deep link
survives a reload. An unknown path under /api/ still gets a JSON 404
rather than the page.
Signing in uses a username and password, because pasting a 64-character
API key into a phone at 3am is not a sign-in flow. Users have no
password until one is set through PUT /api/users/{id}/password, or
optionally at bootstrap. A user without a password is exactly where they
were before this commit and can only use API keys. A login sets an
HttpOnly, SameSite=Lax session cookie. It lasts 30 days and slides
forward while in use, so an on-call phone does not sign itself out.
Only the token's hash is stored, as for API keys.
The cookie needs a CSRF guard where a bearer header does not, because
browsers attach cookies to requests other sites make. So cookie-
authenticated requests go through Go 1.25's http.CrossOriginProtection,
and bearer requests do not. A request carrying an Authorization header
is judged on that header alone and never falls back to the cookie.
Changing a password ends every other session of that user. Changing
your own requires the current password, so a phone left signed in
cannot be used to take the account over.
Failed logins are counted per username and per client address. Ten
failures for one username in 15 minutes refuse that username for the
rest of the window, even with the right password. That makes locking
somebody out possible for anyone who knows their username. It was
accepted because the alternative is unlimited guessing, and during a
lockout the notification's Acknowledge button and API keys keep
working. The address limit reads the first X-Forwarded-For hop, since
behind the gateway RemoteAddr is Envoy. It is looser, because a whole
office behind one NAT shares it.
The Secure flag follows TERDUT_PUBLIC_URL, since TLS terminates at the
gateway and the server itself only ever sees plain HTTP. The chart
already defaults that variable to https://<hostname>.
Schedule editing, statistics and user management stay in terdut-tui for
now. The API they use is unchanged, and bearer authentication behaves
exactly as before.
|