1770e5d94569e9b74c8879cbc39726ad123fc460
3 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
91f03c21e8 |
Web: visual design pass (issue #26)
Addresses the screenshot-review feedback in #26. No framework or build step added — all of this stays within the existing plain HTML/CSS/ vanilla-JS + go:embed architecture. - Nav: re-enable the bottom tab bar that was already built and switched off (Queue/On-call/Alerts/Team + a "More" sheet for Stats/Admin/Account), replacing the hamburger on phone width. - Queue: chip counts, a scroll fade on the filter row, a "Triggered Xh ago" + severity label per row, a chevron on the team switcher so it reads as a dropdown. - On-call: collapse repeated same-person days into shift bars (week view and "your shifts" both), show the week as a date range with the ISO week number as secondary text, split "Current shift" out from "Next shifts" with "ends in Nd", a pill badge + row highlight for "you". - Incident detail: fix the actual bug behind the duplicate "acknowledged" timeline entries (acknowledgeIncident's UPDATE had no guard on the incident's current status, so acknowledging an already-acknowledged incident silently re-logged the event — now idempotent, with regression tests on both the authenticated route and the ntfy ack-button route). Relabel escalation re-pages so they don't look like the same page landing twice. Copy the primary action up near the top. Label the "···" button. Group the timeline by phase (triggered/acknowledged/resolved). Add an "at a glance" summary row (duration/severity/responsible) and collapse the group labels by default. - Team overview: reword the vague copy ("One owner." etc.) into plain labels. - Empty states: fill in missing icons/one-liners across queue, alerts, stats and the incident timeline. - CSS: fix card padding bugs, verify link contrast already passes AA, introduce a --fs-* type-scale token set and migrate the few genuinely isolated cases onto it (left sizes tied to a fixed shape, a deliberately prominent display, or a non-negotiable constraint like the iOS-zoom-prevention input size as documented exceptions rather than guess at a render this change can't see). Verified with the full fmt/lint/test/helm-lint gate, plus a live instance against the test DB with seeded incidents and schedule data to trace the on-call grouping and timeline phase-splitting logic against real API responses. 🤖 Generated with [Claude Code](https://claude.ai/code) Co-Authored-By: Claude <noreply@anthropic.com> |
||
|
|
dc39e3a5d3 |
Move the database to Postgres, before teams need the schema
First step of #1, and it goes first for one reason: #4 adds a team_id to nearly every table, and doing that twice -- once for SQLite, once for Postgres -- is work nobody gets paid for. The teams migrations now only have to be written against one database. The ten SQLite migrations are replaced by a single Postgres baseline rather than ported one by one. They were incremental in a way that has no value on a fresh install: 004 adds columns 008 drops again, and 008's backfill rewrites data a Postgres database never had. The history stays in git; the schema they add up to is now 001_baseline.sql. Timestamps stay BIGINT unix seconds and are NOT converted to timestamptz. Everything in Go already speaks epochs, so converting would have been a second, larger change riding along inside this one. It is worth doing on its own. The JSON columns did move to jsonb, because #4 will want to filter and index on labels. Most of the port is mechanical -- 170 placeholders from ? to $1 -- but four things needed more than a search and replace: * Dynamically built WHERE clauses cannot keep their numbering straight by hand, so they hand out placeholders through sqlArgs instead. A filter can now be added or reordered without renumbering anything. * SUM(resolved_at IS NULL) was SQLite counting a boolean as 0 or 1. Postgres has no sum(boolean), and this was breaking every dead man's switch -- silently, since the sweeper only logs. Now COUNT(*) FILTER. * unixepoch() became FLOOR(EXTRACT(EPOCH FROM now()))::bigint. The FLOOR is load-bearing: a bare cast rounds half up, so a row written at .6 of a second claimed a timestamp a second in the future and disagreed with the time.Now().Unix() the Go side stamps. * The unique-violation check matched SQLite's error text. It matches SQLSTATE 23505 now, so a renamed constraint cannot turn a 409 back into a 500. Tests need a real Postgres, because there is no in-memory Postgres the way there was an in-memory SQLite. Each test gets its own schema on a shared server -- cheaper than a database each, and still isolated. TERDUT_TEST_DSN says where it is; `make test-db` starts one locally and ci.yaml runs one as a service container. An unset DSN fails the suite rather than skipping it: a run that quietly tests nothing is worse than one that does not run. TestMigration_BackfillCarriesAckAndComments is deleted along with the migrations it replayed. What it protected -- an upgrade not losing acknowledgements and comments -- now belongs to scripts/sqlite-to-postgres.go, which is build-tagged so the SQLite driver stays out of the server binary. Both are meant to be deleted once this install has migrated. The chart loses the PVC, the data volume and the python backup sidecar, and requires database.dsnSecret.name: it provisions no database and cannot guess where the credentials live, so a render without it is meant to fail. Backups move to where Postgres actually runs. The other half of that -- the postgresql CR, the k8up pg_dump annotation and the network policy -- is a change to the wrapper chart in Ryuvia/charts and is not in here. Verified rather than assumed: the gate is green with -race against Postgres 17, govulncheck and gitleaks are clean, and the migration script was run end to end against a SQLite database built at the old schema and seeded in every table. Ids survive, so incidents keep their numbers and every foreign key still points where it did; the identity sequences are moved past the copied ids, and a webhook after the migration opened incident 12 rather than colliding at 1. |
||
|
|
bc285799d1 |
Page the on-call person when an incident opens
An incident opened, got assigned to whoever held today's schedule entry,
and then sat there silently until somebody thought to look. The schedule
and the incident model were both built; nothing reached the person
holding the pager.
Notifications go out through ntfy, over plain HTTP with no new
dependencies. Delivery is an outbox rather than an inline call: the pool
is limited to a single connection, so a POST made while holding the
webhook's transaction would stall every other request behind it. The
webhook inserts a row and a notifier goroutine sends it within a tick,
retrying with exponential backoff.
Only opening an incident has to resolve a topic from scratch. Reminders
and all-clears reuse whatever that first notification chose, which keeps
configuration out of resolveIfSettled and gives the right rule for free:
you only hear that something resolved if you were told it started.
Each push carries an Acknowledge button, because the useful thing to do
at 3am is stop the pager without unlocking anything. It POSTs to an
unauthenticated /api/notify/ack/{token} — a notification body lives on
the ntfy server and in the device cache, so a real API key must never
appear in one. The token is minted per delivery, scoped to one incident
and one action, and expires in a day.
Reminders repeat until the incident stops being untouched. The stop
conditions are the states that already mean somebody has it: acknowledged,
snoozed, resolved, archived. Snooze is the mute button, so there is no
separate reminder cap.
Notifications sent to the fallback topic carry no Acknowledge button. The
topic is shared, and a button on it would let any subscriber acknowledge
as somebody else.
|