Add an admin page, and move the behaviour settings into the database #15

Merged
niklas merged 1 commits from admin-settings into main 2026-09-20 16:26:27 +00:00
Owner

Closes #5. Fourth of the six sub-issues in #1.

Three tunables were environment variables, so changing how long an incident waits before being paged again meant editing a chart, merging it, and waiting for a reconcile.

The split is by who owns the value

Stays in the environment — where the server is plugged in: listen address, DSN, ntfy URL and token, public URL. Needed before the database is open, and two are credentials. The settings endpoint reports that ntfy is configured and that a token is set, never what either is.

Moves to the database — how it behaves: notify repeat, stale window, archive window. The environment variable becomes the seed rather than the setting: written once on first start, never overwritten, so a redeploy can't put a chart default back over an administrator's edit — the rule the per-team dead-man switches already follow. The loops read the current value per tick, so a change at 02:00 is obeyed at 02:00.

Key/value rather than a column per knob, because #6 and #7 will both add settings. Unknown keys are refused rather than stored — a typo writing notify_repeat_second would otherwise sit in the table looking like configuration and doing nothing — and each value has bounds loose enough to catch a slipped decimal point without having an opinion about anybody's rota.

Disabling an account is not deleting one

Deleting a user nulls acknowledged_by and assigned_to, quietly rewriting who did what during an incident months later. A disabled user can't authenticate by either credential, loses their sessions immediately, and stays the name on every acknowledgement they made.

The check is part of the lookup in serveAs, not a test afterwards, so there's no path where the row loads and the flag is then forgotten.

The page

A fourth tab, shown only to an administrator — as a courtesy, not a gate: every endpoint under it is refused with 403 regardless, so typing /admin gets an explanation rather than a blank screen. Teams with size and open-incident counts, users with their flags, settings with their bounds, plus the environment half read-only so somebody hunting the ntfy URL learns where it lives.

Delete is disabled rather than offered-and-refused for a team with open incidents, and neither admin action appears on your own account, since the server refuses both.

Verified

make fmt lint test helm-lint green with -race against Postgres 17, plus nine new tests — including the one that matters most: a setting changed through the API takes effect on the very next sweep, with an alert that survives one pass under the seeded six-hour window and expires on the next after the window is shortened to an hour. Also checked that no credential appears in the settings response, and that disabling an account kills its key while the incident it acknowledged still names it.

Driven against a live server too: seeded from that server's own env (TERDUT_STALE_AFTER=26h → 93600s), changed through the API, read back.

https://claude.ai/code/session_01RHPj4ggeFdEjKKfm4SHbD7

Closes #5. Fourth of the six sub-issues in #1. Three tunables were environment variables, so changing how long an incident waits before being paged again meant editing a chart, merging it, and waiting for a reconcile. ### The split is by who owns the value **Stays in the environment** — where the server is plugged in: listen address, DSN, ntfy URL and token, public URL. Needed before the database is open, and two are credentials. The settings endpoint reports *that* ntfy is configured and *that* a token is set, never what either is. **Moves to the database** — how it behaves: notify repeat, stale window, archive window. The environment variable becomes the **seed** rather than the setting: written once on first start, never overwritten, so a redeploy can't put a chart default back over an administrator's edit — the rule the per-team dead-man switches already follow. The loops read the current value per tick, so a change at 02:00 is obeyed at 02:00. Key/value rather than a column per knob, because #6 and #7 will both add settings. Unknown keys are refused rather than stored — a typo writing `notify_repeat_second` would otherwise sit in the table looking like configuration and doing nothing — and each value has bounds loose enough to catch a slipped decimal point without having an opinion about anybody's rota. ### Disabling an account is not deleting one Deleting a user nulls `acknowledged_by` and `assigned_to`, quietly rewriting who did what during an incident months later. A disabled user can't authenticate by either credential, loses their sessions immediately, and stays the name on every acknowledgement they made. The check is part of the lookup in `serveAs`, not a test afterwards, so there's no path where the row loads and the flag is then forgotten. ### The page A fourth tab, shown only to an administrator — as a courtesy, not a gate: every endpoint under it is refused with 403 regardless, so typing `/admin` gets an explanation rather than a blank screen. Teams with size and open-incident counts, users with their flags, settings with their bounds, plus the environment half read-only so somebody hunting the ntfy URL learns where it lives. Delete is *disabled* rather than offered-and-refused for a team with open incidents, and neither admin action appears on your own account, since the server refuses both. ### Verified `make fmt lint test helm-lint` green with `-race` against Postgres 17, plus nine new tests — including the one that matters most: **a setting changed through the API takes effect on the very next sweep**, with an alert that survives one pass under the seeded six-hour window and expires on the next after the window is shortened to an hour. Also checked that no credential appears in the settings response, and that disabling an account kills its key while the incident it acknowledged still names it. Driven against a live server too: seeded from that server's own env (`TERDUT_STALE_AFTER=26h` → 93600s), changed through the API, read back. https://claude.ai/code/session_01RHPj4ggeFdEjKKfm4SHbD7
niklas added 1 commit 2026-09-20 16:24:04 +00:00
Add an admin page, and move the behaviour settings into the database
CI / chart (pull_request) Successful in 1s
CI / security (pull_request) Successful in 13s
CI / test (pull_request) Successful in 2m1s
b0a02c010b
Closes #5. Three of the server's tunables were environment variables,
which meant changing how long an incident waits before being paged again
required editing a chart, merging it and waiting for a reconcile. They
are behaviour rather than infrastructure, and the difference is who needs
to change them and how often.

The split is by who owns the value. What stays in the environment is
where the server is plugged in: the listen address, the DSN, the ntfy URL
and token, the public URL. Those are needed before the database is open
and two of them are credentials -- the settings endpoint reports that
ntfy is configured and that a token is set, and never what either is.

What moves is how it behaves: the notify repeat interval, the stale
window and the archive window. The environment variable becomes the seed
rather than the setting, written once on first start and never
overwritten, so a redeploy cannot put a chart's default back over an
administrator's edit -- the rule the per-team dead man's switches already
follow. The loops read the current value per tick, so a change at 02:00
is obeyed at 02:00.

Key/value rather than a column per knob: #6 and #7 will both add
settings, and a table shaped one-column-per-setting needs a migration for
each. The cost is that values are text and the accessor has to say what
type it wanted, which settings.go does in one place. Unknown keys are
refused rather than stored -- a typo that wrote notify_repeat_second
would otherwise sit in the table looking like configuration and doing
nothing -- and each value has bounds loose enough to catch a slipped
decimal point without having an opinion about anybody's rota.

Disabling an account is new, and is not deleting one. Deleting a user
nulls acknowledged_by and assigned_to, which quietly rewrites who did
what during an incident months after the fact. A disabled user cannot
authenticate by either credential, loses their sessions immediately, and
stays the name on every acknowledgement they made. The check is part of
the lookup in serveAs rather than a test afterwards, so there is no path
where the row is loaded and the flag is then forgotten.

The page itself is a fourth tab, shown only to an administrator and only
as a courtesy: every endpoint under it is refused with 403 regardless, so
somebody who types /admin gets an explanation rather than a blank screen.
It lists teams with their size and open-incident count, users with their
flags, and the settings with their bounds -- plus the environment half,
read-only, so somebody hunting for the ntfy URL learns where it lives
instead of concluding the server has none.

Delete is disabled rather than offered-and-refused for a team with open
incidents, and neither admin action is offered on your own account, since
the server refuses both.

Claude-Session: https://claude.ai/code/session_01RHPj4ggeFdEjKKfm4SHbD7
niklas merged commit 94d23a593c into main 2026-09-20 16:26:27 +00:00
Sign in to join this conversation.