Add an admin page, and move the behaviour settings into the database
CI / chart (pull_request) Successful in 1s
CI / security (pull_request) Successful in 13s
CI / test (pull_request) Successful in 2m1s

Closes #5. Three of the server's tunables were environment variables,
which meant changing how long an incident waits before being paged again
required editing a chart, merging it and waiting for a reconcile. They
are behaviour rather than infrastructure, and the difference is who needs
to change them and how often.

The split is by who owns the value. What stays in the environment is
where the server is plugged in: the listen address, the DSN, the ntfy URL
and token, the public URL. Those are needed before the database is open
and two of them are credentials -- the settings endpoint reports that
ntfy is configured and that a token is set, and never what either is.

What moves is how it behaves: the notify repeat interval, the stale
window and the archive window. The environment variable becomes the seed
rather than the setting, written once on first start and never
overwritten, so a redeploy cannot put a chart's default back over an
administrator's edit -- the rule the per-team dead man's switches already
follow. The loops read the current value per tick, so a change at 02:00
is obeyed at 02:00.

Key/value rather than a column per knob: #6 and #7 will both add
settings, and a table shaped one-column-per-setting needs a migration for
each. The cost is that values are text and the accessor has to say what
type it wanted, which settings.go does in one place. Unknown keys are
refused rather than stored -- a typo that wrote notify_repeat_second
would otherwise sit in the table looking like configuration and doing
nothing -- and each value has bounds loose enough to catch a slipped
decimal point without having an opinion about anybody's rota.

Disabling an account is new, and is not deleting one. Deleting a user
nulls acknowledged_by and assigned_to, which quietly rewrites who did
what during an incident months after the fact. A disabled user cannot
authenticate by either credential, loses their sessions immediately, and
stays the name on every acknowledgement they made. The check is part of
the lookup in serveAs rather than a test afterwards, so there is no path
where the row is loaded and the flag is then forgotten.

The page itself is a fourth tab, shown only to an administrator and only
as a courtesy: every endpoint under it is refused with 403 regardless, so
somebody who types /admin gets an explanation rather than a blank screen.
It lists teams with their size and open-incident count, users with their
flags, and the settings with their bounds -- plus the environment half,
read-only, so somebody hunting for the ntfy URL learns where it lives
instead of concluding the server has none.

Delete is disabled rather than offered-and-refused for a team with open
incidents, and neither admin action is offered on your own account, since
the server refuses both.

Claude-Session: https://claude.ai/code/session_01RHPj4ggeFdEjKKfm4SHbD7
This commit is contained in:
Niklas Ye
2026-09-20 18:23:46 +02:00
parent 303e7a3365
commit b0a02c010b
19 changed files with 1110 additions and 15 deletions
+346
View File
@@ -0,0 +1,346 @@
package api
import (
"context"
"database/sql"
"errors"
"net/http"
"strconv"
"time"
"git.ryuvia.com/niklas/terdut-server/internal/config"
"github.com/go-chi/chi/v5"
)
// The settings an administrator can change at runtime. Each is behaviour rather
// than infrastructure: what the server does, not where it is plugged in.
//
// The values are seconds, stored as text. A duration string would be friendlier
// to read in psql and worse everywhere else — it can be stored unparseable, and
// then the question is what a background loop should do at 02:00 with a
// tuning knob it cannot understand.
const (
SettingNotifyRepeat = "notify_repeat_seconds"
SettingStaleAfter = "stale_after_seconds"
SettingArchiveAfter = "archive_after_seconds"
)
// settingBounds keeps an edit from producing a server that cannot work. The
// ceilings are loose — they exist to catch a slipped decimal point, not to have
// an opinion about anybody's rota.
var settingBounds = map[string]struct {
min, max time.Duration
label string
}{
SettingNotifyRepeat: {0, 24 * time.Hour, "how long an incident may sit unacknowledged before it is paged again; 0 disables reminders"},
SettingStaleAfter: {5 * time.Minute, 30 * 24 * time.Hour, "how long a firing alert may go without a refreshing webhook before the sweeper resolves it"},
SettingArchiveAfter: {time.Minute, 365 * 24 * time.Hour, "how long a resolved alert or incident stays in the default list"},
}
// Settings reads the runtime configuration. It holds no cache: the readers are
// two background loops that tick every 30 seconds and 15 minutes, and handlers
// that run once per request, so a query each time costs nothing measurable and
// means an administrator's change takes effect on the next tick rather than at
// the next restart.
type Settings struct{ db *sql.DB }
// NewSettings returns a reader over db.
func NewSettings(db *sql.DB) *Settings { return &Settings{db: db} }
// Duration reads one setting, falling back to def when the row is missing or
// unreadable. A tuning knob is never worth failing a sweep over: the fallback
// is the value the server started with.
func (s *Settings) Duration(ctx context.Context, key string, def time.Duration) time.Duration {
var raw string
err := s.db.QueryRowContext(ctx, "SELECT value FROM settings WHERE key = $1", key).Scan(&raw)
if err != nil {
return def
}
secs, err := strconv.ParseInt(raw, 10, 64)
if err != nil {
return def
}
return time.Duration(secs) * time.Second
}
// SeedSettings writes each key from the server's environment configuration,
// once. Never overwrites: after the first start the database owns these, and a
// redeploy must not put a chart's default back over an administrator's edit —
// the same rule as the per-team dead man's switches.
func SeedSettings(ctx context.Context, db *sql.DB, cfg config.Config) error {
seeds := map[string]time.Duration{
SettingNotifyRepeat: cfg.NotifyRepeat,
SettingStaleAfter: cfg.StaleAfter,
SettingArchiveAfter: cfg.ArchiveAfter,
}
for key, d := range seeds {
if _, err := db.ExecContext(ctx, `
INSERT INTO settings (key, value) VALUES ($1, $2)
ON CONFLICT (key) DO NOTHING`,
key, strconv.FormatInt(int64(d.Seconds()), 10)); err != nil {
return err
}
}
return nil
}
// settingsResponse is what the admin page renders. The environment half is
// included and marked read-only, so somebody looking for the ntfy URL finds out
// where it lives rather than concluding the server does not have one.
type settingsResponse struct {
Editable map[string]settingValue `json:"editable"`
FromEnv map[string]string `json:"from_env"`
}
type settingValue struct {
Seconds int64 `json:"seconds"`
Description string `json:"description"`
MinSeconds int64 `json:"min_seconds"`
MaxSeconds int64 `json:"max_seconds"`
}
func handleGetSettings(db *sql.DB, cfg config.Config) http.HandlerFunc {
settings := NewSettings(db)
return func(w http.ResponseWriter, r *http.Request) {
out := settingsResponse{
Editable: map[string]settingValue{},
FromEnv: map[string]string{
// Never the ntfy token or the DSN: both are credentials, and an
// admin page that renders them turns a browser tab into a place
// they leak from.
"ntfy_url": cfg.NtfyURL,
"ntfy_configured": strconv.FormatBool(cfg.NtfyURL != ""),
"ntfy_token_set": strconv.FormatBool(cfg.NtfyToken != ""),
"public_url": cfg.PublicURL,
"listen_address": cfg.Addr,
},
}
for key, b := range settingBounds {
def := map[string]time.Duration{
SettingNotifyRepeat: cfg.NotifyRepeat,
SettingStaleAfter: cfg.StaleAfter,
SettingArchiveAfter: cfg.ArchiveAfter,
}[key]
out.Editable[key] = settingValue{
Seconds: int64(settings.Duration(r.Context(), key, def).Seconds()),
Description: b.label,
MinSeconds: int64(b.min.Seconds()),
MaxSeconds: int64(b.max.Seconds()),
}
}
respond(w, http.StatusOK, out)
}
}
// handleSetSettings changes one or more settings. Unknown keys are refused
// rather than stored: a typo that writes notify_repeat_second would otherwise
// sit in the table looking like configuration and doing nothing.
func handleSetSettings(db *sql.DB) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
var req map[string]int64
if err := decodeJSON(r, &req); err != nil {
respond(w, http.StatusBadRequest, errResp("invalid request body"))
return
}
if len(req) == 0 {
respond(w, http.StatusBadRequest, errResp("no settings given"))
return
}
for key, secs := range req {
b, known := settingBounds[key]
if !known {
respond(w, http.StatusBadRequest, errResp("unknown setting: "+key))
return
}
d := time.Duration(secs) * time.Second
if d < b.min || d > b.max {
respond(w, http.StatusBadRequest, errResp(
key+" must be between "+b.min.String()+" and "+b.max.String()))
return
}
}
tx, err := db.BeginTx(r.Context(), nil)
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
defer tx.Rollback() //nolint:errcheck
for key, secs := range req {
if _, err := tx.ExecContext(r.Context(), `
INSERT INTO settings (key, value, updated_at)
VALUES ($1, $2, `+nowEpoch+`)
ON CONFLICT (key) DO UPDATE SET
value = excluded.value, updated_at = excluded.updated_at`,
key, strconv.FormatInt(secs, 10)); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
}
if err := tx.Commit(); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
w.WriteHeader(http.StatusNoContent)
}
}
// handleAdminListTeams lists every team on the server, with its size. The
// ordinary /api/teams answers "what am I in"; this one answers "what exists",
// which only an administrator may ask.
func handleAdminListTeams(db *sql.DB) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
rows, err := db.QueryContext(r.Context(), `
SELECT t.id, t.name, t.created_at,
(SELECT COUNT(*) FROM team_members m WHERE m.team_id = t.id),
(SELECT COUNT(*) FROM incidents i
WHERE i.team_id = t.id AND i.resolved_at IS NULL)
FROM teams t
ORDER BY t.name`)
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
defer rows.Close()
type adminTeam struct {
ID int64 `json:"id"`
Name string `json:"name"`
CreatedAt time.Time `json:"created_at"`
Members int64 `json:"members"`
OpenIncidents int64 `json:"open_incidents"`
}
teams := []adminTeam{}
for rows.Next() {
var t adminTeam
var created int64
if err := rows.Scan(&t.ID, &t.Name, &created, &t.Members, &t.OpenIncidents); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
t.CreatedAt = time.Unix(created, 0).UTC()
teams = append(teams, t)
}
if err := rows.Err(); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
respond(w, http.StatusOK, teams)
}
}
// handleRenameTeam renames a team. An owner's job, and an administrator's when
// a team has nobody left to do it.
func handleRenameTeam(db *sql.DB) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
teamID, ok := teamParam(w, r)
if !ok {
return
}
if !requireTeamOwner(w, r, teamID) {
return
}
var req struct {
Name string `json:"name"`
}
if err := decodeJSON(r, &req); err != nil || req.Name == "" {
respond(w, http.StatusBadRequest, errResp("name is required"))
return
}
res, err := db.ExecContext(r.Context(),
"UPDATE teams SET name = $1 WHERE id = $2", req.Name, teamID)
if err != nil {
if isUniqueViolation(err) {
respond(w, http.StatusConflict, errResp("a team with that name already exists"))
return
}
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
if n, _ := res.RowsAffected(); n == 0 {
respond(w, http.StatusNotFound, errResp("not found"))
return
}
w.WriteHeader(http.StatusNoContent)
}
}
// handleSetUserDisabled takes an account out of use, or puts it back.
//
// Not a delete: the person's acknowledgements, assignments and timeline entries
// stay attached to them. Deleting a user nulls those columns, which rewrites
// what happened during an incident months after the fact.
func handleSetUserDisabled(db *sql.DB) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
id, err := strconv.ParseInt(chi.URLParam(r, "id"), 10, 64)
if err != nil {
respond(w, http.StatusBadRequest, errResp("invalid user id"))
return
}
var req struct {
Disabled *bool `json:"disabled"`
}
if err := decodeJSON(r, &req); err != nil || req.Disabled == nil {
respond(w, http.StatusBadRequest, errResp("disabled is required"))
return
}
if *req.Disabled {
caller, _ := userFromContext(r.Context())
if caller.ID == id {
respond(w, http.StatusConflict, errResp("cannot disable your own account"))
return
}
last, err := isLastAdmin(r.Context(), db, id)
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
if last {
respond(w, http.StatusConflict, errResp("cannot disable the last administrator"))
return
}
}
var res sql.Result
if *req.Disabled {
res, err = db.ExecContext(r.Context(),
"UPDATE users SET disabled_at = "+nowEpoch+" WHERE id = $1 AND disabled_at IS NULL", id)
} else {
res, err = db.ExecContext(r.Context(),
"UPDATE users SET disabled_at = NULL WHERE id = $1", id)
}
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
if n, _ := res.RowsAffected(); n == 0 {
// Either no such user, or already in the state asked for. The
// second is not a failure, so check which before answering.
var exists int
if err := db.QueryRowContext(r.Context(),
"SELECT 1 FROM users WHERE id = $1", id).Scan(&exists); errors.Is(err, sql.ErrNoRows) {
respond(w, http.StatusNotFound, errResp("user not found"))
return
}
}
// Signing back in is the only way to use a re-enabled account, and a
// disabled one must not keep a live session.
if *req.Disabled {
db.ExecContext(r.Context(), "DELETE FROM sessions WHERE user_id = $1", id) //nolint:errcheck
}
user, err := fetchUser(r.Context(), db, id)
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
respond(w, http.StatusOK, user)
}
}