Commit Graph

92 Commits

Author SHA1 Message Date
Niklas Ye 67d68ce058 Show the rota as a month rather than a list of dates
The Team tab printed the next thirty days as thirty rows of date, name and
a Clear button. That is a rota spelled out one day at a time, and it is the
one shape the question cannot be read in: what anybody wants from a rota is
who holds which stretch, and thirty names down a column hides a handover
between two rows that look the same. It was also the longest thing on the
page by a wide margin, so the escalation ladder and the alert sources sat
below a screen of dates.

It is a month now, Monday to Sunday, one coloured initial per day. A shift
becomes a run of one colour, which is the shape the answer actually has; a
gap becomes a hole you can see. The legend underneath says whose colour is
whose, and one line says how many days are left uncovered, counting only
from today -- an empty Tuesday last week is history, not a hole somebody
still has to fill.

Laid out like the on-call page's week, deliberately: heading and arrows
outside the card, days inside it. It is the same rota, and two pages
showing it two ways would be two things to learn.

Colours come from a person's place in the member list, so they hold still
as you page between months, and six of them repeat -- the initial inside
still tells two people apart, and a legend that has to explain nine hues is
not a legend. They are not the severity palette: nothing on a rota is
critical, and a red Thursday would read as one. --teal and --pink are new
in both themes for the two the palette was short.

The per-row Clear button had nowhere left to live, so a day opens the sheet
the app already uses for confirmations: who holds it, a picker, Assign and
Clear. That assign sends replace=true where the range form still asks
first, and the difference is the point -- the sheet has just named whoever
holds the day, so taking it from them is the thing that was asked for
rather than something to warn about. The range form is unchanged and folded
into a details, since filling a whole shift is what it is for; it opens on
the month above it rather than on today, so paging to March to fill March
does not hand you September.

The server is untouched. The month drawn is the month fetched -- the grid's
Monday overhang and its trailing days are real days and are fetched with
it -- so paging is one GET /api/teams/{id}/schedule per month with from and
to, where it used to be one fixed thirty-day window. No new endpoint, no
change to what the API returns, and terdut-tui is unaffected.

Nobody has looked at this in a browser either. What is checked is the
rendering: team.js's own refresh() was run against a stub fetch and a
pocket DOM for September 2026, and it produces 35 cells for a month whose
1st is a Tuesday, the right from/to on the schedule call, today marked on
the 22nd, three people in the legend with "you" on the viewer, the gap
count over a five-day hole, and -- as a member rather than an owner -- the
same grid as plain divs with no sheet and no range form. How it looks at
phone width, and whether the six colours hold up in dark mode, are not
checked.

Claude-Session: https://claude.ai/code/session_01RHPj4ggeFdEjKKfm4SHbD7
2026-09-22 09:13:36 +02:00
Niklas Ye a6fa673e08 Give every team a page of its own
The Admin tab's team list was growing controls the way the user list did
before ac9af8e: a Rename button behind window.prompt, a Delete beside it,
and -- on the Users page, of all places -- an invite form with a team
picker in front of it. The picker was the admission that an invite is a
fact about a team rather than about the server, and a prompt() is the
wrong place to read a 409 about a name already taken.

So a team is now a subject with a page, at /admin/teams/{id}, the mirror
of /admin/users/{id}: when it was created, how many are in it and how
much is open, a field to rename it, the members with their roles, the
invites into it, and deletion. The list goes back to being a list, and
the name in it is the way in.

The member list is the one thing there that needed a new endpoint.
GET /api/teams/{id}/members is requireTeamMember and answers 404 to an
administrator who is not in the team, and that stays exactly as it is:
member means membership and nothing else. Reading a team's shape is a
different question from reading its work, so it gets an endpoint of its
own under AdminOnly -- GET /api/admin/teams/{id}, returning
{"team", "members"} -- rather than an exception carved into that rule. It
is a wrapper and not a team with the members hung off it, because
"members" already means a count on the list endpoint and one name must
not be a number in one answer and an array in the next. The query and its
ordering are copied from handleListTeamMembers so the two answers to "who
is in this team" cannot disagree.

An administrator still sees none of that team's incidents, alerts or
rota. Nothing about what the flag may do changed; it could already rename
and delete any team, and staff one it is not in.

Rename now trims what it is given, as creation has always trimmed. Before
this, " " was a legal name to rename a team to but not to create one
with, which is one rule stated twice and applied once.

Nobody has looked at this in a browser, the caveat ac9af8e and 07914d5
both carried. What is checked is the wiring: admin_test.go covers the new
endpoint for an administrator outside the team, the 404 the member-only
endpoint still gives that same administrator, the 403 for a member who is
not one, a 404 for a team that does not exist, a 400 for an id that is not
a number, and the trim; the module graph evaluates at /admin/teams/{id},
and the server serves index.html there, so a reload survives.

Claude-Session: https://claude.ai/code/session_01RHPj4ggeFdEjKKfm4SHbD7
2026-09-22 09:13:13 +02:00
Niklas Ye ee22eb000c Set the chart's placeholder version to 0.17.0
CI / chart (push) Successful in 1s
CI / security (push) Successful in 13s
CI / test (push) Successful in 2m31s
Release / test (push) Successful in 8s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 19s
Release / image (push) Successful in 53s
Release / scan-image (push) Successful in 3s
Cosmetic, and done anyway for the same reason as 7b9a337, 828cf87 and
8869ac8 before it: .gitea/workflows/release.yaml passes --version and
--app-version to `helm package` from the git tag, so neither line decides
anything about what is published. Being read is all they do, and a tree
heading for v0.17.0 that says 0.16.1 tells its reader something false.

appVersion keeps the v, per APPVERSION_PREFIX in .release.conf, and
image.tag in values.yaml stays "latest" -- that one is what a local
`helm install ./charts/terdut-server` actually pulls.

Claude-Session: https://claude.ai/code/session_01RHPj4ggeFdEjKKfm4SHbD7
v0.17.0
2026-09-21 21:11:29 +02:00
Niklas Ye 07914d5cdb Give the Admin tab sub-sections of its own
Administration was one scrolling page with three cards on it: the teams,
the people, and the settings. There was no way to link somebody to the
settings, no way back to the top of the user list but scrolling, and the
poll loop refetched all three endpoints every tick however little of the
page you were looking at.

Each is now a route -- /admin/teams, /admin/users, /admin/settings --
reached from a strip across the top, with /admin an overview that says
how many of each there are. The three cards themselves are untouched;
they are simply rendered one at a time, so a tab fetches only what it
shows. The Users page is the exception and fetches the teams too, since
its invite form has to offer a team to invite somebody into.

The strip is ordinary links rather than chips. Chips filter what a page
already shows, here and in the queue, and these four go somewhere: the
browser's Back walks them, a reload lands where you were, and the click
is intercepted by the same handler every other link in the app uses.
The current one is marked with aria-current="page", the convention the
tab bar has used since it existed, so the state lives on the attribute
and not in a class.

admin.js owns the table of the four routes, because it also builds the
strip that links to them; app.js parses against that table rather than
keeping a second list to drift from it. Adding a fifth sub-section is
one line.

The bottom tab bar still has six items. 56b8191 made it count-agnostic
when Admin arriving pushed it past four, and the note there records that
six at 420px already leaves 55-65px each -- so the sub-sections went
inside the Admin page rather than beside it.

A person's page keeps its own route at /admin/users/{id}; its back link
now returns to the user list rather than to the top of everything.

Nobody has looked at this in a browser, the same caveat ac9af8e carried.
What is checked is the wiring: the module graph evaluates at every admin
URL, all four tabs render against live server responses with one
aria-current each and the fetches the table above describes, the
non-administrator branch still refuses without fetching, and the server
serves index.html for each new path so a reload survives. The strip's
appearance at phone and desktop width is not checked.

Claude-Session: https://claude.ai/code/session_01RHPj4ggeFdEjKKfm4SHbD7
2026-09-21 21:11:08 +02:00
Niklas Ye 7b9a337d25 Set the chart's placeholder version to 0.16.1
CI / chart (push) Successful in 1s
CI / security (push) Successful in 16s
CI / test (push) Successful in 2m33s
Release / test (push) Successful in 4s
Release / chart (push) Successful in 4s
Release / binaries (push) Successful in 28s
Release / image (push) Successful in 1m5s
Release / scan-image (push) Successful in 2s
Cosmetic, and done anyway. release.yaml passes --version and --app-version
from the git tag when it packages, so neither line decides anything about
what is published; they exist to be read by somebody looking at the tree
before the tag does. A tree heading for v0.16.1 that says 0.16.0 tells that
reader something false.

Its own commit, like 828cf87, 8869ac8 and 9376105 before it, so the feature
commit's diff stays the feature.

Claude-Session: https://claude.ai/code/session_01RHPj4ggeFdEjKKfm4SHbD7
v0.16.1
2026-09-21 15:51:26 +02:00
Niklas Ye fc8b0c8d58 Let people set their own ntfy topic under Account
The first-run checklist's first step is "Set where your pages go", and its
button navigated to /more — which had no field for it. Every new user was
sent to a page that could not do the thing it sent them there for, and the
only ways to actually set a topic were curl or asking an administrator.
That has been true since the checklist shipped in v0.15.0.

Account now has a Notifications section above the password form: the topic,
prefilled and saved through the endpoint that already existed, and a Send a
test push button. The test is offered only once a topic is saved, because
it publishes what the server has stored rather than what is half-typed in
the field, and a button that silently tested the previous value would be
worse than no button.

Saving assigns the response to state.me.user, so the checklist stops asking
and the test button appears without a reload. Clearing works by saving an
empty topic: the server treats that as "no topic of their own" rather than
an error, and returns a user with ntfy_topic absent — it is omitempty — so
the form reads the cleared state from the response rather than assuming it.

The copy says the topic is a shared secret, because people reach for their
own name and it is the only thing between a stranger and their pages. Same
reason the topic stays out of an incident's timeline, which every API key
can read.

No server change: PUT /api/users/{id}/notify has been self-or-admin since
#3 and needed nothing. Only the ntfy topic is per-person — the server is
the install's one TERDUT_NTFY_URL and is not something a user picks.

Also drops a line on that page still sending people to terdut-tui for user
management, which stopped being true one release ago.

Claude-Session: https://claude.ai/code/session_01RHPj4ggeFdEjKKfm4SHbD7
2026-09-21 15:51:16 +02:00
Niklas Ye 828cf87656 Set the chart's placeholder version to 0.16.0
CI / chart (push) Successful in 1s
CI / security (push) Successful in 14s
CI / test (push) Successful in 2m33s
Release / test (push) Successful in 4s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 18s
Release / image (push) Successful in 59s
Release / scan-image (push) Successful in 2s
Cosmetic, and done anyway. release.yaml passes --version and --app-version
from the git tag when it packages, so neither line decides anything about
what is published; they exist to be read by somebody looking at the tree
before the tag does. A tree heading for v0.16.0 that says 0.15.1 tells that
reader something false.

Its own commit, like 8869ac8, 9376105 and 4e8c52c before it, so the feature
commit's diff stays the feature.

Claude-Session: https://claude.ai/code/session_01RHPj4ggeFdEjKKfm4SHbD7
v0.16.0
2026-09-21 12:55:58 +02:00
Niklas Ye ac9af8e4f5 Manage a person's account and teams from one page
The Admin tab could make somebody an administrator and disable them, and
nothing else. Setting a first password, deleting an account and seeing
which teams a person is in all meant curl, and the last one meant opening
each team in turn — the Team tab answers "who is in this team", which is
the wrong way round when the question is about a person.

A name in the user list now opens /admin/users/{id}: their email and when
they joined, where their notifications go, the administrator and disabled
flags, the teams they are in with their role in each, a password field
for a first or forgotten one, and deletion. A section of its own rather
than an expanding row, because memberships and the account actions
together are more than a table row can hold and still be read on a phone.

Adding somebody mints an invite link into a chosen team rather than
creating a bare account. POST /api/users makes a user with no password
and no team, who can sign in nowhere and would see nothing if they did;
the invite machinery from #7 already solves both, and the password is
chosen by the person it belongs to instead of passing through an
administrator.

One new endpoint, GET /api/users/{id}/teams, self or admin. /api/teams is
always about the caller and cannot be asked about anybody else. It 404s
for a user who does not exist, so the page can tell "in no teams" from
"no such person" — an empty list is a real answer and needed to stay one.

No authorisation changed, and the interesting part is why it did not.
requireTeamOwner has accepted the administrator flag since a4fbd60, with
the reason in its own comment: somebody has to be able to repair a team
whose owner has left. It guards nine call sites, so an administrator has
always been able to configure any team on this server — while #1's
decision table and this README both said an admin "is not implicitly in
every team", full stop. The code was right and the prose was wrong in the
safe-sounding direction, which is the worse way round to have it.

So the documentation moved to meet the code. The Teams table marks owner
as owner-or-admin, and the Authentication section states the two
directions separately: an administrator configures any team, and reads
none, because callerTeamIDs is built from real memberships only. Joining
a team to see its queue is a membership change and shows as one.

TestAdmin_ConfiguresATeamTheyAreNotIn pins both halves — the admin
renames, invites, adds and removes on a team they are not in, then sees
zero of its incidents. Nothing tested this from v0.12.0 to here, which is
why four releases of prose could contradict it quietly.

The UI has not been opened in a browser. Its wiring is checked — every
cross-module import resolves, every api.* call exists, every CSS class
has a rule, and the deep link serves index.html — but nobody has clicked
through it, least of all at phone width.

Claude-Session: https://claude.ai/code/session_01RHPj4ggeFdEjKKfm4SHbD7
2026-09-21 12:55:42 +02:00
Niklas Ye 8869ac864f Set the chart's placeholder version to 0.15.1
CI / test (push) Successful in 4s
CI / chart (push) Successful in 1s
CI / security (push) Successful in 10s
Release / test (push) Successful in 4s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 22s
Release / image (push) Successful in 56s
Release / scan-image (push) Successful in 3s
Cosmetic, as in 4c85e76 and 041e159. `make helm-package` passes --version
and --app-version from the tag, so neither field decides anything about
what release.yaml publishes.

Done anyway because a tree heading for v0.15.1 that still says 0.15.0
tells a reader something false. appVersion keeps the v, per
APPVERSION_PREFIX in .release.conf.
v0.15.1
2026-09-21 11:23:56 +02:00
niklas 0677e74cf8 Merge pull request 'Let the tab bar fit however many tabs there are' (#21) from compact-nav into main
CI / chart (push) Successful in 1s
CI / test (push) Successful in 6s
CI / security (push) Successful in 12s
Reviewed-on: #21
2026-09-21 09:23:31 +00:00
Niklas Ye 56b8191a78 Let the tab bar fit however many tabs there are
CI / chart (pull_request) Successful in 2s
CI / security (pull_request) Successful in 17s
CI / test (pull_request) Successful in 2m33s
The bottom tab bar was grid-template-columns: repeat(4, 1fr), written when
there were four tabs. Team and Admin arrived in the last two releases and
nothing updated that number, so six items were being laid into four
columns -- which on a phone is the reported symptom, tabs that do not fit
the width.

grid-auto-flow: column with grid-auto-columns: 1fr makes the count follow
the markup instead. That also handles a case a fixed number cannot: Admin
is only rendered for an administrator, so the tab count genuinely differs
between two people looking at the same install.

Then the compactness. Each link gets min-width: 0 so a column may shrink
below its label's natural width, and the label itself ellipsises rather
than widening the bar. Under 420px the font drops to 10px, the icons to
21px and the badge shrinks to match.

No icon-only breakpoint. The arithmetic says the labels fit: six tabs on
a 320px phone give about 53px each, and the widest label, "On-call", is
about 38px at 10px. A media query that never fires is dead code, and the
ellipsis is the backstop if a future tab is named something longer.

The links gained aria-labels regardless. The icons are aria-hidden, so
the visible text was the accessible name, and it should not be the only
one.

The desktop sidebar is unaffected: it overrides display, padding and
font-size itself, so none of the phone rules reach it.

Not verified on a phone -- I cannot open a browser here, so this is the
cause identified from the CSS and the widths worked out on paper.

Claude-Session: https://claude.ai/code/session_01RHPj4ggeFdEjKKfm4SHbD7
2026-09-21 11:18:30 +02:00
Niklas Ye 93761056eb Set the chart's placeholder version to 0.15.0
CI / chart (push) Successful in 4s
CI / test (push) Successful in 12s
CI / security (push) Successful in 17s
Release / test (push) Successful in 4s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 19s
Release / image (push) Successful in 1m3s
Release / scan-image (push) Successful in 24s
Cosmetic, as in 4c85e76 and 041e159. `make helm-package` passes --version
and --app-version from the tag, so neither field decides anything about
what release.yaml publishes.

Done anyway because a tree heading for v0.15.0 that still says 0.14.0
tells a reader something false. appVersion keeps the v, per
APPVERSION_PREFIX in .release.conf.
v0.15.0
2026-09-21 10:55:02 +02:00
niklas a92da7dcc0 Merge pull request 'Add the sign-up page and the first-run checklist' (#20) from onboarding-ui into main
CI / test (push) Successful in 4s
CI / chart (push) Successful in 1s
CI / security (push) Successful in 11s
Reviewed-on: #20
2026-09-21 08:54:29 +00:00
Niklas Ye b39aac36b7 Add the sign-up page and the first-run checklist
CI / chart (pull_request) Successful in 1s
CI / security (pull_request) Successful in 13s
CI / test (pull_request) Successful in 2m28s
Second half of #7. The API could create accounts from invite links since
the last change; this is the part somebody can actually use.

/signup is the one route that works without a session. It asks the server
what it may offer before showing anything: an invite link that is good
names the team it leads to, a link that is not says so before somebody
picks a password rather than after, and an invite-only server with no
link says that instead of presenting a form it will refuse. The login
card only offers "create one" when sign-up is open, so the door nobody
can walk through is not advertised.

Signing up signs you in and lands on the queue, because the alternative
is a form saying "now go and log in" about the credential just chosen.

The checklist is the other half. Four things have to be true before an
alert reaches a phone -- a notification topic, somebody on the rota, an
alert source, and an alert that has actually arrived -- and on a fresh
install none of them are. It sits above the queue until they are.

It is computed from the data rather than from stored progress: a topic is
set or it is not, an integration exists or it does not. That means it
cannot claim a step is done when it is not, and it comes back by itself
if somebody deletes their integration a month later. The only stored
state is the dismissal, which is per user and not per browser --
finishing on a laptop should not leave the phone nagging.

The topic step is the only one the checklist can finish itself, and the
only proof that counts is a phone buzzing, so there is a test push.
POST /api/me/notify/test publishes directly rather than through the
outbox, which requires an incident this deliberately does not have. Its
failure is the useful part: a wrong topic, a rejected token and an ntfy
that is down all look identical from the phone, which is silence, so the
error comes back to the browser instead.

Verified against a live server with a real ntfy stand-in, the whole path:
an owner mints an invite, the sign-up page reports it valid and names the
team, the invitee signs up and is signed in as a member of that team, the
checklist's four questions answer correctly on a fresh install, a test
push is refused with no topic and delivered with one -- "PAGED
terdut-owner | terdut test" -- and the dismissal survives a reload.

Claude-Session: https://claude.ai/code/session_01RHPj4ggeFdEjKKfm4SHbD7
2026-09-21 10:47:15 +02:00
niklas 19f168ab7e Merge pull request 'Add self-service sign-up and invite links' (#19) from signup-invites into main
CI / test (push) Successful in 4s
CI / chart (push) Successful in 1s
CI / security (push) Successful in 11s
Reviewed-on: #19
2026-09-21 08:42:07 +00:00
Niklas Ye d827ceedff Add self-service sign-up and invite links
CI / chart (pull_request) Successful in 1s
CI / security (pull_request) Successful in 14s
CI / test (pull_request) Successful in 2m30s
First half of #7. Until now the only way to get an account was for
somebody who already had one to create it, and the login page told people
to "ask an admin" -- workable for one operator, impossible for a team.

Two modes, chosen by an administrator in the settings table: invite_only,
which is the default, and open. A third domain-restricted mode was
considered and dropped, because with no email in this server there is
nothing to verify an address against and it would only check the domain
of a string somebody typed.

The default is the closed door. An install that gets a public hostname
before anybody has thought about sign-up should not be collecting
accounts from the internet, and the failure mode of a typo in the setting
is invite_only rather than open.

An invite is a link, not an email. Adding SMTP to send one message would
be a subsystem to run, secure and monitor; the person inviting sends the
link however they already talk to the person they are inviting. A link
carries the team and the role, because an account in no team sees an
empty queue and can be paged by nobody -- that is not a state to invite
somebody into. Links are single-use by default, expire after seven days,
and can be revoked before that: a link that works forever is a credential
nobody remembers issuing, sitting in a chat log.

The uses counter is incremented inside the sign-up transaction and
guarded by `uses < max_uses`, so two people redeeming the last use at
once cannot both get in.

GET /api/signup reports the mode and whether a link is usable, so the
form can say "this link has expired" before somebody picks a password
rather than after. It gives one answer for expired, revoked, used up and
never existed: telling a stranger which it was tells them something about
links they do not hold.

Sign-up signs you in. The alternative is a form that says "now go and log
in", which is the same credential typed twice. login and signup now share
startSession rather than each minting a cookie.

Rate-limited per address on its own limiter, not login's: a burst of
sign-ups must not lock somebody out of logging in.

The settings table grew a second shape for this. It held only durations;
signup_mode is a word from a fixed list, so the admin endpoint now
validates everything before writing anything -- a request that sets two
settings and gets one wrong changes neither.

Still to come in #7: the sign-up and invite-redemption pages, the
first-run checklist, and the in-app integration instructions. The schema
carries onboarding_dismissed_at for the checklist already.

Claude-Session: https://claude.ai/code/session_01RHPj4ggeFdEjKKfm4SHbD7
2026-09-21 09:06:48 +02:00
Niklas Ye 4e8c52c28c Set the chart's placeholder version to 0.14.0
CI / test (push) Successful in 4s
CI / chart (push) Successful in 1s
CI / security (push) Successful in 12s
Release / test (push) Successful in 4s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 24s
Release / image (push) Successful in 53s
Release / scan-image (push) Successful in 2s
Cosmetic, as in 4c85e76 and 041e159. `make helm-package` passes --version
and --app-version from the tag, so neither field decides anything about
what release.yaml publishes.

Done anyway because a tree heading for v0.14.0 that still says 0.13.0
tells a reader something false. appVersion keeps the v, per
APPVERSION_PREFIX in .release.conf.
v0.14.0
2026-09-20 21:29:05 +02:00
niklas fb927aa67b Merge pull request 'Put a team's own settings in the web UI' (#18) from team-settings-ui into main
CI / test (push) Successful in 5s
CI / chart (push) Successful in 1s
CI / security (push) Successful in 11s
Reviewed-on: #18
2026-09-20 19:26:22 +00:00
Niklas Ye d728af53b1 Put a team's own settings in the web UI
CI / chart (pull_request) Successful in 1s
CI / security (pull_request) Successful in 13s
CI / test (pull_request) Successful in 2m19s
Closes #17. Everything a team owner configures was API-only: escalation,
integrations, dead man's switches, membership, and the rota -- which the
on-call view still described as the TUI's job, and the TUI has been broken
against this server since teams landed. Setting up the feature this whole
line of work exists for meant using curl.

A Team tab now holds all of it, one team at a time, with a picker for
somebody in more than one. An owner edits; a member sees the same page
without the controls, because the server refuses their writes anyway --
hiding a button is a courtesy to the reader, not the thing enforcing
anything.

The escalation editor holds a draft and sends the whole ladder, because
the API replaces it wholesale: the levels are an order, and patching one
rung leaves the numbering of the others undecided. Adding a level
defaults to five minutes and the rota, which is the shape almost every
ladder starts as.

An integration key is returned exactly once, so creating one opens a
panel that says so, shows the URL large with a copy button, and renders
the Alertmanager receiver snippet with the URL already in it -- the next
thing anybody does with that key is paste it into a config. The panel
stays until it is dismissed rather than disappearing on the next
re-render.

The incident view gains where an incident is on the ladder and when the
next page is due, which is the question somebody looking at an
unacknowledged incident actually has. The API carries it: the incident
payload now includes escalation_level and escalation_due_at, the latter
computed in the incident SELECT by joining the level's timeout, so a list
costs no extra queries.

Verified against a live server by making every call the page makes,
including the writes: the six reads the Team tab issues, a two-level
ladder saved and read back, an integration created and its key returned
once, three days of rota assigned, switches set, and an incident showing
level 1 with a due time five minutes out.

Claude-Session: https://claude.ai/code/session_01RHPj4ggeFdEjKKfm4SHbD7
2026-09-20 21:22:24 +02:00
Niklas Ye 53e5e03f4e Set the chart's placeholder version to 0.13.0
CI / test (push) Successful in 4s
CI / chart (push) Successful in 1s
CI / security (push) Successful in 11s
Release / test (push) Successful in 4s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 18s
Release / image (push) Successful in 58s
Release / scan-image (push) Successful in 7s
Cosmetic, as in 4c85e76 and 041e159. `make helm-package` passes --version
and --app-version from the tag, so neither field decides anything about
what release.yaml publishes.

Done anyway because a tree heading for v0.13.0 that still says 0.12.0
tells a reader something false. appVersion keeps the v, per
APPVERSION_PREFIX in .release.conf.
v0.13.0
2026-09-20 20:47:52 +02:00
niklas 4d62c1130b Merge pull request 'Page the next person when nobody answers' (#16) from escalation into main
CI / chart (push) Successful in 1s
CI / test (push) Successful in 6s
CI / security (push) Successful in 12s
Reviewed-on: #16
2026-09-20 16:43:03 +00:00
Niklas Ye 3183e7e5c5 Page the next person when nobody answers
CI / chart (pull_request) Successful in 1s
CI / security (pull_request) Successful in 16s
CI / test (pull_request) Successful in 2m21s
Closes #6, and closes the thing this whole line of work was opened for.
Until now an unacknowledged incident re-paged the same topic every
notify_repeat forever, which is a louder version of the same silence: if
the person on call is asleep, out of signal or has left the company,
nothing else happened.

A team can now configure an ordered ladder. Each level has a timeout and
a set of targets; a target is a named person or whoever the team's rota
says is on call today. That second kind is the one that keeps working
when the rota changes and nobody remembers to edit the policy. When a
level's timeout passes with the incident still triggered, the next level
is paged; off the end the chain repeats repeat_count times and then the
team's fallback topic is paged once. The incident stays open throughout,
because running out of people to wake is not somebody answering.

Escalation rides the notifier's existing 30-second tick and its outbox
rather than adding a second scheduler, and runs before delivery so a
level that comes due on a tick is paged on that tick. Each target gets
its own outbox row and therefore its own Acknowledge token: the button in
a notification must acknowledge as the person holding the phone, not as
whoever was paged first.

Acknowledging or resolving takes the incident off the ladder. Snoozing
pauses it -- a deliberate "not now" holds the ladder where it is and it
resumes when the snooze runs out, rather than carrying on without the
person who asked for quiet.

Reminders and escalation never both run. A team with a ladder gets
escalation; a team without keeps today's behaviour exactly. Both would
mean two pages for one silence, which is how a tool gets muted.

A level whose targets cannot be reached -- no topic, a disabled account,
an empty rota -- is entered anyway, recorded as "nobody reachable", and
the ladder moves on. Stalling on a rung that cannot ring would be the
failure this feature exists to prevent, wearing the feature's clothes. A
policy with such a level cannot be created, but an older row could hold
one.

The API replaces the ladder wholesale rather than patching a rung,
because the levels are an order: editing one has to answer what happens
to the numbering of the others, and a whole-ladder PUT makes that the
client's decision and the edit atomic.

Verified against a live server as well as in tests: alice paged, nobody
answers, bob paged, nobody answers, the fallback topic paged once and the
timeline reading "level 2: bob" then "escalation exhausted: paged
terdut-oncall-all" -- and a second incident acknowledged before its
timeout, which woke nobody else.

No UI yet. The team-settings screens for escalation, integrations and
dead man's switches are all still missing, and they are one piece of work
rather than three.

Claude-Session: https://claude.ai/code/session_01RHPj4ggeFdEjKKfm4SHbD7
2026-09-20 18:37:54 +02:00
niklas 94d23a593c Merge pull request 'Add an admin page, and move the behaviour settings into the database' (#15) from admin-settings into main
CI / test (push) Successful in 4s
CI / chart (push) Successful in 1s
CI / security (push) Successful in 11s
Reviewed-on: #15
2026-09-20 16:26:27 +00:00
Niklas Ye b0a02c010b Add an admin page, and move the behaviour settings into the database
CI / chart (pull_request) Successful in 1s
CI / security (pull_request) Successful in 13s
CI / test (pull_request) Successful in 2m1s
Closes #5. Three of the server's tunables were environment variables,
which meant changing how long an incident waits before being paged again
required editing a chart, merging it and waiting for a reconcile. They
are behaviour rather than infrastructure, and the difference is who needs
to change them and how often.

The split is by who owns the value. What stays in the environment is
where the server is plugged in: the listen address, the DSN, the ntfy URL
and token, the public URL. Those are needed before the database is open
and two of them are credentials -- the settings endpoint reports that
ntfy is configured and that a token is set, and never what either is.

What moves is how it behaves: the notify repeat interval, the stale
window and the archive window. The environment variable becomes the seed
rather than the setting, written once on first start and never
overwritten, so a redeploy cannot put a chart's default back over an
administrator's edit -- the rule the per-team dead man's switches already
follow. The loops read the current value per tick, so a change at 02:00
is obeyed at 02:00.

Key/value rather than a column per knob: #6 and #7 will both add
settings, and a table shaped one-column-per-setting needs a migration for
each. The cost is that values are text and the accessor has to say what
type it wanted, which settings.go does in one place. Unknown keys are
refused rather than stored -- a typo that wrote notify_repeat_second
would otherwise sit in the table looking like configuration and doing
nothing -- and each value has bounds loose enough to catch a slipped
decimal point without having an opinion about anybody's rota.

Disabling an account is new, and is not deleting one. Deleting a user
nulls acknowledged_by and assigned_to, which quietly rewrites who did
what during an incident months after the fact. A disabled user cannot
authenticate by either credential, loses their sessions immediately, and
stays the name on every acknowledgement they made. The check is part of
the lookup in serveAs rather than a test afterwards, so there is no path
where the row is loaded and the flag is then forgotten.

The page itself is a fourth tab, shown only to an administrator and only
as a courtesy: every endpoint under it is refused with 403 regardless, so
somebody who types /admin gets an explanation rather than a blank screen.
It lists teams with their size and open-incident count, users with their
flags, and the settings with their bounds -- plus the environment half,
read-only, so somebody hunting for the ntfy URL learns where it lives
instead of concluding the server has none.

Delete is disabled rather than offered-and-refused for a team with open
incidents, and neither admin action is offered on your own account, since
the server refuses both.

Claude-Session: https://claude.ai/code/session_01RHPj4ggeFdEjKKfm4SHbD7
2026-09-20 18:23:46 +02:00
niklas 303e7a3365 Merge pull request 'Remove the unauthenticated webhook and the SQLite migration script' (#14) from cleanup-after-teams into main
CI / chart (push) Successful in 1s
CI / test (push) Successful in 7s
CI / security (push) Successful in 16s
Reviewed-on: #14
2026-09-20 16:14:15 +00:00
Niklas Ye 7c87ae2af8 Remove the unauthenticated webhook and the SQLite migration script
CI / chart (pull_request) Successful in 1s
CI / security (pull_request) Successful in 13s
CI / test (pull_request) Successful in 1m54s
Both existed to carry an upgrade across, and both upgrades are done.

/api/alertmanager/webhook took no credential at all: anything able to
reach the port could open an incident for anybody. v0.12.0 kept it,
deprecated, so the teams release did not stop delivery while the
Alertmanager config was edited, and logged a line per payload asking to
be moved. The cluster's Alertmanager now posts on an integration key --
verified in the log, every two minutes, with no deprecation line since
the rollout -- so the door can be shut rather than left ajar until
somebody remembers. A sender still posting there gets the JSON 404 every
unknown /api path gets.

The tests move with it, which they should have done anyway: the harness
mints an integration key for the default team and posts on that, so they
exercise the path production uses rather than one only they still used.

scripts/sqlite-to-postgres.go goes the same way. It was written to be
temporary, it was the last thing needing modernc.org/sqlite, and this
install migrated on 2026-09-20. `go mod tidy` drops the driver and its
six transitive dependencies with it; the module graph is now chi, pgx,
pgerrcode and x/crypto. Anyone still on v0.10.x can take the script out
of the v0.12.0 tag, which the README now says.

Claude-Session: https://claude.ai/code/session_01RHPj4ggeFdEjKKfm4SHbD7
2026-09-20 18:11:30 +02:00
Niklas Ye 4c85e7646c Set the chart's placeholder version to 0.12.0
CI / test (push) Successful in 5s
CI / chart (push) Successful in 2s
CI / security (push) Successful in 15s
Release / test (push) Successful in 4s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 20s
Release / image (push) Successful in 54s
Release / scan-image (push) Successful in 2s
Cosmetic, as in 041e159 and 989425e. `make helm-package` passes --version
and --app-version from the tag, so neither field decides anything about
what release.yaml publishes.

Done anyway because a tree heading for v0.12.0 that still says 0.11.1
tells a reader something false. appVersion keeps the v, per
APPVERSION_PREFIX in .release.conf.
v0.12.0
2026-09-20 15:36:46 +02:00
niklas 05f82220a6 Merge pull request 'Per-team dead man's switches, and the UI's team badge, filter and cards' (#13) from teams into main
CI / test (push) Successful in 5s
CI / chart (push) Successful in 1s
CI / security (push) Successful in 11s
2026-09-20 13:29:09 +00:00
niklas 5227eb0d5f Merge pull request 'Give each team its own dead man's switches, and the UI a team to show' (#12) from deadman-per-team into teams
CI / test (pull_request) Successful in 5s
CI / chart (pull_request) Successful in 1s
CI / security (pull_request) Successful in 16s
2026-09-20 13:28:30 +00:00
Niklas Ye 74359c72ab Give each team its own dead man's switches, and the UI a team to show
CI / chart (pull_request) Successful in 1s
CI / security (pull_request) Successful in 14s
CI / test (pull_request) Successful in 1m57s
The rest of #4. Two halves that belong together because they are the
same sentence from opposite ends: a team decides which of its alerts are
heartbeats, and the UI has to be able to say which team it is talking
about.

Switches were three environment variables, which made them one setting
for the whole install. That was the last piece of the alerting path a
team could not control: it could take its own alerts on its own key and
still not say which of them were heartbeats, or how long a silence had
to last. They are a row per team now, edited by an owner through
PUT /api/teams/{teamID}/deadman, and the sweeper runs each team against
its own matchers, timeout and severity.

The environment variables become the starting point rather than the
setting. Every team without a configuration is seeded from them at
startup, so an upgrade keeps watching exactly what it was watching, and
SeedDeadmanConfigs never overwrites -- a redeploy must not put the
environment's value back over an owner's edit. A team created later
watches nothing until somebody says otherwise: inheriting an
install-wide heartbeat would page a new team about a source it has never
heard of, and a switch nobody chose is the kind that gets muted rather
than fixed.

A matcher string with no alertname in it is refused at the door instead
of stored. Storing it would produce a switch that watches nothing
silently, which is the exact failure the feature exists to prevent.

NewRouter and Sweep lose their DeadmanConfig parameter -- there is no
longer one answer to hand them. The type stays, because parsing a
matcher string is still parsing a matcher string.

The UI side: rows in the queue carry a team badge, the filter row gains
a team chip per team, and "on call now" shows one card per team. All
three appear only when the viewer is in more than one team -- otherwise
they are the same word repeated down a list, which is noise rather than
information, and the single-team install reads exactly as it did before
teams existed.

Verified against a live two-team server as well as in tests: the
combined queue labelled by team, the team_id filter, a heartbeat that is
a heartbeat in one team and an ordinary alert in another, and a new
team's switches starting empty while the upgraded team keeps the
environment's.

Claude-Session: https://claude.ai/code/session_01RHPj4ggeFdEjKKfm4SHbD7
2026-09-20 15:18:22 +02:00
niklas 43f69272f0 Merge pull request 'Add a system administrator role, and gate account management behind it' (#10) from admin-role into main
CI / test (push) Successful in 4s
CI / chart (push) Successful in 1s
CI / security (push) Successful in 11s
Reviewed-on: #10
2026-09-20 13:10:47 +00:00
niklas 2de5c8412d Merge pull request 'Scope everything to a team, and route alerts by integration key' (#11) from teams into admin-role
CI / test (pull_request) Successful in 4s
CI / chart (pull_request) Successful in 1s
CI / security (pull_request) Successful in 11s
Reviewed-on: #11
2026-09-20 11:39:33 +00:00
Niklas Ye a4fbd60441 Scope everything to a team, and route alerts by integration key
CI / chart (pull_request) Successful in 1s
CI / security (pull_request) Successful in 13s
CI / test (pull_request) Successful in 1m49s
The core of #4, and what #1 is for: terdut stops being one shared space.
A team owns its incidents, alerts, schedule and integrations; a user sees
exactly the teams they are in. Everything that existed moves into one
Default team and every existing user becomes an owner of it, so the
upgrade is a no-op for the people using it.

Ingestion is the load-bearing half. An alert arrives on a team's
integration key, and the key is both the credential and the routing: it
says that the sender may post, and which team the alerts belong to. That
also closes the unauthenticated webhook -- the old path stays for one
release, deprecated and routed to the oldest team, so an upgrade does not
stop delivering while somebody edits the Alertmanager config.

Scoping is enforced in as few places as possible, because the failure
mode is silent. serveAs loads the caller's memberships once; list queries
carry `team_id = ANY(...)`; and every incident route goes through
incidentIDParam, which now parses the id AND checks the team in the same
call, so a new handler cannot remember the first half and forget the
second. Anything in another team is 404, never 403: whether an incident
exists is that team's business.

Two bugs this found, both of which would have been silent:

  * upsertAlerts decided "is this a new occurrence" by looking up the
    fingerprint alone. Across teams that made team B's first alert look
    like a re-send of team A's, so it opened no incident at all. The
    lookups are keyed on (team_id, fingerprint) now, as the index is.

  * Every uniqueness rule was written for one tenant. Two teams watching
    two clusters legitimately see the same fingerprint, the same
    groupKey, and want somebody on call on the same day; all three
    constraints move to include team_id.

Roles inside a team are separate from the system administrator flag: an
owner configures the team, a member works its incidents, and an admin is
NOT implicitly in every team -- administration is about accounts, not
about reading other people's incidents. An admin can still repair a team
whose owner has left, which is why requireTeamOwner lets them through.

A shift can only be given to somebody in the team. Paging a person who
cannot open the incident is worse than paging nobody.

The UI is updated only as far as keeping it working: it loads the
viewer's teams with the session and uses the first one, since nobody has
a second yet. "On call now" shows every team the viewer is in, named only
when there is more than one, so the common case reads exactly as before.
The team switcher, badges and per-team settings pages are the next step.

Breaking for API clients: the schedule endpoints moved under the team,
and /api/schedule/current returns an array rather than an object or a
404. terdut-tui will need a version for that.

Per-team dead-man configuration is deliberately not here. A heartbeat's
incident already opens in the team whose key received it, which is the
part that matters for isolation; moving the matchers out of env into
per-team rows is a change to how deadman.go is configured rather than to
who sees what.

Claude-Session: https://claude.ai/code/session_01RHPj4ggeFdEjKKfm4SHbD7
2026-09-20 13:36:24 +02:00
Niklas Ye 1377d9005b Add a system administrator role, and gate account management behind it
CI / chart (pull_request) Successful in 2s
CI / security (pull_request) Successful in 16s
CI / test (pull_request) Successful in 1m37s
Until now every authenticated caller could create and delete users, set
anybody's password and mint anybody's API keys -- auth.go said so in a
comment. Defensible with one operator and a hand-made account; not once
people sign themselves up (#7), and not in a multi-tenant install (#4),
where the user list is no longer everybody who works here.

users.is_admin is the flag. AdminOnly gates creating and deleting users
and granting the flag itself. The endpoints that are self-service for
your own account and administration for somebody else's -- password,
ntfy topic, API keys -- go through requireSelfOrAdmin instead, because
which rule applies depends on the {id} in the path rather than on the
route.

Minting your own API key stays self-service. A key carries exactly the
rights of the user it belongs to, so issuing one is no more than signing
in again; requiring an admin for it would mean a responder cannot set up
the TUI without somebody else in the room.

/api/users stays readable by everybody. The queue's assignment control
and the on-call schedule both have to name people, and hiding the roster
from the people on it buys nothing.

THE MIGRATION MAKES EVERY EXISTING USER AN ADMINISTRATOR. They already
hold these powers, so nobody's access changes on upgrade: it names what
is already true and leaves demotion as a deliberate act. Promoting only
user 1 would silently strip the others, and could leave an install whose
only administrator is an account nobody has a password for.

Two guards keep an install administrable: the last administrator can be
neither deleted nor demoted, and nobody can delete or demote themselves
-- the likelier accident, where the only admin clears their own flag
while tidying up and locks the door behind them.

No UI changes: there are no account-management screens yet. models.User
carries is_admin (not omitempty, so a client can tell false from an old
server), which is what #5's admin page will render from.
2026-09-20 13:20:04 +02:00
Niklas Ye e3ad19c110 Set the chart's placeholder version to 0.11.1
CI / test (push) Successful in 4s
CI / chart (push) Successful in 1s
CI / security (push) Successful in 11s
Release / test (push) Successful in 9s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 1m1s
Release / image (push) Successful in 1m25s
Release / scan-image (push) Successful in 24s
Cosmetic, as in 041e159 and 989425e. `make helm-package` passes --version
and --app-version from the tag, so neither field decides anything about
what release.yaml publishes.

0.11.0 never reached the registry -- its release run failed at the test
gate -- so the placeholder moves on to the version that will.
v0.11.1
2026-09-20 12:57:20 +02:00
Niklas Ye a8ee742533 Give release.yaml's test job the database ci.yaml already has
CI / test (push) Successful in 4s
CI / chart (push) Successful in 1s
CI / security (push) Successful in 11s
v0.11.0 was tagged and published nothing: release.yaml runs the same
`make fmt lint test` as ci.yaml, and the Postgres service the suite now
needs was added to ci.yaml alone. Its test job failed on the missing
TERDUT_TEST_DSN, and binaries, image, chart and scan-image are all
downstream of it, so they skipped. The tag stays -- a published tag is
immutable and moving one is how this repo ran v0.4.0 for ten days while
every artifact said v0.3.0 -- so the fix is the next version.

The two workflows carry their own copies of this block because a service
container cannot be factored into the Makefile the way the checks are.
That is the second copy, and the reason this failed: the gate is one
target, but what the gate needs to run is declared per workflow.
2026-09-20 11:12:41 +02:00
Niklas Ye 041e159e2a Set the chart's placeholder version to 0.11.0
CI / test (push) Successful in 5s
CI / chart (push) Successful in 1s
CI / security (push) Successful in 12s
Release / test (push) Failing after 4s
Release / binaries (push) Has been skipped
Release / image (push) Has been skipped
Release / chart (push) Has been skipped
Release / scan-image (push) Has been skipped
Cosmetic, as in 989425e and e78f494. `make helm-package` passes --version
and --app-version from the tag, so neither field decides anything about
what release.yaml publishes.

Done anyway because a tree heading for v0.11.0 that still says 0.10.2
tells a reader something false. appVersion keeps the v, per
APPVERSION_PREFIX in .release.conf.
v0.11.0
2026-09-20 11:10:21 +02:00
niklas cede8743a8 Merge pull request 'Take the database password from PGPASSWORD, not the DSN' (#9) from postgres-dsn-password into main
CI / test (push) Successful in 4s
CI / chart (push) Successful in 1s
CI / security (push) Successful in 10s
Reviewed-on: #9
2026-09-20 09:07:32 +00:00
Niklas Ye cc31c993dd Take the database password from PGPASSWORD, not the DSN
CI / test (pull_request) Successful in 4s
CI / chart (pull_request) Successful in 1s
CI / security (pull_request) Successful in 11s
The chart asked for a whole DSN in a Secret. Nothing writes one: the
Zalando postgres operator generates a Secret with `username` and
`password` keys and no connection string, so wiring the wrapper chart up
would have meant hand-maintaining a second copy of a password the
operator owns and rotates on a from-scratch rebuild -- which is charts#176
again, the issue miniflux closed by doing the opposite.

So the DSN becomes a plain value with no password in it, and the password
arrives as PGPASSWORD from a Secret. pgx fills in from libpq's PG*
environment variables whatever the DSN omits, exactly as miniflux's
lib/pq does. Verified rather than assumed, against a real server: a
password-less DSN connects with PGPASSWORD set, and fails with
`password authentication failed` when it is wrong, so the variable is
doing the work rather than being quietly ignored.

It also keeps the credential out of the rendered manifest and out of
`kubectl describe pod`, which a DSN-with-password does not.
2026-09-20 11:00:05 +02:00
niklas 21f0eec807 Merge pull request 'Move the database to Postgres' (#8) from postgres into main
CI / test (push) Successful in 4s
CI / chart (push) Successful in 1s
CI / security (push) Successful in 14s
Reviewed-on: #8
2026-09-20 08:54:24 +00:00
Niklas Ye dc39e3a5d3 Move the database to Postgres, before teams need the schema
CI / chart (pull_request) Successful in 1s
CI / security (pull_request) Successful in 17s
CI / test (pull_request) Successful in 2m5s
First step of #1, and it goes first for one reason: #4 adds a team_id to
nearly every table, and doing that twice -- once for SQLite, once for
Postgres -- is work nobody gets paid for. The teams migrations now only
have to be written against one database.

The ten SQLite migrations are replaced by a single Postgres baseline
rather than ported one by one. They were incremental in a way that has
no value on a fresh install: 004 adds columns 008 drops again, and 008's
backfill rewrites data a Postgres database never had. The history stays
in git; the schema they add up to is now 001_baseline.sql.

Timestamps stay BIGINT unix seconds and are NOT converted to timestamptz.
Everything in Go already speaks epochs, so converting would have been a
second, larger change riding along inside this one. It is worth doing on
its own. The JSON columns did move to jsonb, because #4 will want to
filter and index on labels.

Most of the port is mechanical -- 170 placeholders from ? to $1 -- but
four things needed more than a search and replace:

  * Dynamically built WHERE clauses cannot keep their numbering straight
    by hand, so they hand out placeholders through sqlArgs instead. A
    filter can now be added or reordered without renumbering anything.

  * SUM(resolved_at IS NULL) was SQLite counting a boolean as 0 or 1.
    Postgres has no sum(boolean), and this was breaking every dead man's
    switch -- silently, since the sweeper only logs. Now COUNT(*) FILTER.

  * unixepoch() became FLOOR(EXTRACT(EPOCH FROM now()))::bigint. The
    FLOOR is load-bearing: a bare cast rounds half up, so a row written
    at .6 of a second claimed a timestamp a second in the future and
    disagreed with the time.Now().Unix() the Go side stamps.

  * The unique-violation check matched SQLite's error text. It matches
    SQLSTATE 23505 now, so a renamed constraint cannot turn a 409 back
    into a 500.

Tests need a real Postgres, because there is no in-memory Postgres the
way there was an in-memory SQLite. Each test gets its own schema on a
shared server -- cheaper than a database each, and still isolated.
TERDUT_TEST_DSN says where it is; `make test-db` starts one locally and
ci.yaml runs one as a service container. An unset DSN fails the suite
rather than skipping it: a run that quietly tests nothing is worse than
one that does not run.

TestMigration_BackfillCarriesAckAndComments is deleted along with the
migrations it replayed. What it protected -- an upgrade not losing
acknowledgements and comments -- now belongs to scripts/sqlite-to-postgres.go,
which is build-tagged so the SQLite driver stays out of the server
binary. Both are meant to be deleted once this install has migrated.

The chart loses the PVC, the data volume and the python backup sidecar,
and requires database.dsnSecret.name: it provisions no database and
cannot guess where the credentials live, so a render without it is meant
to fail. Backups move to where Postgres actually runs. The other half of
that -- the postgresql CR, the k8up pg_dump annotation and the network
policy -- is a change to the wrapper chart in Ryuvia/charts and is not in
here.

Verified rather than assumed: the gate is green with -race against
Postgres 17, govulncheck and gitleaks are clean, and the migration script
was run end to end against a SQLite database built at the old schema and
seeded in every table. Ids survive, so incidents keep their numbers and
every foreign key still points where it did; the identity sequences are
moved past the copied ids, and a webhook after the migration opened
incident 12 rather than colliding at 1.
2026-09-20 10:44:12 +02:00
Niklas Ye 989425e550 Set the chart's placeholder version to 0.10.2
CI / chart (push) Successful in 1s
CI / security (push) Successful in 28s
CI / test (push) Successful in 2m13s
Release / test (push) Successful in 1m9s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 1m41s
Release / image (push) Successful in 1m42s
Release / scan-image (push) Successful in 2s
Cosmetic, as in e78f494 and 4cec26e. `make helm-package` passes --version
and --app-version from the tag, so neither field decides anything about
what release.yaml publishes.

Done anyway because a tree heading for v0.10.2 that still says 0.10.1
tells a reader something false. appVersion keeps the v, per
APPVERSION_PREFIX in .release.conf.
v0.10.2
2026-09-19 18:05:06 +02:00
Niklas Ye 5b1ab2c568 Move x/crypto to v0.55.0, clear of the ssh CVEs
v0.10.1's image published, but scan-image refused it. trivy reports ten
HIGH advisories against golang.org/x/crypto v0.49.0, CVE-2026-39828
through CVE-2026-56854, all of them in x/crypto/ssh and its agent and
knownhosts packages. The last is fixed in 0.55.0 and the rest in 0.52.0.

None of them is reachable here. The server imports x/crypto/bcrypt and
nothing else from the module, and the image is FROM scratch, holding a
single binary. But trivy scans at module granularity and cannot tell
that. Shipping a known-vulnerable module version on the strength of a
reachability argument is not the call to make inside a dependency pin,
so the version moves instead.

9abf07f was wrong to call v0.49.0 the newest release that keeps go.mod
at go 1.25.9. v0.55.0 keeps it there too; the directive is unchanged
here. What had held it at 1.26 was the x/sys v0.48.0 that v0.57.0 pulled
in. The `go get golang.org/x/sys@v0.42.0` meant to undo that then
downgraded x/crypto to v0.49.0 to match, without being asked. x/sys now
sits at v0.47.0, the version x/crypto v0.55.0 requires.

Checked before tagging rather than by the pipeline: the gate passes, the
Dockerfile builds on its Go 1.25 builder, and trivy, run locally with
the same flags as `make security-image` (HIGH and CRITICAL, fixed only),
exits 0 on the image that build produced.
2026-09-19 18:05:06 +02:00
Niklas Ye e78f49461a Set the chart's placeholder version to 0.10.1
CI / chart (push) Successful in 1s
CI / security (push) Successful in 27s
CI / test (push) Successful in 2m11s
Release / test (push) Successful in 1m11s
Release / chart (push) Successful in 3s
Release / binaries (push) Successful in 1m32s
Release / image (push) Successful in 1m41s
Release / scan-image (push) Failing after 29s
Cosmetic, as in 4cec26e and 9669b8f. `make helm-package` passes --version
and --app-version from the tag, so neither field decides anything about
what release.yaml publishes.

Done anyway because a tree heading for v0.10.1 that still says 0.10.0
tells a reader something false. appVersion keeps the v, per
APPVERSION_PREFIX in .release.conf.
v0.10.1
2026-09-19 17:56:20 +02:00
Niklas Ye 9abf07f2cb Build on Go 1.25 again, as the Dockerfile does
v0.10.0 never produced an image. Adding bcrypt in dc3879e ran
`go get golang.org/x/crypto`, which took the newest release, v0.57.0.
That version declares `go 1.26.0`, and so does the x/sys v0.48.0 it
pulled along, so go.mod's own directive rose from 1.25.9 to 1.26.0. The
Dockerfile builds on golang:1.25-alpine with GOTOOLCHAIN=local, and
refused:

  go: go.mod requires go >= 1.26.0 (running go 1.25.14; GOTOOLCHAIN=local)

Nothing before the image job could see it. The local toolchain is 1.26,
and the pipeline's test job runs on the runner's Go rather than in the
Dockerfile's builder. So the gate passed on a module the image could not
build.

Pinned to x/crypto v0.49.0 rather than moving the builder to 1.26. That
is the newest release whose own requirements leave go.mod at 1.25.9 and
x/sys where it was. Changing the toolchain the image is built with
belongs in its own change, not inside a dependency addition. go.mod now
differs from v0.9.4 by the one crypto line and nothing else. bcrypt has
not changed in any way that matters here across those releases.

Checked by building the Dockerfile locally, which now succeeds. The
resulting container serves /incidents/3 as the web UI and answers
/api/me with a JSON 401.
2026-09-19 17:56:20 +02:00
Niklas Ye 4cec26edde Set the chart's placeholder version to 0.10.0
CI / chart (push) Successful in 2s
CI / security (push) Successful in 40s
CI / test (push) Successful in 2m47s
Release / test (push) Successful in 1m9s
Release / chart (push) Successful in 2s
Release / image (push) Failing after 42s
Release / scan-image (push) Has been skipped
Release / binaries (push) Successful in 1m57s
Cosmetic, as in 9669b8f and f46e5f5. `make helm-package` passes --version
and --app-version from the tag, so neither field decides anything about
what release.yaml publishes.

Done anyway because a tree heading for v0.10.0 that still says 0.9.4
tells a reader something false, and the tree is what gets read before the
tag exists. appVersion keeps the v, per APPVERSION_PREFIX in .release.conf.
v0.10.0
2026-09-19 17:48:28 +02:00
Niklas Ye dc3879eca6 Serve a web UI for the incident queue, built for phones
Whoever is on call gets paged on a phone, and until now the only ways to
act on a page were the notification's Acknowledge button or a terminal.
Tapping the notification itself opened /api/incidents/{id}, which a
browser can only answer with a 401 in JSON. The server now serves a web
UI at / covering the incident queue, each incident's alerts and timeline
with every action on it, who is on call, the alert feed, and changing
your own password. The notification link now points at /incidents/{id}
in that UI.

It is embedded in the binary and has no build step: plain HTML, CSS and
ES modules under internal/web/static, served with an ETag per file and a
CSP that allows nothing from any other origin. That is how rd-web is
built. It avoids adding a node toolchain to the Dockerfile and the
pipeline for a page this size, and it keeps the page on the same origin
as the API, so no CORS is needed and nothing else has to be deployed.
Paths without a file extension fall back to index.html, so a deep link
survives a reload. An unknown path under /api/ still gets a JSON 404
rather than the page.

Signing in uses a username and password, because pasting a 64-character
API key into a phone at 3am is not a sign-in flow. Users have no
password until one is set through PUT /api/users/{id}/password, or
optionally at bootstrap. A user without a password is exactly where they
were before this commit and can only use API keys. A login sets an
HttpOnly, SameSite=Lax session cookie. It lasts 30 days and slides
forward while in use, so an on-call phone does not sign itself out.
Only the token's hash is stored, as for API keys.

The cookie needs a CSRF guard where a bearer header does not, because
browsers attach cookies to requests other sites make. So cookie-
authenticated requests go through Go 1.25's http.CrossOriginProtection,
and bearer requests do not. A request carrying an Authorization header
is judged on that header alone and never falls back to the cookie.
Changing a password ends every other session of that user. Changing
your own requires the current password, so a phone left signed in
cannot be used to take the account over.

Failed logins are counted per username and per client address. Ten
failures for one username in 15 minutes refuse that username for the
rest of the window, even with the right password. That makes locking
somebody out possible for anyone who knows their username. It was
accepted because the alternative is unlimited guessing, and during a
lockout the notification's Acknowledge button and API keys keep
working. The address limit reads the first X-Forwarded-For hop, since
behind the gateway RemoteAddr is Envoy. It is looser, because a whole
office behind one NAT shares it.

The Secure flag follows TERDUT_PUBLIC_URL, since TLS terminates at the
gateway and the server itself only ever sees plain HTTP. The chart
already defaults that variable to https://<hostname>.

Schedule editing, statistics and user management stay in terdut-tui for
now. The API they use is unchanged, and bearer authentication behaves
exactly as before.
2026-09-19 17:48:21 +02:00
Niklas Ye 9669b8f477 Set the chart's placeholder version to 0.9.4
CI / chart (push) Successful in 0s
CI / security (push) Successful in 24s
CI / test (push) Successful in 28s
Release / test (push) Successful in 28s
Release / chart (push) Successful in 1s
Release / binaries (push) Successful in 28s
Release / image (push) Successful in 1m11s
Release / scan-image (push) Successful in 23s
Cosmetic, as in f46e5f5. `make helm-package` passes --version and
--app-version from the tag, so neither field decides anything about what
release.yaml publishes.

Done anyway because a tree heading for v0.9.4 that still says 0.9.3 tells
a reader something false, and the tree is what gets read before the tag
exists. appVersion keeps the v, per APPVERSION_PREFIX in .release.conf.

Claude-Session: https://claude.ai/code/session_014m2pJdpCTv3mvvUUuBM54Y
v0.9.4
2026-09-04 18:08:21 +02:00
Niklas Ye a7871ed7c6 Stop the bootstrap hook installing curl at run time
The hook's container was alpine:3 and its first line was
`apk add --no-cache curl`. That writes the binary into the container's
writable upper layer, and every exec of it afterwards is, correctly, a
dropped binary: Falco's `Drop and execute new binary in container`
(PCI_DSS_11.5.1, MITRE TA0003) fired twice at Critical on the upgrade to
chart 0.9.3, 65ms after the container started, with
evt.arg.flags=EXE_WRITABLE|EXE_UPPER_LAYER. Ryuvia/charts#100 has the
event lines.

A true positive of the rule and a false positive of intent, and it is not
a one-off: the hook is post-install,post-upgrade, so it recurred on every
release. The cluster is still in the Falco burn-in with detections routed
to a null receiver, which is the only reason nobody was paged for it.

Fixed here rather than with a Falco exception on purpose. An exception
would have to name this container and would then stay in the rule set
forever, blinding it for the one workload that already runs as root with
create-secret RBAC, and it would leave the second problem untouched: this
runs as a post-upgrade hook, a failed hook fails the release, so every
`helm upgrade` of terdut-server depended on dl-cdn.alpinelinux.org
answering. That dependency is now gone.

alpine/curl is still a full Alpine, so sh, cat, sleep, grep, cut, head and
tail are all present -- verified in-cluster before the swap rather than
assumed, since a missing utility would surface as a failed post-upgrade
hook and not as anything visible here. Digest-pinned, as the wrapper
chart's own sidecar images are. The image declares an ENTRYPOINT, which
the Job's `command:` overrides; a comment says so, because rewriting that
to `args:` would silently run curl's entrypoint instead of the script.

No change to the script's logic, to the RBAC, or to when the hook runs.
Nothing on the terdut-tui side of the API moves, and no terdut-tui version
is required or excluded by this.

Worth recording while it is in view, and deliberately not acted on here:
there is no terdut-server-admin-key secret in the namespace, so the POST
returns 403, the hook logs "Server already bootstrapped, nothing to do"
and exits before the secret-creating branch. On an upgrade this hook
currently achieves nothing at all. Narrowing it to post-install would
remove the detection outright, but that changes what the hook is for and
belongs in its own change.

Claude-Session: https://claude.ai/code/session_014m2pJdpCTv3mvvUUuBM54Y
2026-09-04 18:08:05 +02:00
Niklas Ye 79f5db2636 Scan the source and the working tree too, not just the image
CI / chart (push) Successful in 0s
CI / security (push) Successful in 19s
CI / test (push) Successful in 25s
The image scan added yesterday reads the built artifact. It cannot see a
vulnerable dependency the binary never calls into, and it cannot see a
credential in a file that never reaches the image — this one is FROM scratch
and contains a single binary, so almost nothing in the repo is in it. Those
are two different questions and they need two different tools, which is why
riksdata and rd-web have run govulncheck and gitleaks all along.

Both run on every push and pull request rather than only on a tag, since
neither needs anything published.

Checked by hand before wiring in, as with the image scan. govulncheck
reports no vulnerabilities the code can reach, and gitleaks finds nothing in
the tree.

What govulncheck does report is worth writing down, because it is the
argument for having it. It found three advisories in chi and reports none of
them, all three being IP spoofing in middleware.RealIP, which router.go does
not use — it uses Logger and Recoverer. The analysis is symbol-level rather
than dependency-level, so adding middleware.RealIP would turn this red on
the next push. That is precisely when someone should be made to look, and it
is a plausible thing to reach for here, since the API sits behind a gateway
and real client addresses are exactly what RealIP is for. The fourth finding
is an integer overflow in golang.org/x/sys/windows, which a linux/scratch
image will not be calling.

Both gates were checked for the failure direction as well. gitleaks exits 1
on a private key block. Worth knowing when testing it: it allowlists
well-known example credentials, so the AWS key from Amazon's own
documentation does not trip it and proves nothing.

Neither reads git history. gitleaks runs with --no-git, which scans the
working tree, so it stops a secret on the way in and says nothing about what
is already committed.

Claude-Session: https://claude.ai/code/session_01S7R4gWTz5wh5xCY4nCSJjN
2026-09-02 10:11:13 +02:00