.nav-link-secondary's display:none sat before .nav-link's own
display:flex in the file. Both are single-class selectors, so they tie
on specificity, and a tie is broken by which one comes later in the
file -- not by which class the element happens to carry. .nav-link's
declaration, being later, won for every element wearing both classes,
so Stats/Admin/Account rendered as three extra tabs on the phone bar
instead of folding into "More" as intended.
Moved the rule below .nav-link instead of changing either declaration,
since nothing about the values was wrong -- only their order was.
🤖 Generated with [Claude Code](https://claude.ai/code)
Co-Authored-By: Claude <noreply@anthropic.com>
Cosmetic: make helm-package passes --version and --app-version from the
tag, so these two fields decide nothing about what gets published. Kept
in step anyway, the same as 4358e84 and fd26fef before it, so the tree
heading for v0.35.0 does not say 0.34.0 to a reader who hasn't yet seen
the tag.
🤖 Generated with [Claude Code](https://claude.ai/code)
Co-Authored-By: Claude <noreply@anthropic.com>
Addresses the screenshot-review feedback in #26. No framework or build
step added — all of this stays within the existing plain HTML/CSS/
vanilla-JS + go:embed architecture.
- Nav: re-enable the bottom tab bar that was already built and
switched off (Queue/On-call/Alerts/Team + a "More" sheet for
Stats/Admin/Account), replacing the hamburger on phone width.
- Queue: chip counts, a scroll fade on the filter row, a
"Triggered Xh ago" + severity label per row, a chevron on the team
switcher so it reads as a dropdown.
- On-call: collapse repeated same-person days into shift bars (week
view and "your shifts" both), show the week as a date range with
the ISO week number as secondary text, split "Current shift" out
from "Next shifts" with "ends in Nd", a pill badge + row highlight
for "you".
- Incident detail: fix the actual bug behind the duplicate
"acknowledged" timeline entries (acknowledgeIncident's UPDATE had no
guard on the incident's current status, so acknowledging an
already-acknowledged incident silently re-logged the event — now
idempotent, with regression tests on both the authenticated route
and the ntfy ack-button route). Relabel escalation re-pages so they
don't look like the same page landing twice. Copy the primary action
up near the top. Label the "···" button. Group the timeline by
phase (triggered/acknowledged/resolved). Add an "at a glance"
summary row (duration/severity/responsible) and collapse the group
labels by default.
- Team overview: reword the vague copy ("One owner." etc.) into plain
labels.
- Empty states: fill in missing icons/one-liners across queue,
alerts, stats and the incident timeline.
- CSS: fix card padding bugs, verify link contrast already passes AA,
introduce a --fs-* type-scale token set and migrate the few
genuinely isolated cases onto it (left sizes tied to a fixed shape,
a deliberately prominent display, or a non-negotiable constraint
like the iOS-zoom-prevention input size as documented exceptions
rather than guess at a render this change can't see).
Verified with the full fmt/lint/test/helm-lint gate, plus a live
instance against the test DB with seeded incidents and schedule data
to trace the on-call grouping and timeline phase-splitting logic
against real API responses.
🤖 Generated with [Claude Code](https://claude.ai/code)
Co-Authored-By: Claude <noreply@anthropic.com>
ctxUser/ctxTeams (human) and ctxServiceAccount (+ a synthetic ctxTeams
entry, service account) used to be two parallel, un-unified context
representations -- every authz predicate had to remember which one(s) it
needed, and every place that forgot either wrongly 403'd a service account
(terdut-server#23, terdut-operator#3), crashed on an unchecked zero-value
user id, or silently no-op'd. New internal/api/caller.go collapses both
into one Caller, stored under one ctxCaller key by serveAs/serveAsServiceAccount;
every existing predicate (userFromContext, callerTeamIDs, callerRole,
callerIsAdmin, isInstanceServiceAccount, AdminOnly, requireSelfOrAdmin,
requireTeamOwner, OperatorModeBlock) now reads through it, with identical
behavior for every untouched call site (alerts.go, incidents.go,
schedule.go, stats.go, etc.) -- confirmed by the full existing suite
passing unchanged.
Four real fixes land alongside the refactor, not just the restructuring:
1. callerMayManageServiceAccount gains the one load-bearing branch this
exists for: an instance-scoped service account may now manage (mint or
revoke a key on) any team-scoped account, not only a human admin, that
team's human owner, or the account itself. handleCreateServiceAccount
already let an instance-scoped caller *create* a team-scoped account for
any team; adopting or rotating one it didn't just create in the same
call -- terdut-operator's own documented crash-window recovery -- had no
equivalent permission and 403'd forever. Closes terdut-operator#3.
2. handleCreateInvite wrote a service-account caller's zero-value user id
straight into invites.created_by (nullable, but never passed as nil),
which foreign-key-violates against users(id) -- a 500, not success, for
any team-scoped service account minting an invite. Fixed the same way
handleCreateServiceAccount already handles the analogous case. Found
live while verifying this change, not filed separately since it's fixed
in the same place it was found.
3. handleMe and handleTestNotification 500'd for a service-account caller
(fetchUser/the ntfy_topic lookup against a zero-value user id that
matches no row); handleDismissOnboarding silently no-op'd (UPDATE ...
WHERE id = 0). All three now call Caller.AsHuman() and return an
explicit 403 ("this endpoint is for human accounts only").
4. Ratifies, rather than further narrows, two capabilities a team-scoped
service account already had by construction and this document's own
text once called "a gap acknowledged rather than closed": owner-equivalent
reach over membership/invites, and minting another service account for
its own team. terdut-operator's new TerdutTeam invite-minting feature is
about to depend on the first one, so this makes it documented, tested,
intentional behavior instead of an accident nobody was supposed to rely
on.
AdminOnly/requireSelfOrAdmin are unchanged in effect: still human-only,
forever, for every scope of service account -- confirmed by
TestAdminOnly_RefusesEveryServiceAccountScope. terdut-server#23's named
routes (POST /api/users, PUT /api/admin/settings) were never the right
thing to widen; its real fix is the terdut-operator invite feature,
recorded in SERVICE-ACCOUNTS.md's "What this unblocks" and closing that
issue once it ships.
SERVICE-ACCOUNTS.md amended in place (not a new file, its own established
convention) to describe the as-built Caller model, correct its own
aspirational claim about AdminOnly that TEAM-LOOKUP.md had already flagged
as not matching shipped code, and record all of the above.
make helm-package passes --version and --app-version from the tag, so
these fields decide nothing about what is published -- but a tree heading
for v0.33.2 that still says 0.33.1 tells its reader something false.
Same as 4b15079 and 6f8499f before it.
A Deployment created before Postgres has finished its very first boot --
initdb plus Patroni leader election, on a from-scratch postgres-operator
cluster -- crash-looped a few times. db.Open()'s own ping-retry budget
(pingAttempts/pingRetryDelay, internal/db/db.go) is sized for a much
shorter, different race -- NetworkPolicy propagation, a few seconds -- not
for genuine first-time cluster creation, which routinely takes longer, so
it exhausted and the process exited before ever binding its HTTP port. A
startupProbe cannot fix that: the crash happens before there is anything
to probe.
Added a wait-for-postgres init container instead: it loops pg_isready
against database.dsn until Postgres actually answers, before the main
container's own, unchanged retry budget gets a chance to run out.
pg_isready needs no credentials -- it reports PQPING_OK on anything that
amounts to a Postgres backend answering, including an auth challenge --
so no PGPASSWORD is wired into it.
Chart-only; no Go code changed. database.waitForPostgres.enabled defaults
to true and can be turned off if something else already guarantees
Postgres is reachable before this Deployment is created.
Cosmetic: `make helm-package` passes --version/--app-version from the
tag, so this field decides nothing about what gets published. Still done
so the tree doesn't say 0.33.0 while heading for a v0.33.1 release.
Cites fc9f47c, the fix this version actually is.
Closes the race terdut-operator#1 found: /api/bootstrap and OIDC
auto-provisioning both key off the same signal (SELECT COUNT(*) FROM
users) with no coordination between them, so an otherwise-ordinary OIDC
sign-in against a freshly-created, not-yet-bootstrapped install could
create user #1 itself and take the one slot /api/bootstrap expects to win
uncontested (terdut-operator's own design, DESIGN.md §1/§6, assumes it is
the only caller). The operator has no way to recover from losing that
race -- it never gets a credential, and nothing it owns can clear the
occupying user row.
resolveSSOUser now checks the same gate handleBootstrap already does,
right where it's about to create a brand-new user (an identity nobody has
linked yet, that also matches no existing local account by email) -- not
anywhere else, since every other sign-in on an already-bootstrapped
install is unaffected. New sso_error code `not_bootstrapped`: the person
sees "this install is still setting up, try again in a moment" and a
second attempt once something has actually bootstrapped succeeds
normally, same as any other first sign-in.
Does not fix the other half of that issue (BootstrapStateLost's own
"delete and recreate" instructions still don't work once something has
occupied the slot some other way) -- this closes the specific race, not
every path to that state.
Resolves the gap TEAM-LOOKUP.md raised: an instance-scoped service
account had no way to recover a team's id after a 409 on POST
/api/teams, unlike the already-solved equivalent for service accounts
themselves (GET /api/service-accounts?name=).
Same route, extended the same way handleListServiceAccounts already
branches on ?name=: unset behaves exactly as before (the caller's own
teams via team_members); set looks up one team by exact name, open to
any authenticated caller -- not gated by isInstanceServiceAccount or
AdminOnly, since what it discloses (a name is taken, nothing about who's
in it) is the same low sensitivity that lookup already accepts for
service-account names.
Tests cover the exact motivating scenario (create, 409 on a retry,
recover the id via ?name=), the empty-array-not-an-error case, that no
role/source is reported for a non-member match, and that a caller who
isn't a member of the matched team still gets it.
Raised by terdut-operator's TerdutTeam controller (ROADMAP.md Stage 2):
an instance-scoped service account has no way to recover a team's id
after a 409 on POST /api/teams, unlike the equivalent, already-solved
case for service accounts themselves (GET /api/service-accounts?name=).
Also corrects a claim in SERVICE-ACCOUNTS.md's own text that doesn't
match AdminOnly's actual code -- confirmed against source, not assumed.
Cosmetic: make helm-package passes --version and --app-version from the
tag, so these two fields decide nothing about what gets published. Kept
in step anyway, the same as 0ee576f (0.32.0) and 949d659 (0.31.0) before
it, so a tree heading for v0.33.0 does not read as still being on 0.32.0.
The last commit (a4dd60f) shipped the code with no README update, against
this repo's own convention of documenting the whole API surface there.
Adds the Service accounts section and table, the Authentication and Teams
prose covering the new principal and TERDUT_OPERATOR_MODE, and the
/api/version row.
Corrects one thing along the way: a first draft claimed team-scoped
accounts are refused on team membership/invite endpoints. Checked against
teams.go and that is false — nothing server-side carves those two out,
only convention (no sane operator would call them) keeps them out of
automation's hands. README and SERVICE-ACCOUNTS.md now say that plainly
instead of the stronger, incorrect claim.
Service accounts (SERVICE-ACCOUNTS.md) are a scoped, non-human credential:
not a users row, so they never touch OIDC sync, login or the is_admin flag.
Instance scope can create a team and mint a team-scoped account for it;
team scope is owner-equivalent for that one team and nothing else. This is
what unblocks terdut-operator's DESIGN.md §6 — no more impersonating a
human admin, and a real rotation story instead of the unworkable
delete-and-re-bootstrap /api/bootstrap can't actually do.
- migration 014: service_accounts + service_account_keys
- POST /api/service-accounts, POST/DELETE .../keys, GET ?name= self-lookup
- AuthMiddleware resolves a tdsa_-prefixed key to a distinct principal;
a team-scoped account gets a synthetic single membership so
requireTeamMember/requireTeamOwner work on it unmodified
- handleCreateTeam accepts an instance-scoped caller; the team it creates
has no human owner, which is the expected shape for one an operator is
about to hand a team-scoped credential to
Operator mode (TERDUT_OPERATOR_MODE / values.operatorMode) declares an
install gitops-managed: session and user-API-key writes to teams,
escalation policies, dead man's switches and integrations get 403
reason=operator_managed, while a service account's writes still go
through. Team membership/invites and the schedule are deliberately left
out — never gitops-managed by design, and still human day-to-day work.
/api/auth/config reports operator_mode so the web UI can grey these
sections out from the start rather than only after a write fails.
Also: GET /api/version (both terdut-tui and terdut-operator currently
detect server capability by route-probing; this gives them a real answer),
and a PUT for dead man's switches so a reconciler can update one in place
instead of deleting and recreating it.
Proposes a non-human credential type — service_accounts +
service_account_keys, instance- or team-scoped, distinct from both
user API keys (always tied to a human's full rights) and integration
keys (narrow, one-way webhook auth only). Directly unblocks
terdut-operator's DESIGN.md §6, whose bootstrap/rotation plan doesn't
work against /api/bootstrap's actual single-shot-per-install behavior.
Design note only, no implementation yet.
Cosmetic: make helm-package sets the published version and appVersion
from the tag, so these two fields decide nothing (see the comment
above them). Kept in step anyway, same as 949d659 and e5b4df7, so a
tree heading for v0.32.0 doesn't say 0.31.0.
Both sat open by default, competing with the rest of the page for
attention on every visit even though most visits need neither. Team's
rota already has the same problem for its bulk-assign form and solves
it with a native <details>/<summary> disclosure, styled generically in
app.css; this reuses that idiom rather than inventing a JS toggle.
Each section now shows a one-line status (the topic, or whether a
password is set) with the actual form folded under a summary naming
the action ("Set a topic" / "Change topic", "Set a password" /
"Change password"). A successful save closes the fold and confirms
with a toast, since the point of folding is that a saved form goes
back to being just a status line; a validation or API error keeps the
fold open and shows inline, next to the field it's about.
The password section's heading no longer says "Change password" or
"Set a password" itself, since that verb now lives on the summary; it
just says "Password", matching the existing SSO-off case.
Cosmetic: make helm-package sets the published version and appVersion
from the tag, so these two fields decide nothing (see the comment
above them). Kept in step anyway, same as e5b4df7 and 97a4814, so a
tree heading for v0.31.0 doesn't say 0.30.0.
state.js's currentTeam() was hard-coded to teams[0] and never really meant
"the team currently selected" — team.js's settings page and queue.js's
filter chips each kept their own separate, unsynchronized notion of "which
team" instead, so picking one on one page had no effect on the other.
Replaces both with a single state.selectedTeamID, set only through the new
setSelectedTeam (persisted in localStorage, unlike the queue's old per-tab
sessionStorage filter) and broadcast to listeners via onTeamChange. A new
teamselector.js control — a coloured dot plus the team's name, or "All
teams" — sits at the top of both the desktop sidebar and the mobile topbar,
opening the existing bottom-sheet menu to switch. Shown only once someone is
in more than one team, matching every other team-aware control in this app.
Colours come from a new teamColorClass() in format.js, hashing a team's id
into the six-colour rc1..rc6 palette already used for the rota's per-person
chips, so no schema or API change is needed. The queue's team filter chips
pick up the same colours.
Cosmetic: make helm-package sets the published version and appVersion
from the tag, so these two fields decide nothing (see the comment
above them). Kept in step anyway, same as 97a4814 and 155f27c, so a
tree heading for v0.30.0 doesn't say 0.29.1.
Team membership from single sign-on used to come from one env var,
TERDUT_OIDC_GROUP_MAPPINGS, matched against a team by name and creating
the team if none existed. That put the decision in the server's
environment rather than the team's own hands, needed a restart to
change, and let a typo in a team name silently create a stray team.
Each team now carries its own oidc_member_group and oidc_owner_group,
set by its owner (or an administrator) from the Members tab, or PUT
/api/teams/{teamID}/oidc-groups. The "highest role wins" rule
TERDUT_OIDC_GROUP_MAPPINGS used to apply across mappings now applies
across one team's own two fields: being in both makes somebody an
owner. The sync no longer creates a team by name; a group only ever
grants into a team that already exists.
This is a breaking change for anyone already using
TERDUT_OIDC_GROUP_MAPPINGS, deliberately not auto-migrated: an
OIDC-sourced membership is dropped at a user's next sign-in until its
team's owner re-sets the group. The README's OIDC section spells out
the migration and the risk of a visible access gap during it.
TERDUT_OIDC_ADMIN_GROUP and TERDUT_OIDC_ALLOWED_GROUPS are untouched --
only team membership moved. terdut-tui needs no change: it only reads
GET /api/teams and GET /api/teams/{id}/members, and neither response
shape moved.
Cosmetic: make helm-package sets the published version and appVersion
from the tag, so these two fields decide nothing (see the comment
above them). Kept in step anyway, same as 155f27c and c5be55d, so a
tree heading for v0.29.1 doesn't say 0.29.0.
The image is built FROM scratch and carried only the binary, so it had no
trust store, and every HTTPS call failed with "x509: certificate signed by
unknown authority". Nothing needed one until v0.29.0: OIDC discovery and the
token exchange are HTTPS calls to the identity provider, and the first
sign-in against Authentik died in discovery. The tests could not see it,
because they run on the host, whose trust store is fine.
The builder's ca-certificates.crt is copied in by name, so a missing file
fails the build instead of shipping an image that cannot sign anybody in.
Verified by fetching the provider's discovery URL from a scratch image with
and without the bundle: the same x509 error, then 200.
Password login and everything that talks only to Postgres were unaffected.
Cosmetic: make helm-package sets the published version and appVersion
from the tag, so these two fields decide nothing (see the comment
above them). Kept in step anyway, same as c5be55d and 36c00ac, so a
tree heading for v0.29.0 doesn't say 0.28.1.
terdut can now sign people in through any OIDC provider (written against
Authentik), and let groups at the provider decide who may sign in, which
teams they belong to and whether they administer the install. Password
login keeps working alongside it; TERDUT_PASSWORD_LOGIN=false turns it off,
and is refused at startup unless SSO is configured. With no TERDUT_OIDC_*
setting nothing changes, so every existing install behaves as before.
Identity is (issuer, subject), never email or username: those are mutable
at the provider and a recycled address must not inherit an account. An
existing user is linked by email only when the provider marks it verified,
or TERDUT_OIDC_TRUST_EMAIL is set, which Authentik needs.
Group grants are marked source='oidc' on team_members and users, and the
sync changes only those rows. Hand-made memberships and administrators
are left alone, and the sync bypasses the last-owner and last-admin guards
because the provider is the source of truth for what it grants. Editing
managed access by hand is refused with 409, since the next sign-in would
undo it. The web UI badges it as SSO and disables the controls.
Groups are read only at sign-in, so an SSO session carries a hard ceiling
(sessions.max_expires_at, 12h by default) that sliding never extends.
There is no refresh token, which means API keys of somebody removed at the
provider stay valid until an administrator disables the user. That is
accepted and documented, not fixed.
A client with no browser, the TUI over SSH, signs in with a device code
run by terdut itself (POST /api/oidc/device and /device/token), so the
terminal never talks to the provider and ends up with the ordinary
terdut_session cookie. Only a browser session can approve a code; an API
key cannot. /device?code= sends a signed-out visitor through sign-in and
back, which is what oidc_logins.next is for.
oauth2 is pinned to v0.36.0: v0.37 needs Go 1.26 and the Dockerfile
builds on 1.25.
Migrations 011 and 012 add tables and defaulted columns only.
Cosmetic: make helm-package sets the published version and appVersion
from the tag, so these two fields decide nothing (see the comment
above them). Kept in step anyway, same as 36c00ac and 9bf4c92, so a
tree heading for v0.28.1 doesn't say 0.28.0.
format.js defined isoWeek twice. 9d1df2b added a second copy for the
week numbers on the rota without noticing the first, which already
exported the same function. A duplicate function declaration is legal in
a plain script but an early SyntaxError in an ES module, so the browser
refused format.js, every module importing it, and with them app.js. Its
boot() is what hides the spinner, so nothing ever did.
Keeps the first definition, whose comment explains the Thursday rule, and
drops the second. Both give ISO 8601 numbers; the one kept was checked
against 2026-01-01 (week 1), 2026-09-28 (40), 2026-12-31 (53) and
2024-12-30 (1). oncall.js and team.js import the name unchanged.
The suite could not see this: it has no JS, and node --check reads a .js
file as a script, where the redeclaration passes. Checking each file as a
module (.mjs) does catch it.
Cosmetic: make helm-package sets the published version and appVersion
from the tag, so these two fields decide nothing (see the comment
above them). Kept in step anyway, same as 9bf4c92 and 2b396d2, so a
tree heading for v0.28.0 doesn't say 0.27.0.
The rota grid starts each row with its ISO week number, and for an owner
the number is a button: one tap opens a sheet for that week, showing who
holds each of its seven days, and puts one person on all of them. A rota
is usually handed out by the week, and seven taps on seven days was the
only way to do it short of the range form.
ISO 8601 numbering, because the grid already runs Monday to Sunday: week
1 is the one holding the year's first Thursday, which is taken from the
Thursday of the row so the year boundaries come out right (2025-12-29 is
week 1 of 2026, 2020-12-31 is week 53).
Days already past are left alone. Who was on call last Tuesday is a fact,
and "the whole week" should not rewrite it, so a half-elapsed week covers
the days still to come and the sheet says so; a week that is entirely
over has nothing to assign. By default the assignment replaces whoever
holds those days, as the day sheet does and the sheet states, and a
checkbox limits it to the days nobody has yet. The overhang into the
neighbouring month is part of the same week and is included.
Web UI only: the existing schedule endpoint already takes a list of
dates and a replace flag, so nothing changed on the server and there is
nothing to mirror in terdut-tui.
Cosmetic: make helm-package sets the published version and appVersion
from the tag, so these two fields decide nothing (see the comment
above them). Kept in step anyway, same as 2b396d2 and dc92f51, so a
tree heading for v0.27.0 doesn't say 0.26.0.
Team -> Members was a two-column table and an inline add form. It is now
a table in the style of Switches, Sources and Escalation: a status
badge, the member, their role, their next rota day, when they were last
active and when they joined. Adding a member and changing a role moved
into sheets, and removing one asks first.
The badge is the one that matters at 03:00: On call if the rota has them
today, Reachable if they have an ntfy topic, and Can't be paged when a
page to them would go nowhere -- no topic, or a disabled account -- with
the reason under their name. Not being pageable wins over being on call,
since an on-call person nobody can reach is the case worth seeing before
an incident finds it. The rules are the notifier's own. The topic itself
is never in the response, only whether one is set.
Last active is the newer of a member's newest session and API-key use,
and is shown to every member of the team like the rest of the list.
Rota days are UTC dates, and the page formats them as such so a day
cannot show up as the one before.
Removing a member leaves the rota days already assigned to them alone,
which the confirm says, so they are reassigned from the Rota tab rather
than silently dropped.
Demoting the last owner is now refused with 409, as removing them
already was: it was the same outcome by another route, a team with
nobody who can edit it.
API: GET /members gains status, on_call, next_shift, pageable, problem
and last_active_at; additive, no migration, and terdut-tui needs
nothing. POST /members answers 409 for the last-owner demotion.
Cosmetic: make helm-package sets the published version and appVersion
from the tag, so these two fields decide nothing (see the comment
above them). Kept in step anyway, same as dc92f51 and e8d45f9, so a
tree heading for v0.26.0 doesn't say 0.25.0.
Team -> Escalation was the draft form on the page, which showed the
ladder only as inputs. It is now a table in the style of Switches and
Sources: a row per level with a status badge, who it pages, the wait
before the next level, and the open incidents currently waiting on it.
Below it, the repeat count, the fallback topic and when the ladder last
escalated (linking the incident). The editor moved into an "Edit ladder"
sheet, so a poll of the page underneath can no longer throw away half an
edit, and the page-level draft state went with it.
Targets are resolved to who they mean today, and the badge says what
would actually happen: Ready, Escalating (an unanswered incident has
climbed to level 2 or higher), or Pages nobody. The last is the one worth
seeing before an incident finds it: an empty rota, a person with no ntfy
topic or a disabled account each make a rung a silence with a number on
it, and the target says which. The rules are pageLevel's own, so the
page cannot promise a page the notifier would skip.
"Last escalated" comes from the escalated timeline events that already
exist, so there is no migration. Acknowledging or resolving takes an
incident off the ladder, so Escalating clears then while the history
stays.
API: GET /escalation gains status and waiting per level, username,
reachable and problem per target, and last_escalated_at and
last_escalated_incident_id. Output only and additive; PUT is unchanged
and terdut-tui needs nothing.
Cosmetic: make helm-package sets the published version and appVersion
from the tag, so these two fields decide nothing (see the comment
above them). Kept in step anyway, same as e8d45f9 and 3ee8583, so a
tree heading for v0.25.0 doesn't say 0.24.0.
Like Team -> Switches, the page is now a table: a status badge (Active
if the key posted within a day, Quiet if it has but not lately, Never
used), when it last posted a webhook, when an alert last arrived on it,
how many distinct alerts it refreshed in the last 24 hours, and when it
was created. Adding a source moved into a "New source" sheet, and owners
can rename one from its row.
"Last alert" and the count needed alerts to remember which source they
came in on, which they never did, so migration 010 adds
alerts.integration_id and every accepted payload stamps it. Last sender
wins when two sources post the same fingerprint. It is not backfilled: a
NULL says "before this was recorded" rather than guessing, and it heals
by itself as Alertmanager re-sends each alert every repeat_interval.
Revoking a source keeps its alerts, unattributed.
Last webhook and last alert are separate on purpose: a payload with
nothing usable in it stamps the first and not the second. The Quiet
threshold is a fixed day, a colour and not an alarm, since silence that
should page is what dead man's switches are for.
The counts are indexed subqueries (alerts_integration_idx) rather than a
join, which would read every alert a source ever delivered.
API: the integrations list gains status, last_alert_at and alerts_24h,
and PATCH /api/teams/{id}/integrations/{id} renames. Both are additive;
terdut-tui needs nothing.
Cosmetic: make helm-package sets the published version and appVersion
from the tag, so these two fields decide nothing (see the comment
above them). Kept in step anyway, same as 3ee8583 and d2cdcc9, so a
tree heading for v0.24.0 doesn't say 0.23.0.
The page was a bare form: it did not say which switches existed or
whether they were alive. It now lists them, each with a Healthy, Dead or
Dormant badge, when its heartbeat was last heard and when it last opened
an incident (linked while that incident is open). A matcher that several
clusters satisfy is broken down per cluster, since a live cluster must
not hide a dead one. The form moved into a "New switch" sheet, and each
row has a Remove with a confirm.
That needed a switch to be a thing, so switches are rows now
(migration 009) with their own name, matcher, timeout and severity,
instead of one string with one team-wide timeout in deadman_configs.
Existing configuration is split into one row per matcher; a team whose
timeout was zero simply has none. The sweeper and the status endpoint
share one death rule (deadmanAlert.dead), so the page cannot disagree
with the pager. Incident group keys are unchanged, so incidents that
are open across the upgrade keep working.
The environment defaults (TERDUT_DEADMAN_*) are seeded into teams once
per install, recorded in settings, so a team that deletes its last
switch does not get it back on the next restart. Installs that already
had per-team rows are marked as seeded by the migration.
Removing a switch stops the watching but leaves an incident it already
opened open until someone resolves it.
API: GET/PUT /api/teams/{id}/deadman are replaced by
GET/POST /deadman/switches and DELETE /deadman/switches/{switchID}.
terdut-tui does not call them, so nothing to mirror there.
Cosmetic: make helm-package sets the published version and appVersion
from the tag, so these two fields decide nothing (see the comment
above them). Kept in step anyway, same as d2cdcc9 and 71d7e18, so a
tree heading for v0.23.0 doesn't say 0.22.1.
A button in the incident header (also `y`, and "Copy incident" in the
more menu) puts everything the page knows on the clipboard, for pasting
into a chat or an agent prompt with no integration involved.
The text carries the facts, every alert with all its labels and
annotations (the page only shows summary or description), the timeline
with notes in full, and the "Seen before" resolution notes. Times are
ISO 8601 and users are named rather than "you", since relative and
first-person wording is ambiguous once pasted elsewhere.
The async clipboard API needs a secure context and this server is often
reached over plain HTTP, so it falls back to execCommand.
Web UI only: no endpoint or JSON shape changed, so nothing to mirror in
terdut-tui.
Cosmetic: make helm-package sets the published version and appVersion
from the tag, so these two fields decide nothing (see the comment
above them). Kept in step anyway, same as 71d7e18 and 734cd9c, so a
tree heading for v0.22.1 doesn't say 0.22.0.
Claude-Session: https://claude.ai/code/session_01MMados3BD1oSjevHxbmVqU
The list pane is 340-420px wide and its chip row scrolled sideways with the
scrollbar hidden. That works by swipe on a phone, but a mouse has nothing to
grab, so Archived (the last chip) could not be reached on a wide screen. In
the desktop layout the row now wraps instead, and the divider between the
status and team chips is hidden there, since it would sit mid-line.
Phones keep the sideways scroll: the rule is inside the min-width: 900px
block.
Claude-Session: https://claude.ai/code/session_01MMados3BD1oSjevHxbmVqU
Cosmetic: make helm-package sets the published version and appVersion
from the tag, so these two fields decide nothing (see the comment
above them). Kept in step anyway, same as 734cd9c and 43f0044, so a
tree heading for v0.22.0 doesn't say 0.21.0.
Claude-Session: https://claude.ai/code/session_01MMados3BD1oSjevHxbmVqU
Each incident gets a signature: the alert name plus the group labels that
say what is broken, minus the ones that only say where it ran (instance,
pod, container, ...). GET /api/incidents/{id}/similar returns resolved
incidents in the same team with the same signature that have notes.
Notes can be marked as the resolution note, "what fixed it", either with a
resolution field on resolve or pinned on a note. Those lead the similar
list, show on the incident page as "Seen before", and the triggered
notification carries the latest one.
Claude-Session: https://claude.ai/code/session_01MMados3BD1oSjevHxbmVqU
Cosmetic: make helm-package sets the published version and appVersion
from the tag, so these two fields decide nothing (see the comment
above them). Kept in step anyway, same as 43f0044 and e77f04b, so a
tree heading for v0.21.0 doesn't say 0.20.1.
Statistics used to live only in terdut-tui; the account page said so.
The page shows the same figures as the TUI's Stats tab -- incident
counts, MTTA and MTTR, top alerts, and alert frequency by hour (UTC) and
by day of week -- and adds a range picker (Today, 7d, 30d, 90d, All)
that the TUI does not have. The ranges are day-granular because the
server reads from/to as whole UTC dates, so there is no 24h chip.
No server change: the page uses the existing /api/stats/* endpoints,
which already scope to the caller's teams. Charts are inline SVG and
plain elements sized from script, because the CSP forbids inline styles,
inline scripts and CDN libraries.
Removes the "statistics are in terdut-tui" notes from the account page
and the README.
Cosmetic: make helm-package sets the published version and appVersion
from the tag, so these two fields decide nothing (see the comment
above them). Kept in step anyway, same as e77f04b and 429d5fd, so a
tree heading for v0.20.1 doesn't say 0.20.0.
Claude-Session: https://claude.ai/code/session_01RHPj4ggeFdEjKKfm4SHbD7
Every start crashed once or twice before going healthy: kube-router
enforces this namespace's NetworkPolicy per-node, reacting to the new
pod's creation event, and the app's first connection attempt can
reach Postgres's node before that node's allow-set has been updated
with the new pod's IP. The result is "connection refused" -- an
active reject, not a timeout, which is how it was told apart from
Postgres itself not being ready (it had been up for two days in the
run that was diagnosed).
That race resolves within several seconds in practice, so Open now
retries the ping up to five times, two seconds apart, logging each
failure, before giving up with the same wrapped error as before.
Nothing else about Open's behaviour changed: a genuinely absent
database still fails, just after ~8s instead of immediately.
Claude-Session: https://claude.ai/code/session_01RHPj4ggeFdEjKKfm4SHbD7
Cosmetic: make helm-package sets the published version and appVersion
from the tag, so these two fields decide nothing (see the comment
above them). Kept in step anyway, same as 429d5fd and 6a03698, so a
tree heading for v0.20.0 doesn't say 0.19.0.
Claude-Session: https://claude.ai/code/session_01RHPj4ggeFdEjKKfm4SHbD7