155f27ca62734a975f4420c9c53bed001c0d7de2
45 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
155f27ca62 |
Set the chart's placeholder version to 0.29.0
CI / chart (push) Successful in 1s
CI / security (push) Successful in 19s
CI / test (push) Successful in 4m1s
Release / test (push) Successful in 5s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 46s
Release / image (push) Successful in 1m17s
Release / scan-image (push) Successful in 5s
Cosmetic: make helm-package sets the published version and appVersion from the tag, so these two fields decide nothing (see the comment above them). Kept in step anyway, same as |
||
|
|
a27ff49171 |
Sign in through an OpenID Connect provider, and from a terminal
terdut can now sign people in through any OIDC provider (written against Authentik), and let groups at the provider decide who may sign in, which teams they belong to and whether they administer the install. Password login keeps working alongside it; TERDUT_PASSWORD_LOGIN=false turns it off, and is refused at startup unless SSO is configured. With no TERDUT_OIDC_* setting nothing changes, so every existing install behaves as before. Identity is (issuer, subject), never email or username: those are mutable at the provider and a recycled address must not inherit an account. An existing user is linked by email only when the provider marks it verified, or TERDUT_OIDC_TRUST_EMAIL is set, which Authentik needs. Group grants are marked source='oidc' on team_members and users, and the sync changes only those rows. Hand-made memberships and administrators are left alone, and the sync bypasses the last-owner and last-admin guards because the provider is the source of truth for what it grants. Editing managed access by hand is refused with 409, since the next sign-in would undo it. The web UI badges it as SSO and disables the controls. Groups are read only at sign-in, so an SSO session carries a hard ceiling (sessions.max_expires_at, 12h by default) that sliding never extends. There is no refresh token, which means API keys of somebody removed at the provider stay valid until an administrator disables the user. That is accepted and documented, not fixed. A client with no browser, the TUI over SSH, signs in with a device code run by terdut itself (POST /api/oidc/device and /device/token), so the terminal never talks to the provider and ends up with the ordinary terdut_session cookie. Only a browser session can approve a code; an API key cannot. /device?code= sends a signed-out visitor through sign-in and back, which is what oidc_logins.next is for. oauth2 is pinned to v0.36.0: v0.37 needs Go 1.26 and the Dockerfile builds on 1.25. Migrations 011 and 012 add tables and defaulted columns only. |
||
|
|
c5be55dcbc |
Set the chart's placeholder version to 0.28.1
CI / chart (push) Successful in 2s
CI / security (push) Successful in 15s
CI / test (push) Successful in 2m51s
Release / test (push) Successful in 4s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 31s
Release / image (push) Successful in 59s
Release / scan-image (push) Successful in 3s
Cosmetic: make helm-package sets the published version and appVersion from the tag, so these two fields decide nothing (see the comment above them). Kept in step anyway, same as |
||
|
|
36c00acf62 |
Set the chart's placeholder version to 0.28.0
CI / chart (push) Successful in 1s
CI / security (push) Successful in 17s
CI / test (push) Successful in 3m1s
Release / test (push) Successful in 5s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 23s
Release / image (push) Successful in 55s
Release / scan-image (push) Successful in 2s
Cosmetic: make helm-package sets the published version and appVersion from the tag, so these two fields decide nothing (see the comment above them). Kept in step anyway, same as |
||
|
|
9bf4c92bfe |
Set the chart's placeholder version to 0.27.0
CI / chart (push) Successful in 1s
CI / security (push) Successful in 14s
CI / test (push) Successful in 2m59s
Release / test (push) Successful in 8s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 19s
Release / image (push) Successful in 54s
Release / scan-image (push) Successful in 27s
Cosmetic: make helm-package sets the published version and appVersion from the tag, so these two fields decide nothing (see the comment above them). Kept in step anyway, same as |
||
|
|
2b396d22d6 |
Set the chart's placeholder version to 0.26.0
CI / chart (push) Successful in 1s
CI / security (push) Successful in 14s
CI / test (push) Successful in 2m56s
Release / test (push) Successful in 4s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 22s
Release / image (push) Successful in 57s
Release / scan-image (push) Successful in 4s
Cosmetic: make helm-package sets the published version and appVersion from the tag, so these two fields decide nothing (see the comment above them). Kept in step anyway, same as |
||
|
|
dc92f51cf8 |
Set the chart's placeholder version to 0.25.0
CI / chart (push) Successful in 1s
CI / security (push) Successful in 31s
CI / test (push) Successful in 3m20s
Release / test (push) Successful in 4s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 20s
Release / image (push) Successful in 1m2s
Release / scan-image (push) Successful in 2s
Cosmetic: make helm-package sets the published version and appVersion from the tag, so these two fields decide nothing (see the comment above them). Kept in step anyway, same as |
||
|
|
e8d45f9d3d |
Set the chart's placeholder version to 0.24.0
CI / chart (push) Successful in 1s
CI / security (push) Successful in 16s
CI / test (push) Successful in 2m55s
Release / test (push) Successful in 4s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 24s
Release / image (push) Successful in 1m0s
Release / scan-image (push) Successful in 4s
Cosmetic: make helm-package sets the published version and appVersion from the tag, so these two fields decide nothing (see the comment above them). Kept in step anyway, same as |
||
|
|
3ee8583f6f |
Set the chart's placeholder version to 0.23.0
CI / chart (push) Successful in 1s
CI / security (push) Successful in 14s
CI / test (push) Successful in 2m35s
Release / test (push) Successful in 4s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 21s
Release / image (push) Successful in 57s
Release / scan-image (push) Successful in 2s
Cosmetic: make helm-package sets the published version and appVersion from the tag, so these two fields decide nothing (see the comment above them). Kept in step anyway, same as |
||
|
|
d2cdcc9776 |
Set the chart's placeholder version to 0.22.1
CI / chart (push) Successful in 1s
CI / security (push) Successful in 16s
CI / test (push) Successful in 2m45s
Release / test (push) Successful in 4s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 21s
Release / image (push) Successful in 59s
Release / scan-image (push) Successful in 2s
Cosmetic: make helm-package sets the published version and appVersion from the tag, so these two fields decide nothing (see the comment above them). Kept in step anyway, same as |
||
|
|
71d7e1853a |
Set the chart's placeholder version to 0.22.0
CI / chart (push) Successful in 1s
CI / security (push) Successful in 14s
CI / test (push) Successful in 2m41s
Release / test (push) Successful in 4s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 18s
Release / image (push) Successful in 56s
Release / scan-image (push) Successful in 5s
Cosmetic: make helm-package sets the published version and appVersion from the tag, so these two fields decide nothing (see the comment above them). Kept in step anyway, same as |
||
|
|
734cd9c5fd |
Set the chart's placeholder version to 0.21.0
CI / chart (push) Successful in 2s
CI / security (push) Successful in 16s
CI / test (push) Successful in 2m52s
Release / test (push) Successful in 4s
Release / chart (push) Successful in 3s
Release / binaries (push) Successful in 18s
Release / image (push) Successful in 58s
Release / scan-image (push) Successful in 25s
Cosmetic: make helm-package sets the published version and appVersion from the tag, so these two fields decide nothing (see the comment above them). Kept in step anyway, same as |
||
|
|
43f004499b |
Set the chart's placeholder version to 0.20.1
CI / chart (push) Successful in 1s
CI / security (push) Successful in 13s
CI / test (push) Successful in 2m30s
Release / test (push) Successful in 4s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 15s
Release / image (push) Successful in 55s
Release / scan-image (push) Successful in 2s
Cosmetic: make helm-package sets the published version and appVersion from the tag, so these two fields decide nothing (see the comment above them). Kept in step anyway, same as |
||
|
|
e77f04b55e |
Set the chart's placeholder version to 0.20.0
CI / chart (push) Successful in 2s
CI / security (push) Successful in 20s
CI / test (push) Successful in 2m35s
Release / test (push) Successful in 4s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 18s
Release / image (push) Successful in 55s
Release / scan-image (push) Successful in 2s
Cosmetic: make helm-package sets the published version and appVersion from the tag, so these two fields decide nothing (see the comment above them). Kept in step anyway, same as |
||
|
|
429d5fdda3 |
Set the chart's placeholder version to 0.19.0
CI / chart (push) Successful in 1s
CI / security (push) Successful in 13s
CI / test (push) Successful in 2m32s
Release / test (push) Successful in 4s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 24s
Release / image (push) Successful in 55s
Release / scan-image (push) Successful in 2s
Cosmetic, and done anyway for the same reason as |
||
|
|
6a03698f65 |
Set the chart's placeholder version to 0.18.0
CI / chart (push) Successful in 1s
CI / security (push) Successful in 14s
CI / test (push) Successful in 2m40s
Release / test (push) Successful in 4s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 21s
Release / image (push) Successful in 58s
Release / scan-image (push) Successful in 23s
Cosmetic, and done anyway for the same reason as |
||
|
|
ee22eb000c |
Set the chart's placeholder version to 0.17.0
CI / chart (push) Successful in 1s
CI / security (push) Successful in 13s
CI / test (push) Successful in 2m31s
Release / test (push) Successful in 8s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 19s
Release / image (push) Successful in 53s
Release / scan-image (push) Successful in 3s
Cosmetic, and done anyway for the same reason as |
||
|
|
7b9a337d25 |
Set the chart's placeholder version to 0.16.1
CI / chart (push) Successful in 1s
CI / security (push) Successful in 16s
CI / test (push) Successful in 2m33s
Release / test (push) Successful in 4s
Release / chart (push) Successful in 4s
Release / binaries (push) Successful in 28s
Release / image (push) Successful in 1m5s
Release / scan-image (push) Successful in 2s
Cosmetic, and done anyway. release.yaml passes --version and --app-version from the git tag when it packages, so neither line decides anything about what is published; they exist to be read by somebody looking at the tree before the tag does. A tree heading for v0.16.1 that says 0.16.0 tells that reader something false. Its own commit, like |
||
|
|
828cf87656 |
Set the chart's placeholder version to 0.16.0
CI / chart (push) Successful in 1s
CI / security (push) Successful in 14s
CI / test (push) Successful in 2m33s
Release / test (push) Successful in 4s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 18s
Release / image (push) Successful in 59s
Release / scan-image (push) Successful in 2s
Cosmetic, and done anyway. release.yaml passes --version and --app-version from the git tag when it packages, so neither line decides anything about what is published; they exist to be read by somebody looking at the tree before the tag does. A tree heading for v0.16.0 that says 0.15.1 tells that reader something false. Its own commit, like |
||
|
|
8869ac864f |
Set the chart's placeholder version to 0.15.1
CI / test (push) Successful in 4s
CI / chart (push) Successful in 1s
CI / security (push) Successful in 10s
Release / test (push) Successful in 4s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 22s
Release / image (push) Successful in 56s
Release / scan-image (push) Successful in 3s
Cosmetic, as in |
||
|
|
93761056eb |
Set the chart's placeholder version to 0.15.0
CI / chart (push) Successful in 4s
CI / test (push) Successful in 12s
CI / security (push) Successful in 17s
Release / test (push) Successful in 4s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 19s
Release / image (push) Successful in 1m3s
Release / scan-image (push) Successful in 24s
Cosmetic, as in |
||
|
|
4e8c52c28c |
Set the chart's placeholder version to 0.14.0
CI / test (push) Successful in 4s
CI / chart (push) Successful in 1s
CI / security (push) Successful in 12s
Release / test (push) Successful in 4s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 24s
Release / image (push) Successful in 53s
Release / scan-image (push) Successful in 2s
Cosmetic, as in |
||
|
|
53e5e03f4e |
Set the chart's placeholder version to 0.13.0
CI / test (push) Successful in 4s
CI / chart (push) Successful in 1s
CI / security (push) Successful in 11s
Release / test (push) Successful in 4s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 18s
Release / image (push) Successful in 58s
Release / scan-image (push) Successful in 7s
Cosmetic, as in |
||
|
|
4c85e7646c |
Set the chart's placeholder version to 0.12.0
CI / test (push) Successful in 5s
CI / chart (push) Successful in 2s
CI / security (push) Successful in 15s
Release / test (push) Successful in 4s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 20s
Release / image (push) Successful in 54s
Release / scan-image (push) Successful in 2s
Cosmetic, as in |
||
|
|
e3ad19c110 |
Set the chart's placeholder version to 0.11.1
CI / test (push) Successful in 4s
CI / chart (push) Successful in 1s
CI / security (push) Successful in 11s
Release / test (push) Successful in 9s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 1m1s
Release / image (push) Successful in 1m25s
Release / scan-image (push) Successful in 24s
Cosmetic, as in |
||
|
|
041e159e2a |
Set the chart's placeholder version to 0.11.0
CI / test (push) Successful in 5s
CI / chart (push) Successful in 1s
CI / security (push) Successful in 12s
Release / test (push) Failing after 4s
Release / binaries (push) Has been skipped
Release / image (push) Has been skipped
Release / chart (push) Has been skipped
Release / scan-image (push) Has been skipped
Cosmetic, as in |
||
|
|
cc31c993dd |
Take the database password from PGPASSWORD, not the DSN
The chart asked for a whole DSN in a Secret. Nothing writes one: the Zalando postgres operator generates a Secret with `username` and `password` keys and no connection string, so wiring the wrapper chart up would have meant hand-maintaining a second copy of a password the operator owns and rotates on a from-scratch rebuild -- which is charts#176 again, the issue miniflux closed by doing the opposite. So the DSN becomes a plain value with no password in it, and the password arrives as PGPASSWORD from a Secret. pgx fills in from libpq's PG* environment variables whatever the DSN omits, exactly as miniflux's lib/pq does. Verified rather than assumed, against a real server: a password-less DSN connects with PGPASSWORD set, and fails with `password authentication failed` when it is wrong, so the variable is doing the work rather than being quietly ignored. It also keeps the credential out of the rendered manifest and out of `kubectl describe pod`, which a DSN-with-password does not. |
||
|
|
dc39e3a5d3 |
Move the database to Postgres, before teams need the schema
First step of #1, and it goes first for one reason: #4 adds a team_id to nearly every table, and doing that twice -- once for SQLite, once for Postgres -- is work nobody gets paid for. The teams migrations now only have to be written against one database. The ten SQLite migrations are replaced by a single Postgres baseline rather than ported one by one. They were incremental in a way that has no value on a fresh install: 004 adds columns 008 drops again, and 008's backfill rewrites data a Postgres database never had. The history stays in git; the schema they add up to is now 001_baseline.sql. Timestamps stay BIGINT unix seconds and are NOT converted to timestamptz. Everything in Go already speaks epochs, so converting would have been a second, larger change riding along inside this one. It is worth doing on its own. The JSON columns did move to jsonb, because #4 will want to filter and index on labels. Most of the port is mechanical -- 170 placeholders from ? to $1 -- but four things needed more than a search and replace: * Dynamically built WHERE clauses cannot keep their numbering straight by hand, so they hand out placeholders through sqlArgs instead. A filter can now be added or reordered without renumbering anything. * SUM(resolved_at IS NULL) was SQLite counting a boolean as 0 or 1. Postgres has no sum(boolean), and this was breaking every dead man's switch -- silently, since the sweeper only logs. Now COUNT(*) FILTER. * unixepoch() became FLOOR(EXTRACT(EPOCH FROM now()))::bigint. The FLOOR is load-bearing: a bare cast rounds half up, so a row written at .6 of a second claimed a timestamp a second in the future and disagreed with the time.Now().Unix() the Go side stamps. * The unique-violation check matched SQLite's error text. It matches SQLSTATE 23505 now, so a renamed constraint cannot turn a 409 back into a 500. Tests need a real Postgres, because there is no in-memory Postgres the way there was an in-memory SQLite. Each test gets its own schema on a shared server -- cheaper than a database each, and still isolated. TERDUT_TEST_DSN says where it is; `make test-db` starts one locally and ci.yaml runs one as a service container. An unset DSN fails the suite rather than skipping it: a run that quietly tests nothing is worse than one that does not run. TestMigration_BackfillCarriesAckAndComments is deleted along with the migrations it replayed. What it protected -- an upgrade not losing acknowledgements and comments -- now belongs to scripts/sqlite-to-postgres.go, which is build-tagged so the SQLite driver stays out of the server binary. Both are meant to be deleted once this install has migrated. The chart loses the PVC, the data volume and the python backup sidecar, and requires database.dsnSecret.name: it provisions no database and cannot guess where the credentials live, so a render without it is meant to fail. Backups move to where Postgres actually runs. The other half of that -- the postgresql CR, the k8up pg_dump annotation and the network policy -- is a change to the wrapper chart in Ryuvia/charts and is not in here. Verified rather than assumed: the gate is green with -race against Postgres 17, govulncheck and gitleaks are clean, and the migration script was run end to end against a SQLite database built at the old schema and seeded in every table. Ids survive, so incidents keep their numbers and every foreign key still points where it did; the identity sequences are moved past the copied ids, and a webhook after the migration opened incident 12 rather than colliding at 1. |
||
|
|
989425e550 |
Set the chart's placeholder version to 0.10.2
CI / chart (push) Successful in 1s
CI / security (push) Successful in 28s
CI / test (push) Successful in 2m13s
Release / test (push) Successful in 1m9s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 1m41s
Release / image (push) Successful in 1m42s
Release / scan-image (push) Successful in 2s
Cosmetic, as in |
||
|
|
e78f49461a |
Set the chart's placeholder version to 0.10.1
CI / chart (push) Successful in 1s
CI / security (push) Successful in 27s
CI / test (push) Successful in 2m11s
Release / test (push) Successful in 1m11s
Release / chart (push) Successful in 3s
Release / binaries (push) Successful in 1m32s
Release / image (push) Successful in 1m41s
Release / scan-image (push) Failing after 29s
Cosmetic, as in |
||
|
|
4cec26edde |
Set the chart's placeholder version to 0.10.0
CI / chart (push) Successful in 2s
CI / security (push) Successful in 40s
CI / test (push) Successful in 2m47s
Release / test (push) Successful in 1m9s
Release / chart (push) Successful in 2s
Release / image (push) Failing after 42s
Release / scan-image (push) Has been skipped
Release / binaries (push) Successful in 1m57s
Cosmetic, as in |
||
|
|
9669b8f477 |
Set the chart's placeholder version to 0.9.4
CI / chart (push) Successful in 0s
CI / security (push) Successful in 24s
CI / test (push) Successful in 28s
Release / test (push) Successful in 28s
Release / chart (push) Successful in 1s
Release / binaries (push) Successful in 28s
Release / image (push) Successful in 1m11s
Release / scan-image (push) Successful in 23s
Cosmetic, as in
|
||
|
|
a7871ed7c6 |
Stop the bootstrap hook installing curl at run time
The hook's container was alpine:3 and its first line was `apk add --no-cache curl`. That writes the binary into the container's writable upper layer, and every exec of it afterwards is, correctly, a dropped binary: Falco's `Drop and execute new binary in container` (PCI_DSS_11.5.1, MITRE TA0003) fired twice at Critical on the upgrade to chart 0.9.3, 65ms after the container started, with evt.arg.flags=EXE_WRITABLE|EXE_UPPER_LAYER. Ryuvia/charts#100 has the event lines. A true positive of the rule and a false positive of intent, and it is not a one-off: the hook is post-install,post-upgrade, so it recurred on every release. The cluster is still in the Falco burn-in with detections routed to a null receiver, which is the only reason nobody was paged for it. Fixed here rather than with a Falco exception on purpose. An exception would have to name this container and would then stay in the rule set forever, blinding it for the one workload that already runs as root with create-secret RBAC, and it would leave the second problem untouched: this runs as a post-upgrade hook, a failed hook fails the release, so every `helm upgrade` of terdut-server depended on dl-cdn.alpinelinux.org answering. That dependency is now gone. alpine/curl is still a full Alpine, so sh, cat, sleep, grep, cut, head and tail are all present -- verified in-cluster before the swap rather than assumed, since a missing utility would surface as a failed post-upgrade hook and not as anything visible here. Digest-pinned, as the wrapper chart's own sidecar images are. The image declares an ENTRYPOINT, which the Job's `command:` overrides; a comment says so, because rewriting that to `args:` would silently run curl's entrypoint instead of the script. No change to the script's logic, to the RBAC, or to when the hook runs. Nothing on the terdut-tui side of the API moves, and no terdut-tui version is required or excluded by this. Worth recording while it is in view, and deliberately not acted on here: there is no terdut-server-admin-key secret in the namespace, so the POST returns 403, the hook logs "Server already bootstrapped, nothing to do" and exits before the secret-creating branch. On an upgrade this hook currently achieves nothing at all. Narrowing it to post-install would remove the detection outright, but that changes what the hook is for and belongs in its own change. Claude-Session: https://claude.ai/code/session_014m2pJdpCTv3mvvUUuBM54Y |
||
|
|
f46e5f5729 |
Set the chart's placeholder version to 0.9.3
Cosmetic, and done anyway, for the same reason as |
||
|
|
477454ec3c |
Sätt chartets platshållarversion till 0.9.2
Kosmetiskt, och görs ändå. .gitea/workflows/release.yaml stämplar både
version och appVersion från git-taggen när det publicerar (
|
||
|
|
289eca8076 |
Move to Gitea: git.ryuvia.com/niklas/terdut-server
CI / test (push) Successful in 2m15s
The module path, the container image, the Helm chart and the CI pipeline all named GitHub. They now name the Gitea instance everything else already runs on. The workflows are rewritten rather than translated. Gitea's runner image is ubuntu:22.04, whose nodejs is Node 12, so no JS action runs there at all -- actions/checkout@v4 dies with a SyntaxError before it does anything. Every step is shell, checkout is a plain clone (this repo is public, so it needs no credential), and the jobs that need docker or helm run in host mode because the dind bridge a `container:` job gets cannot reach github.com or get.helm.sh. Two consequences worth naming: - upload-artifact/download-artifact are also JS actions, and there is no artifact store here, so the job that builds the binaries is the job that publishes them. Nothing is passed between jobs. - setup-qemu-action is gone with the rest, and the runner has no binfmt registration. The Dockerfile's builder stage now runs on $BUILDPLATFORM and cross-compiles from TARGETARCH instead, which is what keeps the arm64 image buildable -- and makes it native rather than emulated. The chart moves from a GitHub Pages index to an OCI artifact in Gitea's registry. Publishing stays tag-only for the reason recorded in release.yaml: a workflow triggered by the branch push cannot know the version it is about to be tagged with. The GitHub repository is left in place and untouched. Nothing pushes to it any more, but its existing release downloads and chart index keep resolving. |
||
|
|
766f43931c |
chart: publish from the tag only, not from both workflows
Two workflows published the chart and disagreed about its metadata. release.yml stamps version and appVersion from the git tag; chart-release.yml, triggered by any charts/** push to main, took Chart.yaml verbatim, where appVersion is the hardcoded "latest". Both fired for the same commit, both tried to publish the same chart version, and skip_existing turned whichever lost into a no-op — so what a release said about itself came down to which runner was quicker. Chart 0.9.0 went out that way, reading appVersion "latest". Every earlier release got the right answer by accident: Chart.yaml's version lagged the published set, so chart-release.yml always collided with an existing version and skipped, leaving release.yml to win uncontested. Bumping Chart.yaml to match the tag before cutting 0.9.0 removed that accident and the race showed itself. Making the two agree is not possible. The tag is pushed after the branch, so a workflow triggered by the main push cannot know the version it is about to be tagged with — no amount of deriving from git describe fixes that ordering. The fix is one publisher, triggered by the tag, so chart-release.yml is deleted. The chart now only ships with an app release. Nothing is lost: the sed in release.yml ties the chart version to the app version, so a chart-only change never had a version of its own to be released under. Chart fixes ride the next tag. Chart.yaml's version and appVersion are documented as the placeholders they now are, so the next person does not helpfully bump them and reintroduce this. skip_existing stays, for idempotent re-runs of a failed release rather than for the race, and a non-version tag now fails the job instead of silently publishing unstamped metadata. |
||
|
|
14c24f8fda |
Notice when the Watchdog alert stops arriving
Release / release (push) Has been skipped
Release / build (amd64, linux) (push) Has been skipped
Release / build (arm64, darwin) (push) Has been skipped
Release / test (push) Failing after 5s
Release / build (arm64, linux) (push) Has been skipped
Release / docker (push) Has been skipped
Release / chart (push) Has been skipped
Release / build (amd64, darwin) (push) Has been skipped
Everything this server does assumes alerts arrive. If Prometheus stops evaluating, or Alertmanager cannot reach us, nothing arrives — and silence is indistinguishable from everything being fine. The cluster has shipped the alert for exactly this case all along: Watchdog is expr: vector(1), so it fires permanently and is re-sent forever, and it is worth nothing unless something downstream notices it stop. Nothing did. It arrived, opened no incident because a repeat_interval re-send is not a new occurrence, and when the monitoring stack died the sweeper quietly expired it and paged nobody. So the handling is inverted for a configurable set of alerts: receiving one opens no incident, and the absence of one does. TERDUT_DEADMAN_MATCHERS selects them as label matchers, defaulting to alertname=Watchdog. The unit of monitoring is the fingerprint rather than the alert name. Two clusters sending the same Watchdog are two independent switches, so a healthy one can never mask a dead one. Every matcher must name an alertname, which keeps the sweeper's candidate query on alerts_name_idx instead of JSON-extracting labels from every row, and leaves matching with a single implementation. A switch is dormant until its first heartbeat: a matcher nothing has ever sent opens nothing, so a fresh deploy or a restored database does not page. Resolving the incident by hand sticks, exactly as it does for an alert-backed one, so a decommissioned source is a one-time page rather than a nag; the switch re-arms only when the heartbeat comes back, and dying again is a new incident. The incident has no member alerts on purpose. Linking the heartbeat would have the settled-incident cascade close it on the very sweep that opened it, and there is no alert describing the problem anyway — the problem is that no alert arrived. What happened is on the timeline instead, and recovery is the only automatic way out. One narrow exemption in the ingest guard makes recovery possible at all. A heartbeat we declared dead is marked resolved, and the one that proves us wrong carries the unchanged startsAt of an alert that never stopped firing — so "resolution is terminal within an instance" would discard it forever and a switch could die exactly once. The exemption is scoped to resolution_source = 'deadman', which is the only resolution this server infers from silence on a timeout of its own, so nothing another writer set can be undone by a stale retry. Matched alerts are also held back from the generic staleness expiry, which would otherwise resolve a heartbeat as 'expiry' long before its own tighter deadline. The timeout points the opposite way to TERDUT_STALE_AFTER: staleness is a generous grace period around a repeat_interval you do not control, while this is a deadline you set deliberately and configure the heartbeat's route to beat. Inheriting a 4h or 12h repeat_interval gives a dead man's switch with a twelve hour fuse, so the README spells out the route the heartbeat needs. |
||
|
|
17ee290d90 |
docs: the ack token is scoped, not single-use
The handler never deletes the token: it stays valid until expires_at and is purged by the sweeper, so a second tap is an idempotent no-op rather than a rejection. Caught by pressing Acknowledge twice against the live server. What bounds the token is scope -- one incident, one action, one day -- not a use count. |
||
|
|
7caafbaf80 |
chart: back up the database through a python sidecar
Release / test (push) Failing after 7s
Release / build (amd64, darwin) (push) Has been skipped
Release / build (amd64, linux) (push) Has been skipped
Release / build (arm64, darwin) (push) Has been skipped
Release / build (arm64, linux) (push) Has been skipped
Release / docker (push) Has been skipped
Release / chart (push) Has been skipped
Release / release (push) Has been skipped
The image is FROM scratch, so there is no interpreter to run a k8up backupcommand in, and the database runs in WAL mode, where a file-level copy of the volume is not crash-consistent. Also switches to strategy: Recreate. The data PVC is ReadWriteOnce, so a RollingUpdate deadlocks the new pod against the old one holding it. |
||
|
|
bc285799d1 |
Page the on-call person when an incident opens
An incident opened, got assigned to whoever held today's schedule entry,
and then sat there silently until somebody thought to look. The schedule
and the incident model were both built; nothing reached the person
holding the pager.
Notifications go out through ntfy, over plain HTTP with no new
dependencies. Delivery is an outbox rather than an inline call: the pool
is limited to a single connection, so a POST made while holding the
webhook's transaction would stall every other request behind it. The
webhook inserts a row and a notifier goroutine sends it within a tick,
retrying with exponential backoff.
Only opening an incident has to resolve a topic from scratch. Reminders
and all-clears reuse whatever that first notification chose, which keeps
configuration out of resolveIfSettled and gives the right rule for free:
you only hear that something resolved if you were told it started.
Each push carries an Acknowledge button, because the useful thing to do
at 3am is stop the pager without unlocking anything. It POSTs to an
unauthenticated /api/notify/ack/{token} — a notification body lives on
the ntfy server and in the device cache, so a real API key must never
appear in one. The token is minted per delivery, scoped to one incident
and one action, and expires in a day.
Reminders repeat until the incident stops being untouched. The stop
conditions are the states that already mean somebody has it: acknowledged,
snoozed, resolved, archived. Snooze is the mute button, so there is no
separate reminder cap.
Notifications sent to the fallback topic carry no Acknowledge button. The
topic is shared, and a button on it would let any subscriber acknowledge
as somebody else.
|
||
|
|
dcb2a86f9a |
chart: bind the HTTPRoute to a named gateway listener
Release / test (push) Failing after 7s
Release / build (amd64, darwin) (push) Has been skipped
Release / build (amd64, linux) (push) Has been skipped
Release / build (arm64, darwin) (push) Has been skipped
Release / build (arm64, linux) (push) Has been skipped
Release / docker (push) Has been skipped
Release / chart (push) Has been skipped
Release / release (push) Has been skipped
The route carried no sectionName, so it attached to every listener whose hostname matched — including the hostname-less plaintext HTTP listener. On a publicly reachable hostname that means the API accepts bearer tokens over cleartext. networking.listener names the listener to bind to. It defaults to empty, which keeps the previous attach-to-all behaviour. Also document the Kubernetes install path, which the README omitted. |
||
|
|
42e846f876 |
Expire stale firing alerts
Release / build (amd64, darwin) (push) Failing after 12s
Release / build (arm64, darwin) (push) Failing after 11s
Release / build (arm64, linux) (push) Failing after 11s
Release / release (push) Has been skipped
Release / docker (push) Failing after 19s
Release / build (amd64, linux) (push) Failing after 12s
Release / chart (push) Failing after 9s
A resolved webhook was the only path out of the firing state, so a
notification that was dropped, silenced, or lost to a restart pinned an
alert as firing forever — Prometheus showed it resolved while
terdut-server kept listing it. The archiver only ever touched resolved
alerts, and both the list and stats queries compared status with plain
equality, so a stale row was indistinguishable from a live one.
A sweeper pass now resolves firing alerts on either of two signals: the
ends_at watermark Alertmanager sets on outgoing firing notifications has
passed (plus a grace period for clock skew), or no webhook has refreshed
the alert within TERDUT_STALE_AFTER (default 6h, above Alertmanager's 4h
repeat_interval). Such alerts get resolution_source = 'expiry',
distinguishing them from a real 'alertmanager' resolve.
Two related webhook bugs fixed alongside:
- The upsert had no ordering guard, so a retried firing notification
arriving after the resolved one resurrected the alert. Payloads for
an older alert instance are now discarded: a stale retry carries the
same startsAt, a genuine re-fire a newer one.
- archived_at was never cleared on re-fire, leaving a re-fired alert
archived and invisible in the default list.
Stats now exclude archived alerts to match the default list view; this
lowers historical firing/resolved totals.
The chart exposes both sweeper durations via sweeper.staleAfter and
sweeper.archiveAfter.
|
||
|
|
36468a68ed |
chart: add bootstrap job (v0.2.0)
Post-install/post-upgrade Job that calls /api/bootstrap on first deploy and stores the admin API key in a Secret (<release>-admin-key by default). Exits cleanly on subsequent upgrades when bootstrap is already complete. Adds ServiceAccount, Role (secrets:create), and RoleBinding as hook resources. |
||
|
|
591290e4ee |
Add Helm chart and chart-release workflow
charts/terdut-server/ — Helm chart for Kubernetes deployment: - Deployment (replicas=1, /healthz probes, TERDUT_DB_PATH=/data/terdut.db) - Service (ClusterIP :8080) - PVC (1Gi, synology-iscsi) mounted at /data - HTTPRoute via envoy-main gateway .github/workflows/chart-release.yml — packages and publishes the chart to gh-pages branch on any push to main that touches charts/; repo URL will be https://yeniklas.github.io/terdut-server once the repo is made public |