Compare commits

...

9 Commits

Author SHA1 Message Date
Niklas Ye 6047d1a9f7 ci: do not override HELM_REPOSITORY_CACHE alongside the config
CI / test (push) Successful in 4s
Release / test (push) Successful in 3s
Release / chart (push) Successful in 1s
Release / binaries (push) Successful in 2m48s
Release / image (push) Successful in 2m47s
Helm writes a refreshed repository index to the default cache directory and
then looks for it in the overridden one, so setting both is how the chart
tooling breaks on this machine already.
2026-08-19 20:44:06 +02:00
Niklas Ye 289eca8076 Move to Gitea: git.ryuvia.com/niklas/terdut-server
CI / test (push) Successful in 2m15s
The module path, the container image, the Helm chart and the CI pipeline all
named GitHub. They now name the Gitea instance everything else already runs on.

The workflows are rewritten rather than translated. Gitea's runner image is
ubuntu:22.04, whose nodejs is Node 12, so no JS action runs there at all --
actions/checkout@v4 dies with a SyntaxError before it does anything. Every step
is shell, checkout is a plain clone (this repo is public, so it needs no
credential), and the jobs that need docker or helm run in host mode because the
dind bridge a `container:` job gets cannot reach github.com or get.helm.sh.

Two consequences worth naming:

- upload-artifact/download-artifact are also JS actions, and there is no
  artifact store here, so the job that builds the binaries is the job that
  publishes them. Nothing is passed between jobs.
- setup-qemu-action is gone with the rest, and the runner has no binfmt
  registration. The Dockerfile's builder stage now runs on $BUILDPLATFORM and
  cross-compiles from TARGETARCH instead, which is what keeps the arm64 image
  buildable -- and makes it native rather than emulated.

The chart moves from a GitHub Pages index to an OCI artifact in Gitea's
registry. Publishing stays tag-only for the reason recorded in release.yaml: a
workflow triggered by the branch push cannot know the version it is about to be
tagged with.

The GitHub repository is left in place and untouched. Nothing pushes to it any
more, but its existing release downloads and chart index keep resolving.
2026-08-19 20:39:10 +02:00
Niklas Ye 766f43931c chart: publish from the tag only, not from both workflows
Two workflows published the chart and disagreed about its metadata.
release.yml stamps version and appVersion from the git tag;
chart-release.yml, triggered by any charts/** push to main, took
Chart.yaml verbatim, where appVersion is the hardcoded "latest".
Both fired for the same commit, both tried to publish the same chart
version, and skip_existing turned whichever lost into a no-op — so what
a release said about itself came down to which runner was quicker.

Chart 0.9.0 went out that way, reading appVersion "latest". Every
earlier release got the right answer by accident: Chart.yaml's version
lagged the published set, so chart-release.yml always collided with an
existing version and skipped, leaving release.yml to win uncontested.
Bumping Chart.yaml to match the tag before cutting 0.9.0 removed that
accident and the race showed itself.

Making the two agree is not possible. The tag is pushed after the branch,
so a workflow triggered by the main push cannot know the version it is
about to be tagged with — no amount of deriving from git describe fixes
that ordering. The fix is one publisher, triggered by the tag, so
chart-release.yml is deleted.

The chart now only ships with an app release. Nothing is lost: the sed in
release.yml ties the chart version to the app version, so a chart-only
change never had a version of its own to be released under. Chart fixes
ride the next tag.

Chart.yaml's version and appVersion are documented as the placeholders
they now are, so the next person does not helpfully bump them and
reintroduce this. skip_existing stays, for idempotent re-runs of a failed
release rather than for the race, and a non-version tag now fails the job
instead of silently publishing unstamped metadata.
2026-08-08 21:41:44 +02:00
Niklas Ye 14c24f8fda Notice when the Watchdog alert stops arriving
Release / release (push) Has been skipped
Release / build (amd64, linux) (push) Has been skipped
Release / build (arm64, darwin) (push) Has been skipped
Release / test (push) Failing after 5s
Release / build (arm64, linux) (push) Has been skipped
Release / docker (push) Has been skipped
Release / chart (push) Has been skipped
Release / build (amd64, darwin) (push) Has been skipped
Everything this server does assumes alerts arrive. If Prometheus stops
evaluating, or Alertmanager cannot reach us, nothing arrives — and
silence is indistinguishable from everything being fine. The cluster has
shipped the alert for exactly this case all along: Watchdog is
expr: vector(1), so it fires permanently and is re-sent forever, and it
is worth nothing unless something downstream notices it stop. Nothing
did. It arrived, opened no incident because a repeat_interval re-send is
not a new occurrence, and when the monitoring stack died the sweeper
quietly expired it and paged nobody.

So the handling is inverted for a configurable set of alerts: receiving
one opens no incident, and the absence of one does. TERDUT_DEADMAN_MATCHERS
selects them as label matchers, defaulting to alertname=Watchdog.

The unit of monitoring is the fingerprint rather than the alert name. Two
clusters sending the same Watchdog are two independent switches, so a
healthy one can never mask a dead one. Every matcher must name an
alertname, which keeps the sweeper's candidate query on alerts_name_idx
instead of JSON-extracting labels from every row, and leaves matching
with a single implementation.

A switch is dormant until its first heartbeat: a matcher nothing has ever
sent opens nothing, so a fresh deploy or a restored database does not
page. Resolving the incident by hand sticks, exactly as it does for an
alert-backed one, so a decommissioned source is a one-time page rather
than a nag; the switch re-arms only when the heartbeat comes back, and
dying again is a new incident.

The incident has no member alerts on purpose. Linking the heartbeat would
have the settled-incident cascade close it on the very sweep that opened
it, and there is no alert describing the problem anyway — the problem is
that no alert arrived. What happened is on the timeline instead, and
recovery is the only automatic way out.

One narrow exemption in the ingest guard makes recovery possible at all.
A heartbeat we declared dead is marked resolved, and the one that proves
us wrong carries the unchanged startsAt of an alert that never stopped
firing — so "resolution is terminal within an instance" would discard it
forever and a switch could die exactly once. The exemption is scoped to
resolution_source = 'deadman', which is the only resolution this server
infers from silence on a timeout of its own, so nothing another writer
set can be undone by a stale retry. Matched alerts are also held back
from the generic staleness expiry, which would otherwise resolve a
heartbeat as 'expiry' long before its own tighter deadline.

The timeout points the opposite way to TERDUT_STALE_AFTER: staleness is a
generous grace period around a repeat_interval you do not control, while
this is a deadline you set deliberately and configure the heartbeat's
route to beat. Inheriting a 4h or 12h repeat_interval gives a dead man's
switch with a twelve hour fuse, so the README spells out the route the
heartbeat needs.
2026-08-08 21:28:55 +02:00
Niklas Ye e5916d522a Let an on-call day be handed to somebody else
Release / build (amd64, darwin) (push) Has been skipped
Release / build (arm64, darwin) (push) Has been skipped
Release / docker (push) Has been skipped
Release / release (push) Has been skipped
Release / test (push) Failing after 6s
Release / build (amd64, linux) (push) Has been skipped
Release / build (arm64, linux) (push) Has been skipped
Release / chart (push) Has been skipped
A date is held by exactly one person and POST /api/schedule plain-inserts,
so any date that was already taken came back 409. That made reassignment
impossible through the API: the only route was to delete the entry first,
and for a week that meant seven separate deletions. Worse, the reject is
all-or-nothing across the request, so assigning a week where a single day
happened to be taken failed entirely and placed none of the other six.

The refusal itself is worth keeping. Moving a shift off the person
expecting to be paged for it should not be something a plain call does by
accident, so the fix is to make it possible to ask for rather than to
remove the guard: "replace": true takes the dates anyway, and the flag
defaults to off so every existing caller behaves exactly as before.

The delete and the insert share the transaction that was already there.
That matters more than the flag does — a week of free and taken days now
lands as a unit, and a failure part way through leaves the rota as it was
instead of with a shift deleted and nothing put back. A rota with a hole
in it is worse than a rota that refused to change.

One consequence worth naming: under replace a date repeated inside one
request is idempotent rather than a conflict, because the second pass
clears what the first wrote.
2026-08-07 14:04:11 +02:00
Niklas Ye 4224dbe96c Record notification delivery on the incident timeline
Release / test (push) Failing after 8s
Release / build (amd64, darwin) (push) Has been skipped
Release / build (amd64, linux) (push) Has been skipped
Release / build (arm64, darwin) (push) Has been skipped
Release / build (arm64, linux) (push) Has been skipped
Release / docker (push) Has been skipped
Release / chart (push) Has been skipped
Release / release (push) Has been skipped
An incident's history went quiet after "Incident opened": nothing said
that anybody had been paged, reminded, or told it resolved. Delivery
lived only in the notifications outbox, which no API exposes, so when a
page failed to arrive there was nothing in the product that said whether
it had been sent.

The notifier now writes two event types. A notified event once ntfy
accepts the publish, carrying the kind in detail and the paged user in
user_id — absent when the page went to the shared fallback topic, which
belongs to nobody. And a notify_failed event when a notification
exhausts its retries, which is the one worth having: without it a page
that never landed leaves the timeline identical to one that did.

Both are written from the delivery result rather than at enqueue. A
queued notification is an intention, and the timeline is append-only, so
claiming somebody was told before ntfy accepted it would be a lie that
stays there. A failed timeline write is logged rather than returned, so
it cannot make a delivered row look unsent and send the page twice.

The topic is deliberately in neither: it is a shared secret with the
ntfy server, and every API key can read the timeline.

No migration — incident_events.type is free text, unlike
notifications.kind.
2026-08-07 13:31:03 +02:00
Niklas Ye 17ee290d90 docs: the ack token is scoped, not single-use
The handler never deletes the token: it stays valid until expires_at and
is purged by the sweeper, so a second tap is an idempotent no-op rather
than a rejection. Caught by pressing Acknowledge twice against the live
server. What bounds the token is scope -- one incident, one action, one
day -- not a use count.
2026-08-07 11:41:15 +02:00
Niklas Ye 7caafbaf80 chart: back up the database through a python sidecar
Release / test (push) Failing after 7s
Release / build (amd64, darwin) (push) Has been skipped
Release / build (amd64, linux) (push) Has been skipped
Release / build (arm64, darwin) (push) Has been skipped
Release / build (arm64, linux) (push) Has been skipped
Release / docker (push) Has been skipped
Release / chart (push) Has been skipped
Release / release (push) Has been skipped
The image is FROM scratch, so there is no interpreter to run a k8up
backupcommand in, and the database runs in WAL mode, where a file-level
copy of the volume is not crash-consistent.

Also switches to strategy: Recreate. The data PVC is ReadWriteOnce, so a
RollingUpdate deadlocks the new pod against the old one holding it.
2026-08-07 10:57:56 +02:00
Niklas Ye bc285799d1 Page the on-call person when an incident opens
An incident opened, got assigned to whoever held today's schedule entry,
and then sat there silently until somebody thought to look. The schedule
and the incident model were both built; nothing reached the person
holding the pager.

Notifications go out through ntfy, over plain HTTP with no new
dependencies. Delivery is an outbox rather than an inline call: the pool
is limited to a single connection, so a POST made while holding the
webhook's transaction would stall every other request behind it. The
webhook inserts a row and a notifier goroutine sends it within a tick,
retrying with exponential backoff.

Only opening an incident has to resolve a topic from scratch. Reminders
and all-clears reuse whatever that first notification chose, which keeps
configuration out of resolveIfSettled and gives the right rule for free:
you only hear that something resolved if you were told it started.

Each push carries an Acknowledge button, because the useful thing to do
at 3am is stop the pager without unlocking anything. It POSTs to an
unauthenticated /api/notify/ack/{token} — a notification body lives on
the ntfy server and in the device cache, so a real API key must never
appear in one. The token is minted per delivery, scoped to one incident
and one action, and expires in a day.

Reminders repeat until the incident stops being untouched. The stop
conditions are the states that already mean somebody has it: acknowledged,
snoozed, resolved, archived. Snooze is the mute button, so there is no
separate reminder cap.

Notifications sent to the fallback topic carry no Acknowledge button. The
topic is shared, and a button on it would let any subscriber acknowledge
as somebody else.
2026-08-07 08:51:38 +02:00
31 changed files with 3429 additions and 317 deletions
+66
View File
@@ -0,0 +1,66 @@
name: CI
# The release workflow gates a tag, which is late: a broken commit sits green until
# somebody decides to publish. This runs the same checks on the way in.
#
# push is scoped to main rather than all branches so that a branch pushed as part of a
# pull request is not checked twice.
#
# No actions/checkout, deliberately -- same as the letsvisit and charts workflows. The
# runner image is ubuntu:22.04 whose `nodejs` package is Node 12, and actions/checkout@v4
# is built with ES2022 static initialiser blocks, so it dies with
# `SyntaxError: Unexpected token '{'` before running. Cloning with git directly avoids JS
# actions entirely. This repo is public, so the clone needs no credential at all.
#
# `${{ }}` values are passed through `env:` and referenced as quoted shell variables: a
# ref name is attacker-influenced by anyone who can push a branch or open a PR, and
# expanding one straight into `run:` is a shell-injection vector.
on:
push:
branches: [main]
pull_request:
# A rapid series of pushes only needs the last one checked.
concurrency:
group: ci-${{ github.ref }}
cancel-in-progress: true
env:
REPO_URL: https://git.ryuvia.com/niklas/terdut-server.git
jobs:
test:
runs-on: ubuntu-latest
container:
# Runs inside the toolchain image rather than installing Go per job. Note this puts
# the job on the dind bridge, which cannot reach github.com or get.helm.sh --
# proxy.golang.org and git.ryuvia.com are reachable, which is all this job needs.
image: golang:1.26.6-bookworm
# act_runner destroys a job's own volumes when it finishes, so without these every
# run re-downloads the whole module graph. The names must appear in the runner's
# container.valid_volumes allowlist (charts/act-runner in the k8s repo); unlisted
# volumes are dropped silently, so a workflow that looks correct can still be
# running uncached.
volumes:
- go-mod-cache:/go/pkg/mod
- go-build-cache:/root/.cache/go-build
- gobin-cache:/go/bin
steps:
- name: Checkout
env:
REF_NAME: ${{ github.ref_name }}
HEAD_SHA: ${{ github.event.pull_request.head.sha }}
run: |
if [ -n "$HEAD_SHA" ]; then
# A pull_request ref_name is "<n>/merge", which is not a fetchable branch.
git clone "$REPO_URL" .
git checkout -q "$HEAD_SHA"
else
git clone --depth=1 --branch "$REF_NAME" "$REPO_URL" .
fi
- name: Vet
run: go vet ./...
- name: Test
run: go test ./...
+211
View File
@@ -0,0 +1,211 @@
name: Release
# Checkout, interpolation and caching conventions match ci.yaml -- see the header there
# for why there are no JS actions and why every `${{ }}` goes through `env:`.
#
# There is no upload-artifact/download-artifact equivalent here (both are JS actions, and
# this Gitea has no artifact store wired up), so the job that builds the binaries is also
# the job that publishes them. Nothing is handed between jobs at all.
on:
push:
tags:
- 'v*'
workflow_dispatch:
# A tag is not normally re-pushed, so this mostly matters when one is force-moved during
# a botched release -- the superseded run stops holding runner slots.
concurrency:
group: release-${{ github.ref }}
cancel-in-progress: true
env:
REPO_URL: https://git.ryuvia.com/niklas/terdut-server.git
API: https://git.ryuvia.com/api/v1/repos/niklas/terdut-server
REGISTRY: git.ryuvia.com
IMAGE: git.ryuvia.com/niklas/terdut-server
jobs:
# Gates every publishing job below. A tag that fails here publishes nothing: the
# binaries, the image and the chart are all downstream of it.
test:
runs-on: ubuntu-latest
container:
image: golang:1.26.6-bookworm
volumes:
- go-mod-cache:/go/pkg/mod
- go-build-cache:/root/.cache/go-build
- gobin-cache:/go/bin
steps:
- name: Checkout
env:
REF_NAME: ${{ github.ref_name }}
run: git clone --depth=1 --branch "$REF_NAME" "$REPO_URL" .
- name: Vet
run: go vet ./...
- name: Test
run: go test ./...
binaries:
needs: test
runs-on: ubuntu-latest
container:
image: golang:1.26.6-bookworm
volumes:
- go-mod-cache:/go/pkg/mod
- go-build-cache:/root/.cache/go-build
- gobin-cache:/go/bin
steps:
- name: Checkout
env:
REF_NAME: ${{ github.ref_name }}
run: git clone --depth=1 --branch "$REF_NAME" "$REPO_URL" .
- name: Build every target
env:
REF_NAME: ${{ github.ref_name }}
run: |
set -eu
mkdir -p dist
for target in linux/amd64 linux/arm64 darwin/amd64 darwin/arm64; do
GOOS="${target%/*}"
GOARCH="${target#*/}"
out="dist/terdut-${REF_NAME}-${GOOS}-${GOARCH}"
echo "building $out"
GOOS="$GOOS" GOARCH="$GOARCH" go build \
-ldflags "-w -s -X main.version=${REF_NAME}" \
-o "$out" ./cmd/terdut
done
# Creating the release is made idempotent rather than assumed-new: a re-run of a
# failed release must not die on the release that already exists. Assets are
# replaced the same way, so a re-run repairs a partial upload.
- name: Publish the release
env:
REF_NAME: ${{ github.ref_name }}
TOKEN: ${{ secrets.REGISTRY_TOKEN }}
run: |
set -eu
auth="Authorization: token $TOKEN"
body=$(curl -sf -H "$auth" "$API/releases/tags/$REF_NAME" || true)
if [ -z "$body" ]; then
body=$(curl -sf -X POST -H "$auth" -H 'Content-Type: application/json' \
-d "{\"tag_name\":\"$REF_NAME\",\"name\":\"$REF_NAME\"}" \
"$API/releases")
fi
# The release object serialises `id` first, so the first match is the release's
# own id and not one of the nested author/asset ids.
release_id=$(printf '%s' "$body" | grep -o '"id":[0-9]*' | head -1 | cut -d: -f2)
[ -n "$release_id" ] || { echo "::error::could not determine release id"; exit 1; }
echo "release id $release_id"
for f in dist/*; do
name=$(basename "$f")
# Drop an existing asset of the same name first: Gitea happily stores two
# attachments with one name, and the updater matches by name.
old=$(curl -sf -H "$auth" "$API/releases/$release_id/assets" \
| tr '}' '\n' | grep "\"name\":\"$name\"" \
| grep -o '"id":[0-9]*' | head -1 | cut -d: -f2 || true)
if [ -n "$old" ]; then
curl -sf -X DELETE -H "$auth" "$API/releases/$release_id/assets/$old" || true
fi
echo "uploading $name"
curl -sf -X POST -H "$auth" -F "attachment=@$f" \
"$API/releases/$release_id/assets?name=$name" > /dev/null
done
# Host mode on purpose (no `container:`): this is the only context with a Docker CLI
# pointed at the dind daemon. A `container:` job would sit on the dind bridge with no
# docker socket at all.
image:
needs: test
runs-on: ubuntu-latest
steps:
- name: Checkout
env:
REF_NAME: ${{ github.ref_name }}
run: git clone --depth=1 --branch "$REF_NAME" "$REPO_URL" .
- name: Log in to the registry
env:
TOKEN: ${{ secrets.REGISTRY_TOKEN }}
run: echo "$TOKEN" | docker login "$REGISTRY" -u niklas --password-stdin
# The default "docker" driver cannot build more than one platform at a time; the
# docker-container driver can. Reused across runs if it survived the last one.
- name: Prepare buildx
run: docker buildx create --name terdut --use 2>/dev/null || docker buildx use terdut
# No QEMU: the Dockerfile's builder stage runs on $BUILDPLATFORM and cross-compiles
# from TARGETARCH, so both platforms build natively. See the comment in Dockerfile.
- name: Build and push
env:
REF_NAME: ${{ github.ref_name }}
run: |
docker buildx build \
--platform linux/amd64,linux/arm64 \
--build-arg "VERSION=${REF_NAME}" \
--tag "${IMAGE}:latest" \
--tag "${IMAGE}:${REF_NAME}" \
--push .
# Also host mode: helm is baked into the runner image, and a `container:` job could not
# install it -- get.helm.sh is unreachable from the dind bridge.
chart:
needs: test
runs-on: ubuntu-latest
steps:
- name: Checkout
env:
REF_NAME: ${{ github.ref_name }}
run: git clone --depth=1 --branch "$REF_NAME" "$REPO_URL" .
# This job is the only thing that publishes the chart, which is what keeps the
# published metadata honest. There used to be a second publisher on every charts/**
# push to main, and the two raced for the same chart version with different answers:
# this one stamps version and appVersion from the tag, that one took Chart.yaml
# verbatim, where appVersion is the hardcoded "latest". Whichever landed first won,
# so the metadata of a release depended on which runner was quicker -- chart 0.9.0
# went out on 2026-08-08 reading appVersion "latest" that way.
#
# It could not be fixed by making both agree: the tag is pushed after the branch, so
# a workflow triggered by the main push cannot know the version it is about to be
# tagged with. One publisher, triggered by the tag.
#
# The cost is that the chart only ships with an app release. That is no real loss --
# the sed below ties the chart version to the app version, so a chart-only change
# has no version of its own to be released under anyway. Chart fixes ride the next
# tag.
- name: Stamp the chart version from the tag
env:
REF_NAME: ${{ github.ref_name }}
run: |
set -eu
if ! echo "$REF_NAME" | grep -qE '^v[0-9]'; then
echo "::error::refusing to publish a chart for non-version tag ${REF_NAME}"
exit 1
fi
CHART_VERSION="${REF_NAME#v}"
sed -i "s/^version:.*/version: ${CHART_VERSION}/" charts/terdut-server/Chart.yaml
sed -i "s/^appVersion:.*/appVersion: \"${REF_NAME}\"/" charts/terdut-server/Chart.yaml
cat charts/terdut-server/Chart.yaml
- name: Package and push
env:
REF_NAME: ${{ github.ref_name }}
TOKEN: ${{ secrets.REGISTRY_TOKEN }}
run: |
set -eu
echo "$TOKEN" | helm registry login "$REGISTRY" -u niklas --password-stdin
# Isolated repo config: the machine-wide helm repo list is not this job's
# business, and one unreachable entry in it aborts otherwise-fine commands.
# HELM_REPOSITORY_CACHE is deliberately NOT overridden alongside it -- helm
# writes a refreshed index to the default cache and then looks for it in the
# overridden one.
export HELM_REPOSITORY_CONFIG="$PWD/.helm-repos.yaml"
: > "$HELM_REPOSITORY_CONFIG"
helm package charts/terdut-server -d dist
helm push "dist/terdut-server-${REF_NAME#v}.tgz" "oci://${REGISTRY}/niklas"
-38
View File
@@ -1,38 +0,0 @@
name: Release Helm Chart
on:
push:
branches:
- main
paths:
- charts/**
jobs:
release:
runs-on: ubuntu-latest
permissions:
contents: write
pages: write
steps:
- name: Checkout
uses: actions/checkout@v4
with:
fetch-depth: 0
- name: Configure Git
run: |
git config user.name "$GITHUB_ACTOR"
git config user.email "$GITHUB_ACTOR@users.noreply.github.com"
- name: Install Helm
uses: azure/setup-helm@v4
- name: Run chart-releaser
uses: helm/chart-releaser-action@v1.6.0
with:
# A charts/** push without a Chart.yaml version bump would otherwise
# fail trying to re-release the current version. Tagged releases also
# publish the chart from release.yml, so the two can race.
skip_existing: true
env:
CR_TOKEN: "${{ secrets.GITHUB_TOKEN }}"
-34
View File
@@ -1,34 +0,0 @@
name: CI
# The release workflow gates a tag, which is late: a broken commit sits green
# until somebody decides to publish. This runs the same checks on the way in.
#
# push is scoped to main rather than all branches for two reasons: a branch
# pushed as part of a pull request would otherwise be checked twice, and
# gh-pages holds the published Helm chart index with no Go code in it, so
# `go vet ./...` there would fail on a missing go.mod.
on:
push:
branches: [main]
pull_request:
# A rapid series of pushes only needs the last one checked.
concurrency:
group: ci-${{ github.ref }}
cancel-in-progress: true
jobs:
test:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-go@v5
with:
go-version-file: go.mod
- name: Vet
run: go vet ./...
- name: Test
run: go test ./...
-143
View File
@@ -1,143 +0,0 @@
name: Release
on:
push:
tags:
- 'v*'
workflow_dispatch:
jobs:
# Gates every publishing job below. A tag that fails here publishes nothing:
# the binaries, the image and the chart are all downstream of it.
test:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-go@v5
with:
go-version-file: go.mod
- name: Vet
run: go vet ./...
- name: Test
run: go test ./...
build:
needs: test
runs-on: ubuntu-latest
strategy:
matrix:
include:
- goos: linux
goarch: amd64
- goos: linux
goarch: arm64
- goos: darwin
goarch: amd64
- goos: darwin
goarch: arm64
steps:
- uses: actions/checkout@v4
- uses: actions/setup-go@v5
with:
go-version-file: go.mod
- name: Build
env:
GOOS: ${{ matrix.goos }}
GOARCH: ${{ matrix.goarch }}
run: |
go build \
-ldflags "-w -s -X main.version=${{ github.ref_name }}" \
-o terdut-${{ github.ref_name }}-${{ matrix.goos }}-${{ matrix.goarch }} \
./cmd/terdut
- uses: actions/upload-artifact@v4
with:
name: terdut-${{ github.ref_name }}-${{ matrix.goos }}-${{ matrix.goarch }}
path: terdut-${{ github.ref_name }}-${{ matrix.goos }}-${{ matrix.goarch }}
docker:
needs: test
runs-on: ubuntu-latest
permissions:
contents: read
packages: write
steps:
- uses: actions/checkout@v4
- name: Set up QEMU
uses: docker/setup-qemu-action@v3
- name: Set up Docker Buildx
uses: docker/setup-buildx-action@v3
- name: Log in to GHCR
uses: docker/login-action@v3
with:
registry: ghcr.io
username: ${{ github.actor }}
password: ${{ secrets.GITHUB_TOKEN }}
- name: Build and push
uses: docker/build-push-action@v6
with:
context: .
platforms: linux/amd64,linux/arm64
push: true
build-args: VERSION=${{ github.ref_name }}
tags: |
ghcr.io/yeniklas/terdut-server:latest
ghcr.io/yeniklas/terdut-server:${{ github.ref_name }}
chart:
needs: test
runs-on: ubuntu-latest
permissions:
contents: write
steps:
- uses: actions/checkout@v4
with:
fetch-depth: 0
- name: Configure Git
run: |
git config user.name "$GITHUB_ACTOR"
git config user.email "$GITHUB_ACTOR@users.noreply.github.com"
- name: Install Helm
uses: azure/setup-helm@v4
- name: Update chart versions
run: |
VERSION="${{ github.ref_name }}"
if [[ "$VERSION" =~ ^v[0-9] ]]; then
CHART_VERSION="${VERSION#v}"
sed -i "s/^version:.*/version: ${CHART_VERSION}/" charts/terdut-server/Chart.yaml
sed -i "s/^appVersion:.*/appVersion: \"${VERSION}\"/" charts/terdut-server/Chart.yaml
fi
- name: Run chart-releaser
uses: helm/chart-releaser-action@v1.6.0
with:
skip_existing: true
env:
CR_TOKEN: "${{ secrets.GITHUB_TOKEN }}"
release:
needs: build
runs-on: ubuntu-latest
permissions:
contents: write
steps:
- uses: actions/download-artifact@v4
with:
merge-multiple: true
- uses: softprops/action-gh-release@v2
with:
files: 'terdut-*'
+11 -2
View File
@@ -1,10 +1,19 @@
FROM golang:1.25-alpine AS builder
# --platform=$BUILDPLATFORM pins the builder to the machine doing the building, so a
# multi-arch build compiles both targets natively instead of running an emulated arm64
# toolchain under QEMU. Go cross-compiles from TARGETOS/TARGETARCH, which BuildKit fills
# in per platform. The CI runner has no binfmt registration and no way to get one (the
# JS action that used to install it cannot run there), so this is not just an
# optimisation -- it is what makes the arm64 image buildable at all.
FROM --platform=$BUILDPLATFORM golang:1.25-alpine AS builder
WORKDIR /src
COPY go.mod go.sum ./
RUN go mod download
COPY . .
ARG VERSION=dev
RUN CGO_ENABLED=0 GOOS=linux go build -ldflags="-w -s -X main.version=${VERSION}" -o /terdut ./cmd/terdut
ARG TARGETOS
ARG TARGETARCH
RUN CGO_ENABLED=0 GOOS=${TARGETOS} GOARCH=${TARGETARCH} \
go build -ldflags="-w -s -X main.version=${VERSION}" -o /terdut ./cmd/terdut
FROM scratch
COPY --from=builder /terdut /terdut
+234 -11
View File
@@ -17,7 +17,7 @@ Incident management server for teams using Prometheus Alertmanager.
**Prerequisites:** Go 1.21+
```bash
git clone https://github.com/yeniklas/terdut-server
git clone https://git.ryuvia.com/niklas/terdut-server
cd terdut-server
go run ./cmd/terdut
```
@@ -52,11 +52,12 @@ docker run -p 8080:8080 -v $(pwd)/data:/data \
### Kubernetes
A Helm chart is published from this repository:
A Helm chart is published from this repository as an OCI artifact, versioned in lockstep
with the app — chart `x.y.z` is always app `vx.y.z`:
```bash
helm repo add terdut-server https://yeniklas.github.io/terdut-server
helm upgrade --install terdut-server terdut-server/terdut-server \
helm upgrade --install terdut-server oci://git.ryuvia.com/niklas/terdut-server \
--version 0.9.0 \
--namespace terdut-server --create-namespace \
--set networking.hostname=terdut.example.com
```
@@ -71,10 +72,25 @@ than an `Ingress`. TLS is terminated at the gateway, so the server itself never
| `networking.listener` | `""` | Gateway listener (`sectionName`) to bind to. Empty attaches to every matching listener, **including plaintext HTTP** — set it to the HTTPS listener's name to serve TLS only |
| `networking.servicePort` | `8080` | Port the route forwards to; keep in sync with `service.port` |
| `bootstrap.enabled` | `true` | Runs a post-install hook that creates the first user and stores its API key in the `<release>-admin-key` Secret. Already-bootstrapped servers are left alone |
| `backupSidecar.enabled` | `true` | Adds an idle `python` sidecar and the [k8up](https://k8up.io/) annotations that dump the database through it |
The API key travels in an `Authorization: Bearer` header, so set `networking.listener` whenever the
hostname is reachable outside a trusted network.
#### Backups
The server image is `FROM scratch` — the binary and nothing else — so there is no interpreter to
run a database dump in, and the database runs in WAL mode, where a file-level copy of the volume is
not crash-consistent. The chart therefore ships an idle `python:*-alpine` sidecar that shares the
data volume, and points k8up's `backupcommand` at it with `k8up.io/backupcommand-container`. Without
that annotation k8up execs into `.spec.containers[0]` and the dump fails.
The dump is buffered and sanity-checked before its first byte reaches stdout, because k8up streams
stdout straight into Restic: a dump that dies partway is otherwise stored as a silently truncated
snapshot that k8up still reports as successful.
Set `backupSidecar.enabled=false` if you back the volume up some other way.
---
## Configuration
@@ -85,10 +101,22 @@ hostname is reachable outside a trusted network.
| `TERDUT_DB_PATH` | `terdut.db` | Path to the SQLite database file |
| `TERDUT_ARCHIVE_AFTER` | `168h` (7d) | How long a resolved alert or incident stays in the default list before being auto-archived |
| `TERDUT_STALE_AFTER` | `6h` | How long a firing alert may go without a refreshing webhook before it is treated as resolved — **must exceed your Alertmanager `repeat_interval`** |
| `TERDUT_DEADMAN_MATCHERS` | `alertname=Watchdog` | Which alerts are [dead man's switches](#dead-mans-switch). `;` separates matchers, `,` the label conditions within one, `=` is exact equality. Every matcher must name an `alertname` |
| `TERDUT_DEADMAN_TIMEOUT` | `15m` | How long a heartbeat may go unheard before its switch is declared dead — **must be shorter than the `repeat_interval` of the route carrying it**. `0` disables dead man's switch handling |
| `TERDUT_DEADMAN_SEVERITY` | `critical` | Severity a dead man's switch incident opens at |
| `TERDUT_NTFY_URL` | — | ntfy server to publish push notifications to. Empty disables notifications entirely |
| `TERDUT_NTFY_TOKEN` | — | Bearer token for an access-controlled ntfy |
| `TERDUT_NTFY_FALLBACK_TOPIC` | — | Topic used when nobody is on call |
| `TERDUT_PUBLIC_URL` | — | Base URL a phone uses to reach this server, for the link and Acknowledge button inside a notification |
| `TERDUT_NOTIFY_REPEAT` | `15m` | How long an incident may sit unacknowledged before it is paged again. `0` notifies once and never repeats |
Durations use Go syntax (`30m`, `12h`, `168h`). An unparseable value falls back to the default.
In the Helm chart the two sweeper durations are set via `sweeper.staleAfter` and `sweeper.archiveAfter`.
Note that `TERDUT_STALE_AFTER` and `TERDUT_DEADMAN_TIMEOUT` point in opposite directions. Staleness
is a generous grace period around a `repeat_interval` you do not control; a dead man's switch is a
deadline you set deliberately, and the heartbeat's route is configured to beat faster than it.
In the Helm chart the two sweeper durations are set via `sweeper.staleAfter` and `sweeper.archiveAfter`, dead man's switches via the `deadman.*` values, and notifications via the `notify.*` values.
---
@@ -109,6 +137,30 @@ route:
The webhook endpoint requires no authentication.
If you use the [dead man's switch](#dead-mans-switch) — and the default configuration does — give
the heartbeat a route of its own, because the deadline is only as tight as the interval feeding it:
```yaml
route:
receiver: terdut
repeat_interval: 4h
routes:
- matchers: [ 'alertname = "Watchdog"' ]
receiver: terdut
group_wait: 0s
group_interval: 1m
repeat_interval: 1m
```
That delivers a heartbeat every **2 minutes**, not every minute. Alertmanager only reconsiders a
group every `group_interval`, and at exactly one elapsed interval `repeat_interval` has not *quite*
passed, so the send slips to the next tick — equal values give 2×. Two minutes against the 15 minute
default is seven heartbeats per window, which is the point; use `group_interval: 30s` if you want
the numbers to mean what they say.
kube-prometheus-stack users get the `Watchdog` alert (`expr: vector(1)`) for free; it just needs
routing to terdut rather than to `null`.
---
## Alerts and incidents
@@ -159,6 +211,10 @@ alert that never stopped firing.
that group opens a *new* incident rather than reopening this one. If the alert
underneath never stops firing, the incident stays closed — that is what
resolving by hand asserts.
- **On recovery**, for a [dead man's switch](#dead-mans-switch) incident whose
heartbeat started arriving again (`"resolution_source": "recovered"`). These
incidents have no member alerts, so the automatic cascade above cannot reach
them.
To quieten an incident you expect to come back, snooze it instead
(`POST /api/incidents/{id}/snooze`). A snooze hides the incident from the default
@@ -170,6 +226,64 @@ A new incident is assigned to whoever holds today's schedule entry at the moment
it opens (`GET /api/schedule/current`). If nobody is scheduled it opens
unassigned. Reassign with `POST /api/incidents/{id}/assign`.
One person holds a given day, so `POST /api/schedule` refuses a date somebody
already has: taking a shift off the person expecting to be paged for it should
not be something a plain call does by accident. Pass `"replace": true` to take
them anyway. Either way the whole request is one transaction — a week where some
days are free and some are taken moves as a unit, and a failure leaves the rota
exactly as it was rather than with a hole in it.
### Push notifications
With `TERDUT_NTFY_URL` set, an incident that opens is pushed to the on-call
person's phone through [ntfy](https://ntfy.sh). Set each user's topic with
`PUT /api/users/{id}/notify`; a user with no topic falls back to
`TERDUT_NTFY_FALLBACK_TOPIC`, as does an incident that opens with nobody on call.
If neither yields a topic, nothing is queued.
Three things get pushed:
- **triggered** — an incident opened. Priority follows severity (`critical` maps
to ntfy's max priority, the one that overrides the phone's quiet settings).
- **reminder** — the incident is still `triggered` after `TERDUT_NOTIFY_REPEAT`.
Repeats until somebody acts. Acknowledging, snoozing, resolving or archiving
all stop it — snooze is the mute button.
- **resolved** — every alert under the incident stopped firing. Only sent to
whoever was paged in the first place, and only for the automatic cascade:
resolving by hand pushes nothing, since the person who did it already knows.
Notifications carry an **Acknowledge** button that acknowledges the incident
without opening anything. It POSTs to `/api/notify/ack/{token}`, an
unauthenticated route authorised by the 256-bit token in its path — minted fresh
per notification, scoped to one incident and one action, and valid for 24 hours.
A real API key is never put in a notification, because the message is stored on
the ntfy server and cached on the device.
The token is **not** consumed by use. Acknowledging is idempotent, so a token
stays valid for its full 24 hours and a second tap is a no-op that reports the
incident's current state rather than an error — which is what you want when a
tap is retried on a flaky mobile connection. What bounds it is scope, not a use
count: one incident, one action, one day. Expired tokens are purged by the
sweeper.
Two consequences worth planning for:
- `/api/notify/ack/{token}` **must stay publicly reachable**, or the button will
not work when the responder is off your network.
- Notifications sent to the fallback topic carry **no** Acknowledge button. The
topic is shared, and a button on it would let any subscriber acknowledge as
somebody else.
Delivery is a queue, not an inline call: the webhook writes a row and a
background notifier sends it within 30 seconds, retrying with exponential
backoff up to 8 attempts. Nothing about ingestion blocks on ntfy being reachable.
Every delivery is recorded on the incident's timeline: a `notified` event once
ntfy accepts the publish, and a `notify_failed` event when a notification
exhausts its retries. Written from the result rather than at enqueue, so the
timeline says what actually happened — and a page that never landed is visible
instead of looking the same as one that did.
### Stale alert expiry
A resolved webhook is the only signal that an alert has stopped firing, so a
@@ -189,6 +303,72 @@ to distinguish them from a real Alertmanager resolve (`"alertmanager"`).
An expiry cascades: once it leaves an incident with nothing firing under it, the
incident resolves too, in the same sweep.
### Dead man's switch
Everything above assumes alerts arrive. If Prometheus stops evaluating, or
Alertmanager cannot reach this server, nothing arrives — and silence looks
exactly like everything being fine. A dead man's switch inverts the handling for
one designated alert so that silence is the signal:
- **receiving** it opens no incident, and
- the **absence** of it does.
kube-prometheus-stack already ships the alert for this. `Watchdog` is
`expr: vector(1)`, so it fires permanently and is re-sent forever; it is worth
nothing unless something downstream notices it stop. That is what
`TERDUT_DEADMAN_MATCHERS` defaults to.
A matcher is a set of exact label conditions, one of which must be the
`alertname`:
```
TERDUT_DEADMAN_MATCHERS="alertname=Watchdog,cluster=prod; alertname=EdgeHeartbeat"
```
**The unit of monitoring is the fingerprint, not the alert name.** Two clusters
sending the same `Watchdog` are two independent switches, so a healthy one can
never mask a dead one.
#### The lifecycle
A switch is **dormant** until its first heartbeat arrives. A configured matcher
that has never been heard from opens nothing, so a fresh deploy or a restored
database does not page. It also means a matcher that never matches anything is
silently inert — check the startup log line, which lists the matchers that
survived parsing.
Once armed, the sweeper declares it **dead** when either the heartbeat has not
been refreshed within `TERDUT_DEADMAN_TIMEOUT`, or Alertmanager explicitly
resolved it — the sender saying the heartbeat stopped needs no further waiting.
That opens an incident at `TERDUT_DEADMAN_SEVERITY`, assigned and paged like any
other, and marks the heartbeat alert `"resolution_source": "deadman"` so the
alert list stops claiming a dead switch is firing.
It **recovers** when the heartbeat starts arriving again: the incident resolves
with `"resolution_source": "recovered"` and the all-clear goes to whoever was
paged.
Resolving the incident by hand sticks, the same way it does for an alert-backed
one. While the switch stays silent nothing new opens — so a decommissioned
source is a one-time page rather than a nag. The switch **re-arms** on the next
heartbeat: come back and die again, and that is a new incident.
#### Two things to know
`TERDUT_DEADMAN_TIMEOUT` must be **shorter** than the `repeat_interval` of the
route carrying the heartbeat, which is the exact opposite of
`TERDUT_STALE_AFTER`. Inheriting a default `repeat_interval` of 4h gives you a
switch that takes four hours to notice anything, so give the heartbeat
[its own route](#alertmanager-configuration). Matched alerts are exempt from
stale-alert expiry — a heartbeat answers to its own timeout and nothing else.
A dead man's switch incident has **no member alerts**:
`GET /api/incidents/{id}/alerts` returns an empty list. There is no alert
describing the problem, because the problem is that no alert arrived. What
happened is on the timeline instead, as a `deadman_silent` event carrying the age
of the last heartbeat, and the heartbeat's labels are on the incident's
`group_labels`.
---
## API reference
@@ -209,6 +389,7 @@ Authorization: Bearer <api-key>
| `GET` | `/api/users` | List users |
| `POST` | `/api/users` | Create user `{"username","email"}` |
| `DELETE` | `/api/users/{id}` | Delete user (cascades to keys) |
| `PUT` | `/api/users/{id}/notify` | Set push notification target `{"ntfy_topic"}` — empty string clears it |
| `POST` | `/api/users/{id}/api-keys` | Issue API key `{"name"}` — key shown once |
| `DELETE` | `/api/users/{id}/api-keys/{keyID}` | Revoke API key |
@@ -218,6 +399,12 @@ Authorization: Bearer <api-key>
|---|---|---|
| `POST` | `/api/alertmanager/webhook` | Alertmanager v4 webhook receiver (no auth) |
### Notifications
| Method | Path | Description |
|---|---|---|
| `POST` | `/api/notify/ack/{token}` | Acknowledge an incident from a push notification's Acknowledge button. No auth: the token in the path is the credential — one incident, one action, 24 hours, idempotent. Must stay publicly reachable |
### Incidents
| Method | Path | Description |
@@ -260,7 +447,7 @@ only by their author. The rest of the timeline is a record of what happened.
| `assigned_to_id` / `assigned_to` | | *optional* — user id, username |
| `snoozed_until` | timestamp | *optional* — a value in the past reads as not snoozed |
| `resolved_at` | timestamp | *optional* |
| `resolution_source` | string | *optional* — `"alerts"` or `"manual"` |
| `resolution_source` | string | *optional* — `"alerts"`, `"manual"` or `"recovered"` |
| `archived_at` | timestamp | *optional* |
| `alerts` | array | Only on `GET /api/incidents/{id}` |
@@ -281,8 +468,16 @@ name: degrade unknown values to "resolved, reason unknown".
Types written today: `triggered`, `alert_added`, `alert_resolved`,
`acknowledged`, `unacknowledged`, `assigned`, `snoozed`, `unsnoozed`, `resolved`,
`note`. On an `assigned` event `user_id` is the **assignee**, not the actor. New
types may be added; render unknown ones generically rather than dropping them.
`note`, `notified`, `notify_failed`, `deadman_silent`. On an `assigned` event
`user_id` is the **assignee**, not the actor. New types may be added; render
unknown ones generically rather than dropping them.
On `notified` and `notify_failed`, `detail` carries the notification kind
(`triggered` | `reminder` | `resolved`), and on a failure the reason after it.
`user_id` is who was paged — absent means the page went to the shared fallback
topic and so belongs to nobody. The topic itself is never written to the
timeline: it is a shared secret with the ntfy server, and every API key can read
this.
### Alerts
@@ -297,7 +492,8 @@ Archived alerts are hidden from `GET /api/alerts` unless `?archived=true` is
passed; alert archiving is automatic housekeeping by the sweeper, not a user
action. Resolved alerts carry `resolution_source`: `"alertmanager"` for a real
resolved webhook, `"expiry"` when the sweeper inferred it (see
[Stale alert expiry](#stale-alert-expiry)).
[Stale alert expiry](#stale-alert-expiry)), `"deadman"` for a heartbeat declared
dead (see [Dead man's switch](#dead-mans-switch)).
#### The alert object
@@ -318,7 +514,7 @@ when unset, so clients must treat them as nullable.
| `generator_url` | string | Link back to the originating Prometheus |
| `received_at` | timestamp | When the server last accepted a webhook for this alert — see below |
| `incident_id` | integer | *optional* — the most recent incident this alert belongs to |
| `resolution_source` | string | *optional* — `"alertmanager"` or `"expiry"` |
| `resolution_source` | string | *optional* — `"alertmanager"`, `"expiry"` or `"deadman"` |
| `archived_at` | timestamp | *optional* — set while archived |
##### `received_at` is a liveness heartbeat
@@ -373,6 +569,13 @@ which happened. Clients may rely on this:
worthwhile, since `"expiry"` can also mean the alert is still firing and the
notification path broke.
- **`"deadman"` — a heartbeat was declared dead** (see
[Dead man's switch](#dead-mans-switch)). Like `"expiry"`, an inference from
silence rather than an observed end, so `ends_at` is approximate — but a much
tighter one, bounded by `TERDUT_DEADMAN_TIMEOUT`. It is also the one resolution
a re-fire under the same `starts_at` can undo, since the switch coming back is
exactly the evidence that the inference was wrong.
Treat the value as an open set and tolerate ones you do not recognise — new
sources may be added, and unknown values should degrade to "resolved, reason
unknown" rather than being rejected.
@@ -381,7 +584,7 @@ unknown" rather than being rejected.
| Method | Path | Description |
|---|---|---|
| `POST` | `/api/schedule` | Assign user to dates `{"user_id", "dates":["YYYY-MM-DD",...]}` — all-or-nothing |
| `POST` | `/api/schedule` | Assign user to dates `{"user_id", "dates":["YYYY-MM-DD",...], "replace"}` — all-or-nothing |
| `GET` | `/api/schedule` | List entries. Filters: `?from=YYYY-MM-DD`, `?to=YYYY-MM-DD` |
| `GET` | `/api/schedule/current` | Today's on-call user (UTC), 404 if none |
| `DELETE` | `/api/schedule/{id}` | Remove schedule entry |
@@ -429,6 +632,26 @@ Nothing about the two documented alert contracts changes: `received_at` is still
advanced on every accepted webhook, and `resolution_source` still means what it
did.
## Upgrading to dead man's switches
Dead man's switch handling is **on by default**, watching `alertname=Watchdog`
with a 15 minute timeout. If you already route `Watchdog` to this server, the
behaviour of that alert changes on upgrade, in both directions:
- it stops opening incidents when it arrives, and
- it starts opening one when it stops arriving.
**Check your `repeat_interval` before upgrading.** The switch pages whenever a
heartbeat has not been refreshed within `TERDUT_DEADMAN_TIMEOUT`, so a `Watchdog`
route inheriting a 4h or 12h `repeat_interval` will page constantly against the
15 minute default. Either give the heartbeat
[its own fast route](#alertmanager-configuration) — the point of the feature — or
set `TERDUT_DEADMAN_TIMEOUT` above your current `repeat_interval` until you have.
`TERDUT_DEADMAN_TIMEOUT=0` turns the whole thing off.
There is no migration and no schema change. An existing open incident from a
`Watchdog` that arrived under the old behaviour is unaffected; resolve it by hand.
---
## Development
+6 -1
View File
@@ -2,5 +2,10 @@ apiVersion: v2
name: terdut-server
description: A Helm chart for Terminal Duty — on-call alert management server
type: application
version: 0.5.0
# These two are placeholders for a local `helm install ./charts/terdut-server`, not the
# released values. .gitea/workflows/release.yaml rewrites both from the git tag when it
# publishes, so the chart version always equals the app version. Bumping them by hand
# does nothing for a release and is not needed before tagging.
# "latest" is honest here: it matches image.tag in values.yaml.
version: 0.9.0
appVersion: "latest"
@@ -10,10 +10,50 @@ spec:
selector:
matchLabels:
{{- include "terdut-server.selectorLabels" . | nindent 6 }}
# The data PVC is ReadWriteOnce, so a RollingUpdate deadlocks: the new pod
# cannot attach the volume until the old one releases it, and the old one is
# not torn down until the new one is ready.
strategy:
type: Recreate
template:
metadata:
labels:
{{- include "terdut-server.selectorLabels" . | nindent 8 }}
{{- if .Values.backupSidecar.enabled }}
annotations:
# Dumps the whole database: incidents, alerts, users, API key hashes,
# the schedule and the notification outbox.
#
# Runs in the `backup` sidecar, NOT in the app container: the server
# image is FROM scratch and has no interpreter at all. k8up execs into
# .spec.containers[0] unless told otherwise, hence the explicit
# k8up.io/backupcommand-container.
#
# Buffered and sanity-checked before the first byte reaches stdout: k8up
# streams stdout straight into restic, so a dump that dies partway is
# stored as a silently-truncated snapshot that k8up still reports as
# Succeeded. The check counts users rather than incidents -- incidents
# are swept and archived, so an empty incidents table is a legitimate
# state, whereas a database with no users never is.
#
# The connection is read-only but the mount is not: the database runs in
# WAL mode, and opening it mode=ro still needs write access to the -shm
# wal-index.
#
# chr(10), not '\n': k8up parses this annotation with go-shellquote.
k8up.io/backupcommand-container: backup
k8up.io/backupcommand: >-
python3 -c "import sqlite3, sys;
con = sqlite3.connect('file:/data/terdut.db?mode=ro', uri=True);
con.execute('BEGIN');
users = con.execute('SELECT count(*) FROM users').fetchone()[0];
out = chr(10).join(con.iterdump()) + chr(10);
(users > 0 and out.rstrip().endswith('COMMIT;'))
or sys.exit('terdut: db dump failed sanity checks');
sys.stdout.write(out)"
k8up.io/file-extension: ".sql"
k8up.io/backup: "true"
{{- end }}
spec:
enableServiceLinks: false
containers:
@@ -33,6 +73,29 @@ spec:
value: "{{ .Values.sweeper.staleAfter }}"
- name: TERDUT_ARCHIVE_AFTER
value: "{{ .Values.sweeper.archiveAfter }}"
- name: TERDUT_DEADMAN_MATCHERS
value: "{{ .Values.deadman.matchers }}"
- name: TERDUT_DEADMAN_TIMEOUT
value: "{{ .Values.deadman.timeout }}"
- name: TERDUT_DEADMAN_SEVERITY
value: "{{ .Values.deadman.severity }}"
{{- if .Values.notify.ntfyUrl }}
- name: TERDUT_NTFY_URL
value: "{{ .Values.notify.ntfyUrl }}"
- name: TERDUT_NTFY_FALLBACK_TOPIC
value: "{{ .Values.notify.fallbackTopic }}"
- name: TERDUT_NOTIFY_REPEAT
value: "{{ .Values.notify.repeatEvery }}"
- name: TERDUT_PUBLIC_URL
value: "{{ .Values.notify.publicUrl | default (printf "https://%s" .Values.networking.hostname) }}"
{{- if .Values.notify.tokenSecret.name }}
- name: TERDUT_NTFY_TOKEN
valueFrom:
secretKeyRef:
name: {{ .Values.notify.tokenSecret.name }}
key: {{ .Values.notify.tokenSecret.key }}
{{- end }}
{{- end }}
volumeMounts:
- name: data
mountPath: /data
@@ -46,6 +109,25 @@ spec:
path: /healthz
port: http
initialDelaySeconds: 5
{{- if .Values.backupSidecar.enabled }}
# Idle sidecar. It exists only so k8up has a container with a sqlite3
# module to exec the backupcommand in. Mounted read-write on purpose:
# see the note on the backupcommand annotation above.
- name: backup
image: "{{ .Values.backupSidecar.image.repository }}:{{ .Values.backupSidecar.image.tag }}"
imagePullPolicy: {{ .Values.backupSidecar.image.pullPolicy }}
command: ["sleep", "infinity"]
volumeMounts:
- name: data
mountPath: /data
resources:
requests:
memory: "16Mi"
cpu: "10m"
limits:
memory: "64Mi"
{{- end }}
volumes:
- name: data
persistentVolumeClaim:
+75 -1
View File
@@ -7,7 +7,7 @@ networking:
listener: ""
image:
repository: ghcr.io/yeniklas/terdut-server
repository: git.ryuvia.com/niklas/terdut-server
tag: "latest"
pullPolicy: IfNotPresent
@@ -26,6 +26,80 @@ sweeper:
# How long a resolved alert stays in the default list before auto-archiving.
archiveAfter: 168h
# Alerts treated as dead man's switches: receiving one opens no incident, and
# the absence of one does. The Watchdog alert kube-prometheus-stack ships is
# exactly this — an always-firing alert whose only value is something noticing
# when it stops.
deadman:
# Which alerts to treat as heartbeats. ";" separates matchers, "," separates
# the label conditions within one, "=" is exact equality. Every matcher must
# name an alertname:
# alertname=Watchdog,cluster=prod; alertname=EdgeHeartbeat
# Each distinct label set is watched independently, so two clusters sending
# the same alertname are two switches and a live one cannot mask a dead one.
matchers: "alertname=Watchdog"
# How long a heartbeat may go unheard before its switch is declared dead.
#
# This must be SHORTER than the Alertmanager repeat_interval of the route
# carrying the heartbeat — the opposite of sweeper.staleAfter. The default
# repeat_interval of 4h (12h in many setups) makes for a useless dead man's
# switch, so give the heartbeat a route of its own:
#
# - matchers: [ 'alertname = "Watchdog"' ]
# receiver: terdut
# group_wait: 0s
# group_interval: 1m
# repeat_interval: 1m
#
# That delivers every 2m rather than every 1m: a group is only reconsidered
# each group_interval, and at exactly one elapsed interval repeat_interval has
# not quite passed, so equal values give 2x. Fine against 15m; use
# group_interval: 30s if you want a true 1m.
#
# Set to 0 to disable dead man's switch handling entirely.
timeout: 15m
# Severity a dead man's switch incident opens at. These incidents have no
# member alerts to derive one from, and the heartbeat's own severity label is
# meaningless — Watchdog ships as "none". Only "critical" maps to the ntfy
# priority that overrides a phone's quiet hours.
severity: critical
notify:
# ntfy server that push notifications are published to, e.g.
# http://ntfy.ntfy.svc.cluster.local. Empty disables notifications entirely.
ntfyUrl: ""
# Topic used when nobody is on call today. Notifications sent here carry no
# Acknowledge button: the topic is shared, so there is no user to attribute an
# acknowledgement to. Leave empty to send nothing when the schedule is unset.
fallbackTopic: ""
# How long an incident may sit unacknowledged before it is paged again.
# Set to 0 to notify once and never repeat.
repeatEvery: 15m
# Base URL a phone uses to reach this server, for the link and the Acknowledge
# button inside a notification. Defaults to https://<networking.hostname>.
#
# The Acknowledge button is a POST to /api/notify/ack/{token} from the
# responder's phone, so that path has to stay publicly reachable — it is
# authorised by the scoped token in the URL, not by network placement.
publicUrl: ""
# Optional bearer token for an access-controlled ntfy, read from an existing
# Secret. Leave name empty for an open ntfy.
tokenSecret:
name: ""
key: token
# The server image is FROM scratch — just the binary, with no shell, no sqlite3
# and no python — so a k8up backupcommand cannot run in the app container. This
# idle sidecar shares the data volume and is selected with
# k8up.io/backupcommand-container. Only the stdlib sqlite3 module is used, so any
# python image works.
backupSidecar:
enabled: true
image:
repository: python
tag: "3.13-alpine"
pullPolicy: IfNotPresent
bootstrap:
enabled: true
username: admin
+16 -5
View File
@@ -8,9 +8,9 @@ import (
"syscall"
"time"
"github.com/yeniklas/terdut-server/internal/api"
"github.com/yeniklas/terdut-server/internal/config"
"github.com/yeniklas/terdut-server/internal/db"
"git.ryuvia.com/niklas/terdut-server/internal/api"
"git.ryuvia.com/niklas/terdut-server/internal/config"
"git.ryuvia.com/niklas/terdut-server/internal/db"
)
var version = "dev"
@@ -28,7 +28,17 @@ func main() {
log.Fatalf("migrate: %v", err)
}
router := api.NewRouter(database)
notify := api.NotifyConfig{
BaseURL: cfg.NtfyURL,
Token: cfg.NtfyToken,
FallbackTopic: cfg.NtfyFallbackTopic,
PublicURL: cfg.PublicURL,
RepeatEvery: cfg.NotifyRepeat,
}
deadman := api.ParseDeadmanConfig(cfg.DeadmanMatchers, cfg.DeadmanTimeout, cfg.DeadmanSeverity)
router := api.NewRouter(database, notify, deadman)
srv := &http.Server{
Addr: cfg.Addr,
@@ -41,7 +51,8 @@ func main() {
ctx, stop := signal.NotifyContext(context.Background(), syscall.SIGINT, syscall.SIGTERM)
defer stop()
go api.StartArchiver(ctx, database, cfg.ArchiveAfter, cfg.StaleAfter)
go api.StartArchiver(ctx, database, cfg.ArchiveAfter, cfg.StaleAfter, deadman, notify)
go api.StartNotifier(ctx, database, notify)
go func() {
log.Printf("terdut-server %s listening on %s", version, cfg.Addr)
+1 -1
View File
@@ -1,4 +1,4 @@
module github.com/yeniklas/terdut-server
module git.ryuvia.com/niklas/terdut-server
go 1.25.9
+65 -22
View File
@@ -14,6 +14,12 @@ import (
const (
resolutionAlertmanager = "alertmanager"
resolutionExpiry = "expiry"
// resolutionDeadman marks a heartbeat the dead man's switch sweeper declared
// dead. Distinct from expiry because it is load-bearing, not just
// descriptive: it is the one resolution the ingest upsert will let a
// same-instance re-fire undo, so a switch that comes back can be heard.
resolutionDeadman = "deadman"
)
// amPayload mirrors the Alertmanager webhook v4 payload.
@@ -56,9 +62,15 @@ type ingested struct {
// justResolved marks the firing → resolved edge, worth a timeline entry.
justResolved bool
// deadman marks a heartbeat: an alert whose arrival means everything is
// fine. It is stored like any other alert — received_at is the heartbeat —
// but it never reaches an incident. Its absence is what opens one, which
// sweepDeadman decides later and elsewhere.
deadman bool
}
func handleAlertmanagerWebhook(db *sql.DB) http.HandlerFunc {
func handleAlertmanagerWebhook(db *sql.DB, notify NotifyConfig, deadman DeadmanConfig) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
var payload amPayload
if err := decodeJSON(r, &payload); err != nil {
@@ -69,7 +81,7 @@ func handleAlertmanagerWebhook(db *sql.DB) http.HandlerFunc {
// Alertmanager retries anything that is not 2xx, and a retry of a payload
// we failed to store is more useful than an error it cannot act on — so
// failures are logged, not surfaced.
if err := ingest(r.Context(), db, payload); err != nil {
if err := ingest(r.Context(), db, notify, deadman, payload); err != nil {
log.Printf("webhook ingest (group %q): %v", payload.GroupKey, err)
}
@@ -80,14 +92,14 @@ func handleAlertmanagerWebhook(db *sql.DB) http.HandlerFunc {
// ingest stores a payload's alerts and reconciles the incident for its group.
// The whole payload is one transaction: an incident that opened but whose alerts
// failed to link would be a work item nobody could act on.
func ingest(ctx context.Context, db *sql.DB, payload amPayload) error {
func ingest(ctx context.Context, db *sql.DB, notify NotifyConfig, deadman DeadmanConfig, payload amPayload) error {
tx, err := db.BeginTx(ctx, nil)
if err != nil {
return err
}
defer tx.Rollback() //nolint:errcheck
accepted, err := upsertAlerts(ctx, tx, payload.Alerts)
accepted, err := upsertAlerts(ctx, tx, deadman, payload.Alerts)
if err != nil {
return err
}
@@ -96,14 +108,14 @@ func ingest(ctx context.Context, db *sql.DB, payload amPayload) error {
// resolution cascade are recomputed once per incident at the end.
touched := map[int64]bool{}
incidentID, err := incidentForGroup(ctx, tx, payload, accepted)
incidentID, err := incidentForGroup(ctx, tx, notify, payload, accepted)
if err != nil {
return err
}
if incidentID != 0 {
touched[incidentID] = true
for _, a := range accepted {
if !a.firing {
if !a.firing || a.deadman {
continue
}
if err := linkAlert(ctx, tx, incidentID, a.id); err != nil {
@@ -113,7 +125,7 @@ func ingest(ctx context.Context, db *sql.DB, payload amPayload) error {
}
for _, a := range accepted {
if !a.justResolved {
if !a.justResolved || a.deadman {
continue
}
id, err := openIncidentForAlert(ctx, tx, a.id)
@@ -144,7 +156,7 @@ func ingest(ctx context.Context, db *sql.DB, payload amPayload) error {
// upsertAlerts stores each alert of a payload and reports what changed. Payloads
// the ordering guard rejected are left out entirely.
func upsertAlerts(ctx context.Context, tx *sql.Tx, alerts []amAlert) ([]ingested, error) {
func upsertAlerts(ctx context.Context, tx *sql.Tx, deadman DeadmanConfig, alerts []amAlert) ([]ingested, error) {
now := time.Now().Unix()
accepted := make([]ingested, 0, len(alerts))
@@ -187,7 +199,15 @@ func upsertAlerts(ctx context.Context, tx *sql.Tx, alerts []amAlert) ([]ingested
// older than the stored one. Alertmanager retries failed notifications,
// so a stale firing retry can arrive after the resolved one; it carries
// the same startsAt, whereas a genuine re-fire carries a newer one.
// Within a single instance, resolution is terminal.
// Within a single instance, resolution is terminal — with one exception.
//
// A resolution this server synthesised for a dead man's switch is not
// Alertmanager's word that the instance ended; it is our inference from
// silence. The heartbeat that proves us wrong carries the unchanged
// startsAt of an alert that never stopped firing, so without the
// exemption a switch could go dead exactly once and never be heard from
// again. Scoped to 'deadman' so no resolution anybody else wrote can be
// undone by a stale retry.
if _, err := tx.ExecContext(ctx, `
INSERT INTO alerts
(fingerprint, name, status, labels, annotations, starts_at, ends_at,
@@ -211,7 +231,8 @@ func upsertAlerts(ctx context.Context, tx *sql.Tx, alerts []amAlert) ([]ingested
THEN NULL ELSE alerts.archived_at END
WHERE excluded.starts_at > alerts.starts_at
OR (excluded.starts_at = alerts.starts_at
AND NOT (alerts.status = 'resolved' AND excluded.status = 'firing'))`,
AND (alerts.resolution_source = '`+resolutionDeadman+`'
OR NOT (alerts.status = 'resolved' AND excluded.status = 'firing')))`,
a.Fingerprint, name, a.Status,
string(labelsJSON), string(annotationsJSON),
a.StartsAt.Unix(), endsAtUnix,
@@ -244,6 +265,7 @@ func upsertAlerts(ctx context.Context, tx *sql.Tx, alerts []amAlert) ([]ingested
firing: firing,
newOccurrence: firing && (!existed || a.StartsAt.Unix() > prevStartsAt || prevStatus == "resolved"),
justResolved: !firing && existed && prevStatus == "firing",
deadman: deadman.isDeadman(a.Labels),
})
}
@@ -258,10 +280,17 @@ func upsertAlerts(ctx context.Context, tx *sql.Tx, alerts []amAlert) ([]ingested
// something actually started firing. Without that, a manually resolved incident
// would reappear on the next repeat_interval re-send of an alert that never
// stopped, and manual resolution would be meaningless.
func incidentForGroup(ctx context.Context, tx *sql.Tx, payload amPayload, accepted []ingested) (int64, error) {
//
// Heartbeats do not count as anything here. A group of nothing but dead man's
// switch alerts opens no incident at all, and a mixed group gets an incident for
// its real alerts only.
func incidentForGroup(ctx context.Context, tx *sql.Tx, notify NotifyConfig, payload amPayload, accepted []ingested) (int64, error) {
var firstName string
anyFiring, anyNew := false, false
for _, a := range accepted {
if a.deadman {
continue
}
if a.firing {
if !anyFiring {
firstName = a.name
@@ -297,13 +326,20 @@ func incidentForGroup(ctx context.Context, tx *sql.Tx, payload amPayload, accept
if !anyNew {
return 0, nil
}
return openIncident(ctx, tx, groupKey, payload.GroupLabels, firstName)
return openIncident(ctx, tx, notify, groupKey,
incidentTitle(payload.GroupLabels, firstName), payload.GroupLabels, nil)
}
// openIncident creates an incident for a group and assigns it to whoever is on
// call today, which is the point at which the schedule stops being decorative.
func openIncident(ctx context.Context, tx *sql.Tx, groupKey string, groupLabels map[string]string, fallbackName string) (int64, error) {
onCall, err := currentOnCall(ctx, tx)
// openIncident creates an incident and assigns it to whoever is on call today,
// which is the point at which the schedule stops being decorative.
//
// The one place an incident is born, for both of the things that can raise one:
// the webhook, inside its transaction, and the dead man's switch sweeper, inside
// its own. Hence the querier rather than a *sql.Tx. A nil severity leaves the
// column for refreshSeverity to fill from the member alerts; the sweeper passes
// one because its incidents have no members to derive it from.
func openIncident(ctx context.Context, q querier, notify NotifyConfig, groupKey, title string, groupLabels map[string]string, severity *string) (int64, error) {
onCall, err := currentOnCall(ctx, q)
if err != nil {
return 0, err
}
@@ -313,10 +349,10 @@ func openIncident(ctx context.Context, tx *sql.Tx, groupKey string, groupLabels
labelsJSON = []byte("{}")
}
res, err := tx.ExecContext(ctx, `
INSERT INTO incidents (group_key, title, group_labels, status, triggered_at, assigned_to)
VALUES (?, ?, ?, 'triggered', ?, ?)`,
groupKey, incidentTitle(groupLabels, fallbackName), string(labelsJSON),
res, err := q.ExecContext(ctx, `
INSERT INTO incidents (group_key, title, group_labels, status, severity, triggered_at, assigned_to)
VALUES (?, ?, ?, 'triggered', ?, ?, ?)`,
groupKey, title, string(labelsJSON), severity,
time.Now().Unix(), onCall)
if err != nil {
return 0, err
@@ -326,15 +362,22 @@ func openIncident(ctx context.Context, tx *sql.Tx, groupKey string, groupLabels
return 0, err
}
if err := logEvent(ctx, tx, id, evTriggered, nil, nil, nil); err != nil {
if err := logEvent(ctx, q, id, evTriggered, nil, nil, nil); err != nil {
return 0, err
}
if onCall != nil {
// On an "assigned" event user_id is the assignee, not the actor.
if err := logEvent(ctx, tx, id, evAssigned, onCall, nil, nil); err != nil {
if err := logEvent(ctx, q, id, evAssigned, onCall, nil, nil); err != nil {
return 0, err
}
}
// Queue the page, but do not send it here: this runs inside a transaction on
// a single-connection pool, so an HTTP call would hold up every other
// request. The notifier picks the row up within a tick.
if err := enqueueOpened(ctx, q, notify, id, onCall); err != nil {
return 0, err
}
return id, nil
}
+1 -1
View File
@@ -11,7 +11,7 @@ import (
"time"
"github.com/go-chi/chi/v5"
"github.com/yeniklas/terdut-server/internal/models"
"git.ryuvia.com/niklas/terdut-server/internal/models"
)
// alertSelectFrom is the shared SELECT … FROM … clause used by all alert queries.
+148 -9
View File
@@ -12,20 +12,40 @@ import (
"testing"
"time"
"github.com/yeniklas/terdut-server/internal/api"
"github.com/yeniklas/terdut-server/internal/db"
"git.ryuvia.com/niklas/terdut-server/internal/api"
"git.ryuvia.com/niklas/terdut-server/internal/db"
)
// ts wraps httptest.Server with a pre-bootstrapped API key. db is exposed so
// tests can age rows directly — the sweeper's inputs are wall-clock timestamps.
type ts struct {
*httptest.Server
key string
db *sql.DB
key string
db *sql.DB
notify api.NotifyConfig
deadman api.DeadmanConfig
}
func newTS(t *testing.T) *ts {
// newTS builds a server over a fresh in-memory database. Notifications are off
// unless a NotifyConfig is passed, so tests that predate them are unaffected.
// Dead man's switches are off too — see newDeadmanTS.
func newTS(t *testing.T, notify ...api.NotifyConfig) *ts {
t.Helper()
var cfg api.NotifyConfig
if len(notify) > 0 {
cfg = notify[0]
}
return newDeadmanTS(t, api.DeadmanConfig{}, cfg)
}
// newDeadmanTS is newTS with dead man's switch handling configured.
func newDeadmanTS(t *testing.T, deadman api.DeadmanConfig, notify ...api.NotifyConfig) *ts {
t.Helper()
var cfg api.NotifyConfig
if len(notify) > 0 {
cfg = notify[0]
}
database, err := db.Open(":memory:")
if err != nil {
t.Fatalf("open db: %v", err)
@@ -33,7 +53,7 @@ func newTS(t *testing.T) *ts {
if err := db.Migrate(database); err != nil {
t.Fatalf("migrate: %v", err)
}
srv := httptest.NewServer(api.NewRouter(database))
srv := httptest.NewServer(api.NewRouter(database, cfg, deadman))
t.Cleanup(func() { srv.Close(); database.Close() })
body, _ := json.Marshal(map[string]string{"username": "admin", "email": "admin@test.com"})
@@ -49,7 +69,7 @@ func newTS(t *testing.T) *ts {
json.NewDecoder(resp.Body).Decode(&result)
key := result["api_key"].(map[string]any)["key"].(string)
return &ts{Server: srv, key: key, db: database}
return &ts{Server: srv, key: key, db: database, notify: cfg, deadman: deadman}
}
// exec runs a statement against the test database.
@@ -302,6 +322,125 @@ func TestSchedule_MultiDateRollbackOnConflict(t *testing.T) {
}
}
// ---------------------------------------------------------------------------
// Schedule reassignment
// ---------------------------------------------------------------------------
// addUser creates a second person to hand a shift to. The bootstrap user is
// admin, id 1.
func addUser(t *testing.T, s *ts, username string) {
t.Helper()
resp := s.req(t, http.MethodPost, "/api/users",
map[string]any{"username": username, "email": username + "@test.com"})
defer resp.Body.Close()
if resp.StatusCode != http.StatusCreated {
t.Fatalf("create user returned %d", resp.StatusCode)
}
}
// scheduleHolder reports who is on call for one date, or "" for nobody.
func scheduleHolder(t *testing.T, s *ts, date string) string {
t.Helper()
var entries []map[string]any
decode(t, s.req(t, http.MethodGet, "/api/schedule?from="+date+"&to="+date, nil), &entries)
if len(entries) == 0 {
return ""
}
return entries[0]["username"].(string)
}
// Taking a day somebody else holds is possible, but only by asking for it.
func TestSchedule_ReplaceTakesAnAssignedDate(t *testing.T) {
s := newTS(t)
addUser(t, s, "alex")
s.req(t, http.MethodPost, "/api/schedule",
map[string]any{"user_id": 1, "dates": []string{"2026-06-01"}}).Body.Close()
resp := s.req(t, http.MethodPost, "/api/schedule",
map[string]any{"user_id": 2, "dates": []string{"2026-06-01"}, "replace": true})
if resp.StatusCode != http.StatusCreated {
t.Fatalf("expected replace to succeed, got %d", resp.StatusCode)
}
resp.Body.Close()
if got := scheduleHolder(t, s, "2026-06-01"); got != "alex" {
t.Errorf("expected alex to hold the day, got %q", got)
}
// One row, not two: two entries for a date would mean two people believing
// they are on call for it.
var entries []map[string]any
decode(t, s.req(t, http.MethodGet, "/api/schedule?from=2026-06-01&to=2026-06-01", nil), &entries)
if len(entries) != 1 {
t.Errorf("expected exactly one entry for the date, got %d", len(entries))
}
}
// A week where only some days are taken is the case that was impossible before:
// the free days and the taken ones have to land together.
func TestSchedule_ReplaceMixedWeek(t *testing.T) {
s := newTS(t)
addUser(t, s, "alex")
s.req(t, http.MethodPost, "/api/schedule",
map[string]any{"user_id": 1, "dates": []string{"2026-06-02", "2026-06-04"}}).Body.Close()
week := []string{"2026-06-01", "2026-06-02", "2026-06-03", "2026-06-04", "2026-06-05"}
resp := s.req(t, http.MethodPost, "/api/schedule",
map[string]any{"user_id": 2, "dates": week, "replace": true})
if resp.StatusCode != http.StatusCreated {
t.Fatalf("expected the mixed week to succeed, got %d", resp.StatusCode)
}
resp.Body.Close()
for _, d := range week {
if got := scheduleHolder(t, s, d); got != "alex" {
t.Errorf("%s: expected alex, got %q", d, got)
}
}
}
// Without replace the guard stands: nobody loses a shift by accident.
func TestSchedule_ReplaceDefaultsOff(t *testing.T) {
s := newTS(t)
addUser(t, s, "alex")
s.req(t, http.MethodPost, "/api/schedule",
map[string]any{"user_id": 1, "dates": []string{"2026-06-01"}}).Body.Close()
resp := s.req(t, http.MethodPost, "/api/schedule",
map[string]any{"user_id": 2, "dates": []string{"2026-06-01"}})
if resp.StatusCode != http.StatusConflict {
t.Fatalf("expected 409 without replace, got %d", resp.StatusCode)
}
resp.Body.Close()
if got := scheduleHolder(t, s, "2026-06-01"); got != "admin" {
t.Errorf("expected the original holder untouched, got %q", got)
}
}
// Replace makes a repeated date idempotent rather than a conflict: the second
// pass clears what the first wrote and rewrites it. Worth pinning down, because
// the same input without replace is a 409.
func TestSchedule_ReplaceCollapsesRepeatedDates(t *testing.T) {
s := newTS(t)
resp := s.req(t, http.MethodPost, "/api/schedule",
map[string]any{"user_id": 1, "dates": []string{"2026-06-01", "2026-06-01"}, "replace": true})
if resp.StatusCode != http.StatusCreated {
t.Fatalf("expected a repeated date to be accepted under replace, got %d", resp.StatusCode)
}
resp.Body.Close()
var entries []map[string]any
decode(t, s.req(t, http.MethodGet, "/api/schedule?from=2026-06-01&to=2026-06-01", nil), &entries)
if len(entries) != 1 {
t.Errorf("expected one entry for the repeated date, got %d", len(entries))
}
}
// ---------------------------------------------------------------------------
// Stats
// ---------------------------------------------------------------------------
@@ -366,7 +505,7 @@ func TestArchive_AlertListFilter(t *testing.T) {
}
// 2. Let the sweeper archive it: ends_at is already well past archiveAfter.
api.Sweep(context.Background(), s.db, time.Hour, 6*time.Hour)
api.Sweep(context.Background(), s.db, time.Hour, 6*time.Hour, s.deadman, s.notify)
// 3. Default list excludes it.
decode(t, s.req(t, http.MethodGet, "/api/alerts", nil), &alerts)
@@ -407,7 +546,7 @@ func postAlert(t *testing.T, s *ts, fingerprint, status, startsAt, endsAt string
func sweep(t *testing.T, s *ts, staleAfter time.Duration) {
t.Helper()
api.Sweep(context.Background(), s.db, noArchive, staleAfter)
api.Sweep(context.Background(), s.db, noArchive, staleAfter, s.deadman, s.notify)
}
// A firing alert Alertmanager stopped refreshing is resolved via the
+28 -11
View File
@@ -19,31 +19,38 @@ const (
// StartArchiver runs the alert sweeper until ctx is cancelled, starting with an
// immediate pass so a restart reconciles state right away.
func StartArchiver(ctx context.Context, db *sql.DB, archiveAfter, staleAfter time.Duration) {
func StartArchiver(ctx context.Context, db *sql.DB, archiveAfter, staleAfter time.Duration, deadman DeadmanConfig, notify NotifyConfig) {
ticker := time.NewTicker(sweepInterval)
defer ticker.Stop()
Sweep(ctx, db, archiveAfter, staleAfter)
Sweep(ctx, db, archiveAfter, staleAfter, deadman, notify)
for {
select {
case <-ticker.C:
Sweep(ctx, db, archiveAfter, staleAfter)
Sweep(ctx, db, archiveAfter, staleAfter, deadman, notify)
case <-ctx.Done():
return
}
}
}
// Sweep runs a single pass, in dependency order: expire stale firing alerts,
// close the incidents that leaves with nothing firing, then archive whatever has
// been settled long enough. Running them in one pass means an alert can go stale
// and its incident can close and archive without waiting three ticks.
// Sweep runs a single pass, in dependency order: reconcile the dead man's
// switches, expire stale firing alerts, close the incidents that leaves with
// nothing firing, then archive whatever has been settled long enough. Running
// them in one pass means an alert can go stale and its incident can close and
// archive without waiting three ticks.
//
// The switches go first because they hand expireStale the alerts it must not
// touch: a heartbeat answers to its own, much tighter, timeout, and the generic
// staleness rules would otherwise resolve it as 'expiry' long before that.
// Exported so tests can drive a pass without waiting on the ticker.
func Sweep(ctx context.Context, db *sql.DB, archiveAfter, staleAfter time.Duration) {
expireStale(ctx, db, staleAfter)
func Sweep(ctx context.Context, db *sql.DB, archiveAfter, staleAfter time.Duration, deadman DeadmanConfig, notify NotifyConfig) {
heartbeats := sweepDeadman(ctx, db, deadman, notify)
expireStale(ctx, db, staleAfter, heartbeats)
resolveSettledIncidents(ctx, db)
archiveResolved(ctx, db, archiveAfter)
archiveResolvedIncidents(ctx, db, archiveAfter)
purgeAckTokens(ctx, db)
}
// expireStale resolves firing alerts that Alertmanager has stopped refreshing.
@@ -59,16 +66,26 @@ func Sweep(ctx context.Context, db *sql.DB, archiveAfter, staleAfter time.Durati
// notifications every repeat_interval, making received_at a liveness
// heartbeat — provided staleAfter exceeds that interval.
//
// Alerts in skip are left alone: they are dead man's switch heartbeats, whose
// liveness sweepDeadman has already judged against a timeout of its own.
//
// The matching rows are collected before the update rather than updated in bulk,
// because each one owes its incident a timeline entry.
func expireStale(ctx context.Context, db *sql.DB, staleAfter time.Duration) {
func expireStale(ctx context.Context, db *sql.DB, staleAfter time.Duration, skip map[int64]bool) {
now := time.Now()
ids, err := staleAlertIDs(ctx, db, now, staleAfter)
found, err := staleAlertIDs(ctx, db, now, staleAfter)
if err != nil {
log.Printf("sweeper: find stale: %v", err)
return
}
ids := make([]int64, 0, len(found))
for _, id := range found {
if !skip[id] {
ids = append(ids, id)
}
}
if len(ids) == 0 {
return
}
+378
View File
@@ -0,0 +1,378 @@
package api
import (
"context"
"database/sql"
"encoding/json"
"log"
"sort"
"strings"
"time"
)
// deadmanGroupPrefix namespaces the incidents this file opens. Alertmanager
// group keys always contain braces, so this can never collide with one, and the
// partial unique index on open group_key (see 008_incidents.sql) gives one open
// incident per switch for free.
const deadmanGroupPrefix = "deadman:"
// DeadmanMatcher selects the alerts that are heartbeats rather than problems.
// Every condition has to match, and Name — the alertname label — is mandatory:
// it is what lets the sweeper find candidate rows through alerts_name_idx
// instead of JSON-extracting labels from every row in the table.
type DeadmanMatcher struct {
Name string
Labels map[string]string
}
// String renders the matcher the way it was configured, which is also how it
// reads in an incident title.
func (m DeadmanMatcher) String() string {
if len(m.Labels) == 0 {
return m.Name
}
parts := make([]string, 0, len(m.Labels))
for k, v := range m.Labels {
parts = append(parts, k+"="+v)
}
sort.Strings(parts)
return m.Name + " (" + strings.Join(parts, ", ") + ")"
}
// matches reports whether an alert's labels satisfy every condition.
func (m DeadmanMatcher) matches(labels map[string]string) bool {
if labels["alertname"] != m.Name {
return false
}
for k, v := range m.Labels {
if labels[k] != v {
return false
}
}
return true
}
// DeadmanConfig inverts the handling of the alerts it matches: receiving one
// opens nothing, and the absence of one opens an incident.
//
// The unit of monitoring is the fingerprint, not the matcher — two clusters
// sending the same heartbeat alertname are two independent switches, so one
// healthy cluster cannot mask a dead one.
type DeadmanConfig struct {
Matchers []DeadmanMatcher
// Timeout is how long a matched alert may go without a refreshing webhook
// before it is declared dead. It must be shorter than Alertmanager's
// repeat_interval for the heartbeat's route, which is what refreshes it.
// Zero disables dead man's switch handling entirely.
Timeout time.Duration
// Severity is the severity every dead man's switch incident opens at. These
// incidents have no member alerts to derive one from, and the heartbeat's
// own severity label is meaningless — Watchdog ships as "none".
Severity string
}
// enabled reports whether there is anything to watch.
func (c DeadmanConfig) enabled() bool { return c.Timeout > 0 && len(c.Matchers) > 0 }
// match returns the first matcher an alert satisfies.
func (c DeadmanConfig) match(labels map[string]string) (DeadmanMatcher, bool) {
if !c.enabled() {
return DeadmanMatcher{}, false
}
for _, m := range c.Matchers {
if m.matches(labels) {
return m, true
}
}
return DeadmanMatcher{}, false
}
// isDeadman is match without the matcher, for the ingest path.
func (c DeadmanConfig) isDeadman(labels map[string]string) bool {
_, ok := c.match(labels)
return ok
}
// names lists the distinct alertnames worth loading from the database.
func (c DeadmanConfig) names() []string {
seen := map[string]bool{}
out := make([]string, 0, len(c.Matchers))
for _, m := range c.Matchers {
if !seen[m.Name] {
seen[m.Name] = true
out = append(out, m.Name)
}
}
return out
}
// ParseDeadmanConfig reads the matcher list from its configured form:
// ";" separates matchers, "," separates the conditions within one, and "=" is
// exact label equality — `alertname=Watchdog,cluster=prod; alertname=Heartbeat`.
//
// A malformed or alertname-less entry is dropped rather than fatal, following
// config.duration's rule that one bad tuning knob should not take the server
// down. Silence would be worse here than elsewhere, though — a typo that
// disarms the switch is exactly the failure this feature exists to catch — so
// the matchers that survived are logged.
func ParseDeadmanConfig(matchers string, timeout time.Duration, severity string) DeadmanConfig {
cfg := DeadmanConfig{Timeout: timeout, Severity: severity}
for _, entry := range strings.Split(matchers, ";") {
entry = strings.TrimSpace(entry)
if entry == "" {
continue
}
m := DeadmanMatcher{Labels: map[string]string{}}
malformed := false
for _, cond := range strings.Split(entry, ",") {
k, v, ok := strings.Cut(cond, "=")
k, v = strings.TrimSpace(k), strings.TrimSpace(v)
if !ok || k == "" || v == "" {
log.Printf("deadman: ignoring matcher %q: %q is not label=value", entry, strings.TrimSpace(cond))
malformed = true
break
}
if k == "alertname" {
m.Name = v
continue
}
m.Labels[k] = v
}
if malformed {
continue
}
if m.Name == "" {
log.Printf("deadman: ignoring matcher %q: no alertname condition", entry)
continue
}
cfg.Matchers = append(cfg.Matchers, m)
}
switch {
case timeout <= 0:
log.Print("deadman: disabled (timeout is zero)")
case len(cfg.Matchers) == 0:
log.Print("deadman: disabled (no usable matchers)")
default:
rendered := make([]string, 0, len(cfg.Matchers))
for _, m := range cfg.Matchers {
rendered = append(rendered, m.String())
}
log.Printf("deadman: watching %s, timeout %s, severity %s",
strings.Join(rendered, "; "), timeout, severity)
}
return cfg
}
// deadmanAlert is one switch: the alert row carrying its last heartbeat.
type deadmanAlert struct {
id int64
fingerprint string
labels map[string]string
matcher DeadmanMatcher
resolved bool
receivedAt int64
}
// groupKey is the switch's identity as an incident. Per fingerprint, so each
// source is tracked on its own.
func (a deadmanAlert) groupKey() string { return deadmanGroupPrefix + a.fingerprint }
// sweepDeadman is the whole point of the feature: it opens an incident for every
// switch that has stopped chirping, and closes one whose switch came back.
//
// It returns the ids of the alerts it owns, because the generic staleness
// expiry must leave them alone — staleAfter and ends_at would otherwise resolve
// a heartbeat long before its own, much tighter, timeout ever fired.
func sweepDeadman(ctx context.Context, db *sql.DB, cfg DeadmanConfig, notify NotifyConfig) map[int64]bool {
owned := map[int64]bool{}
if !cfg.enabled() {
return owned
}
switches, err := deadmanAlerts(ctx, db, cfg)
if err != nil {
log.Printf("deadman: load switches: %v", err)
return owned
}
now := time.Now()
cutoff := now.Add(-cfg.Timeout).Unix()
for _, sw := range switches {
owned[sw.id] = true
// An explicit resolved from Alertmanager is a stronger death signal than
// mere absence: the sender is telling us the heartbeat stopped, so there
// is nothing left to wait out.
if sw.resolved || sw.receivedAt < cutoff {
if err := deadmanDied(ctx, db, cfg, notify, sw, now); err != nil {
log.Printf("deadman: open incident for %s: %v", sw.matcher.Name, err)
}
continue
}
if err := deadmanRecovered(ctx, db, sw); err != nil {
log.Printf("deadman: resolve incident for %s: %v", sw.matcher.Name, err)
}
}
return owned
}
// deadmanAlerts loads every alert row that a matcher claims. The candidate query
// is narrowed by alertname so it rides alerts_name_idx; the rest of the matching
// happens in Go, which keeps one implementation of the rules. The rows are read
// in full before the caller writes, because the pool holds a single connection.
func deadmanAlerts(ctx context.Context, db *sql.DB, cfg DeadmanConfig) ([]deadmanAlert, error) {
names := cfg.names()
args := make([]any, 0, len(names))
for _, n := range names {
args = append(args, n)
}
rows, err := db.QueryContext(ctx, `
SELECT id, fingerprint, labels, status, received_at
FROM alerts
WHERE name IN (`+placeholders(len(names))+`)
AND archived_at IS NULL`, args...)
if err != nil {
return nil, err
}
defer rows.Close()
var out []deadmanAlert
for rows.Next() {
var a deadmanAlert
var labelsJSON, status string
if err := rows.Scan(&a.id, &a.fingerprint, &labelsJSON, &status, &a.receivedAt); err != nil {
return nil, err
}
json.Unmarshal([]byte(labelsJSON), &a.labels) //nolint:errcheck
m, ok := cfg.match(a.labels)
if !ok {
continue
}
a.matcher = m
a.resolved = status == "resolved"
out = append(out, a)
}
return out, rows.Err()
}
// deadmanDied raises the incident for a switch that has gone quiet.
//
// Two conditions gate it, and both matter. There must be no open incident for
// the switch already — the partial unique index enforces that anyway, but a
// second one would be a wasted page. And the heartbeat must have been seen since
// the last incident was raised, which is the re-arm rule: resolving a dead man's
// switch incident sticks, exactly as resolving an alert-backed one does (see
// incidentForGroup), and a source that is gone for good is a one-time page
// rather than a nag. Only a heartbeat that comes back and dies again earns a new
// incident.
func deadmanDied(ctx context.Context, db *sql.DB, cfg DeadmanConfig, notify NotifyConfig, sw deadmanAlert, now time.Time) error {
var lastTriggered, open int64
if err := db.QueryRowContext(ctx, `
SELECT COALESCE(MAX(triggered_at), 0),
COALESCE(SUM(resolved_at IS NULL), 0)
FROM incidents WHERE group_key = ?`,
sw.groupKey()).Scan(&lastTriggered, &open); err != nil {
return err
}
if open > 0 || sw.receivedAt <= lastTriggered {
return nil
}
tx, err := db.BeginTx(ctx, nil)
if err != nil {
return err
}
defer tx.Rollback() //nolint:errcheck
// A heartbeat nobody has heard from is not firing, and saying otherwise in
// the alert list would be a lie. An Alertmanager-sourced resolution keeps its
// own source: it told us the truth first.
if !sw.resolved {
if _, err := tx.ExecContext(ctx, `
UPDATE alerts
SET status = 'resolved',
resolution_source = ?,
ends_at = COALESCE(ends_at, unixepoch())
WHERE id = ? AND status = 'firing'`, resolutionDeadman, sw.id); err != nil {
return err
}
}
severity := cfg.Severity
var sev *string
if severity != "" {
sev = &severity
}
incidentID, err := openIncident(ctx, tx, notify, sw.groupKey(),
"No heartbeat from "+sw.matcher.String(), sw.labels, sev)
if err != nil {
return err
}
alertID := sw.id
detail := "last heartbeat " + humanDuration(now.Sub(time.Unix(sw.receivedAt, 0))) + " ago"
if err := logEvent(ctx, tx, incidentID, evDeadmanSilent, nil, &alertID, &detail); err != nil {
return err
}
if err := tx.Commit(); err != nil {
return err
}
log.Printf("deadman: %s went silent, opened incident %d", sw.matcher.String(), incidentID)
return nil
}
// deadmanRecovered closes the incident for a switch that started chirping again.
//
// It cannot go through resolveIfSettled: a dead man's switch incident has no
// member alerts (linking the heartbeat would have the settled-incident cascade
// close it on the very same sweep that opened it), so the alert-driven cascade
// ignores it entirely and recovery is the only automatic way out.
func deadmanRecovered(ctx context.Context, db *sql.DB, sw deadmanAlert) error {
var incidentID int64
switch err := db.QueryRowContext(ctx, `
SELECT id FROM incidents
WHERE group_key = ? AND resolved_at IS NULL`, sw.groupKey()).Scan(&incidentID); {
case err == sql.ErrNoRows:
return nil
case err != nil:
return err
}
tx, err := db.BeginTx(ctx, nil)
if err != nil {
return err
}
defer tx.Rollback() //nolint:errcheck
if _, err := tx.ExecContext(ctx, `
UPDATE incidents
SET status = 'resolved', resolved_at = ?, resolution_source = ?
WHERE id = ? AND resolved_at IS NULL`,
time.Now().Unix(), incidentResolutionRecovered, incidentID); err != nil {
return err
}
if err := logEvent(ctx, tx, incidentID, evResolved, nil, nil, nil); err != nil {
return err
}
// The all-clear goes to whoever was paged, which enqueueResolved works out
// from the incident's own notification history.
if err := enqueueResolved(ctx, tx, incidentID); err != nil {
return err
}
if err := tx.Commit(); err != nil {
return err
}
log.Printf("deadman: %s is back, resolved incident %d", sw.matcher.String(), incidentID)
return nil
}
+468
View File
@@ -0,0 +1,468 @@
package api_test
import (
"net/http"
"testing"
"time"
"git.ryuvia.com/niklas/terdut-server/internal/api"
)
// ---------------------------------------------------------------------------
// Harness
// ---------------------------------------------------------------------------
// watchdogGroupKey is what Alertmanager sends for a Watchdog grouped by
// alertname, which is how the deployed route is configured.
const watchdogGroupKey = `{}:{alertname="Watchdog"}`
// deadmanCfg watches Watchdog with a timeout short enough to reason about and
// long enough that a fresh heartbeat is never accidentally stale.
func deadmanCfg() api.DeadmanConfig {
return api.ParseDeadmanConfig("alertname=Watchdog", time.Hour, "critical")
}
// deadmanTS is notifyTS with dead man's switch handling on: notifications
// enabled against a fake ntfy, the admin on call today with a topic.
func deadmanTS(t *testing.T, cfg api.DeadmanConfig) (*ts, *fakeNtfy) {
t.Helper()
f := newFakeNtfy(t)
s := newDeadmanTS(t, cfg, api.NotifyConfig{
BaseURL: f.URL,
PublicURL: "https://terdut.example.com",
})
putOnCall(t, s, 1)
setTopic(t, s, 1, "terdut-admin")
return s, f
}
// heartbeat posts one Watchdog webhook. Its startsAt never changes: a dead man's
// switch alert fires once and is re-sent unchanged forever, which is precisely
// what makes its absence meaningful.
func heartbeat(t *testing.T, s *ts, fingerprint string, labels map[string]string) {
t.Helper()
postWebhook(t, s, []map[string]any{
amAlert(fingerprint, "Watchdog", "firing", "2026-05-20T10:00:00Z", zeroTime, labels),
}, watchdogGroupKey)
}
// silence back-dates a heartbeat's received_at, which is the only clock the
// sweeper reads. There is no fake clock in this package.
func silence(t *testing.T, s *ts, fingerprint string, ago time.Duration) {
t.Helper()
s.exec(t, "UPDATE alerts SET received_at = ? WHERE fingerprint = ?",
time.Now().Add(-ago).Unix(), fingerprint)
}
// ageIncidents back-dates every incident. The re-arm rule compares a heartbeat
// against the last incident raised for its switch, so a test that wants a second
// episode has to put the first one in the past — there is no fake clock here.
func ageIncidents(t *testing.T, s *ts, ago time.Duration) {
t.Helper()
past := time.Now().Add(-ago).Unix()
s.exec(t, `UPDATE incidents
SET triggered_at = ?,
resolved_at = CASE WHEN resolved_at IS NULL THEN NULL ELSE ? END`,
past, past)
}
// incidentByGroup reads the incident for a group key, resolved ones included.
func incidentByGroup(t *testing.T, s *ts, groupKey string) (id int64, status, severity string, source *string) {
t.Helper()
err := s.db.QueryRow(`
SELECT id, status, COALESCE(severity, ''), resolution_source
FROM incidents WHERE group_key = ? ORDER BY id DESC LIMIT 1`,
groupKey).Scan(&id, &status, &severity, &source)
if err != nil {
t.Fatalf("read incident for group %s: %v", groupKey, err)
}
return id, status, severity, source
}
// ---------------------------------------------------------------------------
// Receiving a heartbeat
// ---------------------------------------------------------------------------
// The whole inversion: arrival of a dead man's switch alert is good news, and
// good news is not an incident.
func TestDeadman_HeartbeatOpensNoIncident(t *testing.T) {
s, _ := deadmanTS(t, deadmanCfg())
heartbeat(t, s, "fp-watchdog", nil)
if got := s.countIncidents(t); got != 0 {
t.Fatalf("expected a heartbeat to open no incident, got %d", got)
}
if got := s.countNotifications(t, ""); got != 0 {
t.Errorf("expected no notification for a heartbeat, got %d", got)
}
if status, _, _ := s.alertRow(t, "fp-watchdog"); status != "firing" {
t.Errorf("expected the heartbeat to be stored firing, got %q", status)
}
}
// A heartbeat routed into a group alongside real alerts must not join their
// incident: it is not a symptom of anything.
func TestDeadman_MixedGroupExcludesHeartbeat(t *testing.T) {
s, _ := deadmanTS(t, deadmanCfg())
postWebhook(t, s, []map[string]any{
amAlert("fp-mixed-wd", "Watchdog", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
amAlert("fp-mixed-disk", "DiskFull", "firing", "2026-05-20T10:00:00Z", zeroTime,
map[string]string{"severity": "critical"}),
}, `{}:{namespace="prod"}`)
if got := s.countIncidents(t); got != 1 {
t.Fatalf("expected 1 incident for the real alert, got %d", got)
}
var alerts []map[string]any
decode(t, s.req(t, http.MethodGet, "/api/incidents/1/alerts", nil), &alerts)
if len(alerts) != 1 {
t.Fatalf("expected 1 member alert, got %d", len(alerts))
}
if name := alerts[0]["name"]; name != "DiskFull" {
t.Errorf("expected only the real alert linked, got %v", name)
}
}
// A matcher scoped by label only claims the alerts it names, so a heartbeat from
// somewhere else stays an ordinary alert.
func TestDeadman_LabelScopedMatcherIgnoresOthers(t *testing.T) {
s, _ := deadmanTS(t, api.ParseDeadmanConfig("alertname=Watchdog,cluster=prod", time.Hour, "critical"))
heartbeat(t, s, "fp-dev", map[string]string{"cluster": "dev"})
if got := s.countIncidents(t); got != 1 {
t.Fatalf("expected an unmatched Watchdog to behave like any other alert, got %d incidents", got)
}
}
// ---------------------------------------------------------------------------
// Silence
// ---------------------------------------------------------------------------
func TestDeadman_SilenceOpensIncident(t *testing.T) {
s, f := deadmanTS(t, deadmanCfg())
heartbeat(t, s, "fp-watchdog", nil)
silence(t, s, "fp-watchdog", 2*time.Hour)
sweep(t, s, noArchive)
if got := s.countIncidents(t); got != 1 {
t.Fatalf("expected silence to open 1 incident, got %d", got)
}
id, status, severity, _ := incidentByGroup(t, s, "deadman:fp-watchdog")
if status != "triggered" {
t.Errorf("expected a triggered incident, got %q", status)
}
if severity != "critical" {
t.Errorf("expected the configured severity, got %q", severity)
}
// The alert list must not keep claiming a dead heartbeat is firing.
alertStatus, source, _ := s.alertRow(t, "fp-watchdog")
if alertStatus != "resolved" || source == nil || *source != "deadman" {
t.Errorf("expected the heartbeat resolved as deadman, got %q / %v", alertStatus, source)
}
// Nobody was told anything by an alert here, so the page has to come from
// the switch itself.
s.sweepNotify(t)
msgs := f.messages()
if len(msgs) != 1 {
t.Fatalf("expected 1 page, got %d", len(msgs))
}
if msgs[0].Topic != "terdut-admin" {
t.Errorf("expected the on-call topic, got %q", msgs[0].Topic)
}
if msgs[0].Priority != 5 {
t.Errorf("expected a critical page to override quiet hours (priority 5), got %d", msgs[0].Priority)
}
// The timeline says why, with the age of the last heartbeat.
types := eventTypes(timeline(t, s, int(id)))
found := false
for _, ty := range types {
if ty == "deadman_silent" {
found = true
}
}
if !found {
t.Errorf("expected a deadman_silent event, got %v", types)
}
}
// The generic staleness sweep must keep its hands off heartbeats: they answer to
// their own, much tighter, timeout, and an 'expiry' resolution here would be
// both wrong and unrecoverable.
func TestDeadman_GenericExpiryLeavesHeartbeatAlone(t *testing.T) {
s, _ := deadmanTS(t, deadmanCfg())
heartbeat(t, s, "fp-watchdog", nil)
silence(t, s, "fp-watchdog", 5*time.Minute)
// staleAfter far tighter than the dead man's switch timeout.
sweep(t, s, time.Minute)
status, source, _ := s.alertRow(t, "fp-watchdog")
if status != "firing" || source != nil {
t.Errorf("expected a live heartbeat left alone, got %q / %v", status, source)
}
if got := s.countIncidents(t); got != 0 {
t.Errorf("expected no incident for a heartbeat that is still fresh, got %d", got)
}
}
// An explicit resolved from Alertmanager is the sender telling us the heartbeat
// stopped. There is nothing left to wait out.
func TestDeadman_AlertmanagerResolvedIsImmediateDeath(t *testing.T) {
s, _ := deadmanTS(t, deadmanCfg())
heartbeat(t, s, "fp-watchdog", nil)
postWebhook(t, s, []map[string]any{
amAlert("fp-watchdog", "Watchdog", "resolved", "2026-05-20T10:00:00Z", zeroTime, nil),
}, watchdogGroupKey)
// No ageing: received_at is seconds old, well inside the timeout.
sweep(t, s, noArchive)
if got := s.countIncidents(t); got != 1 {
t.Fatalf("expected a resolved heartbeat to open an incident at once, got %d", got)
}
// Alertmanager told the truth first, so its resolution source stands.
if _, source, _ := s.alertRow(t, "fp-watchdog"); source == nil || *source != "alertmanager" {
t.Errorf("expected the Alertmanager resolution source kept, got %v", source)
}
}
// Each label set is its own switch, so one healthy source cannot mask a dead one.
func TestDeadman_TracksEachFingerprintSeparately(t *testing.T) {
s, _ := deadmanTS(t, deadmanCfg())
heartbeat(t, s, "fp-a", map[string]string{"cluster": "a"})
heartbeat(t, s, "fp-b", map[string]string{"cluster": "b"})
silence(t, s, "fp-b", 2*time.Hour)
sweep(t, s, noArchive)
if got := s.countIncidents(t); got != 1 {
t.Fatalf("expected only the silent switch to page, got %d incidents", got)
}
if _, status, _, _ := incidentByGroup(t, s, "deadman:fp-b"); status != "triggered" {
t.Errorf("expected the incident to belong to the silent switch, got %q", status)
}
if status, _, _ := s.alertRow(t, "fp-a"); status != "firing" {
t.Errorf("expected the live switch untouched, got %q", status)
}
}
// A switch nothing has ever been heard from is dormant. A fresh deploy, a
// restored database or a typo'd alertname must not page.
func TestDeadman_UnheardOfSwitchIsDormant(t *testing.T) {
s, _ := deadmanTS(t, api.ParseDeadmanConfig("alertname=NeverSent", time.Hour, "critical"))
sweep(t, s, noArchive)
if got := s.countIncidents(t); got != 0 {
t.Fatalf("expected a switch that never chirped to be dormant, got %d incidents", got)
}
}
// ---------------------------------------------------------------------------
// Recovery and re-arming
// ---------------------------------------------------------------------------
// The returning heartbeat carries the unchanged startsAt of an alert that never
// stopped firing, so this also covers the ingest guard exemption: without it the
// upsert would discard the payload and the switch could die exactly once.
func TestDeadman_RecoveryResolvesIncident(t *testing.T) {
s, _ := deadmanTS(t, deadmanCfg())
heartbeat(t, s, "fp-watchdog", nil)
silence(t, s, "fp-watchdog", 2*time.Hour)
sweep(t, s, noArchive)
heartbeat(t, s, "fp-watchdog", nil)
if status, source, _ := s.alertRow(t, "fp-watchdog"); status != "firing" || source != nil {
t.Fatalf("expected the returning heartbeat to be accepted, got %q / %v", status, source)
}
sweep(t, s, noArchive)
_, status, _, source := incidentByGroup(t, s, "deadman:fp-watchdog")
if status != "resolved" {
t.Errorf("expected recovery to close the incident, got %q", status)
}
if source == nil || *source != "recovered" {
t.Errorf("expected resolution_source recovered, got %v", source)
}
if got := s.countNotifications(t, "resolved"); got != 1 {
t.Errorf("expected 1 all-clear, got %d", got)
}
}
// Resolving a dead man's switch incident sticks, exactly as it does for an
// alert-backed one. A source that is gone for good is a one-time page.
func TestDeadman_ManualResolveSticksWhileSilent(t *testing.T) {
s, _ := deadmanTS(t, deadmanCfg())
heartbeat(t, s, "fp-watchdog", nil)
silence(t, s, "fp-watchdog", 2*time.Hour)
sweep(t, s, noArchive)
s.req(t, http.MethodPost, "/api/incidents/1/resolve", nil).Body.Close()
// Still silent, several sweeps later.
sweep(t, s, noArchive)
sweep(t, s, noArchive)
if got := s.countIncidents(t); got != 1 {
t.Fatalf("expected a manually resolved incident to stay closed, got %d", got)
}
}
// ...but the switch re-arms, so a heartbeat that comes back and dies again is a
// new incident rather than silence forever.
func TestDeadman_ReArmsAfterHeartbeatReturns(t *testing.T) {
s, _ := deadmanTS(t, deadmanCfg())
heartbeat(t, s, "fp-watchdog", nil)
silence(t, s, "fp-watchdog", 2*time.Hour)
sweep(t, s, noArchive)
s.req(t, http.MethodPost, "/api/incidents/1/resolve", nil).Body.Close()
// That episode is yesterday's news; the heartbeat now returns after it.
ageIncidents(t, s, 10*time.Hour)
heartbeat(t, s, "fp-watchdog", nil)
sweep(t, s, noArchive)
if got := s.countIncidents(t); got != 1 {
t.Fatalf("expected the live switch to open nothing, got %d incidents", got)
}
silence(t, s, "fp-watchdog", 2*time.Hour)
sweep(t, s, noArchive)
if got := s.countIncidents(t); got != 2 {
t.Fatalf("expected a second death to open a second incident, got %d", got)
}
}
// A dead man's switch incident has no member alerts — linking the heartbeat
// would have the settled-incident cascade close it on the very sweep that opened
// it — so the cascade has to leave it alone.
func TestDeadman_SettledCascadeLeavesIncidentOpen(t *testing.T) {
s, _ := deadmanTS(t, deadmanCfg())
heartbeat(t, s, "fp-watchdog", nil)
silence(t, s, "fp-watchdog", 2*time.Hour)
sweep(t, s, noArchive)
sweep(t, s, noArchive)
if _, status, _, _ := incidentByGroup(t, s, "deadman:fp-watchdog"); status != "triggered" {
t.Fatalf("expected the incident to stay open until the switch recovers, got %q", status)
}
}
// ---------------------------------------------------------------------------
// Configuration
// ---------------------------------------------------------------------------
func TestParseDeadmanConfig(t *testing.T) {
tests := []struct {
name string
matchers string
timeout time.Duration
want []api.DeadmanMatcher
enabled bool
}{
{
name: "single alertname",
matchers: "alertname=Watchdog",
timeout: time.Hour,
want: []api.DeadmanMatcher{{Name: "Watchdog", Labels: map[string]string{}}},
enabled: true,
},
{
name: "several matchers with extra labels and whitespace",
matchers: " alertname=Watchdog, cluster=prod ; alertname=EdgeHeartbeat ",
timeout: time.Hour,
want: []api.DeadmanMatcher{
{Name: "Watchdog", Labels: map[string]string{"cluster": "prod"}},
{Name: "EdgeHeartbeat", Labels: map[string]string{}},
},
enabled: true,
},
{
// Mandatory: it is what keeps the sweeper's candidate query on an index.
name: "matcher without alertname is dropped",
matchers: "cluster=prod; alertname=Watchdog",
timeout: time.Hour,
want: []api.DeadmanMatcher{{Name: "Watchdog", Labels: map[string]string{}}},
enabled: true,
},
{
name: "malformed condition drops only its matcher",
matchers: "alertname=Watchdog,garbage; alertname=Other",
timeout: time.Hour,
want: []api.DeadmanMatcher{{Name: "Other", Labels: map[string]string{}}},
enabled: true,
},
{
name: "zero timeout disables",
matchers: "alertname=Watchdog",
timeout: 0,
want: []api.DeadmanMatcher{{Name: "Watchdog", Labels: map[string]string{}}},
enabled: false,
},
{
name: "no usable matchers disables",
matchers: "",
timeout: time.Hour,
want: nil,
enabled: false,
},
}
for _, tc := range tests {
t.Run(tc.name, func(t *testing.T) {
got := api.ParseDeadmanConfig(tc.matchers, tc.timeout, "critical")
if len(got.Matchers) != len(tc.want) {
t.Fatalf("got %d matchers %v, want %d", len(got.Matchers), got.Matchers, len(tc.want))
}
for i, w := range tc.want {
if got.Matchers[i].Name != w.Name {
t.Errorf("matcher %d: name %q, want %q", i, got.Matchers[i].Name, w.Name)
}
if len(got.Matchers[i].Labels) != len(w.Labels) {
t.Errorf("matcher %d: labels %v, want %v", i, got.Matchers[i].Labels, w.Labels)
continue
}
for k, v := range w.Labels {
if got.Matchers[i].Labels[k] != v {
t.Errorf("matcher %d: label %s=%q, want %q", i, k, got.Matchers[i].Labels[k], v)
}
}
}
})
}
}
// A zero config is off, which is what keeps the feature opt-in for anything
// building a router without one.
func TestDeadman_DisabledConfigIsInert(t *testing.T) {
s, _ := deadmanTS(t, api.DeadmanConfig{})
heartbeat(t, s, "fp-watchdog", nil)
silence(t, s, "fp-watchdog", 48*time.Hour)
sweep(t, s, time.Hour)
// Ordinary alert handling: an incident from the arrival, not the absence.
if got := s.countIncidents(t); got != 1 {
t.Fatalf("expected plain alert handling with deadman off, got %d incidents", got)
}
if _, source, _ := s.alertRow(t, "fp-watchdog"); source == nil || *source != "expiry" {
t.Errorf("expected the generic sweeper to own the alert, got %v", source)
}
}
+33 -2
View File
@@ -8,7 +8,7 @@ import (
"strings"
"time"
"github.com/yeniklas/terdut-server/internal/models"
"git.ryuvia.com/niklas/terdut-server/internal/models"
)
// Values for incidents.resolution_source, recording who closed the incident:
@@ -16,6 +16,11 @@ import (
const (
incidentResolutionAlerts = "alerts"
incidentResolutionManual = "manual"
// incidentResolutionRecovered closes a dead man's switch incident whose
// heartbeat started arriving again. It cannot be "alerts": these incidents
// have no member alerts for the cascade to work from.
incidentResolutionRecovered = "recovered"
)
// Incident timeline event types. Stored as free text so adding one later is not
@@ -31,6 +36,7 @@ const (
evUnsnoozed = "unsnoozed"
evResolved = "resolved"
evNote = "note"
evDeadmanSilent = "deadman_silent"
)
// severityLabel is the Alertmanager label an incident's severity is derived from.
@@ -223,7 +229,32 @@ func resolveIfSettled(ctx context.Context, q querier, incidentID int64) (bool, e
if n == 0 {
return false, nil
}
return true, logEvent(ctx, q, incidentID, evResolved, nil, nil, nil)
if err := logEvent(ctx, q, incidentID, evResolved, nil, nil, nil); err != nil {
return false, err
}
// The all-clear goes only to whoever was paged in the first place, which
// enqueueResolved works out from the incident's own notification history.
// Manual resolution sends nothing: the person who closed it already knows.
return true, enqueueResolved(ctx, q, incidentID)
}
// acknowledgeIncident records that userID has picked an incident up, and reports
// whether it changed anything — an already-resolved incident is left alone.
// Shared by the authenticated handler and the Acknowledge button in a push
// notification, so both write the same state and the same timeline entry.
func acknowledgeIncident(ctx context.Context, q querier, incidentID, userID int64) (bool, error) {
res, err := q.ExecContext(ctx, `
UPDATE incidents
SET status = 'acknowledged', acknowledged_by = ?, acknowledged_at = ?
WHERE id = ? AND resolved_at IS NULL`,
userID, time.Now().Unix(), incidentID)
if err != nil {
return false, err
}
if n, _ := res.RowsAffected(); n == 0 {
return false, nil
}
return true, logEvent(ctx, q, incidentID, evAcknowledged, &userID, nil, nil)
}
// openIncidentForAlert returns the open incident an alert currently belongs to,
+9 -6
View File
@@ -9,7 +9,7 @@ import (
"time"
"github.com/go-chi/chi/v5"
"github.com/yeniklas/terdut-server/internal/models"
"git.ryuvia.com/niklas/terdut-server/internal/models"
)
func handleListIncidents(db *sql.DB) http.HandlerFunc {
@@ -189,13 +189,16 @@ func handleIncidentAcknowledge(db *sql.DB) http.HandlerFunc {
return
}
user, _ := userFromContext(r.Context())
if !updateOpenIncident(w, r, db, id,
`UPDATE incidents SET status = 'acknowledged', acknowledged_by = ?, acknowledged_at = ?
WHERE id = ? AND resolved_at IS NULL`, user.ID, time.Now().Unix(), id) {
acked, err := acknowledgeIncident(r.Context(), db, id, user.ID)
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
if err := logEvent(r.Context(), db, id, evAcknowledged, &user.ID, nil, nil); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
if !acked {
if !incidentExists(w, r, db, id) {
return
}
respond(w, http.StatusConflict, errResp("incident is resolved"))
return
}
respondIncident(w, r, db, id)
+3 -3
View File
@@ -10,8 +10,8 @@ import (
"testing"
"time"
"github.com/yeniklas/terdut-server/internal/api"
"github.com/yeniklas/terdut-server/internal/db"
"git.ryuvia.com/niklas/terdut-server/internal/api"
"git.ryuvia.com/niklas/terdut-server/internal/db"
)
// amAlert builds one alert of a webhook payload.
@@ -613,7 +613,7 @@ func TestSweeper_ArchivesResolvedIncidents(t *testing.T) {
s.exec(t, "UPDATE incidents SET resolved_at = ? WHERE id = 1",
time.Now().Add(-30*24*time.Hour).Unix())
api.Sweep(context.Background(), s.db, 7*24*time.Hour, 6*time.Hour)
api.Sweep(context.Background(), s.db, 7*24*time.Hour, 6*time.Hour, s.deadman, s.notify)
if inc := getIncident(t, s, 1); inc["archived_at"] == nil {
t.Error("expected the sweeper to archive a long-resolved incident")
+1 -1
View File
@@ -9,7 +9,7 @@ import (
"strings"
"time"
"github.com/yeniklas/terdut-server/internal/models"
"git.ryuvia.com/niklas/terdut-server/internal/models"
)
type contextKey string
+563
View File
@@ -0,0 +1,563 @@
package api
import (
"bytes"
"context"
"database/sql"
"encoding/json"
"fmt"
"log"
"net/http"
"strings"
"time"
"git.ryuvia.com/niklas/terdut-server/internal/models"
)
const (
// notifyInterval is how often the notifier looks for work. The archiver's
// 15 minute tick is far too coarse for something that has to wake a person.
notifyInterval = 30 * time.Second
// notifyRetryBase and notifyRetryMax bound the delivery backoff. ntfy being
// briefly unreachable should not lose the page.
notifyRetryBase = 30 * time.Second
notifyRetryMax = 15 * time.Minute
// notifyMaxAttempts stops a permanently undeliverable row from being retried
// forever. It keeps last_error so the reason survives.
notifyMaxAttempts = 8
// notifyBatch caps one delivery pass, so a large backlog cannot hold the
// single database connection for an unbounded stretch.
notifyBatch = 100
// ackTokenTTL is how long the Acknowledge button in a notification keeps
// working. Past this the notification is stale enough that the responder
// should look at the incident rather than blind-acknowledge it.
ackTokenTTL = 24 * time.Hour
)
// Notification kinds, recording why a push was sent.
const (
notifyTriggered = "triggered"
notifyReminder = "reminder"
notifyResolved = "resolved"
)
// Timeline event types the notifier writes, so an incident's history says who
// was paged and whether the page landed. Written from the delivery result
// rather than at enqueue: a queued notification is an intention, and claiming
// somebody was told before ntfy accepted it would be a lie the timeline keeps.
//
// The topic is deliberately absent from both. It is a shared secret with the
// ntfy server — anyone holding it can publish to it — and the timeline is
// readable by every API key.
const (
eventNotified = "notified"
eventNotifyFailed = "notify_failed"
)
// NotifyConfig is everything the notifier needs to reach ntfy and to build URLs
// a phone can follow back to this server.
type NotifyConfig struct {
// BaseURL is the ntfy server. Empty disables notifications entirely: no
// goroutine, and nothing is ever enqueued.
BaseURL string
// Token is an optional bearer token for an access-controlled ntfy.
Token string
// FallbackTopic receives incidents that open with nobody on call. Those
// notifications carry no Acknowledge button — there is no user to attribute
// the acknowledgement to, and putting one on a shared topic would let any
// subscriber acknowledge as somebody else.
FallbackTopic string
// PublicURL is the base URL a phone uses to reach this server, for the
// notification's click target and its Acknowledge action. Without it a
// notification is informational only.
PublicURL string
// RepeatEvery is how long an incident may sit unacknowledged before it is
// notified again. Zero disables reminders.
RepeatEvery time.Duration
}
// enabled reports whether notifications are configured at all.
func (c NotifyConfig) enabled() bool { return c.BaseURL != "" }
// notifyClient is shared: a page is small and infrequent, and the timeout is
// what keeps a hung ntfy from stalling the delivery pass.
var notifyClient = &http.Client{Timeout: 10 * time.Second}
// StartNotifier delivers queued notifications until ctx is cancelled, starting
// with an immediate pass so a restart flushes whatever the last one left behind.
func StartNotifier(ctx context.Context, db *sql.DB, cfg NotifyConfig) {
if !cfg.enabled() {
log.Print("notifier: disabled (no ntfy URL configured)")
return
}
log.Printf("notifier: publishing to %s", cfg.BaseURL)
ticker := time.NewTicker(notifyInterval)
defer ticker.Stop()
NotifySweep(ctx, db, cfg)
for {
select {
case <-ticker.C:
NotifySweep(ctx, db, cfg)
case <-ctx.Done():
return
}
}
}
// NotifySweep runs a single pass: queue reminders for incidents nobody has
// picked up, then deliver everything that is due. Reminders are queued first so
// a freshly due one goes out in the same pass rather than a tick later.
// Exported so tests can drive a pass without waiting on the ticker.
func NotifySweep(ctx context.Context, db *sql.DB, cfg NotifyConfig) {
enqueueReminders(ctx, db, cfg)
deliverPending(ctx, db, cfg)
}
// enqueueReminders re-notifies incidents that are still sitting untouched.
//
// The stop conditions are the incident states that already mean "somebody has
// this": acknowledged, snoozed, resolved, archived. Snooze in particular is the
// mute button — a deliberate "not now" that should not keep buzzing — which is
// why there is no separate reminder cap.
//
// The previous notification must have actually been sent before another is
// queued, so an ntfy outage produces a retry backlog rather than a reminder
// backlog that all lands at once when it comes back.
func enqueueReminders(ctx context.Context, db *sql.DB, cfg NotifyConfig) {
if cfg.RepeatEvery <= 0 {
return
}
now := time.Now()
type due struct {
incidentID int64
userID *int64
topic string
}
rows, err := db.QueryContext(ctx, `
SELECT n.incident_id, n.user_id, n.topic
FROM notifications n
JOIN incidents i ON i.id = n.incident_id
WHERE n.id = (SELECT MAX(id) FROM notifications WHERE incident_id = n.incident_id)
AND n.sent_at IS NOT NULL
AND n.created_at <= ?
AND i.resolved_at IS NULL
AND i.archived_at IS NULL
AND i.status = 'triggered'
AND (i.snoozed_until IS NULL OR i.snoozed_until <= ?)`,
now.Add(-cfg.RepeatEvery).Unix(), now.Unix())
if err != nil {
log.Printf("notifier: find reminders: %v", err)
return
}
// Collected before inserting: the pool is limited to a single connection, so
// an open cursor would block the writes behind it.
var pending []due
for rows.Next() {
var d due
if err := rows.Scan(&d.incidentID, &d.userID, &d.topic); err != nil {
rows.Close()
log.Printf("notifier: scan reminder: %v", err)
return
}
pending = append(pending, d)
}
if err := rows.Err(); err != nil {
rows.Close()
log.Printf("notifier: find reminders: %v", err)
return
}
rows.Close()
for _, d := range pending {
if err := enqueueNotification(ctx, db, d.incidentID, d.userID, d.topic, notifyReminder); err != nil {
log.Printf("notifier: queue reminder for incident %d: %v", d.incidentID, err)
}
}
if len(pending) > 0 {
log.Printf("notifier: queued %d reminder(s)", len(pending))
}
}
// outboxRow is one queued notification, read before any HTTP happens.
type outboxRow struct {
id int64
incidentID int64
userID *int64
topic string
kind string
attempts int
}
// deliverPending sends everything that is due and records the outcome.
func deliverPending(ctx context.Context, db *sql.DB, cfg NotifyConfig) {
batch, err := pendingNotifications(ctx, db)
if err != nil {
log.Printf("notifier: find pending: %v", err)
return
}
sent := 0
for _, n := range batch {
if err := deliver(ctx, db, cfg, n); err != nil {
log.Printf("notifier: deliver %d (incident %d): %v", n.id, n.incidentID, err)
markFailed(ctx, db, n, err)
continue
}
if _, err := db.ExecContext(ctx,
"UPDATE notifications SET sent_at = ?, attempts = attempts + 1, last_error = NULL WHERE id = ?",
time.Now().Unix(), n.id); err != nil {
log.Printf("notifier: mark sent %d: %v", n.id, err)
}
// Logged, not returned: the page has already gone out, and treating a
// failed timeline write as a failed delivery would send it again.
if err := logEvent(ctx, db, n.incidentID, eventNotified, n.userID, nil, &n.kind); err != nil {
log.Printf("notifier: log delivery of %d: %v", n.id, err)
}
sent++
}
if sent > 0 {
log.Printf("notifier: delivered %d notification(s)", sent)
}
}
// pendingNotifications reads the due rows and closes the cursor before the
// caller writes, for the same single-connection reason as staleAlertIDs.
func pendingNotifications(ctx context.Context, db *sql.DB) ([]outboxRow, error) {
rows, err := db.QueryContext(ctx, `
SELECT id, incident_id, user_id, topic, kind, attempts
FROM notifications
WHERE sent_at IS NULL
AND send_after <= ?
AND attempts < ?
ORDER BY id
LIMIT ?`, time.Now().Unix(), notifyMaxAttempts, notifyBatch)
if err != nil {
return nil, err
}
defer rows.Close()
var batch []outboxRow
for rows.Next() {
var n outboxRow
if err := rows.Scan(&n.id, &n.incidentID, &n.userID, &n.topic, &n.kind, &n.attempts); err != nil {
return nil, err
}
batch = append(batch, n)
}
return batch, rows.Err()
}
// markFailed bumps the attempt count and pushes the row out to its next retry.
//
// The attempt that exhausts the budget also writes a timeline event. Without it
// a page that never landed leaves the incident's history identical to one that
// did, which is the failure most worth seeing: nobody was told, and nothing
// says so.
func markFailed(ctx context.Context, db *sql.DB, n outboxRow, cause error) {
next := time.Now().Add(retryDelay(n.attempts)).Unix()
if _, err := db.ExecContext(ctx,
"UPDATE notifications SET attempts = attempts + 1, send_after = ?, last_error = ? WHERE id = ?",
next, cause.Error(), n.id); err != nil {
log.Printf("notifier: mark failed %d: %v", n.id, err)
}
if n.attempts+1 < notifyMaxAttempts {
return
}
detail := fmt.Sprintf("%s: %s", n.kind, cause)
if err := logEvent(ctx, db, n.incidentID, eventNotifyFailed, n.userID, nil, &detail); err != nil {
log.Printf("notifier: log failure of %d: %v", n.id, err)
}
}
// retryDelay doubles the wait per attempt, up to notifyRetryMax.
func retryDelay(attempts int) time.Duration {
d := notifyRetryBase << attempts
if d > notifyRetryMax || d <= 0 {
return notifyRetryMax
}
return d
}
// deliver renders one notification against the incident's *current* state and
// publishes it. Rendering happens here rather than at enqueue time so a message
// that waited in the queue while its incident escalated goes out at the
// severity the incident has now.
func deliver(ctx context.Context, db *sql.DB, cfg NotifyConfig, n outboxRow) error {
inc, err := fetchIncident(ctx, db, n.incidentID)
if err != nil {
return fmt.Errorf("load incident: %w", err)
}
var firing int
if err := db.QueryRowContext(ctx, `
SELECT COUNT(*)
FROM incident_alerts ia
JOIN alerts a ON a.id = ia.alert_id
WHERE ia.incident_id = ? AND a.status = 'firing'`, n.incidentID).Scan(&firing); err != nil {
return fmt.Errorf("count firing: %w", err)
}
msg := renderNotification(inc, n, firing, cfg)
// An Acknowledge button needs both a user to attribute the acknowledgement
// to and a URL the phone can reach. Minted per delivery, so every push
// carries its own short-lived token rather than reusing one.
if n.kind != notifyResolved && n.userID != nil && cfg.PublicURL != "" {
raw, err := issueAckToken(ctx, db, n.incidentID, *n.userID)
if err != nil {
return fmt.Errorf("issue ack token: %w", err)
}
msg.Actions = append(msg.Actions, ntfyAction{
Action: "http",
Label: "Acknowledge",
URL: strings.TrimSuffix(cfg.PublicURL, "/") + "/api/notify/ack/" + raw,
Method: "POST",
Clear: true,
})
}
return publish(ctx, cfg, msg)
}
// ntfyMessage is ntfy's JSON publish format. Using it rather than the X-Actions
// header avoids that header's comma and quote escaping rules, which are easy to
// break with a title that happens to contain a comma.
type ntfyMessage struct {
Topic string `json:"topic"`
Title string `json:"title,omitempty"`
Message string `json:"message"`
Priority int `json:"priority,omitempty"`
Tags []string `json:"tags,omitempty"`
Click string `json:"click,omitempty"`
Actions []ntfyAction `json:"actions,omitempty"`
}
type ntfyAction struct {
Action string `json:"action"`
Label string `json:"label"`
URL string `json:"url"`
Method string `json:"method,omitempty"`
Clear bool `json:"clear,omitempty"`
}
// renderNotification builds the message body for one queued notification.
func renderNotification(inc models.Incident, n outboxRow, firing int, cfg NotifyConfig) ntfyMessage {
msg := ntfyMessage{Topic: n.topic}
if cfg.PublicURL != "" {
msg.Click = fmt.Sprintf("%s/api/incidents/%d",
strings.TrimSuffix(cfg.PublicURL, "/"), inc.ID)
}
switch n.kind {
case notifyResolved:
msg.Title = "Resolved: " + inc.Title
msg.Message = "All alerts stopped firing after " +
humanDuration(time.Since(inc.TriggeredAt))
msg.Priority = ntfyPriorityLow
msg.Tags = []string{"white_check_mark"}
return msg
case notifyReminder:
msg.Title = "Still unacknowledged: " + inc.Title
default:
msg.Title = inc.Title
}
severity := derefString(inc.Severity)
parts := []string{fmt.Sprintf("%d alert%s firing", firing, plural(firing))}
if severity != "" {
parts = append(parts, "severity "+severity)
}
if assignee := derefString(inc.AssignedToUser); assignee != "" {
parts = append(parts, "on call: "+assignee)
}
if n.kind == notifyReminder {
parts = append(parts, "open "+humanDuration(time.Since(inc.TriggeredAt)))
}
msg.Message = strings.Join(parts, " · ")
msg.Priority = ntfyPriority(severity)
msg.Tags = []string{severityTag(severity)}
return msg
}
// ntfy's priority scale. Max is the one that overrides the phone's quiet
// settings, which is the whole point of paging on critical.
const (
ntfyPriorityLow = 2
ntfyPriorityDefault = 3
ntfyPriorityHigh = 4
ntfyPriorityMax = 5
)
// ntfyPriority maps an incident's severity onto ntfy's scale, following the
// same ordering severityRank uses. An unrecognised severity gets the default
// rather than being silenced.
func ntfyPriority(severity string) int {
switch severityRank(severity) {
case 4:
return ntfyPriorityMax
case 3:
return ntfyPriorityHigh
case 2:
return ntfyPriorityDefault
case 1:
return ntfyPriorityLow
default:
return ntfyPriorityDefault
}
}
func severityTag(severity string) string {
switch severityRank(severity) {
case 4:
return "rotating_light"
case 3:
return "red_circle"
case 2:
return "warning"
case 1:
return "information_source"
default:
return "bell"
}
}
// publish POSTs one message to ntfy.
func publish(ctx context.Context, cfg NotifyConfig, msg ntfyMessage) error {
body, err := json.Marshal(msg)
if err != nil {
return err
}
req, err := http.NewRequestWithContext(ctx, http.MethodPost,
strings.TrimSuffix(cfg.BaseURL, "/"), bytes.NewReader(body))
if err != nil {
return err
}
req.Header.Set("Content-Type", "application/json")
if cfg.Token != "" {
req.Header.Set("Authorization", "Bearer "+cfg.Token)
}
resp, err := notifyClient.Do(req)
if err != nil {
return err
}
defer resp.Body.Close()
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
return fmt.Errorf("ntfy returned %s", resp.Status)
}
return nil
}
// enqueueNotification adds one row to the outbox, due immediately.
func enqueueNotification(ctx context.Context, q querier, incidentID int64, userID *int64, topic, kind string) error {
now := time.Now().Unix()
_, err := q.ExecContext(ctx, `
INSERT INTO notifications (incident_id, user_id, topic, kind, created_at, send_after)
VALUES (?, ?, ?, ?, ?, ?)`, incidentID, userID, topic, kind, now, now)
return err
}
// notifyTarget decides where a newly opened incident's notification goes.
//
// The on-call user's own topic when they have one, otherwise the fallback
// topic with no user attached. Deliberately not "the fallback topic, attributed
// to the on-call user": the fallback is shared, and an Acknowledge button on a
// shared topic would let any subscriber acknowledge as somebody else.
func notifyTarget(ctx context.Context, q querier, cfg NotifyConfig, onCall *int64) (topic string, userID *int64) {
if onCall != nil {
var t *string
err := q.QueryRowContext(ctx,
"SELECT ntfy_topic FROM users WHERE id = ?", *onCall).Scan(&t)
if err == nil && t != nil && *t != "" {
return *t, onCall
}
}
return cfg.FallbackTopic, nil
}
// enqueueOpened queues the notification for a freshly opened incident. It is the
// only enqueue point that has to resolve a topic from scratch; every later
// notification for the incident reuses what this one chose.
func enqueueOpened(ctx context.Context, q querier, cfg NotifyConfig, incidentID int64, onCall *int64) error {
if !cfg.enabled() {
return nil
}
topic, userID := notifyTarget(ctx, q, cfg, onCall)
if topic == "" {
// Nobody on call has a topic and there is no fallback: there is nowhere
// to send this, and queueing it would only accumulate undeliverable rows.
return nil
}
return enqueueNotification(ctx, q, incidentID, userID, topic, notifyTriggered)
}
// enqueueResolved queues the all-clear, reusing the topic the incident's last
// notification went to. That needs no configuration to reach this function, and
// it gives the right rule for free: you only hear that something resolved if you
// were told it started.
func enqueueResolved(ctx context.Context, q querier, incidentID int64) error {
var topic string
var userID *int64
err := q.QueryRowContext(ctx, `
SELECT topic, user_id FROM notifications
WHERE incident_id = ? ORDER BY id DESC LIMIT 1`, incidentID).Scan(&topic, &userID)
if err == sql.ErrNoRows {
return nil
}
if err != nil {
return err
}
return enqueueNotification(ctx, q, incidentID, userID, topic, notifyResolved)
}
// humanDuration renders an age the way a person reads it at 3am: coarse, and
// never more than two units.
func humanDuration(d time.Duration) string {
if d < time.Minute {
return "less than a minute"
}
if d < time.Hour {
return fmt.Sprintf("%dm", int(d.Minutes()))
}
h := int(d.Hours())
m := int(d.Minutes()) - h*60
if m == 0 {
return fmt.Sprintf("%dh", h)
}
return fmt.Sprintf("%dh%dm", h, m)
}
func plural(n int) string {
if n == 1 {
return ""
}
return "s"
}
// derefString reads a nullable text column as a plain string.
func derefString(s *string) string {
if s == nil {
return ""
}
return *s
}
+98
View File
@@ -0,0 +1,98 @@
package api
import (
"context"
"crypto/sha256"
"database/sql"
"encoding/hex"
"log"
"net/http"
"time"
"github.com/go-chi/chi/v5"
)
// issueAckToken mints the secret behind one notification's Acknowledge button
// and returns the raw value to embed in its URL. Only the hash is stored, the
// same way api_keys works.
//
// A fresh token per delivery rather than one per incident: the raw value only
// exists for as long as it takes to build the message, so there is nothing to
// look up and reuse later, and a reminder that supersedes an earlier page
// carries its own credential.
func issueAckToken(ctx context.Context, q querier, incidentID, userID int64) (string, error) {
raw, hash, err := randomToken()
if err != nil {
return "", err
}
now := time.Now()
if _, err := q.ExecContext(ctx, `
INSERT INTO incident_ack_tokens (token_hash, incident_id, user_id, created_at, expires_at)
VALUES (?, ?, ?, ?, ?)`,
hash, incidentID, userID, now.Unix(), now.Add(ackTokenTTL).Unix()); err != nil {
return "", err
}
return raw, nil
}
// handleNotifyAck acknowledges an incident from the Acknowledge button in a
// push notification.
//
// It is deliberately outside AuthMiddleware: the caller is a phone acting on a
// notification, not a client holding an API key. What stands in for the key is
// the token in the path — 256 bits of entropy, valid for one incident, one
// action, and one day. It must stay publicly reachable for the button to work
// when the responder is off the cluster network.
func handleNotifyAck(db *sql.DB) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
h := sha256.Sum256([]byte(chi.URLParam(r, "token")))
hash := hex.EncodeToString(h[:])
var incidentID, userID int64
err := db.QueryRowContext(r.Context(), `
SELECT incident_id, user_id FROM incident_ack_tokens
WHERE token_hash = ? AND expires_at > ?`,
hash, time.Now().Unix()).Scan(&incidentID, &userID)
if err != nil {
// Unknown and expired get the same answer, so the endpoint cannot be
// used to probe which tokens once existed.
respond(w, http.StatusNotFound, errResp("invalid or expired token"))
return
}
acked, err := acknowledgeIncident(r.Context(), db, incidentID, userID)
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
if !acked {
// The incident closed between the page and the tap. Nothing to do,
// and nothing the responder did wrong — report the state, not an error,
// so ntfy shows a success toast rather than a failure.
respond(w, http.StatusOK, map[string]any{
"incident_id": incidentID,
"status": "resolved",
})
return
}
respond(w, http.StatusOK, map[string]any{
"incident_id": incidentID,
"status": "acknowledged",
})
}
}
// purgeAckTokens drops tokens whose notifications are long past. Nothing else
// deletes them: incidents are archived rather than removed, so the cascade never
// fires in practice.
func purgeAckTokens(ctx context.Context, db *sql.DB) {
res, err := db.ExecContext(ctx,
"DELETE FROM incident_ack_tokens WHERE expires_at < ?", time.Now().Unix())
if err != nil {
log.Printf("sweeper: purge ack tokens: %v", err)
return
}
if n, _ := res.RowsAffected(); n > 0 {
log.Printf("sweeper: purged %d expired ack token(s)", n)
}
}
+724
View File
@@ -0,0 +1,724 @@
package api_test
import (
"context"
"encoding/json"
"fmt"
"net/http"
"net/http/httptest"
"strings"
"sync"
"testing"
"time"
"git.ryuvia.com/niklas/terdut-server/internal/api"
)
// ---------------------------------------------------------------------------
// Fake ntfy
// ---------------------------------------------------------------------------
// pushed is one message the fake ntfy received, in ntfy's JSON publish shape.
type pushed struct {
Topic string `json:"topic"`
Title string `json:"title"`
Message string `json:"message"`
Priority int `json:"priority"`
Tags []string `json:"tags"`
Click string `json:"click"`
Actions []struct {
Action string `json:"action"`
Label string `json:"label"`
URL string `json:"url"`
Method string `json:"method"`
Clear bool `json:"clear"`
} `json:"actions"`
}
// fakeNtfy records what the notifier published. status controls the reply, so a
// test can make delivery fail.
type fakeNtfy struct {
*httptest.Server
mu sync.Mutex
got []pushed
status int
}
func newFakeNtfy(t *testing.T) *fakeNtfy {
t.Helper()
f := &fakeNtfy{status: http.StatusOK}
f.Server = httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
var msg pushed
if err := json.NewDecoder(r.Body).Decode(&msg); err != nil {
http.Error(w, "bad json", http.StatusBadRequest)
return
}
f.mu.Lock()
f.got = append(f.got, msg)
status := f.status
f.mu.Unlock()
w.WriteHeader(status)
}))
t.Cleanup(f.Close)
return f
}
func (f *fakeNtfy) messages() []pushed {
f.mu.Lock()
defer f.mu.Unlock()
return append([]pushed(nil), f.got...)
}
func (f *fakeNtfy) failWith(status int) {
f.mu.Lock()
defer f.mu.Unlock()
f.status = status
}
// ---------------------------------------------------------------------------
// Harness
// ---------------------------------------------------------------------------
// notifyTS builds a server with notifications enabled, the admin on call today,
// and a topic on the admin — the setup every delivery test needs.
func notifyTS(t *testing.T, cfg api.NotifyConfig) (*ts, *fakeNtfy) {
t.Helper()
f := newFakeNtfy(t)
cfg.BaseURL = f.URL
s := newTS(t, cfg)
putOnCall(t, s, 1)
setTopic(t, s, 1, "terdut-admin")
return s, f
}
func putOnCall(t *testing.T, s *ts, userID int) {
t.Helper()
today := time.Now().UTC().Format("2006-01-02")
resp := s.req(t, http.MethodPost, "/api/schedule",
map[string]any{"user_id": userID, "dates": []string{today}})
defer resp.Body.Close()
if resp.StatusCode != http.StatusCreated {
t.Fatalf("schedule assignment returned %d", resp.StatusCode)
}
}
func setTopic(t *testing.T, s *ts, userID int, topic string) {
t.Helper()
resp := s.req(t, http.MethodPut,
fmt.Sprintf("/api/users/%d/notify", userID), map[string]any{"ntfy_topic": topic})
defer resp.Body.Close()
if resp.StatusCode != http.StatusOK {
t.Fatalf("set notify topic returned %d", resp.StatusCode)
}
}
func (s *ts) sweepNotify(t *testing.T) {
t.Helper()
api.NotifySweep(context.Background(), s.db, s.notify)
}
// countNotifications reports how many outbox rows exist, optionally of one kind.
func (s *ts) countNotifications(t *testing.T, kind string) int {
t.Helper()
var n int
query := "SELECT COUNT(*) FROM notifications"
args := []any{}
if kind != "" {
query += " WHERE kind = ?"
args = append(args, kind)
}
if err := s.db.QueryRow(query, args...).Scan(&n); err != nil {
t.Fatalf("count notifications: %v", err)
}
return n
}
// fireCritical posts a single critical alert, which opens one incident.
func fireCritical(t *testing.T, s *ts) {
t.Helper()
postWebhook(t, s, []map[string]any{
amAlert("fp-notify", "DiskFull", "firing", "2026-05-20T10:00:00Z", zeroTime,
map[string]string{"severity": "critical"}),
}, "{}:{alertname=\"DiskFull\"}")
}
// ---------------------------------------------------------------------------
// Delivery
// ---------------------------------------------------------------------------
func TestNotify_TriggeredIncidentPagesOnCall(t *testing.T) {
s, f := notifyTS(t, api.NotifyConfig{PublicURL: "https://terdut.example.com"})
fireCritical(t, s)
if got := s.countNotifications(t, "triggered"); got != 1 {
t.Fatalf("expected 1 queued notification, got %d", got)
}
s.sweepNotify(t)
msgs := f.messages()
if len(msgs) != 1 {
t.Fatalf("expected 1 push, got %d", len(msgs))
}
m := msgs[0]
if m.Topic != "terdut-admin" {
t.Errorf("expected the on-call user's topic, got %q", m.Topic)
}
if m.Priority != 5 {
t.Errorf("expected max priority for a critical incident, got %d", m.Priority)
}
if !strings.Contains(m.Title, "DiskFull") {
t.Errorf("expected the incident title in %q", m.Title)
}
if !strings.Contains(m.Message, "severity critical") {
t.Errorf("expected the severity in %q", m.Message)
}
if m.Click != "https://terdut.example.com/api/incidents/1" {
t.Errorf("unexpected click target %q", m.Click)
}
if len(m.Actions) != 1 || m.Actions[0].Label != "Acknowledge" {
t.Fatalf("expected an Acknowledge action, got %+v", m.Actions)
}
if m.Actions[0].Method != http.MethodPost {
t.Errorf("expected the action to POST, got %q", m.Actions[0].Method)
}
}
// A delivered row must not be delivered again on the next pass.
func TestNotify_DeliveredOnlyOnce(t *testing.T) {
s, f := notifyTS(t, api.NotifyConfig{PublicURL: "https://terdut.example.com"})
fireCritical(t, s)
s.sweepNotify(t)
s.sweepNotify(t)
if got := len(f.messages()); got != 1 {
t.Errorf("expected 1 push across two passes, got %d", got)
}
}
// ---------------------------------------------------------------------------
// Delivery on the timeline
// ---------------------------------------------------------------------------
// notifyEvents picks the notifier's entries out of an incident's timeline.
// Asserted through the API rather than the table: the timeline is what the
// clients read, so its shape is the contract worth covering.
func notifyEvents(t *testing.T, s *ts, id int) []map[string]any {
t.Helper()
var out []map[string]any
for _, e := range timeline(t, s, id) {
if e["type"] == "notified" || e["type"] == "notify_failed" {
out = append(out, e)
}
}
return out
}
func TestNotify_DeliveryIsRecordedOnTheTimeline(t *testing.T) {
s, _ := notifyTS(t, api.NotifyConfig{PublicURL: "https://terdut.example.com"})
fireCritical(t, s)
// Queued is not notified: nothing is on the timeline until ntfy accepts it.
if got := notifyEvents(t, s, 1); len(got) != 0 {
t.Fatalf("expected no event before delivery, got %v", got)
}
s.sweepNotify(t)
events := notifyEvents(t, s, 1)
if len(events) != 1 {
t.Fatalf("expected 1 notification event, got %v", events)
}
e := events[0]
if e["type"] != "notified" {
t.Errorf("expected a notified event, got %v", e["type"])
}
if e["detail"] != "triggered" {
t.Errorf("expected the kind in detail, got %v", e["detail"])
}
if e["username"] != "admin" {
t.Errorf("expected the paged user attached, got %v", e["username"])
}
// The topic is a shared secret with ntfy; the timeline is not the place for it.
for _, v := range e {
if s, ok := v.(string); ok && strings.Contains(s, "terdut-admin") {
t.Errorf("expected the topic kept out of the timeline, found it in %v", e)
}
}
}
// A redelivery-free pass must not double-log either.
func TestNotify_TimelineRecordsOneEventPerDelivery(t *testing.T) {
s, _ := notifyTS(t, api.NotifyConfig{
PublicURL: "https://terdut.example.com",
RepeatEvery: 15 * time.Minute,
})
fireCritical(t, s)
s.sweepNotify(t)
s.sweepNotify(t)
s.ageNotifications(t, 20*time.Minute)
s.sweepNotify(t)
events := notifyEvents(t, s, 1)
if len(events) != 2 {
t.Fatalf("expected one event per delivery, got %v", events)
}
if events[0]["detail"] != "triggered" || events[1]["detail"] != "reminder" {
t.Errorf("expected triggered then reminder, got %v and %v",
events[0]["detail"], events[1]["detail"])
}
}
func TestNotify_AllClearIsRecordedOnTheTimeline(t *testing.T) {
s, _ := notifyTS(t, api.NotifyConfig{PublicURL: "https://terdut.example.com"})
fireCritical(t, s)
s.sweepNotify(t)
postWebhook(t, s, []map[string]any{
amAlert("fp-notify", "DiskFull", "resolved", "2026-05-20T10:00:00Z",
"2026-05-20T11:00:00Z", map[string]string{"severity": "critical"}),
}, "{}:{alertname=\"DiskFull\"}")
s.sweepNotify(t)
events := notifyEvents(t, s, 1)
if len(events) != 2 {
t.Fatalf("expected the all-clear recorded, got %v", events)
}
if events[1]["detail"] != "resolved" {
t.Errorf("expected a resolved event, got %v", events[1]["detail"])
}
}
// A page to the shared fallback belongs to nobody, and the timeline has to say
// so rather than attributing it to whoever happens to be on call now.
func TestNotify_FallbackDeliveryHasNoUser(t *testing.T) {
f := newFakeNtfy(t)
s := newTS(t, api.NotifyConfig{
BaseURL: f.URL,
FallbackTopic: "terdut-oncall",
PublicURL: "https://terdut.example.com",
})
fireCritical(t, s)
s.sweepNotify(t)
events := notifyEvents(t, s, 1)
if len(events) != 1 {
t.Fatalf("expected 1 notification event, got %v", events)
}
if got, ok := events[0]["username"]; ok && got != nil && got != "" {
t.Errorf("expected no user on a fallback-topic page, got %v", got)
}
}
// The failure worth seeing: nobody was paged, and the timeline says so instead
// of looking exactly like a delivery that worked.
func TestNotify_ExhaustedRetriesAreRecordedOnce(t *testing.T) {
s, f := notifyTS(t, api.NotifyConfig{PublicURL: "https://terdut.example.com"})
f.failWith(http.StatusInternalServerError)
fireCritical(t, s)
// One pass per attempt, each made due by clearing the backoff the last one set.
for i := 0; i < 10; i++ {
s.sweepNotify(t)
s.exec(t, "UPDATE notifications SET send_after = ? WHERE sent_at IS NULL",
time.Now().Add(-time.Second).Unix())
}
events := notifyEvents(t, s, 1)
if len(events) != 1 {
t.Fatalf("expected exactly one failure event, got %v", events)
}
if events[0]["type"] != "notify_failed" {
t.Errorf("expected notify_failed, got %v", events[0]["type"])
}
detail, _ := events[0]["detail"].(string)
if !strings.HasPrefix(detail, "triggered: ") || !strings.Contains(detail, "500") {
t.Errorf("expected the kind and the reason in %q", detail)
}
}
// ---------------------------------------------------------------------------
// Acknowledging from the notification
// ---------------------------------------------------------------------------
func TestNotify_AckButtonAcknowledgesIncident(t *testing.T) {
s, f := notifyTS(t, api.NotifyConfig{PublicURL: "https://terdut.example.com"})
fireCritical(t, s)
s.sweepNotify(t)
ackURL := f.messages()[0].Actions[0].URL
// The action URL is built for the public hostname; point it at the test
// server, which is the same handler.
path := ackURL[strings.Index(ackURL, "/api/notify/ack/"):]
resp, err := http.Post(s.URL+path, "application/json", nil)
if err != nil {
t.Fatalf("ack: %v", err)
}
defer resp.Body.Close()
if resp.StatusCode != http.StatusOK {
t.Fatalf("expected 200 from the ack button, got %d", resp.StatusCode)
}
inc := getIncident(t, s, 1)
if inc["status"] != "acknowledged" {
t.Errorf("expected the incident acknowledged, got %v", inc["status"])
}
if inc["acknowledged_by"] != "admin" {
t.Errorf("expected the ack attributed to the token's user, got %v", inc["acknowledged_by"])
}
// The timeline must record it the same way the authenticated route would.
events := timeline(t, s, 1)
if !contains(eventTypes(events), "acknowledged") {
t.Errorf("expected an acknowledged event, got %v", eventTypes(events))
}
for _, e := range events {
if e["type"] == "acknowledged" && e["username"] != "admin" {
t.Errorf("expected the acknowledged event attributed to admin, got %v", e["username"])
}
}
}
func TestNotify_AckRejectsUnknownToken(t *testing.T) {
s, _ := notifyTS(t, api.NotifyConfig{PublicURL: "https://terdut.example.com"})
fireCritical(t, s)
resp, err := http.Post(s.URL+"/api/notify/ack/deadbeef", "application/json", nil)
if err != nil {
t.Fatalf("ack: %v", err)
}
defer resp.Body.Close()
if resp.StatusCode != http.StatusNotFound {
t.Errorf("expected 404 for an unknown token, got %d", resp.StatusCode)
}
if inc := getIncident(t, s, 1); inc["status"] != "triggered" {
t.Errorf("expected the incident untouched, got %v", inc["status"])
}
}
func TestNotify_AckRejectsExpiredToken(t *testing.T) {
s, f := notifyTS(t, api.NotifyConfig{PublicURL: "https://terdut.example.com"})
fireCritical(t, s)
s.sweepNotify(t)
ackURL := f.messages()[0].Actions[0].URL
path := ackURL[strings.Index(ackURL, "/api/notify/ack/"):]
// Age the token past its TTL. The token's inputs are wall-clock timestamps,
// so this is the same trick the sweeper tests use.
s.exec(t, "UPDATE incident_ack_tokens SET expires_at = ?", time.Now().Add(-time.Minute).Unix())
resp, err := http.Post(s.URL+path, "application/json", nil)
if err != nil {
t.Fatalf("ack: %v", err)
}
defer resp.Body.Close()
if resp.StatusCode != http.StatusNotFound {
t.Errorf("expected 404 for an expired token, got %d", resp.StatusCode)
}
if inc := getIncident(t, s, 1); inc["status"] != "triggered" {
t.Errorf("expected the incident untouched, got %v", inc["status"])
}
}
// The sweeper is what stops expired tokens accumulating forever.
func TestNotify_SweepPurgesExpiredAckTokens(t *testing.T) {
s, _ := notifyTS(t, api.NotifyConfig{PublicURL: "https://terdut.example.com"})
fireCritical(t, s)
s.sweepNotify(t)
s.exec(t, "UPDATE incident_ack_tokens SET expires_at = ?", time.Now().Add(-time.Minute).Unix())
api.Sweep(context.Background(), s.db, 168*time.Hour, 6*time.Hour, s.deadman, s.notify)
var n int
if err := s.db.QueryRow("SELECT COUNT(*) FROM incident_ack_tokens").Scan(&n); err != nil {
t.Fatalf("count tokens: %v", err)
}
if n != 0 {
t.Errorf("expected expired tokens purged, %d left", n)
}
}
// ---------------------------------------------------------------------------
// Reminders
// ---------------------------------------------------------------------------
// ageNotifications backdates every sent notification so the next pass sees the
// reminder as due.
func (s *ts) ageNotifications(t *testing.T, by time.Duration) {
t.Helper()
s.exec(t, "UPDATE notifications SET created_at = ? WHERE sent_at IS NOT NULL",
time.Now().Add(-by).Unix())
}
func TestNotify_UnacknowledgedIncidentIsRenotified(t *testing.T) {
s, f := notifyTS(t, api.NotifyConfig{
PublicURL: "https://terdut.example.com",
RepeatEvery: 15 * time.Minute,
})
fireCritical(t, s)
s.sweepNotify(t)
s.ageNotifications(t, 20*time.Minute)
s.sweepNotify(t)
msgs := f.messages()
if len(msgs) != 2 {
t.Fatalf("expected a reminder push, got %d message(s)", len(msgs))
}
if !strings.Contains(msgs[1].Title, "Still unacknowledged") {
t.Errorf("expected the reminder to say so, got %q", msgs[1].Title)
}
if msgs[1].Topic != "terdut-admin" {
t.Errorf("expected the reminder on the same topic, got %q", msgs[1].Topic)
}
// Each page carries its own credential.
if len(msgs[1].Actions) != 1 || msgs[1].Actions[0].URL == msgs[0].Actions[0].URL {
t.Errorf("expected the reminder to carry a fresh ack token")
}
}
func TestNotify_AcknowledgedIncidentStopsReminders(t *testing.T) {
s, f := notifyTS(t, api.NotifyConfig{
PublicURL: "https://terdut.example.com",
RepeatEvery: 15 * time.Minute,
})
fireCritical(t, s)
s.sweepNotify(t)
resp := s.req(t, http.MethodPost, "/api/incidents/1/acknowledge", nil)
resp.Body.Close()
s.ageNotifications(t, 20*time.Minute)
s.sweepNotify(t)
if got := len(f.messages()); got != 1 {
t.Errorf("expected no reminder once acknowledged, got %d message(s)", got)
}
}
// Snooze is the deliberate "not now", and it is what mutes the pager.
func TestNotify_SnoozedIncidentStopsReminders(t *testing.T) {
s, f := notifyTS(t, api.NotifyConfig{
PublicURL: "https://terdut.example.com",
RepeatEvery: 15 * time.Minute,
})
fireCritical(t, s)
s.sweepNotify(t)
resp := s.req(t, http.MethodPost, "/api/incidents/1/snooze", map[string]any{"duration": "1h"})
if resp.StatusCode != http.StatusOK {
t.Fatalf("snooze returned %d", resp.StatusCode)
}
resp.Body.Close()
s.ageNotifications(t, 20*time.Minute)
s.sweepNotify(t)
if got := len(f.messages()); got != 1 {
t.Errorf("expected no reminder while snoozed, got %d message(s)", got)
}
}
func TestNotify_ZeroRepeatDisablesReminders(t *testing.T) {
s, f := notifyTS(t, api.NotifyConfig{PublicURL: "https://terdut.example.com"})
fireCritical(t, s)
s.sweepNotify(t)
s.ageNotifications(t, 24*time.Hour)
s.sweepNotify(t)
if got := len(f.messages()); got != 1 {
t.Errorf("expected reminders off, got %d message(s)", got)
}
}
// ---------------------------------------------------------------------------
// Resolution
// ---------------------------------------------------------------------------
func TestNotify_ResolvedIncidentSendsAllClear(t *testing.T) {
s, f := notifyTS(t, api.NotifyConfig{PublicURL: "https://terdut.example.com"})
fireCritical(t, s)
s.sweepNotify(t)
postWebhook(t, s, []map[string]any{
amAlert("fp-notify", "DiskFull", "resolved", "2026-05-20T10:00:00Z",
"2026-05-20T11:00:00Z", map[string]string{"severity": "critical"}),
}, "{}:{alertname=\"DiskFull\"}")
s.sweepNotify(t)
msgs := f.messages()
if len(msgs) != 2 {
t.Fatalf("expected an all-clear push, got %d message(s)", len(msgs))
}
if !strings.HasPrefix(msgs[1].Title, "Resolved:") {
t.Errorf("expected a resolved title, got %q", msgs[1].Title)
}
if msgs[1].Priority != 2 {
t.Errorf("expected the all-clear at low priority, got %d", msgs[1].Priority)
}
// Nothing to acknowledge on a closed incident.
if len(msgs[1].Actions) != 0 {
t.Errorf("expected no actions on the all-clear, got %+v", msgs[1].Actions)
}
}
// Closing an incident by hand sends nothing: the person who did it knows.
func TestNotify_ManualResolveSendsNothing(t *testing.T) {
s, f := notifyTS(t, api.NotifyConfig{PublicURL: "https://terdut.example.com"})
fireCritical(t, s)
s.sweepNotify(t)
resp := s.req(t, http.MethodPost, "/api/incidents/1/resolve", nil)
resp.Body.Close()
s.sweepNotify(t)
if got := len(f.messages()); got != 1 {
t.Errorf("expected no push for a manual resolve, got %d message(s)", got)
}
}
// ---------------------------------------------------------------------------
// Routing and configuration
// ---------------------------------------------------------------------------
// With nobody on call the page goes to the shared fallback, and carries no
// Acknowledge button — there is no user to attribute the acknowledgement to.
func TestNotify_FallbackTopicHasNoAckButton(t *testing.T) {
f := newFakeNtfy(t)
s := newTS(t, api.NotifyConfig{
BaseURL: f.URL,
FallbackTopic: "terdut-oncall",
PublicURL: "https://terdut.example.com",
})
fireCritical(t, s)
s.sweepNotify(t)
msgs := f.messages()
if len(msgs) != 1 {
t.Fatalf("expected 1 push, got %d", len(msgs))
}
if msgs[0].Topic != "terdut-oncall" {
t.Errorf("expected the fallback topic, got %q", msgs[0].Topic)
}
if len(msgs[0].Actions) != 0 {
t.Errorf("expected no ack button on a shared topic, got %+v", msgs[0].Actions)
}
}
// Nobody on call and no fallback means there is nowhere to send: queueing would
// only pile up rows that can never be delivered.
func TestNotify_NoTargetQueuesNothing(t *testing.T) {
f := newFakeNtfy(t)
s := newTS(t, api.NotifyConfig{BaseURL: f.URL})
fireCritical(t, s)
if got := s.countNotifications(t, ""); got != 0 {
t.Errorf("expected nothing queued without a target, got %d", got)
}
s.sweepNotify(t)
if got := len(f.messages()); got != 0 {
t.Errorf("expected no push, got %d", got)
}
}
// The zero NotifyConfig is what every pre-existing test runs under.
func TestNotify_DisabledQueuesNothing(t *testing.T) {
s := newTS(t)
putOnCall(t, s, 1)
setTopic(t, s, 1, "terdut-admin")
fireCritical(t, s)
if got := s.countNotifications(t, ""); got != 0 {
t.Errorf("expected nothing queued with notifications off, got %d", got)
}
}
// ---------------------------------------------------------------------------
// Retries
// ---------------------------------------------------------------------------
func TestNotify_FailedDeliveryRetriesWithBackoff(t *testing.T) {
s, f := notifyTS(t, api.NotifyConfig{PublicURL: "https://terdut.example.com"})
f.failWith(http.StatusInternalServerError)
fireCritical(t, s)
s.sweepNotify(t)
var attempts int
var sentAt *int64
var sendAfter int64
var lastError *string
if err := s.db.QueryRow(
"SELECT attempts, sent_at, send_after, last_error FROM notifications WHERE id = 1").
Scan(&attempts, &sentAt, &sendAfter, &lastError); err != nil {
t.Fatalf("read notification: %v", err)
}
if attempts != 1 {
t.Errorf("expected 1 attempt recorded, got %d", attempts)
}
if sentAt != nil {
t.Errorf("expected the row unsent, got sent_at %v", *sentAt)
}
if sendAfter <= time.Now().Unix() {
t.Errorf("expected the retry pushed into the future, got %d", sendAfter)
}
if lastError == nil || !strings.Contains(*lastError, "500") {
t.Errorf("expected the failure recorded, got %v", lastError)
}
// Backing off means the next pass leaves it alone until it is due.
s.sweepNotify(t)
if got := len(f.messages()); got != 1 {
t.Errorf("expected no immediate retry, got %d attempt(s)", got)
}
// Once due and once ntfy recovers, it goes out.
f.failWith(http.StatusOK)
s.exec(t, "UPDATE notifications SET send_after = ? WHERE id = 1", time.Now().Add(-time.Second).Unix())
s.sweepNotify(t)
if err := s.db.QueryRow("SELECT sent_at FROM notifications WHERE id = 1").Scan(&sentAt); err != nil {
t.Fatalf("read notification: %v", err)
}
if sentAt == nil {
t.Error("expected the retry to succeed once ntfy recovered")
}
}
// An ntfy outage must not produce a reminder backlog that all lands at once
// when it comes back: the previous page has to have been sent first.
func TestNotify_UnsentNotificationBlocksReminders(t *testing.T) {
s, f := notifyTS(t, api.NotifyConfig{
PublicURL: "https://terdut.example.com",
RepeatEvery: 15 * time.Minute,
})
f.failWith(http.StatusInternalServerError)
fireCritical(t, s)
s.sweepNotify(t)
s.exec(t, "UPDATE notifications SET created_at = ?", time.Now().Add(-time.Hour).Unix())
s.sweepNotify(t)
if got := s.countNotifications(t, "reminder"); got != 0 {
t.Errorf("expected no reminders queued behind an undelivered page, got %d", got)
}
}
+12 -3
View File
@@ -8,7 +8,11 @@ import (
"github.com/go-chi/chi/v5/middleware"
)
func NewRouter(db *sql.DB) http.Handler {
// NewRouter builds the HTTP surface. notify and deadman are passed through to
// the webhook, the only handler that has to decide where a new incident's page
// goes and which arriving alerts are heartbeats rather than problems. A zero
// notify disables notifications; a zero deadman disables dead man's switches.
func NewRouter(db *sql.DB, notify NotifyConfig, deadman DeadmanConfig) http.Handler {
r := chi.NewRouter()
r.Use(middleware.Logger)
r.Use(middleware.Recoverer)
@@ -17,9 +21,13 @@ func NewRouter(db *sql.DB) http.Handler {
respond(w, http.StatusOK, map[string]string{"status": "ok"})
})
// Unauthenticated: bootstrap and Alertmanager webhook receiver.
// Unauthenticated: bootstrap, the Alertmanager webhook receiver, and the
// Acknowledge button in a push notification. The last one is authorised by
// the scoped token in its path rather than an API key, and has to stay
// reachable from outside the cluster for the button to work.
r.Post("/api/bootstrap", handleBootstrap(db))
r.Post("/api/alertmanager/webhook", handleAlertmanagerWebhook(db))
r.Post("/api/alertmanager/webhook", handleAlertmanagerWebhook(db, notify, deadman))
r.Post("/api/notify/ack/{token}", handleNotifyAck(db))
// All other /api routes require a valid API key.
r.Group(func(r chi.Router) {
@@ -28,6 +36,7 @@ func NewRouter(db *sql.DB) http.Handler {
r.Get("/api/users", handleListUsers(db))
r.Post("/api/users", handleCreateUser(db))
r.Delete("/api/users/{id}", handleDeleteUser(db))
r.Put("/api/users/{id}/notify", handleSetNotifyTarget(db))
r.Post("/api/users/{id}/api-keys", handleCreateAPIKey(db))
r.Delete("/api/users/{id}/api-keys/{keyID}", handleDeleteAPIKey(db))
+20 -3
View File
@@ -9,7 +9,7 @@ import (
"time"
"github.com/go-chi/chi/v5"
"github.com/yeniklas/terdut-server/internal/models"
"git.ryuvia.com/niklas/terdut-server/internal/models"
)
func handleCreateSchedule(db *sql.DB) http.HandlerFunc {
@@ -17,6 +17,12 @@ func handleCreateSchedule(db *sql.DB) http.HandlerFunc {
var req struct {
UserID int64 `json:"user_id"`
Dates []string `json:"dates"`
// Replace takes dates that somebody else already holds. It defaults
// to off so that the plain call cannot quietly move a shift off the
// person expecting to be paged for it — reassigning has to be asked
// for.
Replace bool `json:"replace"`
}
if err := decodeJSON(r, &req); err != nil {
respond(w, http.StatusBadRequest, errResp("invalid request body"))
@@ -44,7 +50,10 @@ func handleCreateSchedule(db *sql.DB) http.HandlerFunc {
return
}
// All-or-nothing: if any date already has an assignment, reject the whole request.
// All-or-nothing, in both directions: without replace, one taken date
// rejects the whole request; with it, either every date moves or none
// does. The rota must never be left with a hole where a shift used to
// be, so the delete and the insert share one transaction.
tx, err := db.BeginTx(r.Context(), nil)
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
@@ -53,10 +62,18 @@ func handleCreateSchedule(db *sql.DB) http.HandlerFunc {
defer tx.Rollback()
for _, d := range req.Dates {
if req.Replace {
if _, err := tx.ExecContext(r.Context(),
"DELETE FROM schedule_entries WHERE date = ?", d); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
}
if _, err := tx.ExecContext(r.Context(),
"INSERT INTO schedule_entries (user_id, date) VALUES (?, ?)", req.UserID, d); err != nil {
if strings.Contains(err.Error(), "UNIQUE constraint failed") {
respond(w, http.StatusConflict, errResp("date already assigned: "+d))
respond(w, http.StatusConflict,
errResp("date already assigned: "+d+" (pass replace to take it)"))
return
}
respond(w, http.StatusInternalServerError, errResp("internal error"))
+55 -9
View File
@@ -12,7 +12,7 @@ import (
"time"
"github.com/go-chi/chi/v5"
"github.com/yeniklas/terdut-server/internal/models"
"git.ryuvia.com/niklas/terdut-server/internal/models"
)
func handleBootstrap(db *sql.DB) http.HandlerFunc {
@@ -48,7 +48,7 @@ func handleBootstrap(db *sql.DB) http.HandlerFunc {
}
userID, _ := res.LastInsertId()
raw, hash, err := newAPIKey()
raw, hash, err := randomToken()
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
@@ -70,7 +70,7 @@ func handleBootstrap(db *sql.DB) http.HandlerFunc {
func handleListUsers(db *sql.DB) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
rows, err := db.QueryContext(r.Context(),
"SELECT id, username, email, created_at FROM users ORDER BY id")
"SELECT id, username, email, created_at, ntfy_topic FROM users ORDER BY id")
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
@@ -81,7 +81,7 @@ func handleListUsers(db *sql.DB) http.HandlerFunc {
for rows.Next() {
var u models.User
var ts int64
if err := rows.Scan(&u.ID, &u.Username, &u.Email, &ts); err != nil {
if err := rows.Scan(&u.ID, &u.Username, &u.Email, &ts, &u.NtfyTopic); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
@@ -123,6 +123,50 @@ func handleCreateUser(db *sql.DB) http.HandlerFunc {
}
}
// handleSetNotifyTarget points a user's push notifications at an ntfy topic, or
// clears it with an empty string. The topic is a shared secret with the ntfy
// server — anyone who knows it can publish to it — so pick an unguessable one
// unless your ntfy enforces access control.
func handleSetNotifyTarget(db *sql.DB) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
id, err := strconv.ParseInt(chi.URLParam(r, "id"), 10, 64)
if err != nil {
respond(w, http.StatusBadRequest, errResp("invalid user id"))
return
}
var req struct {
NtfyTopic string `json:"ntfy_topic"`
}
if err := decodeJSON(r, &req); err != nil {
respond(w, http.StatusBadRequest, errResp("invalid request body"))
return
}
var topic *string
if t := strings.TrimSpace(req.NtfyTopic); t != "" {
topic = &t
}
res, err := db.ExecContext(r.Context(),
"UPDATE users SET ntfy_topic = ? WHERE id = ?", topic, id)
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
if n, _ := res.RowsAffected(); n == 0 {
respond(w, http.StatusNotFound, errResp("user not found"))
return
}
user, err := fetchUser(r.Context(), db, id)
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
respond(w, http.StatusOK, user)
}
}
func handleDeleteUser(db *sql.DB) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
id, err := strconv.ParseInt(chi.URLParam(r, "id"), 10, 64)
@@ -170,7 +214,7 @@ func handleCreateAPIKey(db *sql.DB) http.HandlerFunc {
return
}
raw, hash, err := newAPIKey()
raw, hash, err := randomToken()
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
@@ -215,8 +259,9 @@ func handleDeleteAPIKey(db *sql.DB) http.HandlerFunc {
}
}
// newAPIKey generates a random 32-byte key encoded as hex, plus its SHA-256 hash for storage.
func newAPIKey() (raw, hash string, err error) {
// randomToken generates a random 32-byte secret encoded as hex, plus its SHA-256
// hash for storage. Used for API keys and for notification acknowledge tokens.
func randomToken() (raw, hash string, err error) {
b := make([]byte, 32)
if _, err = rand.Read(b); err != nil {
return
@@ -230,8 +275,9 @@ func newAPIKey() (raw, hash string, err error) {
func fetchUser(ctx context.Context, db *sql.DB, id int64) (models.User, error) {
var u models.User
var ts int64
err := db.QueryRowContext(ctx, "SELECT id, username, email, created_at FROM users WHERE id = ?", id).
Scan(&u.ID, &u.Username, &u.Email, &ts)
err := db.QueryRowContext(ctx,
"SELECT id, username, email, created_at, ntfy_topic FROM users WHERE id = ?", id).
Scan(&u.ID, &u.Username, &u.Email, &ts, &u.NtfyTopic)
if err != nil {
return u, err
}
+72 -11
View File
@@ -14,6 +14,44 @@ type Config struct {
// before the sweeper treats it as resolved. It must exceed Alertmanager's
// repeat_interval (default 4h), which is what refreshes the alert.
StaleAfter time.Duration
// DeadmanMatchers selects the alerts that are heartbeats rather than
// problems: receiving one opens no incident, and the absence of one does.
//
// ";" separates matchers, "," the label conditions within one, "=" is exact
// equality — `alertname=Watchdog,cluster=prod; alertname=Heartbeat`. Every
// matcher must name an alertname. See api.ParseDeadmanConfig.
DeadmanMatchers string
// DeadmanTimeout is how long a heartbeat may go unheard before its switch is
// declared dead. It must be *shorter* than the Alertmanager repeat_interval
// of the route carrying the heartbeat — the opposite of StaleAfter, and the
// reason a dead man's switch usually wants a route of its own. Zero disables
// dead man's switch handling entirely.
DeadmanTimeout time.Duration
// DeadmanSeverity is the severity a dead man's switch incident opens at.
// These incidents have no member alerts to derive one from.
DeadmanSeverity string
// NtfyURL is the ntfy server push notifications are published to. Empty
// disables notifications entirely.
NtfyURL string
// NtfyToken is an optional bearer token for an access-controlled ntfy.
NtfyToken string
// NtfyFallbackTopic receives incidents that open with nobody on call.
NtfyFallbackTopic string
// PublicURL is the base URL a phone uses to reach this server, used for the
// link and the Acknowledge button inside a notification. Without it
// notifications carry neither.
PublicURL string
// NotifyRepeat is how long an incident may sit unacknowledged before it is
// notified again. Zero disables reminders.
NotifyRepeat time.Duration
}
func Load() Config {
@@ -25,17 +63,40 @@ func Load() Config {
if dbPath == "" {
dbPath = "terdut.db"
}
archiveAfter := 7 * 24 * time.Hour
if s := os.Getenv("TERDUT_ARCHIVE_AFTER"); s != "" {
if d, err := time.ParseDuration(s); err == nil {
archiveAfter = d
}
deadmanMatchers := os.Getenv("TERDUT_DEADMAN_MATCHERS")
if deadmanMatchers == "" {
deadmanMatchers = "alertname=Watchdog"
}
staleAfter := 6 * time.Hour
if s := os.Getenv("TERDUT_STALE_AFTER"); s != "" {
if d, err := time.ParseDuration(s); err == nil {
staleAfter = d
}
deadmanSeverity := os.Getenv("TERDUT_DEADMAN_SEVERITY")
if deadmanSeverity == "" {
deadmanSeverity = "critical"
}
return Config{
Addr: addr,
DBPath: dbPath,
ArchiveAfter: duration("TERDUT_ARCHIVE_AFTER", 7*24*time.Hour),
StaleAfter: duration("TERDUT_STALE_AFTER", 6*time.Hour),
DeadmanMatchers: deadmanMatchers,
DeadmanTimeout: duration("TERDUT_DEADMAN_TIMEOUT", 15*time.Minute),
DeadmanSeverity: deadmanSeverity,
NtfyURL: os.Getenv("TERDUT_NTFY_URL"),
NtfyToken: os.Getenv("TERDUT_NTFY_TOKEN"),
NtfyFallbackTopic: os.Getenv("TERDUT_NTFY_FALLBACK_TOPIC"),
PublicURL: os.Getenv("TERDUT_PUBLIC_URL"),
NotifyRepeat: duration("TERDUT_NOTIFY_REPEAT", 15*time.Minute),
}
return Config{Addr: addr, DBPath: dbPath, ArchiveAfter: archiveAfter, StaleAfter: staleAfter}
}
// duration reads a time.ParseDuration-formatted env var. An unset or
// unparseable value falls back to def rather than failing startup: a typo in one
// tuning knob should not take the server down.
func duration(env string, def time.Duration) time.Duration {
if s := os.Getenv(env); s != "" {
if d, err := time.ParseDuration(s); err == nil {
return d
}
}
return def
}
@@ -0,0 +1,44 @@
-- Adds push notification delivery, so an incident reaches the person on call
-- instead of waiting to be discovered.
--
-- Delivery is an outbox rather than an inline HTTP call: the pool is limited to
-- a single connection, so a POST made while holding the webhook's transaction
-- would stall every other request behind it. The webhook inserts a row; the
-- notifier goroutine delivers it.
CREATE TABLE notifications (
id INTEGER PRIMARY KEY AUTOINCREMENT,
incident_id INTEGER NOT NULL REFERENCES incidents(id) ON DELETE CASCADE,
-- Nullable: a notification sent to the fallback topic belongs to nobody,
-- because nobody was on call when the incident opened.
user_id INTEGER REFERENCES users(id) ON DELETE SET NULL,
topic TEXT NOT NULL, -- resolved at enqueue: who was on call then
kind TEXT NOT NULL CHECK(kind IN ('triggered', 'reminder', 'resolved')),
created_at INTEGER NOT NULL,
send_after INTEGER NOT NULL, -- retry backoff watermark
attempts INTEGER NOT NULL DEFAULT 0,
sent_at INTEGER,
last_error TEXT -- kept after the last attempt, for debugging
);
-- The delivery loop's only query: what is due and still unsent.
CREATE INDEX notifications_pending_idx ON notifications(send_after) WHERE sent_at IS NULL;
-- Reminders and resolved notices both look up an incident's newest row.
CREATE INDEX notifications_incident_idx ON notifications(incident_id, id DESC);
-- A notification body is stored on the ntfy server and cached on the device, so
-- a real API key must never appear in one. Each delivery mints its own token
-- instead: one incident, one action, one day.
CREATE TABLE incident_ack_tokens (
token_hash TEXT PRIMARY KEY, -- SHA-256 of the raw token, as with api_keys
incident_id INTEGER NOT NULL REFERENCES incidents(id) ON DELETE CASCADE,
user_id INTEGER NOT NULL REFERENCES users(id) ON DELETE CASCADE,
created_at INTEGER NOT NULL,
expires_at INTEGER NOT NULL
);
CREATE INDEX incident_ack_tokens_expires_idx ON incident_ack_tokens(expires_at);
-- Where this user's notifications go. NULL means they get none; incidents
-- assigned to them fall back to the configured fallback topic.
ALTER TABLE users ADD COLUMN ntfy_topic TEXT;
+5
View File
@@ -7,6 +7,11 @@ type User struct {
Username string `json:"username"`
Email string `json:"email"`
CreatedAt time.Time `json:"created_at"`
// NtfyTopic is where this user's push notifications go. Nil means they get
// none of their own; incidents assigned to them fall back to the configured
// fallback topic instead.
NtfyTopic *string `json:"ntfy_topic,omitempty"`
}
type APIKey struct {