Move the database to Postgres, before teams need the schema
CI / chart (pull_request) Successful in 1s
CI / security (pull_request) Successful in 17s
CI / test (pull_request) Successful in 2m5s

First step of #1, and it goes first for one reason: #4 adds a team_id to
nearly every table, and doing that twice -- once for SQLite, once for
Postgres -- is work nobody gets paid for. The teams migrations now only
have to be written against one database.

The ten SQLite migrations are replaced by a single Postgres baseline
rather than ported one by one. They were incremental in a way that has
no value on a fresh install: 004 adds columns 008 drops again, and 008's
backfill rewrites data a Postgres database never had. The history stays
in git; the schema they add up to is now 001_baseline.sql.

Timestamps stay BIGINT unix seconds and are NOT converted to timestamptz.
Everything in Go already speaks epochs, so converting would have been a
second, larger change riding along inside this one. It is worth doing on
its own. The JSON columns did move to jsonb, because #4 will want to
filter and index on labels.

Most of the port is mechanical -- 170 placeholders from ? to $1 -- but
four things needed more than a search and replace:

  * Dynamically built WHERE clauses cannot keep their numbering straight
    by hand, so they hand out placeholders through sqlArgs instead. A
    filter can now be added or reordered without renumbering anything.

  * SUM(resolved_at IS NULL) was SQLite counting a boolean as 0 or 1.
    Postgres has no sum(boolean), and this was breaking every dead man's
    switch -- silently, since the sweeper only logs. Now COUNT(*) FILTER.

  * unixepoch() became FLOOR(EXTRACT(EPOCH FROM now()))::bigint. The
    FLOOR is load-bearing: a bare cast rounds half up, so a row written
    at .6 of a second claimed a timestamp a second in the future and
    disagreed with the time.Now().Unix() the Go side stamps.

  * The unique-violation check matched SQLite's error text. It matches
    SQLSTATE 23505 now, so a renamed constraint cannot turn a 409 back
    into a 500.

Tests need a real Postgres, because there is no in-memory Postgres the
way there was an in-memory SQLite. Each test gets its own schema on a
shared server -- cheaper than a database each, and still isolated.
TERDUT_TEST_DSN says where it is; `make test-db` starts one locally and
ci.yaml runs one as a service container. An unset DSN fails the suite
rather than skipping it: a run that quietly tests nothing is worse than
one that does not run.

TestMigration_BackfillCarriesAckAndComments is deleted along with the
migrations it replayed. What it protected -- an upgrade not losing
acknowledgements and comments -- now belongs to scripts/sqlite-to-postgres.go,
which is build-tagged so the SQLite driver stays out of the server
binary. Both are meant to be deleted once this install has migrated.

The chart loses the PVC, the data volume and the python backup sidecar,
and requires database.dsnSecret.name: it provisions no database and
cannot guess where the credentials live, so a render without it is meant
to fail. Backups move to where Postgres actually runs. The other half of
that -- the postgresql CR, the k8up pg_dump annotation and the network
policy -- is a change to the wrapper chart in Ryuvia/charts and is not in
here.

Verified rather than assumed: the gate is green with -race against
Postgres 17, govulncheck and gitleaks are clean, and the migration script
was run end to end against a SQLite database built at the old schema and
seeded in every table. Ids survive, so incidents keep their numbers and
every foreign key still points where it did; the identity sequences are
moved past the copied ids, and a webhook after the migration opened
incident 12 rather than colliding at 1.
This commit is contained in:
Niklas Ye
2026-09-20 10:44:12 +02:00
parent 989425e550
commit dc39e3a5d3
44 changed files with 1004 additions and 725 deletions
+44 -5
View File
@@ -25,12 +25,45 @@ help: ## Show this help
#
# These three mirror .gitea/workflows/ci.yaml step for step, so a green `make fmt
# lint test` here means the same thing CI means. The one deliberate difference is
# -race below.
# -race below. Both need a Postgres to test against; see test-db.
# The suite needs a Postgres, because the server does: there is no in-memory
# Postgres the way there was an in-memory SQLite. TERDUT_TEST_DSN says where, and
# the tests fail rather than skip without it — a suite that quietly tests nothing
# is worse than one that does not run. `make test-db` starts a local one;
# ci.yaml runs the same thing as a service container.
TEST_DB_CONTAINER ?= terdut-test-db
TEST_DB_PORT ?= 5433
TEST_DB_IMAGE ?= docker.io/library/postgres:17-alpine
export TERDUT_TEST_DSN ?= postgres://terdut:terdut@localhost:$(TEST_DB_PORT)/terdut_test?sslmode=disable
.PHONY: test
test: ## Run the test suite
test: ## Run the test suite (needs TERDUT_TEST_DSN; see test-db)
go test -race ./...
# podman, with docker as the fallback: this is a dev convenience, not part of the
# pipeline, where the database arrives as a service container instead.
.PHONY: test-db
test-db: ## Start a local Postgres for the tests
@runtime=$$(command -v podman || command -v docker); \
if [ -z "$$runtime" ]; then echo "need podman or docker"; exit 1; fi; \
$$runtime run -d --rm --name $(TEST_DB_CONTAINER) \
-e POSTGRES_USER=terdut -e POSTGRES_PASSWORD=terdut -e POSTGRES_DB=terdut_test \
-p $(TEST_DB_PORT):5432 $(TEST_DB_IMAGE) >/dev/null; \
printf 'waiting for postgres'; \
for i in $$(seq 1 60); do \
if $$runtime exec $(TEST_DB_CONTAINER) pg_isready -U terdut -d terdut_test >/dev/null 2>&1; then \
echo " ready: $(TERDUT_TEST_DSN)"; exit 0; \
fi; \
printf '.'; sleep 1; \
done; \
echo " timed out"; exit 1
.PHONY: test-db-stop
test-db-stop: ## Stop the local test Postgres
@runtime=$$(command -v podman || command -v docker); \
$$runtime rm -f $(TEST_DB_CONTAINER) >/dev/null 2>&1 || true
# CI runs a bare `go test ./...`. This is stricter on purpose: the sweeper, the
# notifier goroutine and the deadman sweep all touch the same single-connection
# database, and a race there would surface as a flaky production incident rather
@@ -54,16 +87,22 @@ fmt: ## Report unformatted files
echo "gofmt needed:"; echo "$$unformatted"; gofmt -d .; exit 1; \
fi
# database.dsnSecret.name has no default and the deployment `required`s it: the
# chart provisions no database and cannot guess where the credentials live, so a
# render without it is meant to fail. Setting it here keeps the lint honest about
# what a working install needs.
HELM_LINT_SET = --set image.tag=v0.0.0 --set database.dsnSecret.name=terdut-db
.PHONY: helm-lint
helm-lint: ## Lint and render the chart
helm lint $(HELM_CHART) --set image.tag=v0.0.0
helm lint $(HELM_CHART) $(HELM_LINT_SET)
helm template terdut-server $(HELM_CHART) --namespace terdut-server \
--set image.tag=v0.0.0 >/dev/null
$(HELM_LINT_SET) >/dev/null
@# networking.listener defaults to "", which attaches the route to every
@# matching listener including plaintext HTTP. Production sets it, so the
@# default render proves nothing about the path that actually ships.
helm template terdut-server $(HELM_CHART) --namespace terdut-server \
--set image.tag=v0.0.0 --set networking.listener=https-terdut >/dev/null
$(HELM_LINT_SET) --set networking.listener=https-terdut >/dev/null
## --- release ---