Advisory-lock DB migrations against concurrent replica startup

Migrate's check-then-apply loop against schema_migrations had no
locking: two replicas booting at once against a fresh or
partially-migrated database could both pass the "not yet applied"
check for the same file and race applying it, crashing whichever lost
the duplicate-key insert (confirmed: reverting the lock fails the new
test 10/10 on a duplicate-key violation, racing as early as the
CREATE TABLE IF NOT EXISTS schema_migrations statement itself).

Hold a Postgres advisory lock for Migrate's whole run, on a dedicated
connection reserved via db.Conn so lock and unlock happen on the same
session. Blocking (pg_advisory_lock), unlike the archiver/notifier's
pg_try_advisory_lock: on boot there's no later tick to defer to, so a
second replica should wait for the first to finish migrating rather
than skip ahead.

Adds internal/db's first test file, exercising two concurrent Migrate
calls against a fresh schema.

Still open: the new-incident-insert race on a webhook for a brand-new
groupKey, noted in the chart's updated comment. Login rate limiting
staying in-process, diluted across replicas, is an accepted tradeoff.

Co-authored-by: Claude <noreply@anthropic.com>
This commit is contained in:
Niklas Ye
2026-10-03 11:43:09 +02:00
parent 42180948d1
commit 0050738ca0
3 changed files with 156 additions and 5 deletions
@@ -11,11 +11,10 @@ spec:
matchLabels:
{{- include "terdut-server.selectorLabels" . | nindent 6 }}
# Recreate, not RollingUpdate, even though the PVC that forced it is gone: the
# sweeper and notifier now take a Postgres advisory lock for each pass, so two
# replicas overlapping during a rollout no longer both page for the same
# incident, but DB migrations and new-incident creation on first webhook are
# still unguarded — a second replica starting concurrently with the first can
# still race either of those.
# sweeper, notifier and migration runner now take a Postgres advisory lock
# each, so two replicas overlapping during a rollout no longer both page for
# the same incident or race applying a migration, but new-incident creation
# on the first webhook for a brand-new groupKey is still unguarded.
strategy:
type: Recreate
template: