Guard the archiver and notifier passes with a Postgres advisory lock
Both background loops run unconditionally on every instance with no coordination between them, which the chart's replicas: 1 + strategy: Recreate exists specifically to paper over: with more than one replica, every one of them would sweep and deliver notifications independently, and two overlapping during a rollout would both page for the same incident. Add withAdvisoryLock, which takes a Postgres advisory lock on a dedicated connection and runs a pass only if it gets the lock, otherwise skipping until the next tick. Wire StartArchiver and StartNotifier through it with their own lock keys, so Sweep and NotifySweep themselves are untouched and every existing test calling them directly keeps working unchanged. This also closes the notifier's double-delivery race in passing: two replicas can no longer both be inside deliverPending at once, since only one can hold notifierLockKey at a time. Deliberately not addressed here, and still blocking a replica count above 1: the in-memory login rate limiter, the unlocked migration runner, and the new-incident-insert race on a webhook for a brand-new groupKey. Noted in the updated chart comment. Co-authored-by: Claude <noreply@anthropic.com>
This commit is contained in:
@@ -38,6 +38,13 @@ const (
|
||||
ackTokenTTL = 24 * time.Hour
|
||||
)
|
||||
|
||||
// notifierLockKey is the Postgres advisory lock the notifier takes for the
|
||||
// duration of each pass, so that running more than one replica does not
|
||||
// deliver (or double-deliver) the same notification from more than one of
|
||||
// them at once. Its value has no meaning beyond being distinct from
|
||||
// archiverLockKey.
|
||||
const notifierLockKey int64 = 7265_0002
|
||||
|
||||
// Notification kinds, recording why a push was sent.
|
||||
const (
|
||||
notifyTriggered = "triggered"
|
||||
@@ -97,6 +104,10 @@ var notifyClient = &http.Client{Timeout: 10 * time.Second}
|
||||
|
||||
// StartNotifier delivers queued notifications until ctx is cancelled, starting
|
||||
// with an immediate pass so a restart flushes whatever the last one left behind.
|
||||
//
|
||||
// Each pass runs under notifierLockKey (see withAdvisoryLock), so that on more
|
||||
// than one replica only whichever instance's tick takes the lock first actually
|
||||
// delivers; the rest skip that tick rather than racing the same pass.
|
||||
func StartNotifier(ctx context.Context, db *sql.DB, cfg NotifyConfig) {
|
||||
if !cfg.enabled() {
|
||||
log.Print("notifier: disabled (no ntfy URL configured)")
|
||||
@@ -107,11 +118,17 @@ func StartNotifier(ctx context.Context, db *sql.DB, cfg NotifyConfig) {
|
||||
ticker := time.NewTicker(notifyInterval)
|
||||
defer ticker.Stop()
|
||||
|
||||
NotifySweep(ctx, db, cfg)
|
||||
sweep := func() {
|
||||
withAdvisoryLock(ctx, db, notifierLockKey, "notifier", func() {
|
||||
NotifySweep(ctx, db, cfg)
|
||||
})
|
||||
}
|
||||
|
||||
sweep()
|
||||
for {
|
||||
select {
|
||||
case <-ticker.C:
|
||||
NotifySweep(ctx, db, cfg)
|
||||
sweep()
|
||||
case <-ctx.Done():
|
||||
return
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user