9e5b085d8b
openIncident's INSERT had no ON CONFLICT clause, relying entirely on incidentForGroup's earlier SELECT to avoid a duplicate. On more than one replica, two webhook deliveries for the very first occurrence of a brand-new groupKey can both pass that SELECT before either INSERTs; the loser then hit incidents_open_group_key_idx's unique violation, which rolled back its whole transaction — including that payload's alert upserts, done earlier in the same transaction. ingest's error is only logged and receiveWebhook answers 200 regardless, so nothing retried it: the loser's alerts silently never existed. Add ON CONFLICT (team_id, group_key) WHERE resolved_at IS NULL DO NOTHING to the INSERT, matching the partial unique index. Postgres only resolves that conflict after the winning transaction commits (or rolls back), so by the time RETURNING comes back empty, existingOpenIncident's follow-up SELECT is guaranteed to see the winner's row. The loser attaches to it instead of failing outright, and the rest of its payload commits normally. Covers both callers, since the dead man's switch sweeper shares this same function. New test (package api_test, fires N webhook deliveries for one groupKey from a synchronized start with distinct fingerprints, so they aren't accidentally serialized by upsertAlerts' own per- fingerprint lock) confirmed meaningful: with the ON CONFLICT clause reverted, it fails 10/10 on a missing alert fingerprint; restored, 0/10. Note while building it: "exactly one incident" alone cannot distinguish fixed from broken, since the DB's own unique index already guarantees that either way — the real signal is the loser's payload surviving. Chart comment updated: all three of the chart's original reasons for Recreate are now addressed in code, though replicas stays at 1 and the strategy stays Recreate pending a deliberate decision to raise it. Co-authored-by: Claude <noreply@anthropic.com>