Wait for Postgres to accept connections before the main container starts
A Deployment created before Postgres has finished its very first boot -- initdb plus Patroni leader election, on a from-scratch postgres-operator cluster -- crash-looped a few times. db.Open()'s own ping-retry budget (pingAttempts/pingRetryDelay, internal/db/db.go) is sized for a much shorter, different race -- NetworkPolicy propagation, a few seconds -- not for genuine first-time cluster creation, which routinely takes longer, so it exhausted and the process exited before ever binding its HTTP port. A startupProbe cannot fix that: the crash happens before there is anything to probe. Added a wait-for-postgres init container instead: it loops pg_isready against database.dsn until Postgres actually answers, before the main container's own, unchanged retry budget gets a chance to run out. pg_isready needs no credentials -- it reports PQPING_OK on anything that amounts to a Postgres backend answering, including an auth challenge -- so no PGPASSWORD is wired into it. Chart-only; no Go code changed. database.waitForPostgres.enabled defaults to true and can be turned off if something else already guarantees Postgres is reachable before this Deployment is created.
This commit is contained in:
@@ -21,6 +21,23 @@ spec:
|
|||||||
{{- include "terdut-server.selectorLabels" . | nindent 8 }}
|
{{- include "terdut-server.selectorLabels" . | nindent 8 }}
|
||||||
spec:
|
spec:
|
||||||
enableServiceLinks: false
|
enableServiceLinks: false
|
||||||
|
{{- if .Values.database.waitForPostgres.enabled }}
|
||||||
|
initContainers:
|
||||||
|
- name: wait-for-postgres
|
||||||
|
image: "{{ .Values.database.waitForPostgres.image.repository }}:{{ .Values.database.waitForPostgres.image.tag }}"
|
||||||
|
imagePullPolicy: {{ .Values.database.waitForPostgres.image.pullPolicy }}
|
||||||
|
env:
|
||||||
|
- name: TERDUT_DB_DSN
|
||||||
|
value: {{ required "database.dsn is required" .Values.database.dsn | quote }}
|
||||||
|
command:
|
||||||
|
- sh
|
||||||
|
- -c
|
||||||
|
- |
|
||||||
|
until pg_isready -d "$TERDUT_DB_DSN"; do
|
||||||
|
echo "wait-for-postgres: not ready yet, retrying in 2s"
|
||||||
|
sleep 2
|
||||||
|
done
|
||||||
|
{{- end }}
|
||||||
containers:
|
containers:
|
||||||
- name: terdut-server
|
- name: terdut-server
|
||||||
image: "{{ .Values.image.repository }}:{{ .Values.image.tag }}"
|
image: "{{ .Values.image.repository }}:{{ .Values.image.tag }}"
|
||||||
|
|||||||
@@ -33,6 +33,26 @@ database:
|
|||||||
passwordSecret:
|
passwordSecret:
|
||||||
name: ""
|
name: ""
|
||||||
key: password
|
key: password
|
||||||
|
# Blocks the main container from starting until Postgres accepts
|
||||||
|
# connections. Without this, a Deployment created before Postgres has
|
||||||
|
# finished its very first boot -- initdb plus Patroni leader election, on a
|
||||||
|
# from-scratch postgres-operator cluster -- crash-loops a few times: the
|
||||||
|
# app's own ping-retry budget on startup (pingAttempts/pingRetryDelay in
|
||||||
|
# internal/db/db.go) is sized for a much shorter, different race --
|
||||||
|
# NetworkPolicy propagation, a few seconds -- not for genuine first-time
|
||||||
|
# cluster creation, which routinely takes longer, so it exhausts and the
|
||||||
|
# process exits before ever binding its HTTP port. A startupProbe cannot
|
||||||
|
# help here: the crash happens before there is anything to probe.
|
||||||
|
#
|
||||||
|
# pg_isready needs no credentials -- it reports PQPING_OK on anything that
|
||||||
|
# amounts to "a Postgres backend answered", including an auth challenge --
|
||||||
|
# so no PGPASSWORD is wired into this container.
|
||||||
|
waitForPostgres:
|
||||||
|
enabled: true
|
||||||
|
image:
|
||||||
|
repository: postgres
|
||||||
|
tag: "17-alpine"
|
||||||
|
pullPolicy: IfNotPresent
|
||||||
|
|
||||||
service:
|
service:
|
||||||
type: ClusterIP
|
type: ClusterIP
|
||||||
|
|||||||
Reference in New Issue
Block a user