3bf94a5d7f
replicas and strategy: Recreate were the chart's only guard against the archiver, notifier and migration races; v0.36.0 closed all three with advisory locks and a conflict-resolving incident insert, which made that guard redundant rather than load-bearing. Expose replicaCount (new, no values.yaml key existed before) defaulting to 2, and switch to strategy: RollingUpdate with no explicit maxUnavailable/maxSurge -- the 25%/25% default rounds to 0/1 at replicaCount: 2, which is already zero-downtime. The chart does not gate this on image.tag, so pointing it at a pre-v0.36.0 image with the new default is a foot-gun by omission -- noted in both the values.yaml comment and the deployment.yaml comment, not guarded in code, same as the chart does for every other version-coupled assumption today. Co-authored-by: Claude <noreply@anthropic.com>
204 lines
9.1 KiB
YAML
204 lines
9.1 KiB
YAML
# Safe above 1 since v0.36.0: the sweeper, notifier and migration runner each
|
|
# take a Postgres advisory lock around their own pass, and a webhook that
|
|
# loses the race to open a brand-new incident attaches to the winner's row
|
|
# instead of dropping its payload. An image older than v0.36.0 does not have
|
|
# these guards -- do not raise this against one.
|
|
replicaCount: 2
|
|
|
|
networking:
|
|
hostname: "terdut.example.com"
|
|
servicePort: 8080
|
|
# Gateway listener to bind the HTTPRoute to. Empty attaches to every matching
|
|
# listener, including plaintext HTTP. Set this to the name of the HTTPS
|
|
# listener to serve the API over TLS only.
|
|
listener: ""
|
|
|
|
image:
|
|
repository: git.ryuvia.com/niklas/terdut-server
|
|
tag: "latest"
|
|
pullPolicy: IfNotPresent
|
|
|
|
# Postgres connection. The chart provisions no database; it expects one to exist.
|
|
database:
|
|
# Required. A DSN with no password in it:
|
|
# postgres://terdut@terdut-postgres:5432/terdut?sslmode=require
|
|
#
|
|
# The password is deliberately a separate setting. pgx falls back to libpq's
|
|
# environment variables for anything the DSN omits, so PGPASSWORD supplies it
|
|
# without the credential appearing in values, in the rendered manifest, or in
|
|
# `kubectl describe pod`.
|
|
dsn: ""
|
|
# Where PGPASSWORD comes from. With the Zalando postgres operator this is the
|
|
# Secret it generates for the role — `<user>.<cluster>.credentials.postgresql.acid.zalan.do`,
|
|
# whose keys are `username` and `password` — so a from-scratch rebuild mints a
|
|
# new password and the server picks it up with nothing to keep in sync.
|
|
#
|
|
# Read at process start only: rotating the password needs a pod restart.
|
|
#
|
|
# Leave name empty only if the DSN carries its own password, which puts it in
|
|
# the manifest.
|
|
passwordSecret:
|
|
name: ""
|
|
key: password
|
|
# Blocks the main container from starting until Postgres accepts
|
|
# connections. Without this, a Deployment created before Postgres has
|
|
# finished its very first boot -- initdb plus Patroni leader election, on a
|
|
# from-scratch postgres-operator cluster -- crash-loops a few times: the
|
|
# app's own ping-retry budget on startup (pingAttempts/pingRetryDelay in
|
|
# internal/db/db.go) is sized for a much shorter, different race --
|
|
# NetworkPolicy propagation, a few seconds -- not for genuine first-time
|
|
# cluster creation, which routinely takes longer, so it exhausts and the
|
|
# process exits before ever binding its HTTP port. A startupProbe cannot
|
|
# help here: the crash happens before there is anything to probe.
|
|
#
|
|
# pg_isready needs no credentials -- it reports PQPING_OK on anything that
|
|
# amounts to "a Postgres backend answered", including an auth challenge --
|
|
# so no PGPASSWORD is wired into this container.
|
|
waitForPostgres:
|
|
enabled: true
|
|
image:
|
|
repository: postgres
|
|
tag: "17-alpine"
|
|
pullPolicy: IfNotPresent
|
|
|
|
service:
|
|
type: ClusterIP
|
|
port: 8080
|
|
|
|
sweeper:
|
|
# How long a firing alert may go without a refreshing webhook before it is
|
|
# treated as resolved. Must exceed your Alertmanager repeat_interval.
|
|
staleAfter: 6h
|
|
# How long a resolved alert stays in the default list before auto-archiving.
|
|
archiveAfter: 168h
|
|
|
|
# Alerts treated as dead man's switches: receiving one opens no incident, and
|
|
# the absence of one does. The Watchdog alert kube-prometheus-stack ships is
|
|
# exactly this — an always-firing alert whose only value is something noticing
|
|
# when it stops.
|
|
deadman:
|
|
# Which alerts to treat as heartbeats. ";" separates matchers, "," separates
|
|
# the label conditions within one, "=" is exact equality. Every matcher must
|
|
# name an alertname:
|
|
# alertname=Watchdog,cluster=prod; alertname=EdgeHeartbeat
|
|
# Each distinct label set is watched independently, so two clusters sending
|
|
# the same alertname are two switches and a live one cannot mask a dead one.
|
|
matchers: "alertname=Watchdog"
|
|
# How long a heartbeat may go unheard before its switch is declared dead.
|
|
#
|
|
# This must be SHORTER than the Alertmanager repeat_interval of the route
|
|
# carrying the heartbeat — the opposite of sweeper.staleAfter. The default
|
|
# repeat_interval of 4h (12h in many setups) makes for a useless dead man's
|
|
# switch, so give the heartbeat a route of its own:
|
|
#
|
|
# - matchers: [ 'alertname = "Watchdog"' ]
|
|
# receiver: terdut
|
|
# group_wait: 0s
|
|
# group_interval: 1m
|
|
# repeat_interval: 1m
|
|
#
|
|
# That delivers every 2m rather than every 1m: a group is only reconsidered
|
|
# each group_interval, and at exactly one elapsed interval repeat_interval has
|
|
# not quite passed, so equal values give 2x. Fine against 15m; use
|
|
# group_interval: 30s if you want a true 1m.
|
|
#
|
|
# Set to 0 to disable dead man's switch handling entirely.
|
|
timeout: 15m
|
|
# Severity a dead man's switch incident opens at. These incidents have no
|
|
# member alerts to derive one from, and the heartbeat's own severity label is
|
|
# meaningless — Watchdog ships as "none". Only "critical" maps to the ntfy
|
|
# priority that overrides a phone's quiet hours.
|
|
severity: critical
|
|
|
|
notify:
|
|
# ntfy server that push notifications are published to, e.g.
|
|
# http://ntfy.ntfy.svc.cluster.local. Empty disables notifications entirely.
|
|
ntfyUrl: ""
|
|
# Topic used when nobody is on call today. Notifications sent here carry no
|
|
# Acknowledge button: the topic is shared, so there is no user to attribute an
|
|
# acknowledgement to. Leave empty to send nothing when the schedule is unset.
|
|
fallbackTopic: ""
|
|
# How long an incident may sit unacknowledged before it is paged again.
|
|
# Set to 0 to notify once and never repeat.
|
|
repeatEvery: 15m
|
|
# Base URL a phone uses to reach this server, for the link and the Acknowledge
|
|
# button inside a notification. Defaults to https://<networking.hostname>.
|
|
#
|
|
# The Acknowledge button is a POST to /api/notify/ack/{token} from the
|
|
# responder's phone, so that path has to stay publicly reachable — it is
|
|
# authorised by the scoped token in the URL, not by network placement.
|
|
publicUrl: ""
|
|
# Optional bearer token for an access-controlled ntfy, read from an existing
|
|
# Secret. Leave name empty for an open ntfy.
|
|
tokenSecret:
|
|
name: ""
|
|
key: token
|
|
|
|
# Whether a user may sign in, or sign up, with a password. Turn it off once
|
|
# single sign-on works, to make it the only way in; turn it back on (and
|
|
# redeploy) if the identity provider is down and somebody has to get in.
|
|
passwordLogin: true
|
|
|
|
# Declares this install gitops-managed: writes to teams, escalation policies,
|
|
# dead man's switches and integrations from a session or a user's own API key
|
|
# are refused, while a service account's (see SERVICE-ACCOUNTS.md) are not.
|
|
# Off by default — turning it on is a statement that something like
|
|
# terdut-operator, not a person in the web UI, owns this install's
|
|
# configuration from here on.
|
|
operatorMode: false
|
|
|
|
# Single sign-on through an OpenID Connect provider such as Authentik.
|
|
#
|
|
# At the provider, create an OAuth2/OpenID application whose redirect URI is
|
|
# <notify.publicUrl>/api/oidc/callback
|
|
# (publicUrl defaults to https://<networking.hostname>), a confidential client, and
|
|
# put the client secret in an existing Secret named by clientSecret below.
|
|
#
|
|
# Groups from the provider decide what a person can do. Access it grants is
|
|
# marked as managed by single sign-on and is re-read at every sign-in; anything
|
|
# added by hand in terdut is left alone. Changes in the provider take effect at
|
|
# the person's next sign-in, at most sessionMaxAge later. API keys are NOT
|
|
# revoked when somebody is removed at the provider: disable the user in terdut too.
|
|
oidc:
|
|
enabled: false
|
|
# Issuer URL. For Authentik: https://<authentik>/application/o/<app-slug>/
|
|
issuer: ""
|
|
clientId: ""
|
|
clientSecret:
|
|
name: ""
|
|
key: client-secret
|
|
# What the sign-in button calls the provider.
|
|
name: SSO
|
|
# Authentik puts the groups claim behind the profile scope.
|
|
scopes: "openid profile email"
|
|
usernameClaim: preferred_username
|
|
emailClaim: email
|
|
groupsClaim: groups
|
|
# Link a first sign-in to an existing local user with the same email even when
|
|
# the provider does not mark the address verified. Authentik reports
|
|
# email_verified as false unless configured otherwise.
|
|
trustEmail: false
|
|
# Only people in one of these groups may sign in. Empty admits everybody the
|
|
# provider authenticates, and access control is left to the provider.
|
|
allowedGroups: []
|
|
# Members of this group are system administrators.
|
|
adminGroup: ""
|
|
# Which group grants a team's membership and ownership is each team's own
|
|
# setting now, not chart config: an owner sets it from the Members tab, or
|
|
# PUT /api/teams/{teamID}/oidc-groups. A team must already exist for a group
|
|
# to grant access to it.
|
|
# Hard ceiling on a session made by a single sign-on login.
|
|
sessionMaxAge: 12h
|
|
|
|
# Backups are no longer this chart's business. The SQLite database lived on a PVC
|
|
# beside the app, so it needed a sidecar with a sqlite3 module for k8up to exec a
|
|
# dump in; Postgres is backed up where it runs, through a k8up.io/backupcommand
|
|
# pg_dump annotation on the database pod itself.
|
|
|
|
bootstrap:
|
|
enabled: true
|
|
username: admin
|
|
email: admin@example.com
|
|
# secretName overrides the default of <fullname>-admin-key
|
|
secretName: ""
|