Files
terdut-server/charts/terdut-server/values.yaml
T
Niklas Ye 3bf94a5d7f
CI / chart (push) Successful in 0s
CI / test (push) Successful in 7s
CI / security (push) Successful in 12s
Default to 2 replicas and RollingUpdate now that the singleton jobs are locked
replicas and strategy: Recreate were the chart's only guard against the
archiver, notifier and migration races; v0.36.0 closed all three with
advisory locks and a conflict-resolving incident insert, which made
that guard redundant rather than load-bearing. Expose replicaCount
(new, no values.yaml key existed before) defaulting to 2, and switch
to strategy: RollingUpdate with no explicit maxUnavailable/maxSurge --
the 25%/25% default rounds to 0/1 at replicaCount: 2, which is already
zero-downtime.

The chart does not gate this on image.tag, so pointing it at a
pre-v0.36.0 image with the new default is a foot-gun by omission --
noted in both the values.yaml comment and the deployment.yaml comment,
not guarded in code, same as the chart does for every other
version-coupled assumption today.

Co-authored-by: Claude <noreply@anthropic.com>
2026-10-03 12:30:27 +02:00

204 lines
9.1 KiB
YAML

# Safe above 1 since v0.36.0: the sweeper, notifier and migration runner each
# take a Postgres advisory lock around their own pass, and a webhook that
# loses the race to open a brand-new incident attaches to the winner's row
# instead of dropping its payload. An image older than v0.36.0 does not have
# these guards -- do not raise this against one.
replicaCount: 2
networking:
hostname: "terdut.example.com"
servicePort: 8080
# Gateway listener to bind the HTTPRoute to. Empty attaches to every matching
# listener, including plaintext HTTP. Set this to the name of the HTTPS
# listener to serve the API over TLS only.
listener: ""
image:
repository: git.ryuvia.com/niklas/terdut-server
tag: "latest"
pullPolicy: IfNotPresent
# Postgres connection. The chart provisions no database; it expects one to exist.
database:
# Required. A DSN with no password in it:
# postgres://terdut@terdut-postgres:5432/terdut?sslmode=require
#
# The password is deliberately a separate setting. pgx falls back to libpq's
# environment variables for anything the DSN omits, so PGPASSWORD supplies it
# without the credential appearing in values, in the rendered manifest, or in
# `kubectl describe pod`.
dsn: ""
# Where PGPASSWORD comes from. With the Zalando postgres operator this is the
# Secret it generates for the role — `<user>.<cluster>.credentials.postgresql.acid.zalan.do`,
# whose keys are `username` and `password` — so a from-scratch rebuild mints a
# new password and the server picks it up with nothing to keep in sync.
#
# Read at process start only: rotating the password needs a pod restart.
#
# Leave name empty only if the DSN carries its own password, which puts it in
# the manifest.
passwordSecret:
name: ""
key: password
# Blocks the main container from starting until Postgres accepts
# connections. Without this, a Deployment created before Postgres has
# finished its very first boot -- initdb plus Patroni leader election, on a
# from-scratch postgres-operator cluster -- crash-loops a few times: the
# app's own ping-retry budget on startup (pingAttempts/pingRetryDelay in
# internal/db/db.go) is sized for a much shorter, different race --
# NetworkPolicy propagation, a few seconds -- not for genuine first-time
# cluster creation, which routinely takes longer, so it exhausts and the
# process exits before ever binding its HTTP port. A startupProbe cannot
# help here: the crash happens before there is anything to probe.
#
# pg_isready needs no credentials -- it reports PQPING_OK on anything that
# amounts to "a Postgres backend answered", including an auth challenge --
# so no PGPASSWORD is wired into this container.
waitForPostgres:
enabled: true
image:
repository: postgres
tag: "17-alpine"
pullPolicy: IfNotPresent
service:
type: ClusterIP
port: 8080
sweeper:
# How long a firing alert may go without a refreshing webhook before it is
# treated as resolved. Must exceed your Alertmanager repeat_interval.
staleAfter: 6h
# How long a resolved alert stays in the default list before auto-archiving.
archiveAfter: 168h
# Alerts treated as dead man's switches: receiving one opens no incident, and
# the absence of one does. The Watchdog alert kube-prometheus-stack ships is
# exactly this — an always-firing alert whose only value is something noticing
# when it stops.
deadman:
# Which alerts to treat as heartbeats. ";" separates matchers, "," separates
# the label conditions within one, "=" is exact equality. Every matcher must
# name an alertname:
# alertname=Watchdog,cluster=prod; alertname=EdgeHeartbeat
# Each distinct label set is watched independently, so two clusters sending
# the same alertname are two switches and a live one cannot mask a dead one.
matchers: "alertname=Watchdog"
# How long a heartbeat may go unheard before its switch is declared dead.
#
# This must be SHORTER than the Alertmanager repeat_interval of the route
# carrying the heartbeat — the opposite of sweeper.staleAfter. The default
# repeat_interval of 4h (12h in many setups) makes for a useless dead man's
# switch, so give the heartbeat a route of its own:
#
# - matchers: [ 'alertname = "Watchdog"' ]
# receiver: terdut
# group_wait: 0s
# group_interval: 1m
# repeat_interval: 1m
#
# That delivers every 2m rather than every 1m: a group is only reconsidered
# each group_interval, and at exactly one elapsed interval repeat_interval has
# not quite passed, so equal values give 2x. Fine against 15m; use
# group_interval: 30s if you want a true 1m.
#
# Set to 0 to disable dead man's switch handling entirely.
timeout: 15m
# Severity a dead man's switch incident opens at. These incidents have no
# member alerts to derive one from, and the heartbeat's own severity label is
# meaningless — Watchdog ships as "none". Only "critical" maps to the ntfy
# priority that overrides a phone's quiet hours.
severity: critical
notify:
# ntfy server that push notifications are published to, e.g.
# http://ntfy.ntfy.svc.cluster.local. Empty disables notifications entirely.
ntfyUrl: ""
# Topic used when nobody is on call today. Notifications sent here carry no
# Acknowledge button: the topic is shared, so there is no user to attribute an
# acknowledgement to. Leave empty to send nothing when the schedule is unset.
fallbackTopic: ""
# How long an incident may sit unacknowledged before it is paged again.
# Set to 0 to notify once and never repeat.
repeatEvery: 15m
# Base URL a phone uses to reach this server, for the link and the Acknowledge
# button inside a notification. Defaults to https://<networking.hostname>.
#
# The Acknowledge button is a POST to /api/notify/ack/{token} from the
# responder's phone, so that path has to stay publicly reachable — it is
# authorised by the scoped token in the URL, not by network placement.
publicUrl: ""
# Optional bearer token for an access-controlled ntfy, read from an existing
# Secret. Leave name empty for an open ntfy.
tokenSecret:
name: ""
key: token
# Whether a user may sign in, or sign up, with a password. Turn it off once
# single sign-on works, to make it the only way in; turn it back on (and
# redeploy) if the identity provider is down and somebody has to get in.
passwordLogin: true
# Declares this install gitops-managed: writes to teams, escalation policies,
# dead man's switches and integrations from a session or a user's own API key
# are refused, while a service account's (see SERVICE-ACCOUNTS.md) are not.
# Off by default — turning it on is a statement that something like
# terdut-operator, not a person in the web UI, owns this install's
# configuration from here on.
operatorMode: false
# Single sign-on through an OpenID Connect provider such as Authentik.
#
# At the provider, create an OAuth2/OpenID application whose redirect URI is
# <notify.publicUrl>/api/oidc/callback
# (publicUrl defaults to https://<networking.hostname>), a confidential client, and
# put the client secret in an existing Secret named by clientSecret below.
#
# Groups from the provider decide what a person can do. Access it grants is
# marked as managed by single sign-on and is re-read at every sign-in; anything
# added by hand in terdut is left alone. Changes in the provider take effect at
# the person's next sign-in, at most sessionMaxAge later. API keys are NOT
# revoked when somebody is removed at the provider: disable the user in terdut too.
oidc:
enabled: false
# Issuer URL. For Authentik: https://<authentik>/application/o/<app-slug>/
issuer: ""
clientId: ""
clientSecret:
name: ""
key: client-secret
# What the sign-in button calls the provider.
name: SSO
# Authentik puts the groups claim behind the profile scope.
scopes: "openid profile email"
usernameClaim: preferred_username
emailClaim: email
groupsClaim: groups
# Link a first sign-in to an existing local user with the same email even when
# the provider does not mark the address verified. Authentik reports
# email_verified as false unless configured otherwise.
trustEmail: false
# Only people in one of these groups may sign in. Empty admits everybody the
# provider authenticates, and access control is left to the provider.
allowedGroups: []
# Members of this group are system administrators.
adminGroup: ""
# Which group grants a team's membership and ownership is each team's own
# setting now, not chart config: an owner sets it from the Members tab, or
# PUT /api/teams/{teamID}/oidc-groups. A team must already exist for a group
# to grant access to it.
# Hard ceiling on a session made by a single sign-on login.
sessionMaxAge: 12h
# Backups are no longer this chart's business. The SQLite database lived on a PVC
# beside the app, so it needed a sidecar with a sqlite3 module for k8up to exec a
# dump in; Postgres is backed up where it runs, through a k8up.io/backupcommand
# pg_dump annotation on the database pod itself.
bootstrap:
enabled: true
username: admin
email: admin@example.com
# secretName overrides the default of <fullname>-admin-key
secretName: ""