8 Commits

Author SHA1 Message Date
Niklas Ye b0a431f2a4 Set the chart's placeholder version to 0.5.0
CI / chart (push) Successful in 2s
CI / security (push) Successful in 53s
Release / chart (push) Successful in 3s
CI / test (push) Successful in 2m19s
Release / test (push) Successful in 1m41s
Release / image (push) Successful in 6m30s
Release / scan-image (push) Successful in 4s
Cosmetic: make helm-package passes --version and --app-version from the
tag, so these two fields decide nothing about what gets published. Still
done, as with d502485 (0.4.0) and 88172ad (0.3.0) before it, because a
tree heading for v0.5.0 that still says 0.4.0 tells its reader
something false.

Co-authored-by: Claude <noreply@anthropic.com>
2026-10-08 21:10:27 +02:00
Niklas Ye 62664c93ff Keep the instance credential across a TerdutServer delete, and adopt it on recreate
Deleting a TerdutServer removed the credential Secrets but never touched the
database, so a recreated one found a server that was already bootstrapped and
no key for it: /api/bootstrap answered 403 and the operator stopped at
BootstrapStateLost, whose message and DESIGN.md both said "delete and
recreate". That is how the terdut-demo install on the cluster got stuck on
2026-10-03: Helm's cleanupOnFail deleted its TerdutServer after a failed
upgrade, the recreate found the bootstrapped database, and it sat at Ready:
False for five days until the database was reset by hand. Recreating cannot
fix it, because the finalizer clears Secrets and the database is not its to
reset, so "a fresh create starts clean" was only ever true when the database
went with it.

spec.credentials.deletionPolicy is Retain by default: the finalizer keeps the
instance credential Secret (Delete removes it, as before). The bootstrap
checkpoint is always removed. Before calling /api/bootstrap, reconcile now
looks for the retained Secret and asks the server for the operator's own
service account with its token. Accepted: adopt it and skip bootstrap.
Rejected with 401/403: the Secret outlived a database reset, so ignore it and
bootstrap like a first install, which replaces it. Any other error retries.
terdut-server's own tests already call that endpoint with an instance-scoped
key, so the permission is not new.

BootstrapStateLost is still the answer when the server is bootstrapped and no
credential it accepts survives, but its message now names the Secret to
restore and says that recreating does not clear the database. DESIGN.md §6
says the same, and the chart passes the setting through as
terdutServer.credentials.deletionPolicy.

A retained Secret of a TerdutServer that is gone for good is an orphan to
delete by hand. It is inert: nothing adopts it unless the server accepts the
token.

Checked on the kind demo with a locally built image against the real
terdut-server v0.43.0: deleting the TerdutServer kept the Secret, recreating it
reached Ready with the same credential (identical hash) and both TerdutTeams
came back Ready with their original ids. The controller specs cover adoption,
a rejected token after a reset, the bootstrapped-and-rejected failure, and
both deletion policies.

Co-authored-by: Claude <noreply@anthropic.com>
2026-10-08 21:09:56 +02:00
niklas 6572f63157 Merge pull request 'examples/demo: run terdut-server v0.43.0, with alerts from two clusters' (#6) from demo-matches-v0.4.0-crd into main
CI / chart (push) Successful in 2s
CI / security (push) Successful in 51s
CI / test (push) Successful in 2m26s
Reviewed-on: #6
2026-10-08 17:10:06 +00:00
Niklas Ye 50ce5bcec0 examples/demo: run terdut-server v0.43.0, with alerts from two clusters
CI / test (pull_request) Successful in 6m35s
CI / chart (pull_request) Successful in 2s
CI / security (pull_request) Successful in 1m6s
The demo pinned v0.36.0, the floor for replicas: 2, and so showed none of the
web UI since: the queue and incident layouts, the rota and escalation
pages, the theme toggle, and the cluster chip, filter and page titles
(v0.42.0-v0.43.0). It pins v0.43.0 now; the comment keeps v0.36.0 as the
floor, which is what the replicas setting actually depends on.

fire-alerts.sh takes an optional CLUSTER, standing in for a Prometheus
external label plus `cluster` in Alertmanager's group_by (terdut-server's
README, "Several clusters, one team"). It goes on the alert's labels and
groupLabels, and into the group key and the fingerprint, so the same alert in
two clusters is two incidents and not one. Unset, the payload is exactly what
it was. run-demo.sh fires its alerts across prod-eu and prod-us, high-cpu in
both, so the queue has a chip and a filter to show.

run-demo.sh also failed on its second run, though it says it is safe to
re-run: it expected HTTP 409 when alice already exists, but a spent invite
is answered with 403 "invite link is not usable" before the username is ever
checked. It now tries to log alice in first and skips the signup if that works.

Checked on the kind cluster: the server rolled to v0.43.0, every CR became
Ready and Adopted (server, both teams, both escalation rules, both dead man's
switches, both alert sources), and /api/incidents/clusters,
/api/incidents?cluster=prod-us and the incident titles came back as expected.
No operator code changed, so this needs no operator release.

Co-authored-by: Claude <noreply@anthropic.com>
2026-10-08 18:55:02 +02:00
Niklas Ye 822c80dda6 examples/demo: split the Ready wait so escalation rules wait on alice too
wait_for_ready waited for every demo object at once, including
terdutescalationrule-platform, which names alice as a level-1 target --
but alice does not exist yet at that point in main(): she is created by
redeem_platform_invite, which ran after wait_for_ready. terdut-server
resolves every named username at reconcile time, not just when an
escalation actually fires, so that CR could never reach Ready before
alice did, and main() had no step in between to create her.

Split into wait_for_objects (the shared loop, now taking its object list
as arguments) plus two callers: wait_for_teams_ready, covering just the
server and the two teams redeem_platform_invite/join_payments_team
need, run before alice exists; wait_for_remaining_ready, covering the
escalation rules, dead man's switches and alert sources, run after.

Co-authored-by: Claude <noreply@anthropic.com>
2026-10-03 18:15:42 +02:00
Niklas Ye 0ee7ede648 examples/demo: match v0.4.0's new spec.replicas default and bump terdut-server
replicas: 1 and tag: v0.34.0 were both correct when written, but the CRD's
own default moved to 2 in v0.4.0 (same release this demo is meant to show
off), and v0.34.0 predates v0.36.0's advisory locks that make a second
replica safe instead of racing the first. Left as-is, the demo would have
been the one place in this repo demonstrating the exact unsafe combination
the CRD's own doc comment warns against: more than one replica against an
image that doesn't guard the sweeper/notifier/migration-runner singletons.

replicas is now stated explicitly as 2 rather than dropped to pick up the
default silently, matching every other field in this file's own habit of
spelling out what it depends on. tag moves to v0.36.0 specifically -- the
first version where the lock landed -- with the comment keeping v0.34.0's
original reasoning (the service-account race fix) alongside the new one,
since v0.36.0 still carries that fix forward.

Co-authored-by: Claude <noreply@anthropic.com>
2026-10-03 18:13:09 +02:00
Niklas Ye d50248531c Set the chart's placeholder version to 0.4.0
CI / chart (push) Successful in 1s
CI / security (push) Successful in 59s
CI / test (push) Successful in 2m10s
Release / test (push) Successful in 7m54s
Release / chart (push) Successful in 2s
Release / image (push) Successful in 6m25s
Release / scan-image (push) Successful in 36s
2026-10-03 16:22:15 +02:00
Niklas Ye 4007f54279 Default TerdutServer.spec.replicas to 2 and switch to RollingUpdate
CI / chart (push) Successful in 1s
CI / security (push) Successful in 59s
CI / test (push) Has been cancelled
Mirrors charts/terdut-server's own deployment.yaml change: v0.36.0 put
the sweeper, the notifier and the migration runner each behind a
Postgres advisory lock, and gave incident creation its own conflict
resolution, so the Recreate strategy and replicas-stays-at-1 guidance
this controller carried (explicitly tracking that chart's comment)
are no longer load-bearing.

spec.replicas' +kubebuilder:default moves 1 -> 2 (config/crd/bases and
the chart's CRD template regenerated via controller-gen and
kubebuilder's helm plugin respectively, then hand-verified identical
to the generator's own output rather than trusting a bulk regen --
the plugin's --output-dir charts writes a fresh charts/chart scaffold
rather than updating charts/terdut-operator in place, so only the
diff was taken, not the whole tree). terdutserver_deployment.go's
same-value fallback (reachable only for a TerdutServer stored before
this default existed) moves with it, and its Strategy changes from
Recreate to RollingUpdate with no explicit maxUnavailable/maxSurge --
the 25%/25% default rounds to 0/1 at replicas: 2, already
zero-downtime.

DESIGN.md's three places asserting multi-replica isn't a supported
topology (the illustrative spec.replicas YAML, spec.pod.affinity's
rationale, and the HPA deferred-feature note) are corrected to match;
the HPA note now gives its own standing reason (no scaling metric or
bounds decided yet) rather than a contradiction that no longer holds.

The chart's optional terdutServer.replicas sample value moves 1 -> 2
alongside it. image.tag must be v0.36.0 or newer for any of this to
hold -- stated in both the CRD field's doc comment and the chart
value's comment, not enforced in code, same stance the chart takes on
every other version-coupled assumption.

Co-authored-by: Claude <noreply@anthropic.com>
2026-10-03 12:35:18 +02:00
16 changed files with 533 additions and 127 deletions
+49 -22
View File
@@ -153,7 +153,7 @@ spec:
image:
repository: git.ryuvia.com/niklas/terdut-server
tag: v0.9.3
replicas: 1 # terdut-server is not horizontally-scale-tested; keep the field, default 1
replicas: 2 # default since terdut-server v0.36.0's advisory locks; see TerdutServerSpec.Replicas
networking:
hostname: terdut.example.com
servicePort: 8080
@@ -169,6 +169,8 @@ spec:
matchers: "alertname=Watchdog"
timeout: 15m
severity: critical
credentials:
deletionPolicy: Retain # Retain (default) | Delete -- what deleting this CR does to the instance credential Secret; see §6 point 2
notify:
ntfyURL: "http://ntfy.ntfy.svc.cluster.local"
fallbackTopic: ""
@@ -244,10 +246,10 @@ no custom wrapper buys anything for any of these, matching how
CloudNativePG and the Zalando postgres-operator both expose the same
knobs. `affinity` is pure user-supplied passthrough, not a
toggle-plus-generated-default the way a multi-replica-aware operator's
pod anti-affinity typically is: this operator never auto-generates
affinity of its own, since `replicas` above 1 isn't a supported topology
(the sweeper/notifier singleton constraint, §4.1's own illustrative YAML
comment). `spec.pod.disruptionBudget` is the one field here that isn't a
pod anti-affinity typically is: even though `replicas` now defaults to 2
(terdut-server v0.36.0's advisory locks made that safe, §4.1's own
illustrative YAML comment), this operator still never auto-generates
affinity of its own. `spec.pod.disruptionBudget` is the one field here that isn't a
straight PodTemplateSpec knob — when set, the controller reconciles a
`PodDisruptionBudget` selecting this `TerdutServer`'s pods; clearing it
deletes any it previously created (§7). `minAvailable`/`maxUnavailable`
@@ -529,6 +531,18 @@ when nothing ever crosses into a tenant namespace in the first place.
the same rigor §5's general idempotent-create rule already applies
elsewhere:
- `status.credentialsSecretRef` already set: done, nothing to do.
- Otherwise, look for the instance credential Secret itself
(`<namespace>.<name>-instance-credentials`, point 2 below): a
`TerdutServer` deleted and recreated under the same name leaves it
behind by default (`spec.credentials.deletionPolicy: Retain`), and the
server it logs in to has not changed, because deleting a
`TerdutServer` never touches its database. If it exists, ask the server
for the operator's own service account with that token
(`GET /api/service-accounts?name=terdut-operator`). Accepted: adopt it,
set `status.credentialsSecretRef`, and skip everything below. Rejected
with a `401`/`403`: it is stale (a database reset since), so ignore it
and carry on; the steps below replace it. Any other failure is a retry,
not a guess.
- Otherwise, check for an intermediate
`<namespace>.<name>-bootstrap-admin` Secret in
the operator's own namespace first. If it exists, its key is a still-
@@ -538,15 +552,21 @@ when nothing ever crosses into a tenant namespace in the first place.
on `201`, immediately checkpoint its response's raw admin key
(`{"user": ..., "api_key": {"key": "<raw>", ...}}`) into that Secret
before doing anything else with it. A `403` with neither
`status.credentialsSecretRef` nor this checkpoint Secret present is
the one genuinely pathological case left (the checkpoint deleted out
from under a reconcile already past this point) — handled the same
way the design already handles unrecoverable server-issued material
elsewhere (§5's webhook-Secret-loss rule): fail closed,
`Ready: False, reason: BootstrapStateLost`, with the same recovery as
that case, delete and recreate the `TerdutServer` (its finalizer tears
down the Deployment/database-backing and server-side rows; a fresh
create starts clean) — not a workaround peculiar to this one path.
`status.credentialsSecretRef`, nor an instance credential the server
accepts, nor this checkpoint Secret present is the one genuinely
pathological case left: the server's database is already bootstrapped
and no credential for it survives here (the checkpoint deleted out from
under a reconcile already past this point, `deletionPolicy: Delete`, or
the Secret removed by hand). Fail closed, `Ready: False, reason:
BootstrapStateLost`, with a message that names the instance credential
Secret to restore. Deleting and recreating the `TerdutServer` does **not**
recover from it: the finalizer removes Secrets only, and the database
(the one thing that still says "already bootstrapped") is not the
operator's to reset. Restore the Secret from a copy, or reset the
server's database and then recreate the `TerdutServer`, which then
bootstraps like a first install. (This section used to say a fresh
create "starts clean"; that was only ever true when the database went
with it.)
- With an admin key in hand (fresh or checkpointed): `POST
/api/service-accounts {name: "terdut-operator", scope: "instance"}`.
A `409` here means a prior attempt got this far before being
@@ -563,10 +583,15 @@ when nothing ever crosses into a tenant namespace in the first place.
under a fixed data key, `token`), referenced back from
`TerdutServer.status.credentialsSecretRef: {name, key}` (§4.1). No
`OwnerReference` (those can't cross namespaces, and this Secret doesn't
share a namespace with the `TerdutServer` that caused it); the `TerdutServer`'s
finalizer deletes this Secret directly as part of its own teardown,
the same way it already has to clean up the server-side resources it
created (§5's general finalizer rule extends naturally to this Secret).
share a namespace with the `TerdutServer` that caused it). What the
`TerdutServer`'s finalizer does with it is `spec.credentials.deletionPolicy`:
`Retain` (the default) leaves it for a recreated `TerdutServer` to adopt
(point 1), `Delete` removes it directly as part of the teardown. There are
no server-side resources to undo either way: the bootstrap user and
service account have no delete verb in terdut-server's API. The bootstrap
checkpoint Secret is always removed. A retained Secret of a `TerdutServer`
that is gone for good is an orphan to delete by hand, and it is inert:
adoption asks the server to accept the token first.
3. When a `TerdutTeam` first becomes `Ready` (its `serverRef` resolved,
`allowedTeams` satisfied if cross-namespace), its controller uses the
`TerdutServer`'s instance-scoped credential (read from the operator's own
@@ -832,10 +857,12 @@ what it was, a separate install, until someone deletes it.
- Automatic Deployment restart on upstream Postgres credential rotation.
- `spec.pod.priorityClassName`, pod-label passthrough beyond
`spec.pod.annotations`, and a HorizontalPodAutoscaler for `TerdutServer`
— all considered alongside §4.1's `spec.pod` and explicitly left out of
that round: an HPA in particular would actively contradict
`spec.replicas`'s own stance that this operator doesn't support more
than one replica (the sweeper/notifier singleton constraint).
— all considered alongside §4.1's `spec.pod` and left out of that round.
An HPA no longer contradicts anything now that `spec.replicas` defaults
to 2 (terdut-server v0.36.0's advisory locks), but it is still a
separate, not-yet-made decision: a fixed replica count has no scaling
metric, min/max bounds, or cooldown behaviour to get right, and nobody
has asked for it yet.
- Admission webhooks / CEL-only validation limits (e.g. verifying a
`teamRef` exists at admission time rather than surfacing it as a status
condition after the fact).
+55 -12
View File
@@ -23,6 +23,37 @@ type SecretKeyRef struct {
Key string `json:"key"`
}
// CredentialsDeletionPolicy is what deleting a TerdutServer does to the
// instance credential Secret the operator generated for it.
// +kubebuilder:validation:Enum=Retain;Delete
type CredentialsDeletionPolicy string
const (
// CredentialsRetain keeps the Secret when the TerdutServer is deleted, so
// a TerdutServer recreated with the same name and namespace against the
// same database adopts it again instead of finding a server it cannot
// log in to. Deleting a TerdutServer never touches its database, so the
// operator's service account is still there to be reused. The default.
CredentialsRetain CredentialsDeletionPolicy = "Retain"
// CredentialsDelete removes the Secret with the TerdutServer. Choose it
// when the database goes too, or when the credential must not outlive the
// object.
CredentialsDelete CredentialsDeletionPolicy = "Delete"
)
// CredentialsSpec configures the lifecycle of the generated instance
// credential.
type CredentialsSpec struct {
// deletionPolicy: whether the instance credential Secret is kept
// (Retain, the default) or removed (Delete) when this TerdutServer is
// deleted. A kept Secret is only ever adopted after the server accepts
// its token, so one left over from a database that has since been reset
// is ignored and replaced.
// +kubebuilder:default=Retain
// +optional
DeletionPolicy CredentialsDeletionPolicy `json:"deletionPolicy,omitempty"`
}
// ImageSpec is the terdut-server image to run.
type ImageSpec struct {
// +kubebuilder:validation:MinLength=1
@@ -229,11 +260,11 @@ type PodSpec struct {
Tolerations []corev1.Toleration `json:"tolerations,omitempty"`
// affinity covers node affinity, pod affinity and pod anti-affinity in
// one field -- unlike a multi-replica-aware operator, this one never
// generates a default anti-affinity itself (replicas above 1 isn't a
// supported topology, see TerdutServerSpec.Replicas's own doc comment),
// so this is pure user-supplied passthrough, not a toggle-plus-generated-
// default.
// one field -- even though replicas now defaults to 2 (see
// TerdutServerSpec.Replicas's own doc comment), this operator still
// never generates a default anti-affinity of its own the way a
// multi-replica-aware operator typically would, so this stays pure
// user-supplied passthrough, not a toggle-plus-generated-default.
// +optional
Affinity *corev1.Affinity `json:"affinity,omitempty"`
@@ -301,10 +332,14 @@ type TerdutServerSpec struct {
// +required
Image ImageSpec `json:"image"`
// replicas. terdut-server is not horizontally-scale-tested; keep this
// at its default of 1 unless you've verified otherwise -- the sweeper
// and the notifier are unsynchronised singletons.
// +kubebuilder:default=1
// replicas. Defaults to 2: terdut-server v0.36.0 put the sweeper, the
// notifier and the migration runner each behind a Postgres advisory
// lock, and gave incident creation its own conflict resolution, so
// more than one replica no longer double-pages, races a migration, or
// drops a webhook payload. image.tag must be v0.36.0 or newer for
// that to hold -- an older terdut-server has none of these guards,
// and this field does not check the tag for you.
// +kubebuilder:default=2
// +optional
Replicas int32 `json:"replicas,omitempty"`
@@ -320,6 +355,11 @@ type TerdutServerSpec struct {
// +optional
Deadman DeadmanSpec `json:"deadman,omitempty"`
// credentials: what happens to the instance credential this operator
// generates for the server.
// +optional
Credentials CredentialsSpec `json:"credentials,omitempty"`
// +optional
Notify NotifySpec `json:"notify,omitempty"`
@@ -378,9 +418,12 @@ const (
// ReasonBootstrapStateLost: a checkpointed admin credential
// (DESIGN.md §6) was lost after being used but before the lasting
// credential it was for could be persisted -- the one genuinely
// pathological case in the self-registration flow. Fail-closed, same
// recovery as DESIGN.md §5's webhook-Secret-loss rule: delete and
// recreate this TerdutServer.
// pathological case in the self-registration flow -- or the server's
// database is already bootstrapped and no credential for it survives
// (spec.credentials.deletionPolicy: Delete, or the Secret removed by
// hand). Fail-closed: the operator cannot mint a credential, and
// deleting and recreating the TerdutServer does not clear the database.
// Restore the Secret, or reset the server's database.
ReasonBootstrapStateLost = "BootstrapStateLost"
// ReasonAdopted: the happy path. A working credential is in hand, the
// Deployment has a ready replica, and the database (if postgresClusterRef)
+16
View File
@@ -47,6 +47,21 @@ func (in *AllowedTeamsNamespaces) DeepCopy() *AllowedTeamsNamespaces {
return out
}
// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil.
func (in *CredentialsSpec) DeepCopyInto(out *CredentialsSpec) {
*out = *in
}
// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new CredentialsSpec.
func (in *CredentialsSpec) DeepCopy() *CredentialsSpec {
if in == nil {
return nil
}
out := new(CredentialsSpec)
in.DeepCopyInto(out)
return out
}
// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil.
func (in *DatabaseSpec) DeepCopyInto(out *DatabaseSpec) {
*out = *in
@@ -764,6 +779,7 @@ func (in *TerdutServerSpec) DeepCopyInto(out *TerdutServerSpec) {
in.Database.DeepCopyInto(&out.Database)
out.Sweeper = in.Sweeper
out.Deadman = in.Deadman
out.Credentials = in.Credentials
in.Notify.DeepCopyInto(&out.Notify)
in.OIDC.DeepCopyInto(&out.OIDC)
in.AllowedTeams.DeepCopyInto(&out.AllowedTeams)
+2 -2
View File
@@ -6,8 +6,8 @@ type: application
# These fields decide nothing: `make helm-package` passes --version and
# --app-version from the release tag (same reasoning as terdut-server's own
# chart). They're for whoever reads the tree before a tag exists.
version: 0.3.0
appVersion: "v0.3.0"
version: 0.5.0
appVersion: "v0.5.0"
keywords:
- kubernetes
@@ -128,6 +128,24 @@ spec:
x-kubernetes-map-type: atomic
type: object
type: object
credentials:
description: |-
credentials: what happens to the instance credential this operator
generates for the server.
properties:
deletionPolicy:
default: Retain
description: |-
deletionPolicy: whether the instance credential Secret is kept
(Retain, the default) or removed (Delete) when this TerdutServer is
deleted. A kept Secret is only ever adopted after the server accepts
its token, so one left over from a database that has since been reset
is ignored and replaced.
enum:
- Retain
- Delete
type: string
type: object
database:
description: |-
DatabaseSpec is the Postgres connection this TerdutServer uses. Exactly
@@ -327,11 +345,11 @@ spec:
affinity:
description: |-
affinity covers node affinity, pod affinity and pod anti-affinity in
one field -- unlike a multi-replica-aware operator, this one never
generates a default anti-affinity itself (replicas above 1 isn't a
supported topology, see TerdutServerSpec.Replicas's own doc comment),
so this is pure user-supplied passthrough, not a toggle-plus-generated-
default.
one field -- even though replicas now defaults to 2 (see
TerdutServerSpec.Replicas's own doc comment), this operator still
never generates a default anti-affinity of its own the way a
multi-replica-aware operator typically would, so this stays pure
user-supplied passthrough, not a toggle-plus-generated-default.
properties:
nodeAffinity:
description: Describes node affinity scheduling rules for
@@ -4304,11 +4322,15 @@ spec:
type: array
type: object
replicas:
default: 1
default: 2
description: |-
replicas. terdut-server is not horizontally-scale-tested; keep this
at its default of 1 unless you've verified otherwise -- the sweeper
and the notifier are unsynchronised singletons.
replicas. Defaults to 2: terdut-server v0.36.0 put the sweeper, the
notifier and the migration runner each behind a Postgres advisory
lock, and gave incident creation its own conflict resolution, so
more than one replica no longer double-pages, races a migration, or
drops a webhook payload. image.tag must be v0.36.0 or newer for
that to hold -- an older terdut-server has none of these guards,
and this field does not check the tag for you.
format: int32
type: integer
sweeper:
@@ -38,6 +38,10 @@ spec:
deadman:
{{- toYaml . | nindent 4 }}
{{- end }}
{{- with .Values.terdutServer.credentials }}
credentials:
{{- toYaml . | nindent 4 }}
{{- end }}
{{- with .Values.terdutServer.notify }}
notify:
{{- toYaml . | nindent 4 }}
+12 -1
View File
@@ -233,7 +233,11 @@ terdutServer:
## Required when terdutServer.enabled.
# tag: ""
replicas: 1
## Safe above 1 since terdut-server v0.36.0 (image.tag above must be that or
## newer): the sweeper, notifier and migration runner are each behind a
## Postgres advisory lock, and incident creation resolves its own insert
## conflict, matching this CRD's own spec.replicas default.
replicas: 2
networking:
## Required when terdutServer.enabled -- terdut-server's own public
@@ -266,6 +270,13 @@ terdutServer:
# matchers: "alertname=Watchdog"
# timeout: 15m
# severity: critical
# What happens to the instance credential the operator generates for this
# server when the TerdutServer is deleted. Retain (the default) keeps it, so
# a TerdutServer recreated with the same name against the same database
# adopts it again; Delete removes it with the TerdutServer. A kept Secret
# is only adopted if the server accepts its token.
# credentials:
# deletionPolicy: Retain
# notify: {}
# oidc: {}
@@ -125,6 +125,24 @@ spec:
x-kubernetes-map-type: atomic
type: object
type: object
credentials:
description: |-
credentials: what happens to the instance credential this operator
generates for the server.
properties:
deletionPolicy:
default: Retain
description: |-
deletionPolicy: whether the instance credential Secret is kept
(Retain, the default) or removed (Delete) when this TerdutServer is
deleted. A kept Secret is only ever adopted after the server accepts
its token, so one left over from a database that has since been reset
is ignored and replaced.
enum:
- Retain
- Delete
type: string
type: object
database:
description: |-
DatabaseSpec is the Postgres connection this TerdutServer uses. Exactly
@@ -324,11 +342,11 @@ spec:
affinity:
description: |-
affinity covers node affinity, pod affinity and pod anti-affinity in
one field -- unlike a multi-replica-aware operator, this one never
generates a default anti-affinity itself (replicas above 1 isn't a
supported topology, see TerdutServerSpec.Replicas's own doc comment),
so this is pure user-supplied passthrough, not a toggle-plus-generated-
default.
one field -- even though replicas now defaults to 2 (see
TerdutServerSpec.Replicas's own doc comment), this operator still
never generates a default anti-affinity of its own the way a
multi-replica-aware operator typically would, so this stays pure
user-supplied passthrough, not a toggle-plus-generated-default.
properties:
nodeAffinity:
description: Describes node affinity scheduling rules for
@@ -4301,11 +4319,15 @@ spec:
type: array
type: object
replicas:
default: 1
default: 2
description: |-
replicas. terdut-server is not horizontally-scale-tested; keep this
at its default of 1 unless you've verified otherwise -- the sweeper
and the notifier are unsynchronised singletons.
replicas. Defaults to 2: terdut-server v0.36.0 put the sweeper, the
notifier and the migration runner each behind a Postgres advisory
lock, and gave incident creation its own conflict resolution, so
more than one replica no longer double-pages, races a migration, or
drops a webhook payload. image.tag must be v0.36.0 or newer for
that to hold -- an older terdut-server has none of these guards,
and this field does not check the tag for you.
format: int32
type: integer
sweeper:
+16 -7
View File
@@ -14,13 +14,22 @@ metadata:
spec:
image:
repository: git.ryuvia.com/niklas/terdut-server
# v0.34.0: fixes callerMayManageServiceAccount so an instance-scoped
# service account can adopt/rotate a key on a team-scoped account it
# didn't just create in the same call -- without this, terdutteam-*
# can wedge permanently on exactly the crash-window race this demo
# hit live (niklas/terdut-operator#3).
tag: v0.34.0
replicas: 1
# v0.36.0 is the floor now that replicas below is 2 (this demo pins
# the current release, v0.43.0, so it shows the current web UI too): that
# release put the sweeper, the notifier and the migration runner each
# behind a Postgres advisory lock, and gave incident creation its own
# conflict resolution, which is what makes a second replica safe
# instead of racing the first. (Still carries v0.34.0's fix too --
# callerMayManageServiceAccount, so an instance-scoped service account
# can adopt/rotate a key on a team-scoped account it didn't just create
# in the same call -- without which terdutteam-* can wedge permanently
# on the crash-window race this demo hit live, niklas/terdut-operator#3.)
tag: v0.43.0
# Matches this CRD's own spec.replicas default (v0.4.0) -- stated
# explicitly, like every other field in this file, rather than left to
# the default. RollingUpdate follows automatically; this operator does
# not expose Strategy as a spec field.
replicas: 2
networking:
hostname: terdut-operator-demo.example
servicePort: 8080
+14
View File
@@ -133,6 +133,20 @@ Each `(team, scenario)` pair is one stable fingerprint, so firing the same
one twice updates the same alert (a real re-fire) and `resolve` closes
exactly that one.
Set `CLUSTER` to send the alert as if it came from one of several clusters:
```sh
CLUSTER=prod-eu ./fire-alerts.sh platform high-cpu
CLUSTER=prod-us ./fire-alerts.sh platform high-cpu # a second incident, not a join
CLUSTER=prod-eu ./fire-alerts.sh platform high-cpu resolve
```
It stands in for a Prometheus external label plus `cluster` in Alertmanager's
`group_by` (terdut-server's README, "Several clusters, one team"): the web UI
then shows the cluster chip on each incident and a cluster filter in the
queue. `CLUSTER` is part of the fingerprint, so resolve with the same value you
fired with. `./run-demo.sh` fires its alerts across `prod-eu` and `prod-us`.
### Dead man's switches
`06-deadman-platform.yaml` / `07-deadman-payments.yaml` expect a heartbeat
+18 -6
View File
@@ -16,7 +16,15 @@
# is the one part of that URL still usable here.
#
# Usage:
# ./fire-alerts.sh <platform|payments> <high-cpu|disk-full|pod-crash|heartbeat> [resolve]
# [CLUSTER=prod-eu] ./fire-alerts.sh <platform|payments> <high-cpu|disk-full|pod-crash|heartbeat> [resolve]
#
# CLUSTER stands in for a Prometheus externalLabel plus `cluster` in
# Alertmanager's group_by (terdut-server's README, "Several clusters, one
# team"): it is put on the alert's labels and on groupLabels, so the incident
# carries it and the web UI shows the cluster chip and the queue's cluster
# filter. It is also part of the group key and the fingerprint, which is what
# keeps the same alert in two clusters from joining one incident. Unset, the
# alert is sent exactly as before.
#
# Prerequisites: kubectl context pointed at the demo namespace, jq, curl,
# and (in another terminal) a running:
@@ -24,6 +32,7 @@
set -euo pipefail
NAMESPACE="${NAMESPACE:-}"
CLUSTER="${CLUSTER:-}"
BASE_URL="${BASE_URL:-http://localhost:8080}"
usage() {
@@ -45,6 +54,8 @@ scenarios:
env vars:
NAMESPACE kubectl -n for reading the webhook Secret (required)
BASE_URL where the port-forwarded terdut-server is (default http://localhost:8080)
CLUSTER optional cluster name, e.g. prod-eu: sent as a `cluster` label and
group label, so the UI shows where the incident came from
EOF
exit 1
}
@@ -90,7 +101,7 @@ key="$(kubectl -n "$NAMESPACE" get secret "$secret_name" -o jsonpath='{.data.key
# created: terdut-server correlates on (team_id, fingerprint), not on
# anything else in the payload. Real Alertmanager computes this from the
# alert's label set; a fixed string plays the same role here.
fingerprint="demo-${team}-${scenario}"
fingerprint="demo-${team}-${scenario}${CLUSTER:+-$CLUSTER}"
now="$(date -u +%Y-%m-%dT%H:%M:%SZ)"
if [ "$status" = firing ]; then
@@ -101,7 +112,8 @@ fi
payload="$(jq -n \
--arg status "$status" \
--arg groupKey "demo:${team}:${scenario}" \
--arg groupKey "demo:${team}:${scenario}${CLUSTER:+:$CLUSTER}" \
--arg cluster "$CLUSTER" \
--arg alertname "$alertname" \
--arg team "$team" \
--arg severity "$severity" \
@@ -113,10 +125,10 @@ payload="$(jq -n \
version: "4",
status: $status,
groupKey: $groupKey,
groupLabels: { alertname: $alertname, team: $team },
groupLabels: ({ alertname: $alertname, team: $team } + (if $cluster != "" then { cluster: $cluster } else {} end)),
alerts: [{
status: $status,
labels: { alertname: $alertname, severity: $severity, team: $team, instance: "demo" },
labels: ({ alertname: $alertname, severity: $severity, team: $team, instance: "demo" } + (if $cluster != "" then { cluster: $cluster } else {} end)),
annotations: { summary: $summary },
startsAt: $startsAt,
endsAt: $endsAt,
@@ -126,7 +138,7 @@ payload="$(jq -n \
}')"
url="${BASE_URL}/api/integrations/${key}/alertmanager"
echo "POST $url (team=$team scenario=$scenario status=$status)" >&2
echo "POST $url (team=$team scenario=$scenario status=$status${CLUSTER:+ cluster=$CLUSTER})" >&2
code="$(curl -sS -o /tmp/fire-alerts-response.json -w '%{http_code}' \
-X POST "$url" -H 'Content-Type: application/json' -d "$payload")"
echo "-> HTTP $code" >&2
+51 -19
View File
@@ -107,26 +107,41 @@ apply_demo() {
kubectl apply -n "$NAMESPACE" -k "$SCRIPT_DIR" >/dev/null
}
wait_for_ready() {
local objects=(
"terdutserver/terdut-operator-demo"
"terdutteam/terdutteam-platform"
"terdutteam/terdutteam-payments"
"terdutescalationrule/terdutescalationrule-platform"
"terdutescalationrule/terdutescalationrule-payments"
"terdutdeadmanswitch/terdutdeadmanswitch-platform"
"terdutdeadmanswitch/terdutdeadmanswitch-payments"
"terdutalertsource/terdutalertsource-platform"
"terdutalertsource/terdutalertsource-payments"
)
wait_for_objects() {
local obj
for obj in "${objects[@]}"; do
for obj in "$@"; do
log "waiting for $obj to become Ready"
kubectl wait --for=condition=Ready --timeout "$WAIT_TIMEOUT" -n "$NAMESPACE" "$obj" >/dev/null \
|| die "timed out waiting for $obj -- try: kubectl describe -n $NAMESPACE $obj"
done
}
# Just the server and the two teams -- everything redeem_platform_invite and
# join_payments_team need. Deliberately NOT the escalation rules here: this
# demo kit's own terdutescalationrule-platform names alice as a level-1
# target, and that CR cannot reach Ready until alice actually exists
# (terdut-server resolves every named username at reconcile time, not just
# at escalation time) -- a real dependency this script has to satisfy by
# creating her first, not something kubectl wait can be told to ignore.
wait_for_teams_ready() {
wait_for_objects \
"terdutserver/terdut-operator-demo" \
"terdutteam/terdutteam-platform" \
"terdutteam/terdutteam-payments"
}
# Everything that was waiting on alice (or just on the teams above, now
# already satisfied) to exist.
wait_for_remaining_ready() {
wait_for_objects \
"terdutescalationrule/terdutescalationrule-platform" \
"terdutescalationrule/terdutescalationrule-payments" \
"terdutdeadmanswitch/terdutdeadmanswitch-platform" \
"terdutdeadmanswitch/terdutdeadmanswitch-payments" \
"terdutalertsource/terdutalertsource-platform" \
"terdutalertsource/terdutalertsource-payments"
}
start_port_forward() {
# A stale pidfile from an earlier run would otherwise collide with us on
# $LOCAL_PORT -- if that pid is still alive, stop it first.
@@ -181,6 +196,18 @@ redeem_platform_invite() {
invite_token="${invite_url##*invite=}"
[ -n "$invite_token" ] || die "could not parse an invite token out of $invite_url"
# A re-run: the invite was spent by the first run, and the server answers a
# spent invite with 403 before it ever looks at the username, so the 409
# handled below never arrives. If alice can already sign in, she exists.
local login_code
login_code="$(curl -sS -o /dev/null -w '%{http_code}' \
-X POST "${BASE_URL}/api/login" -H 'Content-Type: application/json' \
-d "$(jq -n --arg u "$ALICE_USERNAME" --arg p "$DEMO_PASSWORD" '{username: $u, password: $p}')")"
if [ "$login_code" = "200" ]; then
log "account ${ALICE_USERNAME} already exists and can sign in, skipping signup (re-run detected)"
return 0
fi
log "signing up ${ALICE_USERNAME} via Platform's invite"
local body resp_file code
body="$(jq -n \
@@ -235,9 +262,13 @@ join_payments_team() {
fire_demo_alerts() {
log "firing representative demo alerts"
NAMESPACE="$NAMESPACE" BASE_URL="$BASE_URL" "$SCRIPT_DIR/fire-alerts.sh" platform high-cpu
NAMESPACE="$NAMESPACE" BASE_URL="$BASE_URL" "$SCRIPT_DIR/fire-alerts.sh" platform disk-full
NAMESPACE="$NAMESPACE" BASE_URL="$BASE_URL" "$SCRIPT_DIR/fire-alerts.sh" payments pod-crash
# Two clusters, so the queue shows the cluster chip and offers its filter.
# high-cpu fires in both: the same alert in two clusters is two incidents.
local fire="$SCRIPT_DIR/fire-alerts.sh"
CLUSTER=prod-eu NAMESPACE="$NAMESPACE" BASE_URL="$BASE_URL" "$fire" platform high-cpu
CLUSTER=prod-us NAMESPACE="$NAMESPACE" BASE_URL="$BASE_URL" "$fire" platform high-cpu
CLUSTER=prod-us NAMESPACE="$NAMESPACE" BASE_URL="$BASE_URL" "$fire" platform disk-full
CLUSTER=prod-eu NAMESPACE="$NAMESPACE" BASE_URL="$BASE_URL" "$fire" payments pod-crash
}
print_summary() {
@@ -254,8 +285,8 @@ terdut demo is up.
Fire more alerts:
export NAMESPACE=${NAMESPACE} BASE_URL=${BASE_URL}
./fire-alerts.sh platform high-cpu
./fire-alerts.sh platform high-cpu resolve
CLUSTER=prod-eu ./fire-alerts.sh platform high-cpu
CLUSTER=prod-eu ./fire-alerts.sh platform high-cpu resolve
./fire-alerts.sh payments heartbeat # send repeatedly (e.g. every minute)
# to keep a dead man's switch alive;
# stop sending it and, 15 minutes
@@ -319,10 +350,11 @@ main() {
ensure_kind_cluster
install_operator
apply_demo
wait_for_ready
wait_for_teams_ready
start_port_forward
redeem_platform_invite
join_payments_team
wait_for_remaining_ready
fire_demo_alerts
print_summary
+75 -10
View File
@@ -16,18 +16,24 @@ import (
)
// bootstrapStateLostError is DESIGN.md §6's one genuinely pathological
// case: a checkpointed admin credential was used and then lost before the
// lasting credential it was for could be persisted. Distinct from a plain
// error so Reconcile can route it to a Ready: False condition (the
// documented recovery is delete-and-recreate, not an automatic retry) rather
// than treating it as a transient reconcile failure.
type bootstrapStateLostError struct{ detail string }
// case: the server's database is already bootstrapped, and no credential for
// it survives here -- a checkpointed admin key was used and then lost before
// the lasting credential could be persisted, or the instance credential
// Secret was deleted (spec.credentials.deletionPolicy: Delete, or by hand).
// Distinct from a plain error so Reconcile can route it to a Ready: False
// condition rather than treating it as a transient reconcile failure: no
// retry can fix it.
type bootstrapStateLostError struct {
detail string
secretName string // the instance credential Secret that would have fixed it
}
func (e *bootstrapStateLostError) Error() string {
return fmt.Sprintf(
"server reports already bootstrapped, but neither status.credentialsSecretRef nor a "+
"checkpointed admin credential exist here: %s. This TerdutServer cannot recover a "+
"credential on its own; delete and recreate it", e.detail)
"server reports already bootstrapped, but there is no credential for it here: %s. "+
"The operator cannot mint one on its own, and deleting and recreating this TerdutServer does "+
"not clear the database. Restore Secret %q in the operator's namespace if you have a copy, "+
"or reset the server's database and recreate this TerdutServer", e.detail, e.secretName)
}
// reconcileBootstrap implements DESIGN.md §6 point 1's self-registration
@@ -36,6 +42,17 @@ func (e *bootstrapStateLostError) Error() string {
// srv.Status.CredentialsSecretRef is nil and the Deployment has a ready
// replica.
func (r *TerdutServerReconciler) reconcileBootstrap(ctx context.Context, srv *terdutv1alpha1.TerdutServer) error {
// A TerdutServer recreated against a database that is already
// bootstrapped: its instance credential may still be here (the default
// deletion policy keeps it), and then there is nothing to bootstrap.
adopted, err := r.adoptRetainedCredentials(ctx, srv)
if err != nil {
return err
}
if adopted {
return nil
}
adminKey, err := r.getOrCreateCheckpointedAdminKey(ctx, srv)
if err != nil {
return err
@@ -64,6 +81,51 @@ func (r *TerdutServerReconciler) reconcileBootstrap(ctx context.Context, srv *te
return nil
}
// adoptRetainedCredentials looks for the instance credential Secret an
// earlier TerdutServer of the same name and namespace left behind
// (spec.credentials.deletionPolicy: Retain), and adopts it if the server
// still accepts its token. Reports whether it did.
//
// The token is tried before it is trusted: a Secret that outlived a database
// reset holds a key the server has never heard of, and adopting that would
// make every later call 401. A rejected token (401/403) is not an error here;
// it just means there is nothing to adopt, and bootstrap proceeds as it would
// for a first install -- which succeeds against a freshly reset database and
// replaces the Secret. Anything else (the server unreachable, a 5xx) is
// returned for a retry rather than guessed at.
func (r *TerdutServerReconciler) adoptRetainedCredentials(ctx context.Context, srv *terdutv1alpha1.TerdutServer) (bool, error) {
name := credentialsSecretName(srv)
var sec corev1.Secret
if err := r.Get(ctx, client.ObjectKey{Namespace: r.OperatorNamespace, Name: name}, &sec); err != nil {
if apierrors.IsNotFound(err) {
return false, nil
}
return false, err
}
token := string(sec.Data[credentialsSecretDataKey])
if token == "" {
return false, nil
}
// The operator's own service account is the one thing this credential
// is for, so asking the server for it by name checks the token and that
// it belongs to that account in the same call.
sa, err := r.NewClient(serviceURL(srv)).WithToken(token).GetServiceAccountByName(ctx, serviceAccountName)
if err != nil {
if statusErr, ok := errors.AsType[*tdclient.StatusError](err); ok &&
(statusErr.Code == http.StatusUnauthorized || statusErr.Code == http.StatusForbidden) {
return false, nil
}
return false, fmt.Errorf("checking the retained credential in Secret %q: %w", name, err)
}
if sa == nil {
return false, nil
}
srv.Status.CredentialsSecretRef = &terdutv1alpha1.SecretKeyRef{Name: name, Key: credentialsSecretDataKey}
return true, nil
}
// getOrCreateCheckpointedAdminKey returns a usable admin key: from the
// checkpoint Secret if an earlier, interrupted attempt already got one, or
// freshly from /api/bootstrap, immediately checkpointed before it's used
@@ -89,7 +151,10 @@ func (r *TerdutServerReconciler) getOrCreateCheckpointedAdminKey(ctx context.Con
// status.credentialsSecretRef) means a prior reconcile already
// won this exact race and its checkpoint was lost afterward --
// the one case §6 doesn't try to paper over.
return "", &bootstrapStateLostError{detail: "/api/bootstrap returned 403"}
return "", &bootstrapStateLostError{
detail: "/api/bootstrap returned 403",
secretName: credentialsSecretName(srv),
}
}
return "", fmt.Errorf("POST /api/bootstrap: %w", err)
}
+23 -10
View File
@@ -39,10 +39,10 @@ const resyncInterval = 5 * time.Minute
// than "someone edited something out of band."
const waitInterval = 15 * time.Second
// finalizerName cleans up the credentials Secret(s) this controller
// generates in the operator's own namespace on delete — the Deployment and
// Service are owned (OwnerReference, DESIGN.md §7) and need no finalizer of
// their own.
// finalizerName cleans up the Secret(s) this controller generates in the
// operator's own namespace on delete (which of them, spec.credentials.
// deletionPolicy decides) — the Deployment and Service are owned
// (OwnerReference, DESIGN.md §7) and need no finalizer of their own.
const finalizerName = "terdut.ryuvia.com/terdutserver"
// serviceAccountName is the name the operator registers itself under
@@ -215,18 +215,31 @@ func (r *TerdutServerReconciler) setNotReady(
return ctrl.Result{RequeueAfter: d}, nil
}
// reconcileDelete cleans up the credentials Secret(s) this controller
// generated in the operator's own namespace. The Deployment and Service are
// owned (OwnerReference, DESIGN.md §7) and need no attention here — normal
// GC handles them. There is no server-side "delete this install" call to
// reconcileDelete cleans up the Secrets this controller generated in the
// operator's own namespace. The Deployment and Service are owned
// (OwnerReference, DESIGN.md §7) and need no attention here — normal GC
// handles them. There is no server-side "delete this install" call to
// make: bootstrap created a user and a service account, and terdut-server's
// API has no way to delete either (only to revoke individual keys), so
// there is nothing meaningful to undo there either.
// there is nothing meaningful to undo there either, and the database is
// never touched.
//
// That is why the instance credential is kept by default
// (spec.credentials.deletionPolicy: Retain). Everything it logs in to
// outlives the TerdutServer, so a recreated one finds a bootstrapped server
// it has no key for -- unless the key is still here to be adopted (see
// adoptRetainedCredentials). The bootstrap checkpoint is always removed: it
// is a short-lived admin key, and the instance credential is all that is
// needed afterwards.
func (r *TerdutServerReconciler) reconcileDelete(ctx context.Context, srv *terdutv1alpha1.TerdutServer) (ctrl.Result, error) {
if !controllerutil.ContainsFinalizer(srv, finalizerName) {
return ctrl.Result{}, nil
}
for _, name := range []string{checkpointSecretName(srv), credentialsSecretName(srv)} {
names := []string{checkpointSecretName(srv)}
if srv.Spec.Credentials.DeletionPolicy == terdutv1alpha1.CredentialsDelete {
names = append(names, credentialsSecretName(srv))
}
for _, name := range names {
sec := &corev1.Secret{ObjectMeta: metav1.ObjectMeta{Name: name, Namespace: r.OperatorNamespace}}
if err := r.Delete(ctx, sec); err != nil && !apierrors.IsNotFound(err) {
return ctrl.Result{}, err
@@ -54,6 +54,10 @@ type fakeTerdutServer struct {
mu sync.Mutex
bootstrapped bool
bootstrap403 bool // force every /api/bootstrap call to 403, even the first
// rejectedTokens: bearer tokens the fake answers 401 on GET
// /api/service-accounts, the way a server that never issued the key
// (a database reset since) would.
rejectedTokens map[string]bool
nextID int64
accounts map[string]int64 // name -> id
keyMints map[int64]int // id -> number of keys minted so far
@@ -170,6 +174,10 @@ func (f *fakeTerdutServer) ServeHTTP(w http.ResponseWriter, r *http.Request) {
})
case r.URL.Path == "/api/service-accounts" && r.Method == http.MethodGet:
if f.rejectedTokens[strings.TrimPrefix(r.Header.Get("Authorization"), "Bearer ")] {
writeJSON(w, http.StatusUnauthorized, map[string]string{errJSONKey: "invalid or expired API key"})
return
}
name := r.URL.Query().Get("name")
id, exists := f.accounts[name]
if !exists {
@@ -875,13 +883,14 @@ var _ = Describe("TerdutServer Controller", func() {
})
Describe("deletion", func() {
It("removes the credentials and checkpoint Secrets and the finalizer", func(ctx SpecContext) {
fake, fakeSrv := newFakeTerdutServer()
_ = fake
// runToReady brings a server to Ready against a fresh fake and
// returns the name of the instance credential Secret it minted.
runToReady := func(ctx SpecContext, spec terdutv1alpha1.TerdutServerSpec) string {
_, fakeSrv := newFakeTerdutServer()
DeferCleanup(fakeSrv.Close)
reconciler.NewClient = func(string) *tdclient.Client { return tdclient.New(fakeSrv.URL) }
createServer(ctx, dsnSpec())
createServer(ctx, spec)
reconcileOnce(ctx)
reconcileOnce(ctx)
markDeploymentReady(ctx)
@@ -889,17 +898,114 @@ var _ = Describe("TerdutServer Controller", func() {
srv := &terdutv1alpha1.TerdutServer{}
Expect(k8sClient.Get(ctx, objKey, srv)).To(Succeed())
credsName := srv.Status.CredentialsSecretRef.Name
return srv.Status.CredentialsSecretRef.Name
}
deleteAndFinalize := func(ctx SpecContext) {
srv := &terdutv1alpha1.TerdutServer{}
Expect(k8sClient.Get(ctx, objKey, srv)).To(Succeed())
Expect(k8sClient.Delete(ctx, srv)).To(Succeed())
reconcileOnce(ctx) // runs the finalizer
Expect(k8sClient.Get(ctx, objKey, srv)).NotTo(Succeed(),
"the TerdutServer itself should be gone once the finalizer clears")
}
secretExists := func(ctx SpecContext, secretName string) bool {
var s corev1.Secret
err := k8sClient.Get(ctx, types.NamespacedName{Name: secretName, Namespace: operatorNamespace}, &s)
return err == nil
}
err := k8sClient.Get(ctx, objKey, srv)
Expect(err).To(HaveOccurred(), "the TerdutServer itself should be gone once the finalizer clears")
It("keeps the instance credential by default and removes the checkpoint", func(ctx SpecContext) {
credsName := runToReady(ctx, dsnSpec())
// A checkpoint left behind, as if the best-effort delete had failed.
Expect(k8sClient.Create(ctx, &corev1.Secret{
ObjectMeta: metav1.ObjectMeta{Name: checkpointSecretNameFor(name), Namespace: operatorNamespace},
Data: map[string][]byte{credentialsSecretDataKey: []byte("admin-key-raw")},
})).To(Succeed())
var leftover corev1.Secret
err = k8sClient.Get(ctx, types.NamespacedName{Name: credsName, Namespace: operatorNamespace}, &leftover)
Expect(err).To(HaveOccurred(), "the credentials Secret should have been cleaned up by the finalizer")
deleteAndFinalize(ctx)
Expect(secretExists(ctx, credsName)).To(BeTrue(),
"the credential is kept so a recreated TerdutServer can adopt it")
Expect(secretExists(ctx, checkpointSecretNameFor(name))).To(BeFalse(),
"the checkpoint is a short-lived admin key and is always removed")
})
It("removes the credentials and checkpoint Secrets and the finalizer when the policy is Delete", func(ctx SpecContext) {
spec := dsnSpec()
spec.Credentials.DeletionPolicy = terdutv1alpha1.CredentialsDelete
credsName := runToReady(ctx, spec)
deleteAndFinalize(ctx)
Expect(secretExists(ctx, credsName)).To(BeFalse(),
"the credentials Secret should have been cleaned up by the finalizer")
})
})
Describe("recreating a TerdutServer against an already-bootstrapped server", func() {
// retainedSecret stands in for what the previous TerdutServer left.
retainedSecret := func(ctx SpecContext, token string) {
Expect(k8sClient.Create(ctx, &corev1.Secret{
ObjectMeta: metav1.ObjectMeta{Name: credentialsSecretNameFor(name), Namespace: operatorNamespace},
Data: map[string][]byte{credentialsSecretDataKey: []byte(token)},
})).To(Succeed())
}
bringUp := func(ctx SpecContext, fakeSrv *httptest.Server) {
DeferCleanup(fakeSrv.Close)
reconciler.NewClient = func(string) *tdclient.Client { return tdclient.New(fakeSrv.URL) }
createServer(ctx, dsnSpec())
reconcileOnce(ctx) // finalizer
reconcileOnce(ctx) // Deployment/Service
markDeploymentReady(ctx)
reconcileOnce(ctx)
}
credsToken := func(ctx SpecContext) string {
var s corev1.Secret
Expect(k8sClient.Get(ctx, types.NamespacedName{Name: credentialsSecretNameFor(name), Namespace: operatorNamespace}, &s)).To(Succeed())
return string(s.Data[credentialsSecretDataKey])
}
It("adopts the retained credential when the server still accepts it, without bootstrapping", func(ctx SpecContext) {
fake, fakeSrv := newFakeTerdutServer()
fake.bootstrapped = true // every /api/bootstrap call 403s
fake.accounts[serviceAccountName] = 7
retainedSecret(ctx, "retained-key")
bringUp(ctx, fakeSrv)
Expect(readyCondition(ctx).Status).To(Equal(metav1.ConditionTrue))
srv := &terdutv1alpha1.TerdutServer{}
Expect(k8sClient.Get(ctx, objKey, srv)).To(Succeed())
Expect(srv.Status.CredentialsSecretRef).NotTo(BeNil())
Expect(srv.Status.CredentialsSecretRef.Name).To(Equal(credentialsSecretNameFor(name)))
Expect(credsToken(ctx)).To(Equal("retained-key"), "the retained key is reused, not replaced")
})
It("ignores a retained credential the server rejects and bootstraps afresh after a database reset", func(ctx SpecContext) {
fake, fakeSrv := newFakeTerdutServer() // a fresh database: bootstrap succeeds
fake.rejectedTokens = map[string]bool{"stale-key": true}
retainedSecret(ctx, "stale-key")
bringUp(ctx, fakeSrv)
Expect(readyCondition(ctx).Status).To(Equal(metav1.ConditionTrue))
Expect(credsToken(ctx)).NotTo(Equal("stale-key"), "the stale credential is replaced by a freshly minted one")
})
It("fails closed, naming the Secret to restore, when the server is bootstrapped and the retained credential is rejected", func(ctx SpecContext) {
fake, fakeSrv := newFakeTerdutServer()
fake.bootstrapped = true
fake.rejectedTokens = map[string]bool{"stale-key": true}
retainedSecret(ctx, "stale-key")
bringUp(ctx, fakeSrv)
cond := readyCondition(ctx)
Expect(cond.Status).To(Equal(metav1.ConditionFalse))
Expect(cond.Reason).To(Equal(terdutv1alpha1.ReasonBootstrapStateLost))
Expect(cond.Message).To(ContainSubstring(credentialsSecretNameFor(name)),
"the message names the Secret that would restore it")
Expect(cond.Message).To(ContainSubstring("does not clear the database"))
})
})
+16 -6
View File
@@ -37,17 +37,27 @@ func (r *TerdutServerReconciler) reconcileDeployment(
_, err := controllerutil.CreateOrUpdate(ctx, r.Client, deploy, func() error {
replicas := srv.Spec.Replicas
if replicas == 0 {
replicas = 1
// Only reachable for a TerdutServer stored before the
// +kubebuilder:default=2 marker existed -- the API server's own
// CRD defaulting fills this in for anything created or updated
// through it, so a fresh zero value here means a pre-existing
// object that predates the default, not a deliberate "none"
// (there is no way to request zero replicas).
replicas = 2
}
labels := labelsFor(srv)
deploy.Spec.Replicas = &replicas
deploy.Spec.Selector = &metav1.LabelSelector{MatchLabels: labels}
// Recreate, not RollingUpdate: the sweeper and the notifier are
// unsynchronised singletons inside terdut-server, and two replicas
// overlapping during a rollout would both page for the same
// incident (matches the chart's own deployment.yaml comment).
deploy.Spec.Strategy = appsv1.DeploymentStrategy{Type: appsv1.RecreateDeploymentStrategyType}
// RollingUpdate, not Recreate: terdut-server v0.36.0 put the sweeper,
// the notifier and the migration runner each behind a Postgres
// advisory lock, and gave incident creation its own conflict
// resolution, so two replicas overlapping during a rollout no longer
// double-page, race a migration, or drop a webhook payload (matches
// the chart's own deployment.yaml comment). No explicit
// maxUnavailable/maxSurge: left at the 25%/25% default, which rounds
// to 0/1 at the default replicas: 2 -- already zero-downtime.
deploy.Spec.Strategy = appsv1.DeploymentStrategy{Type: appsv1.RollingUpdateDeploymentStrategyType}
pod := srv.Spec.Pod
deploy.Spec.Template = corev1.PodTemplateSpec{
// pod.Annotations is assigned directly, not merged -- nothing