10 Commits

Author SHA1 Message Date
Niklas Ye e2c4475867 Build and scan with Go 1.26.9
govulncheck in the security job reports ten standard-library
vulnerabilities (net/http, mime/multipart, crypto/tls), all fixed in
1.26.9. The workflows pinned golang:1.26.6-bookworm.

Claude-Session: https://claude.ai/code/session_016mBLURvJoMuUEr9cB2RpUN
2026-10-09 15:17:03 +02:00
Niklas Ye e1103f2b7d Authenticate with a seeded operator key; fold escalation and switches into TerdutTeam
Credentials: the TerdutServer controller generates <name>-operator-key in
the server's own namespace (owned by it) and hands it to the pods as
TERDUT_OPERATOR_KEY; the server creates its instance-scoped account from it
at every start. A replaced Secret rolls the pods. The bootstrap handshake,
the checkpoint Secret, per-team service accounts and credentials Secrets,
BootstrapStateLost and credentials.deletionPolicy are gone.

CRDs: TerdutServer, TerdutTeam and TerdutAlertSource. TerdutEscalationRule
and TerdutDeadmanSwitch become spec.escalation and spec.deadmanSwitches[]
on the team (matched by name, extras removed); team invites are removed.
A team is created under the identity <namespace>/<name> (external_id), so a
retry, a lost status or a deleted team heal by repeating the same call, and
a display name owned by another team is TeamNameTaken instead of an
adoption. The server resolves escalation usernames (UnknownUser condition).
OIDC claim names and trustEmail are spec fields.

Fixes: query values are URL-escaped; every delete treats 404 as success;
deleting a team no longer depends on allowedTeams consent; a switch or
integration deleted on the server is recreated; unnamed switches take the
CR's name.

Cleanup: scaffold e2e test, AGENTS.md, devcontainer, unused config/ pieces
and Client.Version() removed; DESIGN.md, README, ROADMAP and the demo
(run-demo.sh, manifests) rewritten for the new design. Secret RBAC stays
cluster-wide, now stated in DESIGN.md section 9.

Claude-Session: https://claude.ai/code/session_016mBLURvJoMuUEr9cB2RpUN
2026-10-09 14:56:22 +02:00
Niklas Ye b0a431f2a4 Set the chart's placeholder version to 0.5.0
CI / chart (push) Successful in 2s
CI / security (push) Successful in 53s
Release / chart (push) Successful in 3s
CI / test (push) Successful in 2m19s
Release / test (push) Successful in 1m41s
Release / image (push) Successful in 6m30s
Release / scan-image (push) Successful in 4s
Cosmetic: make helm-package passes --version and --app-version from the
tag, so these two fields decide nothing about what gets published. Still
done, as with d502485 (0.4.0) and 88172ad (0.3.0) before it, because a
tree heading for v0.5.0 that still says 0.4.0 tells its reader
something false.

Co-authored-by: Claude <noreply@anthropic.com>
2026-10-08 21:10:27 +02:00
Niklas Ye 62664c93ff Keep the instance credential across a TerdutServer delete, and adopt it on recreate
Deleting a TerdutServer removed the credential Secrets but never touched the
database, so a recreated one found a server that was already bootstrapped and
no key for it: /api/bootstrap answered 403 and the operator stopped at
BootstrapStateLost, whose message and DESIGN.md both said "delete and
recreate". That is how the terdut-demo install on the cluster got stuck on
2026-10-03: Helm's cleanupOnFail deleted its TerdutServer after a failed
upgrade, the recreate found the bootstrapped database, and it sat at Ready:
False for five days until the database was reset by hand. Recreating cannot
fix it, because the finalizer clears Secrets and the database is not its to
reset, so "a fresh create starts clean" was only ever true when the database
went with it.

spec.credentials.deletionPolicy is Retain by default: the finalizer keeps the
instance credential Secret (Delete removes it, as before). The bootstrap
checkpoint is always removed. Before calling /api/bootstrap, reconcile now
looks for the retained Secret and asks the server for the operator's own
service account with its token. Accepted: adopt it and skip bootstrap.
Rejected with 401/403: the Secret outlived a database reset, so ignore it and
bootstrap like a first install, which replaces it. Any other error retries.
terdut-server's own tests already call that endpoint with an instance-scoped
key, so the permission is not new.

BootstrapStateLost is still the answer when the server is bootstrapped and no
credential it accepts survives, but its message now names the Secret to
restore and says that recreating does not clear the database. DESIGN.md §6
says the same, and the chart passes the setting through as
terdutServer.credentials.deletionPolicy.

A retained Secret of a TerdutServer that is gone for good is an orphan to
delete by hand. It is inert: nothing adopts it unless the server accepts the
token.

Checked on the kind demo with a locally built image against the real
terdut-server v0.43.0: deleting the TerdutServer kept the Secret, recreating it
reached Ready with the same credential (identical hash) and both TerdutTeams
came back Ready with their original ids. The controller specs cover adoption,
a rejected token after a reset, the bootstrapped-and-rejected failure, and
both deletion policies.

Co-authored-by: Claude <noreply@anthropic.com>
2026-10-08 21:09:56 +02:00
niklas 6572f63157 Merge pull request 'examples/demo: run terdut-server v0.43.0, with alerts from two clusters' (#6) from demo-matches-v0.4.0-crd into main
CI / chart (push) Successful in 2s
CI / security (push) Successful in 51s
CI / test (push) Successful in 2m26s
Reviewed-on: #6
2026-10-08 17:10:06 +00:00
Niklas Ye 50ce5bcec0 examples/demo: run terdut-server v0.43.0, with alerts from two clusters
CI / test (pull_request) Successful in 6m35s
CI / chart (pull_request) Successful in 2s
CI / security (pull_request) Successful in 1m6s
The demo pinned v0.36.0, the floor for replicas: 2, and so showed none of the
web UI since: the queue and incident layouts, the rota and escalation
pages, the theme toggle, and the cluster chip, filter and page titles
(v0.42.0-v0.43.0). It pins v0.43.0 now; the comment keeps v0.36.0 as the
floor, which is what the replicas setting actually depends on.

fire-alerts.sh takes an optional CLUSTER, standing in for a Prometheus
external label plus `cluster` in Alertmanager's group_by (terdut-server's
README, "Several clusters, one team"). It goes on the alert's labels and
groupLabels, and into the group key and the fingerprint, so the same alert in
two clusters is two incidents and not one. Unset, the payload is exactly what
it was. run-demo.sh fires its alerts across prod-eu and prod-us, high-cpu in
both, so the queue has a chip and a filter to show.

run-demo.sh also failed on its second run, though it says it is safe to
re-run: it expected HTTP 409 when alice already exists, but a spent invite
is answered with 403 "invite link is not usable" before the username is ever
checked. It now tries to log alice in first and skips the signup if that works.

Checked on the kind cluster: the server rolled to v0.43.0, every CR became
Ready and Adopted (server, both teams, both escalation rules, both dead man's
switches, both alert sources), and /api/incidents/clusters,
/api/incidents?cluster=prod-us and the incident titles came back as expected.
No operator code changed, so this needs no operator release.

Co-authored-by: Claude <noreply@anthropic.com>
2026-10-08 18:55:02 +02:00
Niklas Ye 822c80dda6 examples/demo: split the Ready wait so escalation rules wait on alice too
wait_for_ready waited for every demo object at once, including
terdutescalationrule-platform, which names alice as a level-1 target --
but alice does not exist yet at that point in main(): she is created by
redeem_platform_invite, which ran after wait_for_ready. terdut-server
resolves every named username at reconcile time, not just when an
escalation actually fires, so that CR could never reach Ready before
alice did, and main() had no step in between to create her.

Split into wait_for_objects (the shared loop, now taking its object list
as arguments) plus two callers: wait_for_teams_ready, covering just the
server and the two teams redeem_platform_invite/join_payments_team
need, run before alice exists; wait_for_remaining_ready, covering the
escalation rules, dead man's switches and alert sources, run after.

Co-authored-by: Claude <noreply@anthropic.com>
2026-10-03 18:15:42 +02:00
Niklas Ye 0ee7ede648 examples/demo: match v0.4.0's new spec.replicas default and bump terdut-server
replicas: 1 and tag: v0.34.0 were both correct when written, but the CRD's
own default moved to 2 in v0.4.0 (same release this demo is meant to show
off), and v0.34.0 predates v0.36.0's advisory locks that make a second
replica safe instead of racing the first. Left as-is, the demo would have
been the one place in this repo demonstrating the exact unsafe combination
the CRD's own doc comment warns against: more than one replica against an
image that doesn't guard the sweeper/notifier/migration-runner singletons.

replicas is now stated explicitly as 2 rather than dropped to pick up the
default silently, matching every other field in this file's own habit of
spelling out what it depends on. tag moves to v0.36.0 specifically -- the
first version where the lock landed -- with the comment keeping v0.34.0's
original reasoning (the service-account race fix) alongside the new one,
since v0.36.0 still carries that fix forward.

Co-authored-by: Claude <noreply@anthropic.com>
2026-10-03 18:13:09 +02:00
Niklas Ye d50248531c Set the chart's placeholder version to 0.4.0
CI / chart (push) Successful in 1s
CI / security (push) Successful in 59s
CI / test (push) Successful in 2m10s
Release / test (push) Successful in 7m54s
Release / chart (push) Successful in 2s
Release / image (push) Successful in 6m25s
Release / scan-image (push) Successful in 36s
2026-10-03 16:22:15 +02:00
Niklas Ye 4007f54279 Default TerdutServer.spec.replicas to 2 and switch to RollingUpdate
CI / chart (push) Successful in 1s
CI / security (push) Successful in 59s
CI / test (push) Has been cancelled
Mirrors charts/terdut-server's own deployment.yaml change: v0.36.0 put
the sweeper, the notifier and the migration runner each behind a
Postgres advisory lock, and gave incident creation its own conflict
resolution, so the Recreate strategy and replicas-stays-at-1 guidance
this controller carried (explicitly tracking that chart's comment)
are no longer load-bearing.

spec.replicas' +kubebuilder:default moves 1 -> 2 (config/crd/bases and
the chart's CRD template regenerated via controller-gen and
kubebuilder's helm plugin respectively, then hand-verified identical
to the generator's own output rather than trusting a bulk regen --
the plugin's --output-dir charts writes a fresh charts/chart scaffold
rather than updating charts/terdut-operator in place, so only the
diff was taken, not the whole tree). terdutserver_deployment.go's
same-value fallback (reachable only for a TerdutServer stored before
this default existed) moves with it, and its Strategy changes from
Recreate to RollingUpdate with no explicit maxUnavailable/maxSurge --
the 25%/25% default rounds to 0/1 at replicas: 2, already
zero-downtime.

DESIGN.md's three places asserting multi-replica isn't a supported
topology (the illustrative spec.replicas YAML, spec.pod.affinity's
rationale, and the HPA deferred-feature note) are corrected to match;
the HPA note now gives its own standing reason (no scaling metric or
bounds decided yet) rather than a contradiction that no longer holds.

The chart's optional terdutServer.replicas sample value moves 1 -> 2
alongside it. image.tag must be v0.36.0 or newer for any of this to
hold -- stated in both the CRD field's doc comment and the chart
value's comment, not enforced in code, same stance the chart takes on
every other version-coupled assumption.

Co-authored-by: Claude <noreply@anthropic.com>
2026-10-03 12:35:18 +02:00
95 changed files with 2231 additions and 7304 deletions
-35
View File
@@ -1,35 +0,0 @@
{
"name": "Kubebuilder DevContainer",
"image": "golang:1.26",
"features": {
"ghcr.io/devcontainers/features/docker-in-docker:2": {
"moby": false,
"dockerDefaultAddressPool": "base=172.30.0.0/16,size=24"
},
"ghcr.io/devcontainers/features/git:1": {},
"ghcr.io/devcontainers/features/common-utils:2": {
"upgradePackages": true
}
},
"runArgs": ["--privileged", "--init"],
"customizations": {
"vscode": {
"settings": {
"terminal.integrated.shell.linux": "/bin/bash"
},
"extensions": [
"ms-kubernetes-tools.vscode-kubernetes-tools",
"ms-azuretools.vscode-docker"
]
}
},
"remoteEnv": {
"GO111MODULE": "on"
},
"onCreateCommand": "bash .devcontainer/post-install.sh"
}
-153
View File
@@ -1,153 +0,0 @@
#!/bin/bash
set -euo pipefail
echo "===================================="
echo "Kubebuilder DevContainer Setup"
echo "===================================="
# Verify running as root (required for installing to /usr/local/bin and /etc)
if [ "$(id -u)" -ne 0 ]; then
echo "ERROR: This script must be run as root"
exit 1
fi
echo ""
echo "Detecting system architecture..."
# Detect architecture using uname
MACHINE=$(uname -m)
case "${MACHINE}" in
x86_64)
ARCH="amd64"
;;
aarch64|arm64)
ARCH="arm64"
;;
*)
echo "WARNING: Unsupported architecture ${MACHINE}, defaulting to amd64"
ARCH="amd64"
;;
esac
echo "Architecture: ${ARCH}"
echo ""
echo "------------------------------------"
echo "Setting up bash completion..."
echo "------------------------------------"
BASH_COMPLETIONS_DIR="/usr/share/bash-completion/completions"
# Enable bash-completion in root's .bashrc (devcontainer runs as root)
if ! grep -q "source /usr/share/bash-completion/bash_completion" ~/.bashrc 2>/dev/null; then
echo 'source /usr/share/bash-completion/bash_completion' >> ~/.bashrc
echo "Added bash-completion to .bashrc"
fi
echo ""
echo "------------------------------------"
echo "Installing development tools..."
echo "------------------------------------"
# Install kind
if ! command -v kind &> /dev/null; then
echo "Installing kind..."
curl -Lo /usr/local/bin/kind "https://kind.sigs.k8s.io/dl/latest/kind-linux-${ARCH}"
chmod +x /usr/local/bin/kind
echo "kind installed successfully"
fi
# Generate kind bash completion
if command -v kind &> /dev/null; then
if kind completion bash > "${BASH_COMPLETIONS_DIR}/kind" 2>/dev/null; then
echo "kind completion installed"
else
echo "WARNING: Failed to generate kind completion"
fi
fi
# Install kubebuilder
if ! command -v kubebuilder &> /dev/null; then
echo "Installing kubebuilder..."
curl -Lo /usr/local/bin/kubebuilder "https://go.kubebuilder.io/dl/latest/linux/${ARCH}"
chmod +x /usr/local/bin/kubebuilder
echo "kubebuilder installed successfully"
fi
# Generate kubebuilder bash completion
if command -v kubebuilder &> /dev/null; then
if kubebuilder completion bash > "${BASH_COMPLETIONS_DIR}/kubebuilder" 2>/dev/null; then
echo "kubebuilder completion installed"
else
echo "WARNING: Failed to generate kubebuilder completion"
fi
fi
# Install kubectl
if ! command -v kubectl &> /dev/null; then
echo "Installing kubectl..."
KUBECTL_VERSION=$(curl -Ls https://dl.k8s.io/release/stable.txt)
curl -Lo /usr/local/bin/kubectl "https://dl.k8s.io/release/${KUBECTL_VERSION}/bin/linux/${ARCH}/kubectl"
chmod +x /usr/local/bin/kubectl
echo "kubectl installed successfully"
fi
# Generate kubectl bash completion
if command -v kubectl &> /dev/null; then
if kubectl completion bash > "${BASH_COMPLETIONS_DIR}/kubectl" 2>/dev/null; then
echo "kubectl completion installed"
else
echo "WARNING: Failed to generate kubectl completion"
fi
fi
# Generate Docker bash completion
if command -v docker &> /dev/null; then
if docker completion bash > "${BASH_COMPLETIONS_DIR}/docker" 2>/dev/null; then
echo "docker completion installed"
else
echo "WARNING: Failed to generate docker completion"
fi
fi
echo ""
echo "------------------------------------"
echo "Configuring Docker environment..."
echo "------------------------------------"
# Wait for Docker to be ready
echo "Waiting for Docker to be ready..."
for i in {1..30}; do
if docker info >/dev/null 2>&1; then
echo "Docker is ready"
break
fi
if [ "$i" -eq 30 ]; then
echo "WARNING: Docker not ready after 30s"
fi
sleep 1
done
# Create kind network (ignore if already exists)
if ! docker network inspect kind >/dev/null 2>&1; then
if docker network create kind >/dev/null 2>&1; then
echo "Created kind network"
else
echo "WARNING: Failed to create kind network (may already exist)"
fi
fi
echo ""
echo "------------------------------------"
echo "Verifying installations..."
echo "------------------------------------"
kind version
kubebuilder version
kubectl version --client
docker --version
go version
echo ""
echo "===================================="
echo "DevContainer ready!"
echo "===================================="
echo "All development tools installed successfully."
echo "You can now start building Kubernetes operators."
+2 -2
View File
@@ -55,7 +55,7 @@ jobs:
test:
runs-on: ubuntu-latest
container:
image: golang:1.26.6-bookworm
image: golang:1.26.9-bookworm
volumes:
- go-mod-cache:/go/pkg/mod
- go-build-cache:/root/.cache/go-build
@@ -85,7 +85,7 @@ jobs:
security:
runs-on: ubuntu-latest
container:
image: golang:1.26.6-bookworm
image: golang:1.26.9-bookworm
volumes:
- go-mod-cache:/go/pkg/mod
- go-build-cache:/root/.cache/go-build
+1 -1
View File
@@ -26,7 +26,7 @@ jobs:
test:
runs-on: ubuntu-latest
container:
image: golang:1.26.6-bookworm
image: golang:1.26.9-bookworm
volumes:
- go-mod-cache:/go/pkg/mod
- go-build-cache:/root/.cache/go-build
-320
View File
@@ -1,320 +0,0 @@
# terdut-operator - AI Agent Guide
## Project Structure
**Single-group layout (default):**
```
cmd/main.go Manager entry (registers controllers/webhooks)
api/<version>/*_types.go CRD schemas (+kubebuilder markers)
api/<version>/zz_generated.* Auto-generated (DO NOT EDIT)
internal/controller/* Reconciliation logic
internal/webhook/* Validation/defaulting (if present)
config/crd/bases/* Generated CRDs (DO NOT EDIT)
config/rbac/role.yaml Generated RBAC (DO NOT EDIT)
config/samples/* Example CRs (edit these)
Makefile Build/test/deploy commands
PROJECT Kubebuilder metadata Auto-generated (DO NOT EDIT)
```
**Multi-group layout** (for projects with multiple API groups):
```
api/<group>/<version>/*_types.go CRD schemas by group
internal/controller/<group>/* Controllers by group
internal/webhook/<group>/<version>/* Webhooks by group and version (if present)
```
Multi-group layout organizes APIs by group name (e.g., `batch`, `apps`). Check the `PROJECT` file for `multigroup: true`.
**To convert to multi-group layout:**
1. Run: `kubebuilder edit --multigroup=true`
2. Move APIs: `mkdir -p api/<group> && mv api/<version> api/<group>/`
3. Move controllers: `mkdir -p internal/controller/<group> && mv internal/controller/*.go internal/controller/<group>/`
4. Move webhooks (if present): `mkdir -p internal/webhook/<group> && mv internal/webhook/<version> internal/webhook/<group>/`
5. Update import paths in all files
6. Fix `path` in `PROJECT` file for each resource
7. Update test suite CRD paths (add one more `..` to relative paths)
## Critical Rules
### Never Edit These (Auto-Generated)
- `config/crd/bases/*.yaml` - from `make manifests`
- `config/rbac/role.yaml` - from `make manifests`
- `config/webhook/manifests.yaml` - from `make manifests`
- `**/zz_generated.*.go` - from `make generate`
- `PROJECT` - from `kubebuilder [OPTIONS]`
### Never Remove Scaffold Markers
Do NOT delete `// +kubebuilder:scaffold:*` comments. CLI injects code at these markers.
### Keep Project Structure
Do not move files around. The CLI expects files in specific locations.
### Always Use CLI Commands
Always use `kubebuilder create api` and `kubebuilder create webhook` to scaffold. Do NOT create files manually.
### E2E Tests Require an Isolated Kind Cluster
The e2e tests are designed to validate the solution in an isolated environment (similar to GitHub Actions CI).
Ensure you run them against a dedicated [Kind](https://kind.sigs.k8s.io/) cluster (not your “real” dev/prod cluster).
## After Making Changes
**After editing `*_types.go` or markers:**
```
make manifests # Regenerate CRDs/RBAC from markers
make generate # Regenerate DeepCopy methods
```
**After editing `*.go` files:**
```
make lint-fix # Auto-fix code style
make test # Run unit tests
```
## CLI Commands Cheat Sheet
### Create API (your own types)
```bash
kubebuilder create api --group <group> --version <version> --kind <Kind>
```
### Deploy Image Plugin (scaffold to deploy/manage ANY container image)
Generate a controller that deploys and manages a container image (nginx, redis, memcached, your app, etc.):
```bash
# Example: deploying memcached
kubebuilder create api --group example.com --version v1alpha1 --kind Memcached \
--image=memcached:alpine \
--plugins=deploy-image.go.kubebuilder.io/v1-alpha
```
Scaffolds good-practice code: reconciliation logic, status conditions, finalizers, RBAC. Use as a reference implementation.
### Create Webhooks
```bash
# Validation + defaulting
kubebuilder create webhook --group <group> --version <version> --kind <Kind> \
--defaulting --programmatic-validation
# Conversion webhook (for multi-version APIs)
kubebuilder create webhook --group <group> --version v1 --kind <Kind> \
--conversion --spoke v2
```
### Controller for Core Kubernetes Types
```bash
# Watch Pods
kubebuilder create api --group core --version v1 --kind Pod \
--controller=true --resource=false
# Watch Deployments
kubebuilder create api --group apps --version v1 --kind Deployment \
--controller=true --resource=false
```
### Controller for External Types (e.g., from other operators)
Watch resources from external APIs (cert-manager, Argo CD, Istio, etc.):
```bash
# Example: watching cert-manager Certificate resources
kubebuilder create api \
--group cert-manager --version v1 --kind Certificate \
--controller=true --resource=false \
--external-api-path=github.com/cert-manager/cert-manager/pkg/apis/certmanager/v1 \
--external-api-domain=io \
--external-api-module=github.com/cert-manager/cert-manager
```
**Note:** Use `--external-api-module=<module>@<version>` only if you need a specific version. Otherwise, omit `@<version>` to use what's in go.mod.
### Webhook for External Types
```bash
# Example: validating external resources
kubebuilder create webhook \
--group cert-manager --version v1 --kind Issuer \
--defaulting \
--external-api-path=github.com/cert-manager/cert-manager/pkg/apis/certmanager/v1 \
--external-api-domain=io \
--external-api-module=github.com/cert-manager/cert-manager
```
## Testing & Development
```bash
make test # Run unit tests (uses envtest: real K8s API + etcd)
make run # Run locally (uses current kubeconfig context)
```
Tests use **Ginkgo + Gomega** (BDD style). Check `suite_test.go` for setup.
## Deployment Workflow
```bash
# 1. Regenerate manifests
make manifests generate
# 2. Build & deploy
export IMG=<registry>/<project>:tag
make docker-build docker-push IMG=$IMG # Or: kind load docker-image $IMG --name <cluster>
make deploy IMG=$IMG
# 3. Test
kubectl apply -k config/samples/
# 4. Debug
kubectl logs -n <project>-system deployment/<project>-controller-manager -c manager -f
```
### API Design
**Key markers for** `api/<version>/*_types.go`:
```go
// +kubebuilder:object:root=true
// +kubebuilder:subresource:status
// +kubebuilder:resource:scope=Namespaced
// +kubebuilder:printcolumn:name="Status",type=string,JSONPath=".status.conditions[?(@.type=='Ready')].status"
// On fields:
// +kubebuilder:validation:Required
// +kubebuilder:validation:Minimum=1
// +kubebuilder:validation:MaxLength=100
// +kubebuilder:validation:Pattern="^[a-z]+$"
// +kubebuilder:default="value"
```
- **Use** `metav1.Condition` for status (not custom string fields)
- **Use predefined types**: `metav1.Time` instead of `string` for dates
- **Follow K8s API conventions**: Standard field names (`spec`, `status`, `metadata`)
### Controller Design
**RBAC markers in** `internal/controller/*_controller.go`:
```go
// +kubebuilder:rbac:groups=mygroup.example.com,resources=mykinds,verbs=get;list;watch;create;update;patch;delete
// +kubebuilder:rbac:groups=mygroup.example.com,resources=mykinds/status,verbs=get;update;patch
// +kubebuilder:rbac:groups=mygroup.example.com,resources=mykinds/finalizers,verbs=update
// +kubebuilder:rbac:groups=events.k8s.io,resources=events,verbs=create;patch
// +kubebuilder:rbac:groups=apps,resources=deployments,verbs=get;list;watch;create;update;patch;delete
```
**Implementation rules:**
- **Idempotent reconciliation**: Safe to run multiple times
- **Re-fetch before updates**: `r.Get(ctx, req.NamespacedName, obj)` before `r.Update` to avoid conflicts
- **Structured logging**: `log := log.FromContext(ctx); log.Info("msg", "key", val)`
- **Owner references**: Enable automatic garbage collection (`SetControllerReference`)
- **Watch secondary resources**: Use `.Owns()` or `.Watches()`, not just `RequeueAfter`
- **Finalizers**: Clean up external resources (buckets, VMs, DNS entries)
### Logging
**Follow Kubernetes logging message style guidelines:**
- Start from a capital letter
- Do not end the message with a period
- Active voice: subject present (`"Deployment could not create Pod"`) or omitted (`"Could not create Pod"`)
- Past tense: `"Could not delete Pod"` not `"Cannot delete Pod"`
- Specify object type: `"Deleted Pod"` not `"Deleted"`
- Balanced key-value pairs
```go
log.Info("Starting reconciliation")
log.Info("Created Deployment", "name", deploy.Name)
log.Error(err, "Failed to create Pod", "name", name)
```
**Reference:** https://github.com/kubernetes/community/blob/master/contributors/devel/sig-instrumentation/logging.md#message-style-guidelines
### Webhooks
- **Create all types together**: `--defaulting --programmatic-validation --conversion`
- **When`--force`is used**: Backup custom logic first, then restore after scaffolding
- **For multi-version APIs**: Use hub-and-spoke pattern (`--conversion --spoke v2`)
- Hub version: Usually oldest stable version (v1)
- Spoke versions: Newer versions that convert to/from hub (v2, v3)
- Example: `--group crew --version v1 --kind Captain --conversion --spoke v2` (v1 is hub, v2 is spoke)
### Learning from Examples
The **deploy-image plugin** scaffolds a complete controller following good practices. Use it as a reference implementation:
```bash
kubebuilder create api --group example --version v1alpha1 --kind MyApp \
--image=<your-image> --plugins=deploy-image.go.kubebuilder.io/v1-alpha
```
Generated code includes: status conditions (`metav1.Condition`), finalizers, owner references, events, idempotent reconciliation.
## Distribution Options
### Option 1: YAML Bundle (Kustomize)
```bash
# Generate dist/install.yaml from Kustomize manifests
make build-installer IMG=<registry>/<project>:tag
```
**Key points:**
- The `dist/install.yaml` is generated from Kustomize manifests (CRDs, RBAC, Deployment)
- Commit this file to your repository for easy distribution
- Users only need `kubectl` to install (no additional tools required)
**Example:** Users install with a single command:
```bash
kubectl apply -f https://raw.githubusercontent.com/<org>/<repo>/<tag>/dist/install.yaml
```
### Option 2: Helm Chart
```bash
kubebuilder edit --plugins=helm/v2-alpha # Generates dist/chart/ (default)
kubebuilder edit --plugins=helm/v2-alpha --output-dir=charts # Generates charts/chart/
```
**For development:**
```bash
make helm-deploy IMG=<registry>/<project>:<tag> # Deploy manager via Helm
make helm-deploy IMG=$IMG HELM_EXTRA_ARGS="--set ..." # Deploy with custom values
make helm-status # Show release status
make helm-uninstall # Remove release
make helm-history # View release history
make helm-rollback # Rollback to previous version
```
**For end users/production:**
```bash
helm install my-release ./<output-dir>/chart/ --namespace <ns> --create-namespace
```
**Important:** If you add webhooks or modify manifests after initial chart generation:
1. Backup any customizations in `<output-dir>/chart/values.yaml` and `<output-dir>/chart/manager/manager.yaml`
2. Re-run: `kubebuilder edit --plugins=helm/v2-alpha --force` (use same `--output-dir` if customized)
3. Manually restore your custom values from the backup
### Publish Container Image
```bash
export IMG=<registry>/<project>:<version>
make docker-build docker-push IMG=$IMG
```
## References
### Essential Reading
- **Kubebuilder Book**: https://book.kubebuilder.io (comprehensive guide)
- **controller-runtime FAQ**: https://github.com/kubernetes-sigs/controller-runtime/blob/main/FAQ.md (common patterns and questions)
- **Good Practices**: https://book.kubebuilder.io/reference/good-practices.html (why reconciliation is idempotent, status conditions, etc.)
- **Logging Conventions**: https://github.com/kubernetes/community/blob/master/contributors/devel/sig-instrumentation/logging.md#message-style-guidelines (message style, verbosity levels)
### API Design & Implementation
- **API Conventions**: https://github.com/kubernetes/community/blob/master/contributors/devel/sig-architecture/api-conventions.md
- **Operator Pattern**: https://kubernetes.io/docs/concepts/extend-kubernetes/operator/
- **Markers Reference**: https://book.kubebuilder.io/reference/markers.html
### Tools & Libraries
- **controller-runtime**: https://github.com/kubernetes-sigs/controller-runtime
- **controller-tools**: https://github.com/kubernetes-sigs/controller-tools
- **Kubebuilder Repo**: https://github.com/kubernetes-sigs/kubebuilder
+5 -8
View File
@@ -1,8 +1,8 @@
# terdut-operator
Kubebuilder/controller-runtime operator for terdut-server. See `DESIGN.md` for the
settled design (CRD catalog, reconciliation semantics, bootstrap/auth, RBAC) and
`ROADMAP.md` for the staged build plan this repo is following. `README.md` stays the
design (start with its revision section: credentials and the CRD catalog) and
`ROADMAP.md` for status and deferred work. `README.md` stays the
short pitch.
## Checks
@@ -16,15 +16,12 @@ required beyond Go itself and network access to `proxy.golang.org`/`storage.goog
`security` runs `make security-go`/`security-secrets` (govulncheck/gitleaks), same as
terdut-server's own `security` job.
The kubebuilder-scaffolded `make test-e2e` (a disposable, generic smoke test) is separate
from the real golden-path `kind` e2e pass ROADMAP.md's Stage 5 describes (create every CRD
kind, verify against terdut-server's own API, delete, verify gone) — the latter is a
manual pass run and recorded in ROADMAP.md, not a CI job, matching Stage 1-4's own
precedent of validating against a real cluster outside CI.
There is no CI end-to-end job: the golden-path `kind` pass is manual, via
`examples/demo/run-demo.sh` (see ROADMAP.md).
## Release
Wired as of Stage 5 (ROADMAP.md): `.release.conf`, `make release-vars`/`helm-lint`/
Wired: `.release.conf`, `make release-vars`/`helm-lint`/
`push`/`helm-package`/`helm-push`/`release`, and `.gitea/workflows/release.yaml`
(`test` → `image`/`chart` → `scan-image`) all follow terdut-server's established shape —
see that repo's Makefile/`.release.conf` for the shared reasoning, not restated here.
+185 -594
View File
@@ -1,146 +1,83 @@
# terdut-operator design
This is the design reference for implementing terdut-operator. It exists so
implementation can start from settled decisions instead of re-litigating them
mid-PR. It is deliberately more detailed than the README; the README stays as
the short pitch and now points here instead of carrying open questions.
Written against terdut-server as of the Postgres-only, per-team-resources
version (teams, escalation policies, dead man's switches and integrations are
all rows scoped to a team, managed over `internal/api/*` — not env/config-file
driven; see that repo's `charts/terdut-server` for the current deploy story
this operator supersedes).
The design of terdut-operator as built. terdut-server is the system of record; this operator
makes one install of it — the server, its teams, their escalation ladders, dead man's switches
and alert-source integrations — describable as Kubernetes objects and manageable through gitops.
## 1. Goals & non-goals
**Goal:** let a terdut-server install — the server itself, its teams, their
escalation policies, dead man's switches and alert-source integrations — be
fully described as Kubernetes objects and managed through gitops, following
**Goal:** a terdut-server install fully described as Kubernetes objects, following
controller-runtime / Kubebuilder conventions.
**The operator creates and owns every `TerdutServer` it manages. It never
adopts a pre-existing, independently-deployed terdut-server** — whether
deployed by hand or by `charts/terdut-server`. There is no migration path
from an existing chart-based install, and none is planned (§10): starting
with the operator means applying a fresh `TerdutServer` CR, not converting
one. This is the root a few things downstream hang off of — notably §6's
bootstrap flow, which only has to handle the operator bootstrapping a server
it just created, never a server something else already bootstrapped first.
**The operator creates and owns every `TerdutServer` it manages. It never adopts a pre-existing,
independently-deployed terdut-server**, whether deployed by hand or by `charts/terdut-server`.
There is no migration path from a chart-based install (§10): starting with the operator means
applying a fresh `TerdutServer`.
**Non-goals (v1):**
- Not a general-purpose Postgres operator. It *consumes* a database that
either the Zalando `postgres-operator` or something else already provides.
- Not managing Alertmanager itself, or the routing rules that decide which
alerts reach which integration webhook — only the terdut-server side
(creating the integration and handing back its URL/key).
- Not OLM packaging. Plain Kubebuilder manifests + Helm chart for
install, matching how terdut-server itself ships.
- Cross-namespace references are limited to exactly one edge:
`TerdutTeam.spec.serverRef` may name a `TerdutServer` in a different
namespace, gated by that `TerdutServer`'s own `spec.allowedTeams` consent
field (§4.1, §4.2) — this is the multi-tenant shape the operator exists
for (one platform team owns a `TerdutServer`; other teams self-service a
`TerdutTeam` against it without needing write access to the server's
namespace). Every other reference (`teamRef` on the escalation
rule/dead-man-switch/alert-source CRDs) stays same-namespace-as-its-`TerdutTeam`
only, in v1 — those manage a specific team's own resources and are
expected to live alongside it.
- No validating/mutating admission webhooks in v1. CEL validation rules on
the CRDs (OpenAPI `x-kubernetes-validations`) cover what they can; anything
that needs a live look at another object (e.g. "does this teamRef exist")
is a status condition, not an admission rejection — keeps v1 to a
controller-only deployment with no cert-manager/webhook dependency.
- **Never manages external exposure/ingress for `TerdutServer`, in any
form — a permanent non-goal, not a staged one.** The operator creates a
plain `ClusterIP` Service (§4.1) and stops there: no `Ingress`, no
Gateway API `HTTPRoute`, no Istio `VirtualService`, nothing. Some
installs won't expose `TerdutServer` outside the cluster at all; others
will use whichever of those mechanisms already fits their cluster. That
choice belongs to whoever deploys it, not to this operator — see
`examples/networking` for worked (but not operator-managed) examples.
- Not a Postgres operator. It consumes a database that the Zalando `postgres-operator` or
something else provides (§8).
- Not managing Alertmanager or its routing, only the terdut-server side (creating the
integration and handing back its URL/key).
- Not OLM packaging: plain Kubebuilder manifests and a Helm chart, like terdut-server.
- Cross-namespace references are limited to one edge: `TerdutTeam.spec.serverRef` may name a
`TerdutServer` in another namespace, gated by that server's `spec.allowedTeams` (§4.6). A
`TerdutAlertSource` lives beside its `TerdutTeam`.
- No admission webhooks. CEL validation covers what it can; anything needing a live look at
another object is a status condition, not an admission rejection.
- **Never manages external exposure for a `TerdutServer`** (Ingress, HTTPRoute, VirtualService) —
a permanent non-goal. The operator creates a plain `ClusterIP` Service; see `examples/networking`.
## 2. The README's open questions, resolved
## 2. Credentials and CRD shape (the 2026-10 redesign)
> How do team-crd connect with server-crd?
What the design is, and why — each point replaced something heavier:
Explicit `spec.serverRef: {name, namespace}` on `TerdutTeam` — same reasoning
as before (explicit, greppable, trivially validated), but **`namespace` is
deliberately part of the reference**: one team can run and own a
`TerdutServer`, and other teams — in their own namespaces, without any write
access to the server-owning team's namespace — self-service a `TerdutTeam`
against it. `namespace` defaults to the `TerdutTeam`'s own namespace when
omitted, so the common single-tenant case (`serverRef: {name: terdut}`) is
unchanged.
- **No bootstrap handshake.** The operator generates a key into a Secret named
`<TerdutServer>-operator-key`, in the TerdutServer's own namespace and owned by
it (a pod can only mount Secrets of its own namespace, and an owner reference
replaces the old finalizer and `credentials.deletionPolicy`). The Deployment
hands it to terdut-server as `TERDUT_OPERATOR_KEY`; the server creates or
re-keys its instance-scoped service account `terdut-operator` from it at every
start. A replaced Secret rolls the pods (key-hash annotation). `/api/bootstrap`
stays free for the first human administrator. Gone with it: the checkpoint
Secret, `BootstrapStateLost`, `DatabaseReady`/`Bootstrapped` conditions.
- **One credential per server.** An instance-scoped account acts as owner of every
team's configuration (not a member, so it reads no incidents). The per-team
service accounts, Secrets and `status.credentialsSecretRef`/`serverEndpoint` on
`TerdutTeam` are gone.
- **Team identity is `external_id`.** The operator creates a team with
`external_id: <namespace>/<name>` of its CR; the server returns the existing team
for a known id (200) instead of creating one, so crash recovery, a lost status
and a deleted team all heal by repeating the same call, and a display name that
belongs to another team is a 409 (`TeamNameTaken`) instead of an adoption.
- **Three CRDs.** `TerdutServer`, `TerdutTeam` and `TerdutAlertSource`.
`TerdutEscalationRule` and `TerdutDeadmanSwitch` are `spec.escalation` and
`spec.deadmanSwitches[]` (matched by name, unique per team server-side; ones not
listed are removed) on the team: they were one-to-one children with the team's
lifecycle, and folding them removes `teamRef`, the two-rules-clobber footgun and
three controllers. `TerdutAlertSource` stays separate because it owns a webhook
Secret in its own namespace.
- **Usernames are resolved by the server** (`PUT .../escalation` accepts
`username`), so an unknown user is a `UnknownUser` condition, not a list-and-match.
- **Invites are removed.** Membership is not modelled; people get in through the
server's own signup/OIDC.
- **Env mirrors the chart's where it matters:** OIDC claim names and `trustEmail`
are spec fields; `TERDUT_DEADMAN_*` no longer exist on the server.
This cross-namespace edge needs the target namespace's explicit consent —
otherwise any namespace in the cluster could point a `TerdutTeam` at
someone else's `TerdutServer` and have the operator provision a team on
its behalf, which is a namespace-boundary violation, not a gitops
convenience. (This consent gate is about which `TerdutTeam`s the operator
will act on, not about credential exposure — no terdut-server credential
is ever placed in a `TerdutTeam`'s own namespace regardless of this
setting; see §6.) Kubernetes has two established patterns for this kind
of consent, and Gateway API itself uses both, for two different
relationships:
## 3. CRD catalog
- **`ReferenceGrant`** (used for a Route reaching into an arbitrary
Service/Secret): a separate object, living in the *target* namespace,
enumerating exact `{fromNamespace, fromKind} → {toKind, toName}` pairs.
No wildcard, no selector — every permitted namespace is spelled out.
- **An inline field on the parent** (used for a `ListenerSet` attaching to
a shared `Gateway`, GA since Gateway API v1.5): the parent carries
`spec.allowedListeners.namespaces: {from: None|Same|All|Selector,
selector}` directly, no separate CRD.
Group `terdut.ryuvia.com`, version `v1alpha1`; module `git.ryuvia.com/niklas/terdut-operator`.
`TerdutTeam` attaching to a shared `TerdutServer` is structurally the
second case, not the first — a bounded set of expected children attaching
to a parent they were deliberately made shareable, not an arbitrary
backend reference — so this design follows the `ListenerSet` precedent:
`TerdutServer.spec.allowedTeams` (§4.1), no extra CRD. §4.2 covers how
`TerdutTeam` resolves against it. No other reference in this design
(`teamRef` on the child CRDs) crosses a namespace boundary, so this is the
only place cross-namespace consent is needed at all (§1, §5, §6, §9).
> How does escalationrules, switches and alertsources connect to a team?
Explicit `spec.teamRef: {name}` on each of `TerdutEscalationRule`,
`TerdutDeadmanSwitch`, `TerdutAlertSource` — same reasoning, and it mirrors
terdut-server's own data model, where every one of these rows carries a
`team_id` foreign key already. A matcher/selector on `TerdutTeam` would be
inventing a second source of truth for an association the server already
models as a plain reference.
> Support for both postgres-operator (Zalando) and bring-your-own, how do we
> design that to be user friendly?
`TerdutServer.spec.database` is a oneOf, mirroring the chart's existing
`database.dsn` / `database.passwordSecret` contract (see §8):
- `dsn` + `passwordSecretRef` — bring-your-own, exactly today's chart inputs.
- `postgresClusterRef` — points at a Zalando `postgresql.acid.zalan.do` CR;
the operator derives the DSN and resolves the generated credentials Secret
itself (see §8). CloudNativePG support is a natural follow-up using the
same shape and is called out as deferred (§13), not designed in detail now.
## 3. API group, versions, CRD catalog
- Group: `terdut.ryuvia.com`, version: `v1alpha1` (matches the `ryuvia.com`
domain terdut-server already uses; bump to `v1beta1`/`v1` per the normal
Kubernetes API graduation criteria once the shapes below have proven
stable against a real install).
- Module: `git.ryuvia.com/niklas/terdut-operator`, scaffolded with
Kubebuilder (controller-runtime), matching terdut-server's Go toolchain
and house style.
| Kind | Scope | Purpose |
|---|---|---|
| `TerdutServer` | Namespaced | One terdut-server install: Deployment, Service, database wiring, bootstrap, operator credentials, cross-namespace team consent. |
| `TerdutTeam` | Namespaced | One team on a `TerdutServer`, possibly in another namespace: name, OIDC group mapping. |
| `TerdutEscalationRule` | Namespaced | A team's escalation policy (levels, targets, repeat). |
| `TerdutDeadmanSwitch` | Namespaced | One dead man's switch on a team. |
| `TerdutAlertSource` | Namespaced | One alert-ingest integration on a team (currently: Alertmanager webhook). |
| Kind | Purpose |
|---|---|
| `TerdutServer` | One terdut-server install: Deployment, Service, database wiring, operator key, cross-namespace team consent. |
| `TerdutTeam` | One team on a server, possibly in another namespace: name, OIDC groups, escalation ladder, dead man's switches. |
| `TerdutAlertSource` | One alert-ingest integration on a team; owns a webhook Secret. |
## 4. Per-CRD spec
(The escalation and dead man's switch specs that used to be §4.3 and §4.4 are part of §4.2; the section numbers are kept because code comments cite them.)
### 4.1 `TerdutServer`
```yaml
@@ -153,7 +90,7 @@ spec:
image:
repository: git.ryuvia.com/niklas/terdut-server
tag: v0.9.3
replicas: 1 # terdut-server is not horizontally-scale-tested; keep the field, default 1
replicas: 2 # default since terdut-server v0.36.0's advisory locks; see TerdutServerSpec.Replicas
networking:
hostname: terdut.example.com
servicePort: 8080
@@ -165,10 +102,6 @@ spec:
sweeper:
staleAfter: 6h
archiveAfter: 168h
deadman:
matchers: "alertname=Watchdog"
timeout: 15m
severity: critical
notify:
ntfyURL: "http://ntfy.ntfy.svc.cluster.local"
fallbackTopic: ""
@@ -183,6 +116,8 @@ spec:
scopes: "openid profile email"
allowedGroups: []
adminGroup: ""
usernameClaim: preferred_username # also emailClaim, groupsClaim
trustEmail: false
sessionMaxAge: 12h
passwordLogin: true
# Consent for TerdutTeams in OTHER namespaces to set serverRef at this
@@ -210,10 +145,9 @@ spec:
disruptionBudget:
minAvailable: 1 # mutually exclusive with maxUnavailable
status:
conditions: [...] # Ready, DatabaseReady, Bootstrapped
conditions: [...] # Ready
observedGeneration: 3
serviceName: terdut
credentialsSecretRef: {name: terdut.platform-oncall-instance-credentials, key: token} # see §6; pure output -- generated by the controller's own self-registration flow, always in the OPERATOR's namespace (always, implicitly -- not stored here, since it's never anything else), under a fixed key ("token").
credentialsSecretRef: {name: terdut-operator-key, key: token} # see §6; pure output, in this TerdutServer's own namespace
```
Field-for-field this is the chart's `values.yaml` reshaped as a spec — not
@@ -244,10 +178,10 @@ no custom wrapper buys anything for any of these, matching how
CloudNativePG and the Zalando postgres-operator both expose the same
knobs. `affinity` is pure user-supplied passthrough, not a
toggle-plus-generated-default the way a multi-replica-aware operator's
pod anti-affinity typically is: this operator never auto-generates
affinity of its own, since `replicas` above 1 isn't a supported topology
(the sweeper/notifier singleton constraint, §4.1's own illustrative YAML
comment). `spec.pod.disruptionBudget` is the one field here that isn't a
pod anti-affinity typically is: even though `replicas` now defaults to 2
(terdut-server v0.36.0's advisory locks made that safe, §4.1's own
illustrative YAML comment), this operator still never auto-generates
affinity of its own. `spec.pod.disruptionBudget` is the one field here that isn't a
straight PodTemplateSpec knob — when set, the controller reconciles a
`PodDisruptionBudget` selecting this `TerdutServer`'s pods; clearing it
deletes any it previously created (§7). `minAvailable`/`maxUnavailable`
@@ -258,89 +192,37 @@ same way as `spec.database`'s own `dsn`/`postgresClusterRef` rule.
```yaml
spec:
serverRef:
name: terdut
namespace: platform-oncall # optional; defaults to this TerdutTeam's own namespace.
# Cross-namespace requires that TerdutServer's spec.allowedTeams
# (§4.1) to admit this namespace — otherwise Ready: False, reason: RefNotPermitted.
displayName: "Platform" # -> POST /api/teams {"name": ...}; server assigns the ID
oidc:
memberGroup: "terdut-platform-members"
ownerGroup: "terdut-platform-owners"
serverRef: {name: terdut, namespace: platform-oncall} # namespace optional; cross-namespace needs allowedTeams
displayName: Platform
oidc: {memberGroup: terdut-platform-members, ownerGroup: terdut-platform-owners}
escalation:
repeatCount: 2
fallbackTopic: platform-oncall
levels:
- timeout: 5m
targets: [{kind: oncall}, {kind: user, username: alice}]
- timeout: 15m
targets: [{kind: oncall}]
deadmanSwitches:
- {name: watchdog, matcher: "alertname=Watchdog,cluster=prod", timeout: 15m, severity: critical}
status:
conditions: [...]
teamID: 42 # the server-side ID; needed by every child object's controller
credentialsSecretRef: {name: platform-oncall.platform-team-credentials, key: token} # see §6; this team's own scoped key, always in the OPERATOR's own namespace (not stored here — same reasoning as TerdutServer's, §4.1)
serverEndpoint: "http://terdut.platform-oncall.svc:8080" # resolved once, here, from spec.serverRef -- see §5: this is what actually makes "a child never needs to chain up to TerdutServer" true, not just a stated intent. Set alongside teamID/credentialsSecretRef, same reconcile.
conditions: [...] # Ready; reasons: ServerRefNotFound, RefNotPermitted, WaitingForServer,
# TeamNameTaken, UnknownUser, InvalidSpec, Adopted
teamID: 42
observedGeneration: 1
```
Team *membership* (which users belong, `team_members`) is explicitly **not**
modeled as a CRD field in v1: terdut-server already manages membership via
OIDC group sync at login for SSO installs, and manual membership for
password-login installs is a people-management action, not infrastructure —
forcing it through gitops would mean a human's team change goes through a PR
review. Flagged in §13 as revisitable if a real gitops-membership need shows up.
The team is created under the identity `<namespace>/<name>` of its CR (`external_id`), so a retry,
a lost status or a team deleted behind the operator's back all heal by repeating the same call; a
`displayName` that belongs to a different team is `TeamNameTaken`, never an adoption. Every
reconcile applies, in order: create-or-find, rename, OIDC groups, the escalation ladder (one PUT
of the whole ladder; absent `spec.escalation` clears it), and the switches. Switches are matched
by name (unique per team on the server): missing ones are created, changed ones updated in
place, and any not listed are deleted, since in operator mode nobody else can add one. Durations
are validated by a CRD pattern and re-checked (`InvalidSpec`). A username the server does not
know is `UnknownUser` until that person exists.
### 4.3 `TerdutEscalationRule`
```yaml
spec:
teamRef: {name: platform-team}
repeatCount: 2
fallbackTopic: "platform-oncall"
levels:
- timeout: 5m
targets:
- kind: oncall # "oncall" or "user"
- kind: user
username: alice # resolved to a user ID by the controller at apply time
- timeout: 15m
targets:
- kind: oncall
status:
conditions: [...]
observedGeneration: 1
```
One `TerdutEscalationRule` per team — the server itself models a policy as
one row (`escalation_policies`) with an owned list of levels, so a
one-CRD-to-one-policy mapping (not one-CRD-per-level) matches the server's
own aggregate and lets the whole thing be reconciled with the single
`PUT /api/teams/{teamID}/escalation` the API actually exposes (see §5). Not
enforced at admission (no webhooks in v1, §1) — two `TerdutEscalationRule`s
naming the same team would both `PUT` it and clobber each other every
reconcile; a footgun worth this one sentence, not a technical guard.
A `username` target is resolved to the `user_id` `PUT /api/teams/{teamID}/escalation`
actually requires (confirmed against source: `escalationTargetJSON.UserID
*int64`, no username field at all) via `GET /api/users` — confirmed open to
any authenticated caller, not gated by team membership or admin
(`internal/api/router.go`'s own comment: "readable by anyone signed in"), so
the team-scoped credential this controller already holds is enough; no
extra RBAC-equivalent server-side needed. Re-resolved every reconcile rather
than cached, in case a username is renamed. An unresolvable username is
`Ready: False, reason: UnknownUser`, naming which one.
### 4.4 `TerdutDeadmanSwitch`
```yaml
spec:
teamRef: {name: platform-team}
name: "prod-watchdog"
matcher: "alertname=Watchdog,cluster=prod"
timeout: 15m
severity: critical
status:
conditions: [...]
switchID: 7
```
No unique-name constraint server-side (§5's table row, confirmed against
source) — the controller's own idempotent-create step is a `GET`-list and
name match, not a conflict to recover from. `name` is optional, same as the
API: left empty, terdut-server derives it from `matcher`'s own canonical
form, and that's what the lookup matches against too.
Team *membership* is not modelled: the server manages it through OIDC group sync and its own UI.
### 4.5 `TerdutAlertSource`
@@ -357,7 +239,7 @@ status:
The integration key is shown by the API exactly once, at creation
(`Integration.Key`/`URL` in terdut-server's own model) — never re-readable,
same shape as the bootstrap admin key. The controller writes it straight into
a one-shot value. The controller writes it straight into
a generated, owner-referenced Secret on the create it caused and never logs
or stores it anywhere else; the CR's `status` carries only the Secret
reference, matching how e.g. cert-manager's `Certificate` exposes
@@ -377,8 +259,7 @@ of both patterns).
`serverRef` into this `TerdutServer`. Same-namespace `TerdutTeam`s are
always allowed regardless of this field.
- `from: Same` — equivalent to `None` in effect (same-namespace is already
unrestricted) but kept for parity with the upstream enum and to make the
policy self-documenting in a diff.
unrestricted); kept for parity with the upstream enum.
- `from: All` — any namespace in the cluster may reference in. Appropriate
for a genuinely shared, cluster-wide `TerdutServer`; the audit trail is
"check `allowedTeams` plus who has RBAC to create a `TerdutTeam`
@@ -401,257 +282,64 @@ of both patterns).
it used to authorize finds itself no longer permitted, flips
`Ready: False, reason: RefNotPermitted`, and — deliberately — does
**not** delete the team server-side on revocation alone; it stops
reconciling further changes until access is restored or the `TerdutTeam`
CR itself is deleted (whose finalizer still needs the credentials
Secret described in §6 to clean up, so blocking *new* changes rather
than forcing an immediate, possibly credential-less deletion is the
safer failure mode).
reconciling further changes until access is restored. Deleting the
`TerdutTeam` CR still deletes the team: consent gates what the operator
starts acting on, not whether it may clean up after itself.
## 5. Reconciliation semantics
terdut-server's REST surface (`internal/api/router.go`) does not give every
resource a full update verb, so reconciliation strategy is per-resource:
| Resource | Verbs available | Strategy |
| Resource | Server verbs | Strategy |
|---|---|---|
| Team | POST create, PUT rename, DELETE, PUT oidc-groups, GET by name | Real update-in-place: diff spec vs. last-applied, PUT the changed pieces. `GET /api/teams?name=` (terdut-server's `TEAM-LOOKUP.md`, landed 2026-10-01) is what makes the idempotent-create general rule below actually true for Team — confirmed by checking: until that endpoint existed, an instance-scoped service account had no way to recover a team's id after a 409, unlike every other resource in this table, where the adopt-on-conflict rule had a real lookup to call. |
| Escalation policy | GET/PUT whole-policy | Update-in-place: PUT the full desired policy every reconcile that finds drift; cheap because whole-policy is small and already loaded whole server-side. |
| Dead man's switch | POST create, **PUT update-in-place** (added in terdut-server v0.33.0), DELETE, GET-list | Real update-in-place, same shape as Team/Escalation: PUT the whole switch every reconcile once its id is known. No unique-name constraint server-side (confirmed against source — `handleCreateTeamDeadman` has no conflict handling at all, unlike Team/service-account creation), so idempotent-create here can't rely on a 409 to adopt from: before POSTing, `GET /api/teams/{teamID}/deadman/switches` and match by `name` first: found → adopt its id; not found → POST. `handleUpdateTeamDeadman`'s own doc comment (terdut-server) confirms the motivation directly: "Added alongside create/delete so an automated caller (terdut-operator) can reconcile a spec change without deleting and recreating the switch, which would otherwise ... needlessly rotate its id for no reason a reconciler's diff should ever manufacture." An earlier draft of this row, written before that endpoint existed, described delete-and-recreate; corrected here, confirmed against source rather than left stale. |
| Integration (alert source) | POST create, PATCH rename, DELETE | Rename via PATCH; any other spec change (kind) is delete-and-recreate, which **rotates the webhook key** — called out loudly in the CRD's field docs and in a `Warning` event, since it breaks whatever sends to the old URL/key until the new Secret is picked up. (See the general rule below for what happens if the Secret is lost with *no* spec change.) |
| Team | POST (idempotent on `external_id`), PUT rename, PUT oidc-groups, DELETE | Repeat the idempotent POST, then PUT the rest, every reconcile. Delete runs in a finalizer; the server refuses (409) while the team has open incidents, which is retried. |
| Escalation ladder | PUT whole ladder | PUT the full desired ladder every reconcile; the server resolves usernames. |
| Dead man's switch | list, POST, PUT, DELETE; names unique per team | Diff by name against the list. |
| Integration (alert source) | POST, PATCH rename, DELETE | Rename via PATCH; a kind change is delete-and-recreate, which **rotates the webhook key** (Warning event). An integration deleted on the server is recreated with a new key (Warning event). |
General rules for every controller:
- **Idempotent create**: before POSTing, check `status.<serverSideID>` is
unset; if the server already has a same-named object from a previous
partial reconcile (e.g. after a crash between POST and status-write), treat
a 409/name-conflict as "adopt" — GET-by-name and populate status, rather
than erroring forever. terdut-server's list endpoints in each of these
areas return objects by name, so this is a straightforward correlation.
- **Periodic resync** in addition to watch-triggered reconciles (Kubebuilder
default `RequeueAfter` on success, e.g. every 5–10 minutes) to catch drift
from **someone changing state directly against the server's API/UI**,
since gitops correctness means the CR wins, not "first write wins".
- **Finalizers** on every CRD that has a server-side counterpart, so deletion
calls the corresponding DELETE before the Kubernetes object disappears.
Failure to delete server-side (e.g. server unreachable) blocks finalizer
removal and surfaces as a `Degraded` condition + event, rather than
silently orphaning a row.
- **Generated Secrets holding unrecoverable server-issued material are
watched, and their loss is fail-closed, not self-healed.** Currently this
is just `TerdutAlertSource`'s webhook Secret (§4.5): the controller adds
it to its `Owns()` watches, not just the CR. If it disappears while
`status.integrationID` is still set, the controller does **not** attempt
to recreate it — the key is genuinely gone (§4.5: never stored anywhere
but that one Secret), so silently minting a replacement would rotate a
live production webhook URL with no corresponding spec change to explain
why. Instead it flips `Ready: False, reason: WebhookSecretLost` and fires
a `Warning` event telling the operator to delete and recreate the
`TerdutAlertSource`. No new mechanism is needed for recovery: deleting the
CR runs the existing finalizer (DELETE the still-live integration
server-side, above), and recreating it runs the existing idempotent-create
path (this same section) — a fresh POST, a new key, a new Secret. This is
deliberately the same recovery motion as the kind-change rotation above,
just human-triggered instead of spec-triggered.
- **Owner chain for status resolution, not API calls**: `TerdutTeam`'s
controller does not call any other controller; every child CRD's
controller independently resolves its own `teamRef` → `TerdutTeam.status`
for the `teamID`, `credentialsSecretRef`, and `serverEndpoint` it needs to
call the API (§6) — it never needs to chain further up to `TerdutServer`
at all, not just as a stated intent but literally: `serverEndpoint` is
resolved once, by `TerdutTeam`'s own controller, and stored in its status
specifically so no child ever needs its own `TerdutServer` RBAC (`get`/
`list`/`watch` on `terdutservers`) to find out where to send a request —
the team's own scoped credential plus that one status field is everything
a child resource's controller requires. If the referenced `TerdutTeam`
isn't `Ready` yet (which includes not having a `credentialsSecretRef` set),
the child requeues with backoff and reports `Ready: False, reason:
WaitingForTeam` — no cross-controller RPC.
- **Cross-namespace `serverRef` is re-checked every reconcile, not just at
creation**: `TerdutTeam`'s controller reads the target `TerdutServer`'s
`spec.allowedTeams` (and, under `Selector`, a `Get` on its own `Namespace`
object for labels) on every pass before touching a cross-namespace
`TerdutServer` — revocation (§4.6) takes effect on the team's very next
reconcile, not just when the CR is first applied.
General rules:
- **Periodic resync** (5 minutes) besides watch-triggered reconciles, to catch someone changing
state directly against the server: the CR wins.
- **Finalizers** on `TerdutTeam` and `TerdutAlertSource` call the server's DELETE first. A 404 is
success; any other failure blocks removal and surfaces as an event rather than orphaning a row.
`TerdutServer` needs none: its Secret is owned and its database is never touched.
- **A webhook Secret that is lost fails closed.** The key is never re-readable from the server, so
the controller does not mint a replacement for a live URL with no spec change to explain it:
`Ready: False, reason: WebhookSecretLost`; delete and recreate the `TerdutAlertSource`.
- **Children resolve through the team.** A `TerdutAlertSource` finds its `TerdutTeam` (same
namespace), requires it Ready, then reads the server named by the team's `serverRef` and that
server's operator key. A `TerdutTeam` re-reconciles when its `TerdutServer` changes.
- **Cross-namespace consent is re-checked every reconcile** (§4.6), so revocation takes effect on
the team's next pass.
## 6. Bootstrap & authentication to terdut-server's API
## 6. Authentication to terdut-server's API
terdut-server's scoped service-account credential type
(`terdut-server`'s `SERVICE-ACCOUNTS.md`) has shipped — confirmed against
source: `internal/api/service_accounts.go`, migration
`014_service_accounts.sql`, and `internal/api/router.go` wiring it in under
`AuthMiddleware`. This section is no longer blocked on it; the "v1-blocking"
framing here was accurate when this section was first written and is stale
now.
The operator talks to the server over HTTP with one bearer key per `TerdutServer`:
**`GET /api/service-accounts?name=` is not an unauthenticated lookup —
confirmed against `internal/api/middleware.go`'s `AuthMiddleware`, which
hard-rejects any request carrying neither a Bearer token nor a session
cookie with `401` before any handler ever runs.** This would matter a great
deal if the operator's `/api/bootstrap` call could ever lose a race to
something else bootstrapping the same server first — a credential-less
loser would have no authenticated way to recover. It doesn't matter here,
by construction (§1): **the operator only ever calls `/api/bootstrap`
against a `TerdutServer` it just created**, so there is nothing else in a
position to race it. An earlier draft of this section added a
`spec.credentialsSecretRef` bring-your-own input specifically to work around
that race, for a world where the operator might adopt a server something
else had already bootstrapped. That world doesn't exist (§1), so the field
was removed rather than kept as unused flexibility — self-registration
(point 1, below) is simply the only path, not one of two.
1. The `TerdutServer` controller creates the Secret `<name>-operator-key` (data key `token`,
`tdsa_` + 48 hex characters) in the server's own namespace, owned by the `TerdutServer`. It
is created once and never overwritten while it exists: the running server was seeded with it.
2. The Deployment hands it to the pods as `TERDUT_OPERATOR_KEY` through a `secretKeyRef` (the key
never appears in the pod spec), plus an annotation holding a hash of it so a replaced Secret
rolls the pods.
3. At every start the server creates or re-keys its instance-scoped service account
`terdut-operator` from that value. An instance-scoped account acts as owner of every team's
configuration but is not a member of any team, so it reads no incidents and is never an
administrator.
4. `status.credentialsSecretRef` points at the Secret; `TerdutTeam` and `TerdutAlertSource`
controllers read it from the server's namespace.
**Every credential the operator holds — the one instance-scoped key per
`TerdutServer`, and one team-scoped key per `TerdutTeam` — lives in a Secret
in the *operator's own* namespace, never in the namespace of the CR it
authenticates for.** Reconciliation happens entirely inside the operator's
controller loop, which is a single Deployment/ServiceAccount already
watching every namespace it's granted (§9); nothing about calling
terdut-server's API on a CR's behalf requires the credential to be
physically located near that CR, and no CR owner (human or otherwise) ever
needs to see, hold, or have RBAC to read a terdut-server credential. This is
a straight simplification of an earlier draft of this section, which mirrored
a shared credential into each consenting namespace instead — that version
conflated "the CR's owner never needs to see this" (true, and preserved
here) with "so the credential must live in the CR's namespace" (a
non-sequitur once you don't need to grant *anyone else* namespace-local
read access). Dropping that assumption also removes an entire class of
complexity: no on-demand mirroring, no garbage-collecting an orphaned copy
when `allowedTeams` narrows, no "OwnerReferences can't cross namespaces so
track it in status instead" workaround — none of that machinery is needed
when nothing ever crosses into a tenant namespace in the first place.
1. **First reconcile, confirmed against source**
(`internal/api/users.go`'s `handleBootstrap`): `/api/bootstrap` is
single-shot *per install*, gated on `SELECT COUNT(*) FROM users` — once
non-zero, every call `403`s regardless of identity. There are two crash
windows between "get an admin key" and "have a lasting, usable
credential" — getting from `/api/bootstrap`'s key to a minted
service-account key, and getting from that key to a persisted Secret —
and both get a checkpoint rather than being left as a theoretical gap,
the same rigor §5's general idempotent-create rule already applies
elsewhere:
- `status.credentialsSecretRef` already set: done, nothing to do.
- Otherwise, check for an intermediate
`<namespace>.<name>-bootstrap-admin` Secret in
the operator's own namespace first. If it exists, its key is a still-
valid admin credential from an earlier, interrupted attempt — skip
`/api/bootstrap` entirely and reuse it. If not, call `/api/bootstrap`
once the Deployment this `TerdutServer` created has a ready replica;
on `201`, immediately checkpoint its response's raw admin key
(`{"user": ..., "api_key": {"key": "<raw>", ...}}`) into that Secret
before doing anything else with it. A `403` with neither
`status.credentialsSecretRef` nor this checkpoint Secret present is
the one genuinely pathological case left (the checkpoint deleted out
from under a reconcile already past this point) — handled the same
way the design already handles unrecoverable server-issued material
elsewhere (§5's webhook-Secret-loss rule): fail closed,
`Ready: False, reason: BootstrapStateLost`, with the same recovery as
that case, delete and recreate the `TerdutServer` (its finalizer tears
down the Deployment/database-backing and server-side rows; a fresh
create starts clean) — not a workaround peculiar to this one path.
- With an admin key in hand (fresh or checkpointed): `POST
/api/service-accounts {name: "terdut-operator", scope: "instance"}`.
A `409` here means a prior attempt got this far before being
interrupted — adopt rather than error, per §5's general rule:
`GET /api/service-accounts?name=terdut-operator` (authenticated with
the checkpointed admin key, not an unauthenticated lookup) to find its
id, then `POST /api/service-accounts/{id}/keys` to mint a fresh key —
an orphaned first key some interrupted attempt minted and never used
is inert, not a cleanup obligation.
2. The resulting instance-scoped key — not the checkpointed admin key, which
is deleted once this step succeeds — is written to a generated Secret in
the **operator's own namespace** (e.g.
`<serverRef.namespace>.<serverRef.name>-instance-credentials`,
under a fixed data key, `token`), referenced back from
`TerdutServer.status.credentialsSecretRef: {name, key}` (§4.1). No
`OwnerReference` (those can't cross namespaces, and this Secret doesn't
share a namespace with the `TerdutServer` that caused it); the `TerdutServer`'s
finalizer deletes this Secret directly as part of its own teardown,
the same way it already has to clean up the server-side resources it
created (§5's general finalizer rule extends naturally to this Secret).
3. When a `TerdutTeam` first becomes `Ready` (its `serverRef` resolved,
`allowedTeams` satisfied if cross-namespace), its controller uses the
`TerdutServer`'s instance-scoped credential (read from the operator's own
namespace, resolved via the owner chain in §5) to mint a **team-scoped**
service account for itself: `POST /api/service-accounts` with
`scope: team, teamID: <status.teamID>`. The resulting key is written to
its own generated Secret, again in the **operator's own namespace**
(e.g. `<teamNamespace>.<teamName>-team-credentials`), referenced from
`TerdutTeam.status.credentialsSecretRef` (§4.2). Same finalizer pattern as
point 2: the `TerdutTeam`'s finalizer deletes this Secret as part of its
own teardown.
4. Every child controller (`TerdutEscalationRule`, `TerdutDeadmanSwitch`,
`TerdutAlertSource`) reads its team's `credentialsSecretRef` — resolved
through its `teamRef` → `TerdutTeam.status` (§5) — and never touches the
instance-scoped credential at all. Since child CRDs stay same-namespace-
as-their-`TerdutTeam` in v1 (§1), and the credential itself lives in the
operator's namespace regardless of where the `TerdutTeam` or its children
are, this works identically whether the `TerdutTeam` is same-namespace or
cross-namespace relative to its `TerdutServer` — there is no separate
cross-namespace case to handle here at all, unlike the mirroring design
this replaced.
5. **Blast radius**: a team-scoped key can only touch its own `team_id`'s
escalation policy, dead-man switches, integrations, schedule and OIDC
group bindings server-side (enforced by terdut-server itself, per
`SERVICE-ACCOUNTS.md`) — compromising one such Secret (e.g. a bug that
leaks operator-namespace Secrets, or an overly broad RBAC grant on that
one namespace) exposes exactly one team, never the whole server. This is
the real fix for what an earlier draft of this section called out as its
weak point (every mirrored copy being server-admin-equivalent); it falls
out of team-scoped credentials existing at all, independent of where
they're stored — the operator-private storage described above closes the
RBAC-footprint half of the problem, team scoping closes the credential-
privilege half.
6. **Rotation**: `POST /api/service-accounts/{id}/keys` mints a new key on
the existing account without recreating it; the old key is revoked via
`DELETE /api/service-accounts/{id}/keys/{keyID}`; the operator's local
Secret is updated in place. No DB-level workaround, no re-triggering a
single-shot endpoint that can't fire twice (which is what made rotation
unworkable under the old `/api/bootstrap`-only design).
7. **Operator mode** (`TERDUT_OPERATOR_MODE`, terdut-server's own
deploy-time flag, off by default) is the complementary half of this
trust model: it makes terdut-server itself refuse a *human* write (a
session or a user's own API key) on a route, while a service account's
— this operator's — still goes through. Confirmed against source
(`internal/api/middleware.go`'s `OperatorModeBlock`,
`internal/api/router.go`'s `opMode` wrapper): the blocked set is exactly
team create/rename/delete, a team's OIDC-group binding, its escalation
policy, its dead man's switches, and its integrations — precisely the
resources `TerdutTeam`, `TerdutEscalationRule`, `TerdutDeadmanSwitch` and
`TerdutAlertSource` manage, and nothing more. **Deliberately not
blocked, confirmed against the same router**: team membership and
invites (the router's own comment: "membership is deliberately never
gitops-managed"), and the on-call schedule/rota
(`/api/teams/{teamID}/schedule`, `/api/schedule/current`) — neither
route carries the `opMode` wrapper at all. A human can still add or
remove a team member, or assign who's on call, on a server running in
operator mode; only the CRD-shaped resources above are locked to
GitOps. This isn't a gap to close — it's the same boundary §4.2 already
draws for membership, confirmed to hold on the server side too, not
just stated as an intent here.
There is no bootstrap handshake: `/api/bootstrap` stays free for the first human administrator.
Deleting the `TerdutServer` deletes the key with it; a recreated one gets a new key and the server
re-seeds on its next start. Rotating by hand means deleting the Secret: the next reconcile makes
a new one and rolls the pods.
## 7. Ownership, status, garbage collection
- Every generated object that lives in the *same* namespace as the CR that
caused it (Deployment, Service, webhook Secret, and `TerdutServer`'s own
PodDisruptionBudget) carries a standard `metav1.OwnerReference` — GC
handles these, no finalizer needed. PodDisruptionBudget is the one
member of that list that's conditionally created/deleted rather than
always present: it exists only while `spec.pod.disruptionBudget` is set,
and the controller deletes it itself the moment that field is cleared
(it doesn't wait on GC for that case, only for the `TerdutServer` being
deleted outright). The two credential Secrets from §6 are the one
exception to OwnerReference-based cleanup generally: they live in the
operator's own namespace regardless of where their owning CR lives, so
`OwnerReference` doesn't apply (cross-namespace) and cleanup instead runs
through that CR's finalizer directly, alongside the server-side DELETE
it already has to issue (§5).
- Status conditions follow the standard `metav1.Condition` shape with at
least `Ready` on every kind, plus kind-specific ones (`TerdutServer`:
`DatabaseReady`, `Bootstrapped`; children: `Synced`).
- `status.observedGeneration` on every kind, bumped only after a successful
reconcile against that generation's spec — the standard way a client
(or `kubectl wait`) tells "applied" from "seen".
- No cluster-scoped aggregation object (e.g. no cluster-wide "all servers"
status) in v1 — `kubectl get terdutservers -A` is the aggregate view.
- Everything the controllers generate in a CR's namespace carries an `OwnerReference` (Deployment,
Service, PodDisruptionBudget, the operator key Secret, the webhook Secret), so GC cleans up
and no finalizer is needed. The PDB exists only while `spec.pod.disruptionBudget` is set; the
controller deletes it itself when the field is cleared.
- Every kind has a `Ready` condition and `status.observedGeneration`.
- No cluster-scoped aggregate object: `kubectl get terdutservers -A` is the overview.
## 8. Postgres integration
@@ -672,85 +360,43 @@ documented and tested operationally:
(`<user>.<cluster>.credentials.postgresql.acid.zalan.do`) the same way
the chart's comment already documents, and wires it in as
`PGPASSWORD` the same way.
- Watches that Secret (not just the `postgresql` CR) so a credential
rotation triggers a requeue — the chart today requires a manual pod
restart for this; the operator can at least detect and report it via a
condition even if restarting on rotation is left as a §13 follow-up
rather than done automatically (a rolling restart on credential change
is a behavior change worth its own design pass, not folded in here).
- Does not watch that Secret: a rotated credential is noticed at the next
5-minute resync, and restarting pods on rotation is a §13 follow-up.
- Requires read RBAC on `postgresql.acid.zalan.do` (optional CRD — the
operator's ClusterRole/Role should not hard-fail if the CRD isn't
installed and a given `TerdutServer` uses BYO DSN instead).
## 9. RBAC
- The operator's own ServiceAccount needs, per namespace it's granted:
`get/list/watch/create/update/patch/delete` on `Deployments`, `Services`
and `PodDisruptionBudgets` it owns, and `get/list/watch` on
`postgresql.acid.zalan.do` (optional, degrade gracefully if absent per
§8), plus cluster-wide `get/list` on `Namespace` (labels only, for
`allowedTeams: {from: Selector}` evaluation — §4.1, §4.6).
- **Two different `Secret` scopes, not one — corrected from an earlier draft
of this section.** That earlier draft said `Secret` access was "scoped to
the operator's own namespace only... nowhere else," reasoning that with
every credential held privately in the operator's own namespace (§6)
there was no legitimate reason to touch a `Secret` anywhere else. That was
wrong once §4.5 existed:
- The §6 credential Secrets (one instance-scoped key per `TerdutServer`,
one team-scoped key per `TerdutTeam`) do live in, and are only ever
touched from, the operator's own namespace —
`get/list/watch/create/update/patch/delete` there, nowhere else. The
rest of the original reasoning stands for *these* Secrets specifically:
no human or team's own RBAC is ever granted access to a terdut-server
credential by this design, and the operator itself never needs
cross-namespace access to reach them.
- The §4.5 webhook Secret is different: it's owned by and lives beside
its `TerdutAlertSource`, in that CR's own tenant namespace, not the
operator's. The per-namespace `Role` already granted for
`Deployments`/`Services` in each watched namespace (below) must carry
the same `Secret` verbs there too, or the controller cannot create,
watch, or even detect the loss of (§5) that Secret at all.
- This necessarily widens the operator's footprint in each watched
tenant namespace to "any `Secret` in that namespace," not just the ones
it created — Kubernetes RBAC has no owner-scoped grant finer than the
namespace itself, and the design already accepts this same granularity
for Deployments/Services there. Flagged as an accepted trade-off, not a
silent gap (§13).
- No cluster-scoped resources are created by this operator (namespaced CRDs
only, per §1) — a `Role` + `RoleBinding` per watched (tenant) namespace is
sufficient for Deployments/Services/the webhook `Secret`/the optional
Zalando CRD, plus a separate `Role` + `RoleBinding` in the operator's own
namespace for the §6 credential Secrets; a `ClusterRole` is only needed
for watching CRDs across all namespaces (the normal Kubebuilder
multi-tenant-operator default) — none of the `Secret` access above needs
to be cluster-scoped.
- terdut-server's own RBAC is unaffected — the operator talks to it purely
over HTTP with service-account API keys (§6), never via the Kubernetes
API for app-level state.
- The operator needs `get/list/watch/create/update/patch/delete` on `Deployments`, `Services`,
`PodDisruptionBudgets` and `Secrets` it owns, `get/list/watch` on `postgresql.acid.zalan.do`
(optional; degraded gracefully when the CRD is absent), and `get/list/watch` on `Namespaces`
(labels only, for `allowedTeams: {from: Selector}`).
- Secrets live in the `TerdutServer`'s namespace (the operator key) and in each
`TerdutAlertSource`'s namespace (the webhook Secret), so the operator needs Secret access in
every tenant namespace. Kubernetes RBAC has no owner-scoped grant finer than the namespace.
- **As built (2026-10): broader than the above.** The shipped default is a
`ClusterRole` with full verbs on `Secrets` in every namespace (the chart's
`rbac.namespaced: true` gives a `Role` in the release namespace only, which
cannot serve tenant namespaces), and the manager's cache is not restricted
to watched namespaces. A per-namespace `Role` split (a Role and RoleBinding per watched namespace, with the cache
restricted to them) is the intended end state, not implemented, and the decision (2026-10) is to stay
cluster-wide for now: treat this operator as able to read every Secret in the cluster.
- terdut-server's own RBAC is unaffected: the operator uses only its HTTP API (§6), never the
Kubernetes API for app-level state.
## 10. Relationship to `charts/terdut-server`
The chart's Deployment/Service/bootstrap-job templates are redundant once
`TerdutServer` exists — running both would mean two controllers (Helm and
this operator) reconciling the same Deployment, which is exactly the
conflict Kubernetes operators exist to avoid. The chart is repurposed into
an **installer chart**: it installs the operator + CRDs (and optionally one
`TerdutServer` CR from `values.yaml`, for users who want "helm install and
get a server" without hand-writing a CR) rather than templating the
Deployment directly.
The server chart's Deployment, Service and bootstrap Job are redundant once `TerdutServer`
exists: running both would have two controllers reconciling the same Deployment. The operator's
chart (`charts/terdut-operator`) installs the operator, CRDs and RBAC, and optionally one
`TerdutServer` from `values.yaml` (`terdutServer.enabled`). There is no migration from a
chart-based install and none is planned (§1); whatever the server chart deployed stays a separate
install until someone deletes it.
**No migration path from an existing chart-based install, and none is
planned (§1).** An earlier draft of this section spent most of its length on
one — `helm template` the current release's `values.yaml` into an equivalent
`TerdutServer` CR, uninstall or shrink the old release, and a whole
sub-question about who gets to call `/api/bootstrap` first, the chart's Job
or the operator — all of which presupposed the operator might end up
managing a server the chart had already deployed and bootstrapped. §1 rules
that out: the operator only ever manages servers it created itself, so
there's nothing to migrate and no bootstrap race to settle (§6 covers why
that race doesn't exist either). Adopting the operator means applying a
fresh `TerdutServer` CR; whatever the chart deployed before stays exactly
what it was, a separate install, until someone deletes it.
The operator's Deployment builder mirrors the chart's env block for the knobs both expose (the
chart's `deployment.yaml` and `buildEnv` in `internal/controller/terdutserver_deployment.go`);
a new server setting is added in `config.go`, the chart, and the operator, in that order.
## 11. Testing strategy
@@ -761,82 +407,27 @@ what it was, a separate install, until someone deletes it.
request/response shapes (already well-documented in
`internal/api/*_test.go` on the server side) — no real Postgres or real
terdut-server binary needed for controller unit tests.
- A smaller number of true end-to-end tests (`kind` cluster + real
terdut-server image + real Postgres) covering the golden path per CRD:
create `TerdutServer` → `TerdutTeam` → one of each child kind → verify via
terdut-server's own API that the objects exist with the right shape →
delete the CR → verify the server-side object is gone.
- The golden path (`kind` cluster + real terdut-server image + real Postgres:
create `TerdutServer` → `TerdutTeam` → `TerdutAlertSource`, verify through
terdut-server's own API, delete, verify it is gone) is a manual pass via
`examples/demo/run-demo.sh`, not a CI job.
## 12. Observability
- Standard controller-runtime metrics (reconcile duration/error counts) are
enough for v1 — no custom metrics.
- Every externally-visible action (bootstrap, key rotation-needed, delete-
and-recreate on the no-PUT resources, adopt-on-conflict) emits a
- Every externally-visible action (a team created or renamed, an integration recreated, a delete
that failed) emits a
Kubernetes `Event` on the CR, since that's what shows up in `kubectl
describe` and gitops tooling (Argo CD/Flux) surfaces without extra wiring.
## 13. Deferred / explicitly out of scope for this design
## 13. Deferred / out of scope
- **A real scoped service-account/token type in terdut-server — not merely
deferred, this is v1-blocking for §6 as written** (verified: without it,
§6's bootstrap flow has no working credential-rotation path and no clean
answer to the chart-vs-operator bootstrap race; see §6 points 1, 5, 6 and
§10). Sequence this server-side change *before* implementing the
`TerdutServer` controller's bootstrap logic, not after.
- **A version-discovery endpoint on terdut-server** (e.g. `GET /api/version`).
Neither this operator nor terdut-tui has one today — both independently
detect capability by probing specific routes (terdut-tui via `GET
/api/teams` 404-checking; this operator would otherwise need to invent
its own equivalent probe). An unattended reconciler is more exposed to a
silent breaking API change than an interactive TUI a human is watching;
raising this alongside the service-account request rather than inventing
another route-probe here.
- CloudNativePG support — same `spec.database` shape as Zalando should
extend to it, but the concrete field/Secret-naming conventions need their
own look.
- Cross-namespace `teamRef` on the child CRDs (`TerdutEscalationRule`,
`TerdutDeadmanSwitch`, `TerdutAlertSource`) — only `TerdutTeam.serverRef`
crosses namespaces in v1 (§2, §4.2, §4.6); these stay same-namespace as
their `TerdutTeam` until a real need for splitting them out shows up.
- Narrower-than-namespace RBAC for the §4.5 webhook Secret (Kubernetes RBAC
has no owner-scoped grant below the namespace itself, per §9) — revisit
if the widened per-tenant-namespace `Secret` access proves too broad in
practice.
- A mid-life `spec.teamRef` change on a child CRD (`TerdutEscalationRule`,
`TerdutDeadmanSwitch`, `TerdutAlertSource`) isn't specially detected —
noticed while grounding Stage 4 against source, not newly introduced by
it: all three controllers always resolve `spec.teamRef` fresh every
reconcile and act against whatever `TerdutTeam` that currently names,
trusting the server-side id already stored in `status` remains valid
there. There's no server-side verb that could move an existing
integration/policy/switch to a different team in place regardless, so
retargeting one onto a live child isn't a supported operation in v1 —
delete and recreate the CR instead.
- `TerdutServerSpec.OIDC` has no `trustEmail` field (nor `usernameClaim`,
`emailClaim`, `groupsClaim` — the "rather than being added here
speculatively" fields its own doc comment already names), unlike
`charts/terdut-server`'s own chart, which sets `oidc.trustEmail: true`
for the production install specifically because Authentik reports
`email_verified: false` and without it a user's first SSO sign-in
creates a second, empty account instead of linking to their existing
one (`terdut-server/README.md`'s own account of this). Found by
actually trying to stand up a second real `TerdutServer` against the
same Authentik provider (`terdut-demo`, `Ryuvia/charts#275`), not by
inspection: that install's OIDC config is otherwise a straight copy of
production's and runs with terdut-server's own default (`trustEmail:
false`) regardless, since the CRD has nowhere to put the override.
Worth closing if a second real OIDC install becomes routine rather than
a one-off exercise.
- Gitops-managed team *membership* (see §4.2).
- CloudNativePG support, alongside the Zalando `postgresClusterRef` (same `spec.database` shape).
- Gitops-managed team membership (§4.2).
- Per-namespace RBAC with a restricted cache (§9).
- Automatic Deployment restart on upstream Postgres credential rotation.
- `spec.pod.priorityClassName`, pod-label passthrough beyond
`spec.pod.annotations`, and a HorizontalPodAutoscaler for `TerdutServer`
— all considered alongside §4.1's `spec.pod` and explicitly left out of
that round: an HPA in particular would actively contradict
`spec.replicas`'s own stance that this operator doesn't support more
than one replica (the sweeper/notifier singleton constraint).
- Admission webhooks / CEL-only validation limits (e.g. verifying a
`teamRef` exists at admission time rather than surfacing it as a status
condition after the fact).
- `spec.pod.priorityClassName`, pod labels beyond annotations, and an HPA for `TerdutServer`.
- Admission webhooks beyond CEL (for example, checking a `teamRef` exists at admission time).
- A shared API types module or generated client for terdut-server, terdut-operator and terdut-tui.
- OLM packaging.
+1 -33
View File
@@ -61,38 +61,7 @@ vet: ## Run go vet against code.
.PHONY: test
test: manifests generate fmt vet setup-envtest ## Run tests.
KUBEBUILDER_ASSETS="$(shell "$(ENVTEST)" use $(ENVTEST_K8S_VERSION) --bin-dir "$(LOCALBIN)" -p path)" go test $$(go list ./... | grep -v /e2e) -coverprofile cover.out
# TODO(user): To use a different vendor for e2e tests, modify the setup under 'tests/e2e'.
# The default setup assumes Kind is pre-installed and builds/loads the Manager Docker image locally.
# kubectl kuberc is disabled by default for test isolation; enable with:
# - KUBECTL_KUBERC=true
# CertManager is installed by default; skip with:
# - CERT_MANAGER_INSTALL_SKIP=true
KIND_CLUSTER ?= terdut-operator-test-e2e
.PHONY: setup-test-e2e
setup-test-e2e: ## Set up a Kind cluster for e2e tests if it does not exist
@command -v $(KIND) >/dev/null 2>&1 || { \
echo "Kind is not installed. Please install Kind manually."; \
exit 1; \
}
@case "$$($(KIND) get clusters)" in \
*"$(KIND_CLUSTER)"*) \
echo "Kind cluster '$(KIND_CLUSTER)' already exists. Skipping creation." ;; \
*) \
echo "Creating Kind cluster '$(KIND_CLUSTER)'..."; \
$(KIND) create cluster --name $(KIND_CLUSTER) ;; \
esac
.PHONY: test-e2e
test-e2e: setup-test-e2e manifests generate fmt vet ## Run the e2e tests. Expected an isolated environment using Kind.
KIND=$(KIND) KIND_CLUSTER=$(KIND_CLUSTER) go test -tags=e2e ./test/e2e/ -v -ginkgo.v
$(MAKE) cleanup-test-e2e
.PHONY: cleanup-test-e2e
cleanup-test-e2e: ## Tear down the Kind cluster used for e2e tests
@$(KIND) delete cluster --name $(KIND_CLUSTER)
KUBEBUILDER_ASSETS="$(shell "$(ENVTEST)" use $(ENVTEST_K8S_VERSION) --bin-dir "$(LOCALBIN)" -p path)" go test $$(go list ./...) -coverprofile cover.out
.PHONY: lint
lint: golangci-lint ## Run golangci-lint linter
@@ -186,7 +155,6 @@ $(LOCALBIN):
## Tool Binaries
KUBECTL ?= kubectl
KIND ?= kind
KUSTOMIZE ?= $(LOCALBIN)/kustomize
CONTROLLER_GEN ?= $(LOCALBIN)/controller-gen
ENVTEST ?= $(LOCALBIN)/setup-envtest
-18
View File
@@ -31,24 +31,6 @@ resources:
kind: TerdutTeam
path: git.ryuvia.com/niklas/terdut-operator/api/v1alpha1
version: v1alpha1
- api:
crdVersion: v1
namespaced: true
controller: true
domain: ryuvia.com
group: terdut
kind: TerdutEscalationRule
path: git.ryuvia.com/niklas/terdut-operator/api/v1alpha1
version: v1alpha1
- api:
crdVersion: v1
namespaced: true
controller: true
domain: ryuvia.com
group: terdut
kind: TerdutDeadmanSwitch
path: git.ryuvia.com/niklas/terdut-operator/api/v1alpha1
version: v1alpha1
- api:
crdVersion: v1
namespaced: true
+22 -35
View File
@@ -1,44 +1,31 @@
# Terdut operator
Aims to expose most config as CRD's, so end users can self-service over gitops.
Exposes terdut-server's configuration as Kubernetes objects, so teams can self-service it over
gitops. See [DESIGN.md](./DESIGN.md) for the design (read its revision section first: it is the
current shape of credentials and the CRD catalog), and [examples/demo](./examples/demo) for a
working install.
See [DESIGN.md](./DESIGN.md) for the full design: CRD catalog and specs,
reconciliation semantics, bootstrap/auth, Postgres integration, RBAC, and the
relationship to `charts/terdut-server`. This README stays a short pitch; the
open questions it used to carry are now resolved decisions there (§2).
## CRDs
## CRD's
### TerdutServer
A terdut-server install: Deployment, Service, database wiring (a DSN, or a Zalando
`postgresClusterRef`), and an operator key Secret it hands to the server so the operator can
authenticate. `allowedTeams` consents to `TerdutTeam`s in other namespaces.
### terdutServers
Creates a server — Deployment, Service, database wiring, bootstrap, operator
credentials, and `allowedTeams` consent for cross-namespace teams. See
DESIGN.md §4.1, §4.6.
### TerdutTeam
One team on a server, possibly in another namespace (`serverRef`, gated by that server's
`allowedTeams`). It carries the team's whole configuration:
- `displayName` and OIDC group bindings
- `escalation` — the escalation ladder
- `deadmanSwitches` — dead man's switches, by name (switches not listed are removed)
### terdutTeams
- team name
- oidc groups
- `serverRef` — explicit reference to its `TerdutServer`, may be in a
different namespace (one team owns the server, others self-service a
team against it), gated by that `TerdutServer`'s own `allowedTeams`
field (DESIGN.md §2, §4.1, §4.2, §4.6)
### terdutEscalationrules
- rule
- `teamRef` — explicit reference to its `TerdutTeam` (DESIGN.md §2, §4.3)
### terdutDeadmansswitches
- rule
- `teamRef` (DESIGN.md §4.4)
### terdutAlertSources
- `teamRef` (DESIGN.md §4.5)
- URL/key are generated by the server at creation and surfaced only via a
generated Secret, never set explicitly
### TerdutAlertSource
An alert-ingest integration on a team (`teamRef`). The server shows the webhook key once; it is
surfaced only through a generated Secret next to the object, never set explicitly.
## Demo
[`examples/demo`](./examples/demo) wires one of every CRD above together
— two teams, each with an escalation rule, a dead man's switch and an
alert source — plus a script that fires synthetic Alertmanager webhooks
at it, so you can watch real incidents open, escalate and resolve without
a real Alertmanager anywhere in the picture.
[`examples/demo`](./examples/demo) wires one of each together — two teams, each with an
escalation ladder, a dead man's switch and an alert source — plus a script that fires synthetic
Alertmanager webhooks at it, so you can watch incidents open, escalate and resolve without a real
Alertmanager.
+20 -235
View File
@@ -1,240 +1,25 @@
# terdut-operator build roadmap
# terdut-operator status and deferred work
This is the staging plan for implementing the operator against `DESIGN.md`'s
settled decisions. It exists for the same reason `DESIGN.md` and
`SERVICE-ACCOUNTS.md` do: so each stage starts from an agreed sequencing
instead of re-litigating "what do we build first" mid-PR.
The staged build plan this file used to hold (scaffolding, TerdutServer, TerdutTeam, the child
kinds, the installer chart and first release) is done and shipped; the history is in git. The
design it implemented is in `DESIGN.md`, and its revision section at the top is the current
shape of credentials and the CRD catalog.
## Sequencing call this roadmap makes
## Validation
**The operator creates and owns every `TerdutServer` it manages — it never
adopts one deployed independently, by hand or by `charts/terdut-server`**
(`DESIGN.md` §1). An earlier version of this roadmap staged `TerdutServer`'s
Deployment/Service/bootstrap takeover separately (old Stage 5), behind a
hand-deployed server the simpler CRDs could be proven against first.
That staging existed only because a credential-less operator couldn't
`/api/bootstrap` its way into a server something else had already
bootstrapped (`DESIGN.md` §6's original gap). With no server to adopt at
all, that split has nothing left to justify it: `TerdutServer` now builds
its full lifecycle — Deployment, Service, database wiring, bootstrap,
credentials — in one stage, Stage 1, since bootstrap only has something to
bootstrap once the Deployment exists.
There is no CI end-to-end job. The golden path (create every kind against a real terdut-server on
`kind`, check the server's own API, delete, check it is gone) is a manual pass, using
`examples/demo/run-demo.sh`. It has not been re-run since the credential and CRD redesign
(2026-10): do that before the next release.
## Stage 0 — Scaffolding & CI
## Deferred
- `go.mod` (`git.ryuvia.com/niklas/terdut-operator`) + Kubebuilder v4
scaffold (`cmd/main.go`, `config/`, `Makefile`, `PROJECT`), matching
terdut-server's Go toolchain and house style (§3). Kubebuilder's own
scaffolded `Makefile` already wires `manifests`/`generate`
(`controller-gen`) and `setup-envtest` into `test`, and `golangci-lint`
into `lint`, all fetched on demand into `bin/` — no separate install
step needed beyond what `make test`/`make lint` already do.
- `.release.conf` deliberately **not** added yet: it names a `HELM_CHART`
this repo doesn't have until Stage 5. Adding it now would either be a
stub that lies about what's releasable or dead config nobody can run —
it lands in Stage 5, alongside the chart it describes.
- Gitea Actions CI calling `fmt lint test`, mirroring terdut-server's
`ci.yaml` convention (its `CLAUDE.md`: "a green gate here and a green
pipeline are the same code, not two descriptions of it") minus the
`chart`/`security` jobs, which need a chart (Stage 5) and real controller
code (Stage 1+) respectively to have anything to check.
- Drop kubebuilder's default `.github/workflows/*` scaffold — this org
runs on Gitea, not GitHub; `.gitea/workflows/ci.yaml` is the only CI this
repo has.
- Housekeeping: drop the stray `.DESIGN.md.swp` (leftover vim swapfile,
shouldn't be committed); correct `DESIGN.md` §6/§13's "v1-blocking, not
v1-shippable" language — the service-account feature it was blocking on
has since shipped in terdut-server.
**Done when:** CI is green on an otherwise-empty scaffold.
## Stage 1 — `TerdutServer`, full lifecycle
Supersedes the Stage 1 shipped before this redesign (commit `1be7cf2`)
outright — that `TerdutServerSpec`/`Status`/controller/tests implemented the
now-removed bring-your-own path and get replaced wholesale, not extended.
New commits build forward over the old ones; no git history rewrite.
- Full §4.1 spec: `image`, `replicas`, `networking`, `database`, `sweeper`,
`deadman`, `notify`, `oidc`, `passwordLogin`, `allowedTeams`, all together
— no narrowing, since bootstrap needs the Deployment it's narrowed away
from in the version this replaces.
- Controller manages the Deployment + Service, both Postgres paths from §8
at once (bring-your-own DSN *and* the Zalando `postgres-operator`
`postgresClusterRef` integration — not sequenced, per the user's call),
and bootstrap/credentials per §6's self-registration flow: `/api/bootstrap`
once the Deployment has a ready replica, checkpoint the admin key, mint
the instance-scoped service account, generated credentials Secret in the
operator's own namespace, `status.credentialsSecretRef`. Finalizer cleans
up that Secret (and the checkpoint, if one's still there) on delete —
there's no server-side row to clean up alongside it: terdut-server's API
has no way to delete a user or a service account, only to revoke
individual keys, so there's nothing to undo there regardless.
- RBAC: read-only watch on `postgresql.acid.zalan.do`, degrading gracefully
if that CRD isn't installed (§8, §9).
- Shipped, scoped down from §8's full ambition in one way, called out in
code rather than silently dropped: no live watch on the Zalando-
generated credentials Secret for rotation (relies on the periodic resync
to notice eventually, higher latency than a watch). A near-term
follow-up, not deferred to a later stage.
- External exposure (a Gateway API `HTTPRoute` from
`spec.networking.hostname`) was originally sketched here too, as a
second near-term follow-up alongside the one above. It's since become an
explicit, permanent non-goal instead (DESIGN.md §1): the operator will
never manage ingress/exposure for `TerdutServer` in any form. See
`examples/networking` for how to do that yourself.
- `envtest` covering Deployment/Service reconciliation and both database
paths — the Zalando path needs that CRD's schema vendored into the test
environment (there's no real `postgres-operator` controller in `envtest`,
only the CRD shape to create fixture objects against) — plus the
self-registration flow against an `httptest.Server` fake of
`/api/bootstrap` and `/api/service-accounts` (§11), including the
adopt-on-409 recovery path and the one fail-closed case
(`BootstrapStateLost`), not just the happy path.
- **Done, 2026-10-01**: a real `kind` end-to-end pass (bring-your-own DSN,
real terdut-server `v0.33.0` image, operator built into a real image and
deployed as a real Pod, not `go run` against the cluster). `TerdutServer`
went `Ready`; the generated credential authenticated and exercised its
real capability against the actual server
(`GET`/`POST /api/teams` → `200`/`201`, confirmed from terdut-server's own
access log). Caught two real bugs no `envtest` suite could have (its
client bypasses RBAC): `.dockerignore`'s `!**/*.go` not working under
podman, and missing RBAC for `events.k8s.io` (the new events API
`GetEventRecorder` uses) — both fixed. The Zalando path can additionally
be validated for real against the org's own cluster later, where
`postgres-operator` already runs, rather than only in a disposable `kind`
stand-in — not done in this pass.
## Stage 2 — `TerdutTeam`
- §4.2: `serverRef` resolution, real update-in-place (POST create / PUT
rename / PUT oidc-groups), team-scoped service-account minting once
`Ready` (§6 point 3), finalizer that DELETEs the team server-side and its
credential Secret.
- First place the "every child resolves its own `teamRef` →
`TerdutTeam.status`, never chains up to `TerdutServer`" pattern (§5) gets
proven end to end.
## Stage 3 — `TerdutEscalationRule` + `TerdutDeadmanSwitch`
- Built together: both stay same-namespace-as-their-`TerdutTeam` (§1), so
neither exercises cross-namespace complexity, but together they cover the
two different reconciliation shapes §5's table calls out — whole-policy
PUT-upsert for the escalation policy (no separate create step at all),
real create/update-in-place/delete for the dead man's switch (PUT added
in terdut-server `v0.33.0` specifically for this operator) — against the
same shared create/finalizer/resync scaffolding Stage 2 already built.
- New shared `resolveTeamAndClient` helper (`childref.go`) implements §5's
"every child resolves its own `teamRef` → `TerdutTeam.status`, never
chains up to `TerdutServer`" rule once, for both controllers —
`TerdutTeam.status.serverEndpoint`, added in this stage, is what makes
that literally true rather than just a stated intent.
- `TerdutEscalationRule` resolves each "user" target's username to a
user_id via `GET /api/users` (confirmed open to any authenticated
caller) and reports `Ready: False, reason: UnknownUser` if it doesn't
resolve. No `DELETE` exists for this resource, so its delete path `PUT`s
an empty policy as the closest available undo.
- `TerdutDeadmanSwitch` has no unique-name constraint server-side, so its
idempotent-create is `GET`-list-and-match-by-name rather than
adopt-on-409 (unlike every other resource in this operator).
- **Done, 2026-10-01**: `envtest` coverage for both controllers' happy
path, `TeamRefNotFound`/`WaitingForTeam`, `UnknownUser`, list-and-match
adoption, update-in-place on spec drift, and deletion. `make fmt lint
test build` all clean; `internal/controller` envtest coverage
50.5% → 71.7%. No `kind` e2e pass for this stage — Stage 1's already
proved the real-cluster mechanics (RBAC, image, bootstrap) these two
controllers reuse unchanged, and neither introduces a new mechanism that
pass would exercise differently (same reasoning Stage 2 used to skip
one).
## Stage 4 — `TerdutAlertSource`
- Last of the children on purpose: it has the subtlest failure mode of the
four. Covers webhook Secret generation/ownership (§4.5, §7), the
`WebhookSecretLost` fail-closed condition + `Warning` event (§5, added
2026-09-30), and the kind-change delete-and-recreate rotation path — all
easier to get right with the other three controllers' patterns already
in place to build on.
- Idempotent-create here is neither adopt-on-409 (Team/service-account) nor
list-and-match-by-name (`TerdutDeadmanSwitch`): terdut-server shows the
webhook key exactly once, at creation, and never again, so no server-side
lookup could ever recover it after a crash. The generated webhook Secret
itself — written immediately after the POST, before `status` is ever
touched — is this CR's only durable record that a create already
succeeded; found on a later reconcile with `status.integrationID` still
unset, it's read back directly rather than POSTing again. Found missing
with `status.integrationID` *set*, that's the already-designed
`WebhookSecretLost` fail-closed case instead.
- **Done, 2026-10-01**: `envtest` coverage for the happy path, rename
(PATCH, no key rotation), a kind change (delete-and-recreate, new id and
key), crash recovery between POST and the Secret write, `WebhookSecretLost`,
`TeamRefNotFound`/`WaitingForTeam`, and deletion. `make fmt lint test
build` all clean; `internal/controller` envtest coverage holds at 71.6%.
No `kind` e2e pass for this stage, same reasoning as Stage 3 (reuses
Stage 1's already-proven real-cluster mechanics unchanged).
## Stage 5 — Installer chart + real release
- Package CRDs + the operator's own Deployment/RBAC into the installer
chart §10 describes; wire `.release.conf`/release-vars the same way
terdut-server does; run it through the `release` skill for a real first
cut.
- Full `kind` end-to-end test per §11: create `TerdutServer` → `TerdutTeam`
→ one of each child kind → verify via terdut-server's own API that each
object exists with the right shape → delete the CR → verify the
server-side object is gone.
- Chart built via kubebuilder's own `helm/v2-alpha` plugin from `config/`'s
kustomize output (`charts/terdut-operator`), not hand-rolled -- CRDs +
manager Deployment/RBAC come from the same markers/manifests every other
stage already generates, so there's exactly one source of truth for
them. Hand-added on top: the optional `terdutServer` values block (§10's
"helm install and get a server" path), `.release.conf`, and the
`release-vars`/`helm-lint`/`push`/`helm-package`/`helm-push`/`release`
Makefile targets `.gitea/workflows/release.yaml` calls, mirroring
terdut-server's own shape end to end (same registry/namespace
convention, same multi-arch buildx push, same trivy/govulncheck/gitleaks
scans). Also fixed while wiring this: the Dockerfile's builder stage
didn't pin `--platform=$BUILDPLATFORM`, which would have made a
multi-arch release build fail outright on this org's runners (no binfmt
registration) -- caught before it ever shipped, not discovered mid-release;
and govulncheck surfaced one real, reachable finding (`google.golang.org/grpc`
v1.82.1, transitive via controller-runtime's otel exporter), fixed by
bumping to v1.83.1.
- **Done, 2026-10-01**: the full golden-path pass above, run for real
against a `kind` cluster, installed via `helm install` (not raw
kustomize/kubectl apply -- the first time the chart itself, not just
`config/`, was exercised): `TerdutServer` (real terdut-server `v0.33.0`
image, bring-your-own DSN against a throwaway in-cluster Postgres) →
`TerdutTeam` → one `TerdutEscalationRule` + `TerdutDeadmanSwitch` +
`TerdutAlertSource`, each confirmed `Ready` and then confirmed a second
way, independent of the operator's own status: a `curl` pod inside the
cluster, authenticated with the generated team credential, hit
terdut-server's real API directly (`GET /api/teams/{id}/escalation`,
`.../deadman/switches`, `.../integrations`) and got back exactly the
policy/switch/integration each spec declared. Deleting every CR in
reverse order was verified the same way: the escalation policy came back
empty (its only available "undo"), the switch and the integration were
both gone from their list endpoints, the team no longer resolved by
name, and the Deployment/Service/every generated Secret were gone from
the cluster. No new bugs found this pass -- Stage 1's own kind e2e pass
already caught the two issues (`events.k8s.io` RBAC, the podman
`.dockerignore` fix) a real cluster catches and `envtest` can't, and
nothing since has touched that surface.
- Not done in this pass, deliberately: an actual tagged release. `make
release-vars`/`helm-lint`/`push`/`helm-package`/`helm-push` all work
locally and `.gitea/workflows/release.yaml` is wired, but
`release-preflight` found there is no `terdut-operator/` entry under
`Ryuvia/charts` yet to bump -- every other onboarded repo had that
one-time wrapper-chart bootstrap done for it before its own first
release, and this one doesn't, since deploying this operator for real is
a decision for whoever runs the cluster, not a side effect of finishing
this stage. Cutting the first real release (and creating that wrapper
entry) is therefore the next action, not yet taken.
## Deferred (§13, unchanged by this roadmap)
Cross-namespace `allowedTeams` exercised against a real second namespace,
CloudNativePG support, narrower-than-namespace Secret RBAC for the webhook
Secret, gitops-managed team membership, automatic Deployment restart on
upstream Postgres credential rotation, admission webhooks/CEL-only
validation limits, OLM packaging.
- Narrower Secret RBAC: per-namespace Roles and a restricted cache (today a ClusterRole with
Secret access cluster-wide, DESIGN.md §9).
- CloudNativePG support alongside the Zalando `postgresClusterRef`.
- Gitops-managed team membership.
- Automatic Deployment restart on upstream Postgres credential rotation.
- Admission webhooks beyond CEL validation.
- OLM packaging.
- A shared API types module (or generated client) between terdut-server, terdut-operator and
terdut-tui, so contract drift is a compile error and not a manual mirror.
-98
View File
@@ -1,98 +0,0 @@
package v1alpha1
import (
metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
"k8s.io/apimachinery/pkg/runtime"
)
// TerdutDeadmanSwitchSpec defines the desired state of TerdutDeadmanSwitch.
//
// One per switch (DESIGN.md §4.4). Reconciled with real update-in-place
// (terdut-server v0.33.0 added PUT specifically for this, §5) -- but with no
// unique-name constraint server-side, idempotent-create here means
// GET-list-and-match-by-name, not adopt-on-409.
type TerdutDeadmanSwitchSpec struct {
// +required
TeamRef TerdutTeamRef `json:"teamRef"`
// name is optional, same as the API: left empty, terdut-server derives
// it from matcher's own canonical form, and that's what the
// idempotent-create lookup matches against too.
// +optional
Name string `json:"name,omitempty"`
// matcher names the alerts this switch watches, e.g.
// "alertname=Watchdog,cluster=prod". One matcher per switch -- add
// another TerdutDeadmanSwitch instead of separating with ";"
// (terdut-server's own restriction, mirrored here so a bad spec is
// rejected at apply time).
// +required
// +kubebuilder:validation:MinLength=1
// +kubebuilder:validation:XValidation:rule="!self.contains(';')",message="one matcher per switch: add another TerdutDeadmanSwitch instead of separating with ;"
Matcher string `json:"matcher"`
// timeout is a Go duration string, e.g. "15m".
// +required
// +kubebuilder:validation:MinLength=1
Timeout string `json:"timeout"`
// +kubebuilder:validation:Enum=critical;error;warning;info
// +kubebuilder:default=critical
// +optional
Severity string `json:"severity,omitempty"`
}
// TerdutDeadmanSwitchStatus defines the observed state of TerdutDeadmanSwitch.
type TerdutDeadmanSwitchStatus struct {
// +listType=map
// +listMapKey=type
// +optional
Conditions []metav1.Condition `json:"conditions,omitempty"`
// switchID is the server-side id.
// +optional
SwitchID int64 `json:"switchID,omitempty"`
// +optional
ObservedGeneration int64 `json:"observedGeneration,omitempty"`
}
// +kubebuilder:object:root=true
// +kubebuilder:subresource:status
// +kubebuilder:printcolumn:name="Team",type=string,JSONPath=`.spec.teamRef.name`
// +kubebuilder:printcolumn:name="SwitchID",type=integer,JSONPath=`.status.switchID`
// +kubebuilder:printcolumn:name="Ready",type=string,JSONPath=`.status.conditions[?(@.type=="Ready")].status`
// +kubebuilder:printcolumn:name="Reason",type=string,JSONPath=`.status.conditions[?(@.type=="Ready")].reason`
// TerdutDeadmanSwitch is the Schema for the terdutdeadmanswitches API
type TerdutDeadmanSwitch struct {
metav1.TypeMeta `json:",inline"`
// metadata is a standard object metadata
// +optional
metav1.ObjectMeta `json:"metadata,omitzero"`
// spec defines the desired state of TerdutDeadmanSwitch
// +required
Spec TerdutDeadmanSwitchSpec `json:"spec"`
// status defines the observed state of TerdutDeadmanSwitch
// +optional
Status TerdutDeadmanSwitchStatus `json:"status,omitzero"`
}
// +kubebuilder:object:root=true
// TerdutDeadmanSwitchList contains a list of TerdutDeadmanSwitch
type TerdutDeadmanSwitchList struct {
metav1.TypeMeta `json:",inline"`
metav1.ListMeta `json:"metadata,omitzero"`
Items []TerdutDeadmanSwitch `json:"items"`
}
func init() {
SchemeBuilder.Register(func(s *runtime.Scheme) error {
s.AddKnownTypes(SchemeGroupVersion, &TerdutDeadmanSwitch{}, &TerdutDeadmanSwitchList{})
return nil
})
}
-137
View File
@@ -1,137 +0,0 @@
package v1alpha1
import (
metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
"k8s.io/apimachinery/pkg/runtime"
)
// TerdutTeamRef names the TerdutTeam this resource belongs to. Always
// same-namespace as the CR itself (DESIGN.md §1: only TerdutTeam.spec.serverRef
// crosses namespaces in v1) -- no namespace field, unlike TerdutServerRef.
type TerdutTeamRef struct {
// +kubebuilder:validation:MinLength=1
Name string `json:"name"`
}
// EscalationTargetKind is who one rung of the ladder pages.
// +kubebuilder:validation:Enum=oncall;user
type EscalationTargetKind string
const (
EscalationTargetOncall EscalationTargetKind = "oncall"
EscalationTargetUser EscalationTargetKind = "user"
)
// EscalationTarget is one page within a level. username is required iff
// kind is "user" (terdut-server's own validation, internal/api/escalation.go's
// handleSetEscalation -- mirrored here as a CEL rule so a bad spec is
// rejected at apply time, not discovered on the next failed PUT).
// +kubebuilder:validation:XValidation:rule="self.kind != 'user' || has(self.username)",message="username is required when kind is user"
// +kubebuilder:validation:XValidation:rule="self.kind != 'oncall' || !has(self.username)",message="username must not be set when kind is oncall"
type EscalationTarget struct {
// +required
Kind EscalationTargetKind `json:"kind"`
// +optional
Username string `json:"username,omitempty"`
}
// EscalationLevel is one rung of the ladder: how long to wait, and who to
// page if nobody's acknowledged by then.
type EscalationLevel struct {
// timeout is a Go duration string, e.g. "5m".
// +required
// +kubebuilder:validation:MinLength=1
Timeout string `json:"timeout"`
// +required
// +kubebuilder:validation:MinItems=1
Targets []EscalationTarget `json:"targets"`
}
// TerdutEscalationRuleSpec defines the desired state of TerdutEscalationRule.
//
// One per team (DESIGN.md §4.3) -- terdut-server models a policy as one row
// with an owned list of levels, reconciled with a single whole-policy PUT.
// Not enforced at admission if two CRs name the same team (no webhooks in
// v1, §1); they would simply clobber each other every reconcile.
type TerdutEscalationRuleSpec struct {
// +required
TeamRef TerdutTeamRef `json:"teamRef"`
// +kubebuilder:validation:Minimum=0
// +kubebuilder:validation:Maximum=10
// +optional
RepeatCount int64 `json:"repeatCount,omitempty"`
// +optional
FallbackTopic string `json:"fallbackTopic,omitempty"`
// +required
// +kubebuilder:validation:MinItems=1
Levels []EscalationLevel `json:"levels"`
}
// Condition reasons shared by TerdutEscalationRule and TerdutDeadmanSwitch
// (both resolve a teamRef the same way, DESIGN.md §5).
const (
// ReasonTeamRefNotFound: spec.teamRef names no TerdutTeam (yet).
ReasonTeamRefNotFound = "TeamRefNotFound"
// ReasonWaitingForTeam: the referenced TerdutTeam exists but isn't
// Ready yet (no status.credentialsSecretRef to read).
ReasonWaitingForTeam = "WaitingForTeam"
// ReasonUnknownUser: an escalation target's username doesn't resolve to
// any user server-side (TerdutEscalationRule only).
ReasonUnknownUser = "UnknownUser"
// ReasonChildAdopted: the happy path, shared by both child kinds.
ReasonChildAdopted = "Adopted"
)
// TerdutEscalationRuleStatus defines the observed state of TerdutEscalationRule.
type TerdutEscalationRuleStatus struct {
// +listType=map
// +listMapKey=type
// +optional
Conditions []metav1.Condition `json:"conditions,omitempty"`
// +optional
ObservedGeneration int64 `json:"observedGeneration,omitempty"`
}
// +kubebuilder:object:root=true
// +kubebuilder:subresource:status
// +kubebuilder:printcolumn:name="Team",type=string,JSONPath=`.spec.teamRef.name`
// +kubebuilder:printcolumn:name="Ready",type=string,JSONPath=`.status.conditions[?(@.type=="Ready")].status`
// +kubebuilder:printcolumn:name="Reason",type=string,JSONPath=`.status.conditions[?(@.type=="Ready")].reason`
// TerdutEscalationRule is the Schema for the terdutescalationrules API
type TerdutEscalationRule struct {
metav1.TypeMeta `json:",inline"`
// metadata is a standard object metadata
// +optional
metav1.ObjectMeta `json:"metadata,omitzero"`
// spec defines the desired state of TerdutEscalationRule
// +required
Spec TerdutEscalationRuleSpec `json:"spec"`
// status defines the observed state of TerdutEscalationRule
// +optional
Status TerdutEscalationRuleStatus `json:"status,omitzero"`
}
// +kubebuilder:object:root=true
// TerdutEscalationRuleList contains a list of TerdutEscalationRule
type TerdutEscalationRuleList struct {
metav1.TypeMeta `json:",inline"`
metav1.ListMeta `json:"metadata,omitzero"`
Items []TerdutEscalationRule `json:"items"`
}
func init() {
SchemeBuilder.Register(func(s *runtime.Scheme) error {
s.AddKnownTypes(SchemeGroupVersion, &TerdutEscalationRule{}, &TerdutEscalationRuleList{})
return nil
})
}
+37 -61
View File
@@ -7,12 +7,11 @@ import (
"k8s.io/apimachinery/pkg/util/intstr"
)
// SecretKeyRef names one data key inside a Secret. Every use of this type in
// TerdutServerSpec resolves in the TerdutServer's own namespace (it's wired
// straight into the Deployment's pod spec as a secretKeyRef env source,
// which Kubernetes itself only allows same-namespace) -- unlike the
// generated credentials Secret (DESIGN.md §6), which always lives in the
// operator's own namespace and is never referenced through this type.
// SecretKeyRef names one data key inside a Secret in the TerdutServer's own
// namespace. Every use of this type is wired into the Deployment's pod spec as
// a secretKeyRef env source, which Kubernetes only allows same-namespace --
// including status.credentialsSecretRef, the operator key the controller
// generates there.
type SecretKeyRef struct {
// name is the Secret's name.
// +kubebuilder:validation:MinLength=1
@@ -104,18 +103,6 @@ type SweeperSpec struct {
ArchiveAfter string `json:"archiveAfter,omitempty"`
}
// DeadmanSpec controls dead man's switch alerts. Matchers/Timeout/Severity
// map straight to TERDUT_DEADMAN_MATCHERS/TERDUT_DEADMAN_TIMEOUT/
// TERDUT_DEADMAN_SEVERITY.
type DeadmanSpec struct {
// +optional
Matchers string `json:"matchers,omitempty"`
// +optional
Timeout string `json:"timeout,omitempty"`
// +optional
Severity string `json:"severity,omitempty"`
}
// NotifySpec controls push notifications via ntfy. Empty ntfyURL disables
// notifications entirely (matches the chart's own default).
type NotifySpec struct {
@@ -131,11 +118,7 @@ type NotifySpec struct {
TokenSecretRef *SecretKeyRef `json:"tokenSecretRef,omitempty"`
}
// OIDCSpec controls single sign-on. Fields the chart also exposes but
// DESIGN.md's spec doesn't (usernameClaim, emailClaim, groupsClaim,
// trustEmail) use terdut-server's own defaults
// (preferred_username/email/groups/false) rather than being added here
// speculatively.
// OIDCSpec controls single sign-on.
type OIDCSpec struct {
// +optional
Enabled bool `json:"enabled,omitempty"`
@@ -158,6 +141,21 @@ type OIDCSpec struct {
// +kubebuilder:default="12h"
// +optional
SessionMaxAge string `json:"sessionMaxAge,omitempty"`
// usernameClaim, emailClaim and groupsClaim name the ID token claims read.
// +kubebuilder:default="preferred_username"
// +optional
UsernameClaim string `json:"usernameClaim,omitempty"`
// +kubebuilder:default="email"
// +optional
EmailClaim string `json:"emailClaim,omitempty"`
// +kubebuilder:default="groups"
// +optional
GroupsClaim string `json:"groupsClaim,omitempty"`
// trustEmail links a sign-in to an existing local user by email even when
// the provider does not vouch the address is verified (Authentik reports
// email_verified false unless told otherwise).
// +optional
TrustEmail bool `json:"trustEmail,omitempty"`
}
// AllowedTeamsNamespaces gates which namespaces a TerdutTeam may resolve a
@@ -229,11 +227,11 @@ type PodSpec struct {
Tolerations []corev1.Toleration `json:"tolerations,omitempty"`
// affinity covers node affinity, pod affinity and pod anti-affinity in
// one field -- unlike a multi-replica-aware operator, this one never
// generates a default anti-affinity itself (replicas above 1 isn't a
// supported topology, see TerdutServerSpec.Replicas's own doc comment),
// so this is pure user-supplied passthrough, not a toggle-plus-generated-
// default.
// one field -- even though replicas now defaults to 2 (see
// TerdutServerSpec.Replicas's own doc comment), this operator still
// never generates a default anti-affinity of its own the way a
// multi-replica-aware operator typically would, so this stays pure
// user-supplied passthrough, not a toggle-plus-generated-default.
// +optional
Affinity *corev1.Affinity `json:"affinity,omitempty"`
@@ -301,10 +299,14 @@ type TerdutServerSpec struct {
// +required
Image ImageSpec `json:"image"`
// replicas. terdut-server is not horizontally-scale-tested; keep this
// at its default of 1 unless you've verified otherwise -- the sweeper
// and the notifier are unsynchronised singletons.
// +kubebuilder:default=1
// replicas. Defaults to 2: terdut-server v0.36.0 put the sweeper, the
// notifier and the migration runner each behind a Postgres advisory
// lock, and gave incident creation its own conflict resolution, so
// more than one replica no longer double-pages, races a migration, or
// drops a webhook payload. image.tag must be v0.36.0 or newer for
// that to hold -- an older terdut-server has none of these guards,
// and this field does not check the tag for you.
// +kubebuilder:default=2
// +optional
Replicas int32 `json:"replicas,omitempty"`
@@ -317,9 +319,6 @@ type TerdutServerSpec struct {
// +optional
Sweeper SweeperSpec `json:"sweeper,omitempty"`
// +optional
Deadman DeadmanSpec `json:"deadman,omitempty"`
// +optional
Notify NotifySpec `json:"notify,omitempty"`
@@ -349,15 +348,6 @@ const (
// ConditionReady is the standard top-level condition every CRD carries
// (DESIGN.md §7).
ConditionReady = "Ready"
// ConditionDatabaseReady reflects whether the configured database is
// usable -- for postgresClusterRef, whether the Zalando CR and its
// generated credentials Secret both resolved; for a plain dsn, always
// true once set (DESIGN.md §8: "no connectivity check beyond what the
// Deployment's own readiness probe already gives").
ConditionDatabaseReady = "DatabaseReady"
// ConditionBootstrapped reflects whether a working credential has been
// acquired via self-registration (DESIGN.md §6).
ConditionBootstrapped = "Bootstrapped"
)
// Condition reasons this controller sets.
@@ -375,13 +365,6 @@ const (
// instead -- but a TerdutServer that explicitly asks for it still needs
// to say clearly that it can't be satisfied).
ReasonPostgresOperatorCRDNotInstalled = "PostgresOperatorCRDNotInstalled"
// ReasonBootstrapStateLost: a checkpointed admin credential
// (DESIGN.md §6) was lost after being used but before the lasting
// credential it was for could be persisted -- the one genuinely
// pathological case in the self-registration flow. Fail-closed, same
// recovery as DESIGN.md §5's webhook-Secret-loss rule: delete and
// recreate this TerdutServer.
ReasonBootstrapStateLost = "BootstrapStateLost"
// ReasonAdopted: the happy path. A working credential is in hand, the
// Deployment has a ready replica, and the database (if postgresClusterRef)
// resolved.
@@ -402,16 +385,9 @@ type TerdutServerStatus struct {
// +optional
ObservedGeneration int64 `json:"observedGeneration,omitempty"`
// serviceName is the Service this controller created for the
// Deployment, so other objects can reference it without recomputing the
// naming convention.
// +optional
ServiceName string `json:"serviceName,omitempty"`
// credentialsSecretRef is the generated instance-scoped credential
// (DESIGN.md §6) -- pure output, always in the operator's own
// namespace, under a fixed data key ("token"). Set only once
// Bootstrapped is True.
// credentialsSecretRef is the operator key this controller generated for
// the server (TERDUT_OPERATOR_KEY): pure output, in the TerdutServer's own
// namespace and owned by it, under the data key "token".
// +optional
CredentialsSecretRef *SecretKeyRef `json:"credentialsSecretRef,omitempty"`
}
+120 -75
View File
@@ -27,52 +27,15 @@ type TerdutTeamOIDC struct {
OwnerGroup string `json:"ownerGroup,omitempty"`
}
// TerdutTeamInvite requests a standing invite link into this team, minted
// with the team's own team-scoped credential — requireTeamOwner already
// treats that credential as owner-equivalent for every /invites route
// (ratified, not a gap, as of terdut-server's SERVICE-ACCOUNTS.md). This is
// the real answer to "how does a human ever get a first login on a
// password-only, operator-managed install" (terdut-server#23): no signup_mode
// flip, no admin token, just a link redeemed the same way anyone else's
// invite would be.
type TerdutTeamInvite struct {
// enabled mints (and keeps refreshed ahead of terdut-server's own fixed
// 7-day TTL) an invite link while true. Flipping it back to false
// revokes the current one server-side rather than leaving it to expire
// on its own.
// +optional
Enabled bool `json:"enabled,omitempty"`
// role is what the invite grants: member or owner. Defaults to member —
// owner by default would make every invite link a standing
// administrative credential for the team, a much bigger blast radius
// than "let a human see the queue".
// +optional
// +kubebuilder:validation:Enum=member;owner
// +kubebuilder:default=member
Role string `json:"role,omitempty"`
// maxUses bounds how many times this link may be redeemed before it
// stops working, mirroring terdut-server's own 1-100 range
// (POST /api/teams/{teamID}/invites). Defaults to 1: a link meant for
// one specific person, not a standing door.
// +optional
// +kubebuilder:validation:Minimum=1
// +kubebuilder:validation:Maximum=100
// +kubebuilder:default=1
MaxUses int64 `json:"maxUses,omitempty"`
}
// TerdutTeamSpec defines the desired state of TerdutTeam.
type TerdutTeamSpec struct {
// serverRef names the TerdutServer this team belongs to.
// +required
ServerRef TerdutServerRef `json:"serverRef"`
// displayName is this team's name, both in terdut-server's own data
// (POST /api/teams {"name": ...}) and as the identity POST /api/teams
// and GET /api/teams?name= correlate on (DESIGN.md §5's idempotent-
// create rule, via TEAM-LOOKUP.md).
// displayName is this team's name on the server. It can be changed freely:
// the team is found by the CR's own identity (<namespace>/<name>, sent as
// external_id), not by this name.
// +required
// +kubebuilder:validation:MinLength=1
DisplayName string `json:"displayName"`
@@ -80,8 +43,107 @@ type TerdutTeamSpec struct {
// +optional
OIDC TerdutTeamOIDC `json:"oidc,omitempty"`
// escalation is this team's escalation ladder. Omitted, the team has none
// (the server's plain reminder behaviour applies).
// +optional
Invite TerdutTeamInvite `json:"invite,omitempty"`
Escalation *EscalationSpec `json:"escalation,omitempty"`
// deadmanSwitches are this team's dead man's switches, by name. Switches on
// the server that are not listed here are removed: in operator mode this
// list is the whole truth.
// +listType=map
// +listMapKey=name
// +kubebuilder:validation:MaxItems=50
// +optional
DeadmanSwitches []DeadmanSwitchSpec `json:"deadmanSwitches,omitempty"`
}
// EscalationTargetKind is who one rung of the ladder pages.
// +kubebuilder:validation:Enum=oncall;user
type EscalationTargetKind string
const (
EscalationTargetOncall EscalationTargetKind = "oncall"
EscalationTargetUser EscalationTargetKind = "user"
)
// EscalationTarget is one page within a level. username is required iff kind is
// "user".
// +kubebuilder:validation:XValidation:rule="self.kind != 'user' || has(self.username)",message="username is required when kind is user"
// +kubebuilder:validation:XValidation:rule="self.kind != 'oncall' || !has(self.username)",message="username must not be set when kind is oncall"
type EscalationTarget struct {
// +required
Kind EscalationTargetKind `json:"kind"`
// +kubebuilder:validation:MaxLength=255
// +optional
Username string `json:"username,omitempty"`
}
// EscalationLevel is one rung of the ladder: how long to wait, and who to page
// if nobody has acknowledged by then.
type EscalationLevel struct {
// timeout is a Go duration string, e.g. "5m".
// +required
// +kubebuilder:validation:MaxLength=32
// +kubebuilder:validation:Pattern=`^([0-9]+(\.[0-9]+)?(ns|us|µs|ms|s|m|h))+$`
Timeout string `json:"timeout"`
// +required
// +kubebuilder:validation:MinItems=1
// +kubebuilder:validation:MaxItems=20
Targets []EscalationTarget `json:"targets"`
}
// EscalationSpec is a team's escalation ladder.
type EscalationSpec struct {
// +kubebuilder:validation:Minimum=0
// +kubebuilder:validation:Maximum=10
// +optional
RepeatCount int64 `json:"repeatCount,omitempty"`
// +optional
FallbackTopic string `json:"fallbackTopic,omitempty"`
// +required
// +kubebuilder:validation:MinItems=1
// +kubebuilder:validation:MaxItems=10
Levels []EscalationLevel `json:"levels"`
}
// DeadmanSwitchSpec is one dead man's switch: the absence of an alert matching
// matcher for longer than timeout opens an incident.
type DeadmanSwitchSpec struct {
// name identifies the switch within the team.
// +required
// +kubebuilder:validation:MinLength=1
// +kubebuilder:validation:MaxLength=100
Name string `json:"name"`
// matcher names the alerts this switch watches, e.g.
// "alertname=Watchdog,cluster=prod". One matcher per switch.
// +required
// +kubebuilder:validation:MinLength=1
// +kubebuilder:validation:MaxLength=512
// +kubebuilder:validation:XValidation:rule="!self.contains(';')",message="one matcher per switch: add another entry instead of separating with ;"
Matcher string `json:"matcher"`
// timeout is a Go duration string, e.g. "15m".
// +required
// +kubebuilder:validation:MaxLength=32
// +kubebuilder:validation:Pattern=`^([0-9]+(\.[0-9]+)?(ns|us|µs|ms|s|m|h))+$`
Timeout string `json:"timeout"`
// +kubebuilder:validation:Enum=critical;error;warning;info
// +kubebuilder:default=critical
// +optional
Severity string `json:"severity,omitempty"`
}
// TerdutTeamRef names the TerdutTeam a TerdutAlertSource belongs to. Always
// same-namespace as the CR itself.
type TerdutTeamRef struct {
// +kubebuilder:validation:MinLength=1
Name string `json:"name"`
}
// Condition reasons this controller sets.
@@ -97,21 +159,28 @@ const (
// DESIGN.md §5's "every child requeues with backoff, no cross-
// controller RPC" rule.
ReasonWaitingForServer = "WaitingForServer"
// ReasonTeamNameTaken: the server already has a team with spec.displayName
// that belongs to a different TerdutTeam (or to a person). The name is global
// to the server, so the operator waits for one of them to change.
ReasonTeamNameTaken = "TeamNameTaken"
// ReasonUnknownUser: an escalation target's username matches no user on the
// server (yet).
ReasonUnknownUser = "UnknownUser"
// ReasonInvalidSpec: a duration in the spec does not parse.
ReasonInvalidSpec = "InvalidSpec"
// ReasonTeamAdopted: the happy path.
ReasonTeamAdopted = "Adopted"
)
// Condition reasons for spec.invite reconciliation (TerdutTeamInvite). Not
// surfaced on the Ready condition itself — an invite is a convenience, not
// a dependency anything else in this team's own readiness waits on — but
// recorded as Events and readable via `kubectl describe`.
// Condition reasons shared by the resources that hang off a TerdutTeam
// (currently TerdutAlertSource).
const (
// ReasonInviteMinted: spec.invite.enabled is true and status.inviteSecretRef
// is populated and live.
ReasonInviteMinted = "InviteMinted"
// ReasonInviteRevoked: spec.invite.enabled flipped back to false and the
// server-side invite was revoked (or there was nothing to revoke).
ReasonInviteRevoked = "InviteRevoked"
// ReasonTeamRefNotFound: spec.teamRef names no TerdutTeam (yet).
ReasonTeamRefNotFound = "TeamRefNotFound"
// ReasonWaitingForTeam: the referenced TerdutTeam exists but is not Ready.
ReasonWaitingForTeam = "WaitingForTeam"
// ReasonChildAdopted: the happy path.
ReasonChildAdopted = "Adopted"
)
// TerdutTeamStatus defines the observed state of TerdutTeam.
@@ -126,30 +195,6 @@ type TerdutTeamStatus struct {
// +optional
TeamID int64 `json:"teamID,omitempty"`
// credentialsSecretRef is this team's own scoped credential
// (DESIGN.md §6 point 3) -- pure output, always in the operator's own
// namespace, under a fixed data key ("token").
// +optional
CredentialsSecretRef *SecretKeyRef `json:"credentialsSecretRef,omitempty"`
// serverEndpoint is the resolved TerdutServer's base URL, resolved once
// here so no child controller (TerdutEscalationRule, TerdutDeadmanSwitch,
// TerdutAlertSource) ever needs its own RBAC on terdutservers just to
// find out where to send a request (DESIGN.md §5).
// +optional
ServerEndpoint string `json:"serverEndpoint,omitempty"`
// inviteSecretRef is this team's current invite link, if spec.invite.enabled.
// Unlike credentialsSecretRef, this lives in the TerdutTeam's OWN
// namespace, not the operator's: an invite is bounded, limited-use, and
// meant for this namespace's own human operators to read and hand out,
// not a durable high-privilege credential — same shape as
// TerdutAlertSource's status.webhookURLSecretRef, not TerdutServer's
// cross-namespace credentialsSecretRef. Nil whenever spec.invite.enabled
// is false or unset.
// +optional
InviteSecretRef *LocalSecretRef `json:"inviteSecretRef,omitempty"`
// +optional
ObservedGeneration int64 `json:"observedGeneration,omitempty"`
}
+37 -233
View File
@@ -73,16 +73,16 @@ func (in *DatabaseSpec) DeepCopy() *DatabaseSpec {
}
// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil.
func (in *DeadmanSpec) DeepCopyInto(out *DeadmanSpec) {
func (in *DeadmanSwitchSpec) DeepCopyInto(out *DeadmanSwitchSpec) {
*out = *in
}
// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new DeadmanSpec.
func (in *DeadmanSpec) DeepCopy() *DeadmanSpec {
// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new DeadmanSwitchSpec.
func (in *DeadmanSwitchSpec) DeepCopy() *DeadmanSwitchSpec {
if in == nil {
return nil
}
out := new(DeadmanSpec)
out := new(DeadmanSwitchSpec)
in.DeepCopyInto(out)
return out
}
@@ -107,6 +107,28 @@ func (in *EscalationLevel) DeepCopy() *EscalationLevel {
return out
}
// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil.
func (in *EscalationSpec) DeepCopyInto(out *EscalationSpec) {
*out = *in
if in.Levels != nil {
in, out := &in.Levels, &out.Levels
*out = make([]EscalationLevel, len(*in))
for i := range *in {
(*in)[i].DeepCopyInto(&(*out)[i])
}
}
}
// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new EscalationSpec.
func (in *EscalationSpec) DeepCopy() *EscalationSpec {
if in == nil {
return nil
}
out := new(EscalationSpec)
in.DeepCopyInto(out)
return out
}
// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil.
func (in *EscalationTarget) DeepCopyInto(out *EscalationTarget) {
*out = *in
@@ -481,207 +503,6 @@ func (in *TerdutAlertSourceStatus) DeepCopy() *TerdutAlertSourceStatus {
return out
}
// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil.
func (in *TerdutDeadmanSwitch) DeepCopyInto(out *TerdutDeadmanSwitch) {
*out = *in
out.TypeMeta = in.TypeMeta
in.ObjectMeta.DeepCopyInto(&out.ObjectMeta)
out.Spec = in.Spec
in.Status.DeepCopyInto(&out.Status)
}
// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new TerdutDeadmanSwitch.
func (in *TerdutDeadmanSwitch) DeepCopy() *TerdutDeadmanSwitch {
if in == nil {
return nil
}
out := new(TerdutDeadmanSwitch)
in.DeepCopyInto(out)
return out
}
// DeepCopyObject is an autogenerated deepcopy function, copying the receiver, creating a new runtime.Object.
func (in *TerdutDeadmanSwitch) DeepCopyObject() runtime.Object {
if c := in.DeepCopy(); c != nil {
return c
}
return nil
}
// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil.
func (in *TerdutDeadmanSwitchList) DeepCopyInto(out *TerdutDeadmanSwitchList) {
*out = *in
out.TypeMeta = in.TypeMeta
in.ListMeta.DeepCopyInto(&out.ListMeta)
if in.Items != nil {
in, out := &in.Items, &out.Items
*out = make([]TerdutDeadmanSwitch, len(*in))
for i := range *in {
(*in)[i].DeepCopyInto(&(*out)[i])
}
}
}
// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new TerdutDeadmanSwitchList.
func (in *TerdutDeadmanSwitchList) DeepCopy() *TerdutDeadmanSwitchList {
if in == nil {
return nil
}
out := new(TerdutDeadmanSwitchList)
in.DeepCopyInto(out)
return out
}
// DeepCopyObject is an autogenerated deepcopy function, copying the receiver, creating a new runtime.Object.
func (in *TerdutDeadmanSwitchList) DeepCopyObject() runtime.Object {
if c := in.DeepCopy(); c != nil {
return c
}
return nil
}
// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil.
func (in *TerdutDeadmanSwitchSpec) DeepCopyInto(out *TerdutDeadmanSwitchSpec) {
*out = *in
out.TeamRef = in.TeamRef
}
// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new TerdutDeadmanSwitchSpec.
func (in *TerdutDeadmanSwitchSpec) DeepCopy() *TerdutDeadmanSwitchSpec {
if in == nil {
return nil
}
out := new(TerdutDeadmanSwitchSpec)
in.DeepCopyInto(out)
return out
}
// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil.
func (in *TerdutDeadmanSwitchStatus) DeepCopyInto(out *TerdutDeadmanSwitchStatus) {
*out = *in
if in.Conditions != nil {
in, out := &in.Conditions, &out.Conditions
*out = make([]v1.Condition, len(*in))
for i := range *in {
(*in)[i].DeepCopyInto(&(*out)[i])
}
}
}
// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new TerdutDeadmanSwitchStatus.
func (in *TerdutDeadmanSwitchStatus) DeepCopy() *TerdutDeadmanSwitchStatus {
if in == nil {
return nil
}
out := new(TerdutDeadmanSwitchStatus)
in.DeepCopyInto(out)
return out
}
// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil.
func (in *TerdutEscalationRule) DeepCopyInto(out *TerdutEscalationRule) {
*out = *in
out.TypeMeta = in.TypeMeta
in.ObjectMeta.DeepCopyInto(&out.ObjectMeta)
in.Spec.DeepCopyInto(&out.Spec)
in.Status.DeepCopyInto(&out.Status)
}
// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new TerdutEscalationRule.
func (in *TerdutEscalationRule) DeepCopy() *TerdutEscalationRule {
if in == nil {
return nil
}
out := new(TerdutEscalationRule)
in.DeepCopyInto(out)
return out
}
// DeepCopyObject is an autogenerated deepcopy function, copying the receiver, creating a new runtime.Object.
func (in *TerdutEscalationRule) DeepCopyObject() runtime.Object {
if c := in.DeepCopy(); c != nil {
return c
}
return nil
}
// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil.
func (in *TerdutEscalationRuleList) DeepCopyInto(out *TerdutEscalationRuleList) {
*out = *in
out.TypeMeta = in.TypeMeta
in.ListMeta.DeepCopyInto(&out.ListMeta)
if in.Items != nil {
in, out := &in.Items, &out.Items
*out = make([]TerdutEscalationRule, len(*in))
for i := range *in {
(*in)[i].DeepCopyInto(&(*out)[i])
}
}
}
// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new TerdutEscalationRuleList.
func (in *TerdutEscalationRuleList) DeepCopy() *TerdutEscalationRuleList {
if in == nil {
return nil
}
out := new(TerdutEscalationRuleList)
in.DeepCopyInto(out)
return out
}
// DeepCopyObject is an autogenerated deepcopy function, copying the receiver, creating a new runtime.Object.
func (in *TerdutEscalationRuleList) DeepCopyObject() runtime.Object {
if c := in.DeepCopy(); c != nil {
return c
}
return nil
}
// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil.
func (in *TerdutEscalationRuleSpec) DeepCopyInto(out *TerdutEscalationRuleSpec) {
*out = *in
out.TeamRef = in.TeamRef
if in.Levels != nil {
in, out := &in.Levels, &out.Levels
*out = make([]EscalationLevel, len(*in))
for i := range *in {
(*in)[i].DeepCopyInto(&(*out)[i])
}
}
}
// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new TerdutEscalationRuleSpec.
func (in *TerdutEscalationRuleSpec) DeepCopy() *TerdutEscalationRuleSpec {
if in == nil {
return nil
}
out := new(TerdutEscalationRuleSpec)
in.DeepCopyInto(out)
return out
}
// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil.
func (in *TerdutEscalationRuleStatus) DeepCopyInto(out *TerdutEscalationRuleStatus) {
*out = *in
if in.Conditions != nil {
in, out := &in.Conditions, &out.Conditions
*out = make([]v1.Condition, len(*in))
for i := range *in {
(*in)[i].DeepCopyInto(&(*out)[i])
}
}
}
// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new TerdutEscalationRuleStatus.
func (in *TerdutEscalationRuleStatus) DeepCopy() *TerdutEscalationRuleStatus {
if in == nil {
return nil
}
out := new(TerdutEscalationRuleStatus)
in.DeepCopyInto(out)
return out
}
// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil.
func (in *TerdutServer) DeepCopyInto(out *TerdutServer) {
*out = *in
@@ -763,7 +584,6 @@ func (in *TerdutServerSpec) DeepCopyInto(out *TerdutServerSpec) {
out.Networking = in.Networking
in.Database.DeepCopyInto(&out.Database)
out.Sweeper = in.Sweeper
out.Deadman = in.Deadman
in.Notify.DeepCopyInto(&out.Notify)
in.OIDC.DeepCopyInto(&out.OIDC)
in.AllowedTeams.DeepCopyInto(&out.AllowedTeams)
@@ -812,7 +632,7 @@ func (in *TerdutTeam) DeepCopyInto(out *TerdutTeam) {
*out = *in
out.TypeMeta = in.TypeMeta
in.ObjectMeta.DeepCopyInto(&out.ObjectMeta)
out.Spec = in.Spec
in.Spec.DeepCopyInto(&out.Spec)
in.Status.DeepCopyInto(&out.Status)
}
@@ -834,21 +654,6 @@ func (in *TerdutTeam) DeepCopyObject() runtime.Object {
return nil
}
// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil.
func (in *TerdutTeamInvite) DeepCopyInto(out *TerdutTeamInvite) {
*out = *in
}
// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new TerdutTeamInvite.
func (in *TerdutTeamInvite) DeepCopy() *TerdutTeamInvite {
if in == nil {
return nil
}
out := new(TerdutTeamInvite)
in.DeepCopyInto(out)
return out
}
// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil.
func (in *TerdutTeamList) DeepCopyInto(out *TerdutTeamList) {
*out = *in
@@ -916,7 +721,16 @@ func (in *TerdutTeamSpec) DeepCopyInto(out *TerdutTeamSpec) {
*out = *in
out.ServerRef = in.ServerRef
out.OIDC = in.OIDC
out.Invite = in.Invite
if in.Escalation != nil {
in, out := &in.Escalation, &out.Escalation
*out = new(EscalationSpec)
(*in).DeepCopyInto(*out)
}
if in.DeadmanSwitches != nil {
in, out := &in.DeadmanSwitches, &out.DeadmanSwitches
*out = make([]DeadmanSwitchSpec, len(*in))
copy(*out, *in)
}
}
// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new TerdutTeamSpec.
@@ -939,16 +753,6 @@ func (in *TerdutTeamStatus) DeepCopyInto(out *TerdutTeamStatus) {
(*in)[i].DeepCopyInto(&(*out)[i])
}
}
if in.CredentialsSecretRef != nil {
in, out := &in.CredentialsSecretRef, &out.CredentialsSecretRef
*out = new(SecretKeyRef)
**out = **in
}
if in.InviteSecretRef != nil {
in, out := &in.InviteSecretRef, &out.InviteSecretRef
*out = new(LocalSecretRef)
**out = **in
}
}
// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new TerdutTeamStatus.
+2 -2
View File
@@ -6,8 +6,8 @@ type: application
# These fields decide nothing: `make helm-package` passes --version and
# --app-version from the release tag (same reasoning as terdut-server's own
# chart). They're for whoever reads the tree before a tag exists.
version: 0.3.0
appVersion: "v0.3.0"
version: 0.5.0
appVersion: "v0.5.0"
keywords:
- kubernetes
@@ -1,184 +0,0 @@
{{- if .Values.crd.enabled }}
apiVersion: apiextensions.k8s.io/v1
kind: CustomResourceDefinition
metadata:
annotations:
{{- if .Values.crd.keep }}
"helm.sh/resource-policy": keep
{{- end }}
controller-gen.kubebuilder.io/version: v0.22.0
name: terdutdeadmanswitches.terdut.ryuvia.com
spec:
group: terdut.ryuvia.com
names:
kind: TerdutDeadmanSwitch
listKind: TerdutDeadmanSwitchList
plural: terdutdeadmanswitches
singular: terdutdeadmanswitch
scope: Namespaced
versions:
- additionalPrinterColumns:
- jsonPath: .spec.teamRef.name
name: Team
type: string
- jsonPath: .status.switchID
name: SwitchID
type: integer
- jsonPath: .status.conditions[?(@.type=="Ready")].status
name: Ready
type: string
- jsonPath: .status.conditions[?(@.type=="Ready")].reason
name: Reason
type: string
name: v1alpha1
schema:
openAPIV3Schema:
description: TerdutDeadmanSwitch is the Schema for the terdutdeadmanswitches
API
properties:
apiVersion:
description: |-
APIVersion defines the versioned schema of this representation of an object.
Servers should convert recognized schemas to the latest internal value, and
may reject unrecognized values.
More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#resources
type: string
kind:
description: |-
Kind is a string value representing the REST resource this object represents.
Servers may infer this from the endpoint the client submits requests to.
Cannot be updated.
In CamelCase.
More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds
type: string
metadata:
type: object
spec:
description: spec defines the desired state of TerdutDeadmanSwitch
properties:
matcher:
description: |-
matcher names the alerts this switch watches, e.g.
"alertname=Watchdog,cluster=prod". One matcher per switch -- add
another TerdutDeadmanSwitch instead of separating with ";"
(terdut-server's own restriction, mirrored here so a bad spec is
rejected at apply time).
minLength: 1
type: string
x-kubernetes-validations:
- message: 'one matcher per switch: add another TerdutDeadmanSwitch
instead of separating with ;'
rule: '!self.contains('';'')'
name:
description: |-
name is optional, same as the API: left empty, terdut-server derives
it from matcher's own canonical form, and that's what the
idempotent-create lookup matches against too.
type: string
severity:
default: critical
enum:
- critical
- error
- warning
- info
type: string
teamRef:
description: |-
TerdutTeamRef names the TerdutTeam this resource belongs to. Always
same-namespace as the CR itself (DESIGN.md §1: only TerdutTeam.spec.serverRef
crosses namespaces in v1) -- no namespace field, unlike TerdutServerRef.
properties:
name:
minLength: 1
type: string
required:
- name
type: object
timeout:
description: timeout is a Go duration string, e.g. "15m".
minLength: 1
type: string
required:
- matcher
- teamRef
- timeout
type: object
status:
description: status defines the observed state of TerdutDeadmanSwitch
properties:
conditions:
items:
description: Condition contains details for one aspect of the current
state of this API Resource.
properties:
lastTransitionTime:
description: |-
lastTransitionTime is the last time the condition transitioned from one status to another.
This should be when the underlying condition changed. If that is not known, then using the time when the API field changed is acceptable.
format: date-time
type: string
message:
description: |-
message is a human readable message indicating details about the transition.
This may be an empty string.
maxLength: 32768
type: string
observedGeneration:
description: |-
observedGeneration represents the .metadata.generation that the condition was set based upon.
For instance, if .metadata.generation is currently 12, but the .status.conditions[x].observedGeneration is 9, the condition is out of date
with respect to the current state of the instance.
format: int64
minimum: 0
type: integer
reason:
description: |-
reason contains a programmatic identifier indicating the reason for the condition's last transition.
Producers of specific condition types may define expected values and meanings for this field,
and whether the values are considered a guaranteed API.
The value should be a CamelCase string.
This field may not be empty.
maxLength: 1024
minLength: 1
pattern: ^[A-Za-z]([A-Za-z0-9_,:]*[A-Za-z0-9_])?$
type: string
status:
description: status of the condition, one of True, False, Unknown.
enum:
- "True"
- "False"
- Unknown
type: string
type:
description: type of condition in CamelCase or in foo.example.com/CamelCase.
maxLength: 316
pattern: ^([a-z0-9]([-a-z0-9]*[a-z0-9])?(\.[a-z0-9]([-a-z0-9]*[a-z0-9])?)*/)?(([A-Za-z0-9][-A-Za-z0-9_.]*)?[A-Za-z0-9])$
type: string
required:
- lastTransitionTime
- message
- reason
- status
- type
type: object
type: array
x-kubernetes-list-map-keys:
- type
x-kubernetes-list-type: map
observedGeneration:
format: int64
type: integer
switchID:
description: switchID is the server-side id.
format: int64
type: integer
type: object
required:
- spec
type: object
served: true
storage: true
subresources:
status: {}
{{- end }}
@@ -1,195 +0,0 @@
{{- if .Values.crd.enabled }}
apiVersion: apiextensions.k8s.io/v1
kind: CustomResourceDefinition
metadata:
annotations:
{{- if .Values.crd.keep }}
"helm.sh/resource-policy": keep
{{- end }}
controller-gen.kubebuilder.io/version: v0.22.0
name: terdutescalationrules.terdut.ryuvia.com
spec:
group: terdut.ryuvia.com
names:
kind: TerdutEscalationRule
listKind: TerdutEscalationRuleList
plural: terdutescalationrules
singular: terdutescalationrule
scope: Namespaced
versions:
- additionalPrinterColumns:
- jsonPath: .spec.teamRef.name
name: Team
type: string
- jsonPath: .status.conditions[?(@.type=="Ready")].status
name: Ready
type: string
- jsonPath: .status.conditions[?(@.type=="Ready")].reason
name: Reason
type: string
name: v1alpha1
schema:
openAPIV3Schema:
description: TerdutEscalationRule is the Schema for the terdutescalationrules
API
properties:
apiVersion:
description: |-
APIVersion defines the versioned schema of this representation of an object.
Servers should convert recognized schemas to the latest internal value, and
may reject unrecognized values.
More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#resources
type: string
kind:
description: |-
Kind is a string value representing the REST resource this object represents.
Servers may infer this from the endpoint the client submits requests to.
Cannot be updated.
In CamelCase.
More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds
type: string
metadata:
type: object
spec:
description: spec defines the desired state of TerdutEscalationRule
properties:
fallbackTopic:
type: string
levels:
items:
description: |-
EscalationLevel is one rung of the ladder: how long to wait, and who to
page if nobody's acknowledged by then.
properties:
targets:
items:
description: |-
EscalationTarget is one page within a level. username is required iff
kind is "user" (terdut-server's own validation, internal/api/escalation.go's
handleSetEscalation -- mirrored here as a CEL rule so a bad spec is
rejected at apply time, not discovered on the next failed PUT).
properties:
kind:
description: EscalationTargetKind is who one rung of the
ladder pages.
enum:
- oncall
- user
type: string
username:
type: string
required:
- kind
type: object
x-kubernetes-validations:
- message: username is required when kind is user
rule: self.kind != 'user' || has(self.username)
- message: username must not be set when kind is oncall
rule: self.kind != 'oncall' || !has(self.username)
minItems: 1
type: array
timeout:
description: timeout is a Go duration string, e.g. "5m".
minLength: 1
type: string
required:
- targets
- timeout
type: object
minItems: 1
type: array
repeatCount:
format: int64
maximum: 10
minimum: 0
type: integer
teamRef:
description: |-
TerdutTeamRef names the TerdutTeam this resource belongs to. Always
same-namespace as the CR itself (DESIGN.md §1: only TerdutTeam.spec.serverRef
crosses namespaces in v1) -- no namespace field, unlike TerdutServerRef.
properties:
name:
minLength: 1
type: string
required:
- name
type: object
required:
- levels
- teamRef
type: object
status:
description: status defines the observed state of TerdutEscalationRule
properties:
conditions:
items:
description: Condition contains details for one aspect of the current
state of this API Resource.
properties:
lastTransitionTime:
description: |-
lastTransitionTime is the last time the condition transitioned from one status to another.
This should be when the underlying condition changed. If that is not known, then using the time when the API field changed is acceptable.
format: date-time
type: string
message:
description: |-
message is a human readable message indicating details about the transition.
This may be an empty string.
maxLength: 32768
type: string
observedGeneration:
description: |-
observedGeneration represents the .metadata.generation that the condition was set based upon.
For instance, if .metadata.generation is currently 12, but the .status.conditions[x].observedGeneration is 9, the condition is out of date
with respect to the current state of the instance.
format: int64
minimum: 0
type: integer
reason:
description: |-
reason contains a programmatic identifier indicating the reason for the condition's last transition.
Producers of specific condition types may define expected values and meanings for this field,
and whether the values are considered a guaranteed API.
The value should be a CamelCase string.
This field may not be empty.
maxLength: 1024
minLength: 1
pattern: ^[A-Za-z]([A-Za-z0-9_,:]*[A-Za-z0-9_])?$
type: string
status:
description: status of the condition, one of True, False, Unknown.
enum:
- "True"
- "False"
- Unknown
type: string
type:
description: type of condition in CamelCase or in foo.example.com/CamelCase.
maxLength: 316
pattern: ^([a-z0-9]([-a-z0-9]*[a-z0-9])?(\.[a-z0-9]([-a-z0-9]*[a-z0-9])?)*/)?(([A-Za-z0-9][-A-Za-z0-9_.]*)?[A-Za-z0-9])$
type: string
required:
- lastTransitionTime
- message
- reason
- status
- type
type: object
type: array
x-kubernetes-list-map-keys:
- type
x-kubernetes-list-type: map
observedGeneration:
format: int64
type: integer
type: object
required:
- spec
type: object
served: true
storage: true
subresources:
status: {}
{{- end }}
@@ -178,19 +178,6 @@ spec:
- message: exactly one of dsn or postgresClusterRef must be set
rule: '(has(self.dsn) ? 1 : 0) + (has(self.postgresClusterRef) ?
1 : 0) == 1'
deadman:
description: |-
DeadmanSpec controls dead man's switch alerts. Matchers/Timeout/Severity
map straight to TERDUT_DEADMAN_MATCHERS/TERDUT_DEADMAN_TIMEOUT/
TERDUT_DEADMAN_SEVERITY.
properties:
matchers:
type: string
severity:
type: string
timeout:
type: string
type: object
image:
description: ImageSpec is the terdut-server image to run.
properties:
@@ -262,12 +249,7 @@ spec:
type: object
type: object
oidc:
description: |-
OIDCSpec controls single sign-on. Fields the chart also exposes but
DESIGN.md's spec doesn't (usernameClaim, emailClaim, groupsClaim,
trustEmail) use terdut-server's own defaults
(preferred_username/email/groups/false) rather than being added here
speculatively.
description: OIDCSpec controls single sign-on.
properties:
adminGroup:
type: string
@@ -279,12 +261,11 @@ spec:
type: string
clientSecretRef:
description: |-
SecretKeyRef names one data key inside a Secret. Every use of this type in
TerdutServerSpec resolves in the TerdutServer's own namespace (it's wired
straight into the Deployment's pod spec as a secretKeyRef env source,
which Kubernetes itself only allows same-namespace) -- unlike the
generated credentials Secret (DESIGN.md §6), which always lives in the
operator's own namespace and is never referenced through this type.
SecretKeyRef names one data key inside a Secret in the TerdutServer's own
namespace. Every use of this type is wired into the Deployment's pod spec as
a secretKeyRef env source, which Kubernetes only allows same-namespace --
including status.credentialsSecretRef, the operator key the controller
generates there.
properties:
key:
description: key is the data key inside the Secret holding
@@ -299,8 +280,14 @@ spec:
- key
- name
type: object
emailClaim:
default: email
type: string
enabled:
type: boolean
groupsClaim:
default: groups
type: string
issuer:
type: string
name:
@@ -312,6 +299,17 @@ spec:
sessionMaxAge:
default: 12h
type: string
trustEmail:
description: |-
trustEmail links a sign-in to an existing local user by email even when
the provider does not vouch the address is verified (Authentik reports
email_verified false unless told otherwise).
type: boolean
usernameClaim:
default: preferred_username
description: usernameClaim, emailClaim and groupsClaim name the
ID token claims read.
type: string
type: object
passwordLogin:
default: true
@@ -327,11 +325,11 @@ spec:
affinity:
description: |-
affinity covers node affinity, pod affinity and pod anti-affinity in
one field -- unlike a multi-replica-aware operator, this one never
generates a default anti-affinity itself (replicas above 1 isn't a
supported topology, see TerdutServerSpec.Replicas's own doc comment),
so this is pure user-supplied passthrough, not a toggle-plus-generated-
default.
one field -- even though replicas now defaults to 2 (see
TerdutServerSpec.Replicas's own doc comment), this operator still
never generates a default anti-affinity of its own the way a
multi-replica-aware operator typically would, so this stays pure
user-supplied passthrough, not a toggle-plus-generated-default.
properties:
nodeAffinity:
description: Describes node affinity scheduling rules for
@@ -4304,11 +4302,15 @@ spec:
type: array
type: object
replicas:
default: 1
default: 2
description: |-
replicas. terdut-server is not horizontally-scale-tested; keep this
at its default of 1 unless you've verified otherwise -- the sweeper
and the notifier are unsynchronised singletons.
replicas. Defaults to 2: terdut-server v0.36.0 put the sweeper, the
notifier and the migration runner each behind a Postgres advisory
lock, and gave incident creation its own conflict resolution, so
more than one replica no longer double-pages, races a migration, or
drops a webhook payload. image.tag must be v0.36.0 or newer for
that to hold -- an older terdut-server has none of these guards,
and this field does not check the tag for you.
format: int32
type: integer
sweeper:
@@ -4395,10 +4397,9 @@ spec:
x-kubernetes-list-type: map
credentialsSecretRef:
description: |-
credentialsSecretRef is the generated instance-scoped credential
(DESIGN.md §6) -- pure output, always in the operator's own
namespace, under a fixed data key ("token"). Set only once
Bootstrapped is True.
credentialsSecretRef is the operator key this controller generated for
the server (TERDUT_OPERATOR_KEY): pure output, in the TerdutServer's own
namespace and owned by it, under the data key "token".
properties:
key:
description: key is the data key inside the Secret holding the
@@ -4420,12 +4421,6 @@ spec:
tells "applied" from "seen" (DESIGN.md §7).
format: int64
type: integer
serviceName:
description: |-
serviceName is the Service this controller created for the
Deployment, so other objects can reference it without recomputing the
naming convention.
type: string
type: object
required:
- spec
@@ -55,54 +55,121 @@ spec:
spec:
description: spec defines the desired state of TerdutTeam
properties:
deadmanSwitches:
description: |-
deadmanSwitches are this team's dead man's switches, by name. Switches on
the server that are not listed here are removed: in operator mode this
list is the whole truth.
items:
description: |-
DeadmanSwitchSpec is one dead man's switch: the absence of an alert matching
matcher for longer than timeout opens an incident.
properties:
matcher:
description: |-
matcher names the alerts this switch watches, e.g.
"alertname=Watchdog,cluster=prod". One matcher per switch.
maxLength: 512
minLength: 1
type: string
x-kubernetes-validations:
- message: 'one matcher per switch: add another entry instead
of separating with ;'
rule: '!self.contains('';'')'
name:
description: name identifies the switch within the team.
maxLength: 100
minLength: 1
type: string
severity:
default: critical
enum:
- critical
- error
- warning
- info
type: string
timeout:
description: timeout is a Go duration string, e.g. "15m".
maxLength: 32
pattern: ^([0-9]+(\.[0-9]+)?(ns|us|µs|ms|s|m|h))+$
type: string
required:
- matcher
- name
- timeout
type: object
maxItems: 50
type: array
x-kubernetes-list-map-keys:
- name
x-kubernetes-list-type: map
displayName:
description: |-
displayName is this team's name, both in terdut-server's own data
(POST /api/teams {"name": ...}) and as the identity POST /api/teams
and GET /api/teams?name= correlate on (DESIGN.md §5's idempotent-
create rule, via TEAM-LOOKUP.md).
displayName is this team's name on the server. It can be changed freely:
the team is found by the CR's own identity (<namespace>/<name>, sent as
external_id), not by this name.
minLength: 1
type: string
invite:
escalation:
description: |-
TerdutTeamInvite requests a standing invite link into this team, minted
with the team's own team-scoped credential — requireTeamOwner already
treats that credential as owner-equivalent for every /invites route
(ratified, not a gap, as of terdut-server's SERVICE-ACCOUNTS.md). This is
the real answer to "how does a human ever get a first login on a
password-only, operator-managed install" (terdut-server#23): no signup_mode
flip, no admin token, just a link redeemed the same way anyone else's
invite would be.
escalation is this team's escalation ladder. Omitted, the team has none
(the server's plain reminder behaviour applies).
properties:
enabled:
description: |-
enabled mints (and keeps refreshed ahead of terdut-server's own fixed
7-day TTL) an invite link while true. Flipping it back to false
revokes the current one server-side rather than leaving it to expire
on its own.
type: boolean
maxUses:
default: 1
description: |-
maxUses bounds how many times this link may be redeemed before it
stops working, mirroring terdut-server's own 1-100 range
(POST /api/teams/{teamID}/invites). Defaults to 1: a link meant for
one specific person, not a standing door.
format: int64
maximum: 100
minimum: 1
type: integer
role:
default: member
description: |-
role is what the invite grants: member or owner. Defaults to member —
owner by default would make every invite link a standing
administrative credential for the team, a much bigger blast radius
than "let a human see the queue".
enum:
- member
- owner
fallbackTopic:
type: string
levels:
items:
description: |-
EscalationLevel is one rung of the ladder: how long to wait, and who to page
if nobody has acknowledged by then.
properties:
targets:
items:
description: |-
EscalationTarget is one page within a level. username is required iff kind is
"user".
properties:
kind:
description: EscalationTargetKind is who one rung
of the ladder pages.
enum:
- oncall
- user
type: string
username:
maxLength: 255
type: string
required:
- kind
type: object
x-kubernetes-validations:
- message: username is required when kind is user
rule: self.kind != 'user' || has(self.username)
- message: username must not be set when kind is oncall
rule: self.kind != 'oncall' || !has(self.username)
maxItems: 20
minItems: 1
type: array
timeout:
description: timeout is a Go duration string, e.g. "5m".
maxLength: 32
pattern: ^([0-9]+(\.[0-9]+)?(ns|us|µs|ms|s|m|h))+$
type: string
required:
- targets
- timeout
type: object
maxItems: 10
minItems: 1
type: array
repeatCount:
format: int64
maximum: 10
minimum: 0
type: integer
required:
- levels
type: object
oidc:
description: |-
@@ -193,53 +260,9 @@ spec:
x-kubernetes-list-map-keys:
- type
x-kubernetes-list-type: map
credentialsSecretRef:
description: |-
credentialsSecretRef is this team's own scoped credential
(DESIGN.md §6 point 3) -- pure output, always in the operator's own
namespace, under a fixed data key ("token").
properties:
key:
description: key is the data key inside the Secret holding the
raw value.
minLength: 1
type: string
name:
description: name is the Secret's name.
minLength: 1
type: string
required:
- key
- name
type: object
inviteSecretRef:
description: |-
inviteSecretRef is this team's current invite link, if spec.invite.enabled.
Unlike credentialsSecretRef, this lives in the TerdutTeam's OWN
namespace, not the operator's: an invite is bounded, limited-use, and
meant for this namespace's own human operators to read and hand out,
not a durable high-privilege credential — same shape as
TerdutAlertSource's status.webhookURLSecretRef, not TerdutServer's
cross-namespace credentialsSecretRef. Nil whenever spec.invite.enabled
is false or unset.
properties:
name:
description: name is the Secret's name.
minLength: 1
type: string
required:
- name
type: object
observedGeneration:
format: int64
type: integer
serverEndpoint:
description: |-
serverEndpoint is the resolved TerdutServer's base URL, resolved once
here so no child controller (TerdutEscalationRule, TerdutDeadmanSwitch,
TerdutAlertSource) ever needs its own RBAC on terdutservers just to
find out where to send a request (DESIGN.md §5).
type: string
teamID:
description: |-
teamID is the server-side id -- needed by every child object's
@@ -74,8 +74,6 @@ rules:
- terdut.ryuvia.com
resources:
- terdutalertsources
- terdutdeadmanswitches
- terdutescalationrules
- terdutservers
- terdutteams
verbs:
@@ -90,8 +88,6 @@ rules:
- terdut.ryuvia.com
resources:
- terdutalertsources/finalizers
- terdutdeadmanswitches/finalizers
- terdutescalationrules/finalizers
- terdutservers/finalizers
- terdutteams/finalizers
verbs:
@@ -100,8 +96,6 @@ rules:
- terdut.ryuvia.com
resources:
- terdutalertsources/status
- terdutdeadmanswitches/status
- terdutescalationrules/status
- terdutservers/status
- terdutteams/status
verbs:
@@ -1,31 +0,0 @@
{{- if .Values.rbac.helpers.enabled }}
apiVersion: rbac.authorization.k8s.io/v1
{{- if .Values.rbac.namespaced }}
kind: Role
{{- else }}
kind: ClusterRole
{{- end }}
metadata:
{{- if .Values.rbac.namespaced }}
namespace: {{ .Release.Namespace }}
{{- end }}
labels:
app.kubernetes.io/managed-by: {{ .Release.Service }}
app.kubernetes.io/name: {{ include "terdut-operator.name" . }}
helm.sh/chart: {{ .Chart.Name }}-{{ .Chart.Version | replace "+" "_" }}
app.kubernetes.io/instance: {{ .Release.Name }}
name: {{ include "terdut-operator.resourceName" (dict "suffix" "terdutdeadmanswitch-admin-role" "context" $) }}
rules:
- apiGroups:
- terdut.ryuvia.com
resources:
- terdutdeadmanswitches
verbs:
- '*'
- apiGroups:
- terdut.ryuvia.com
resources:
- terdutdeadmanswitches/status
verbs:
- get
{{- end }}
@@ -1,37 +0,0 @@
{{- if .Values.rbac.helpers.enabled }}
apiVersion: rbac.authorization.k8s.io/v1
{{- if .Values.rbac.namespaced }}
kind: Role
{{- else }}
kind: ClusterRole
{{- end }}
metadata:
{{- if .Values.rbac.namespaced }}
namespace: {{ .Release.Namespace }}
{{- end }}
labels:
app.kubernetes.io/managed-by: {{ .Release.Service }}
app.kubernetes.io/name: {{ include "terdut-operator.name" . }}
helm.sh/chart: {{ .Chart.Name }}-{{ .Chart.Version | replace "+" "_" }}
app.kubernetes.io/instance: {{ .Release.Name }}
name: {{ include "terdut-operator.resourceName" (dict "suffix" "terdutdeadmanswitch-editor-role" "context" $) }}
rules:
- apiGroups:
- terdut.ryuvia.com
resources:
- terdutdeadmanswitches
verbs:
- create
- delete
- get
- list
- patch
- update
- watch
- apiGroups:
- terdut.ryuvia.com
resources:
- terdutdeadmanswitches/status
verbs:
- get
{{- end }}
@@ -1,33 +0,0 @@
{{- if .Values.rbac.helpers.enabled }}
apiVersion: rbac.authorization.k8s.io/v1
{{- if .Values.rbac.namespaced }}
kind: Role
{{- else }}
kind: ClusterRole
{{- end }}
metadata:
{{- if .Values.rbac.namespaced }}
namespace: {{ .Release.Namespace }}
{{- end }}
labels:
app.kubernetes.io/managed-by: {{ .Release.Service }}
app.kubernetes.io/name: {{ include "terdut-operator.name" . }}
helm.sh/chart: {{ .Chart.Name }}-{{ .Chart.Version | replace "+" "_" }}
app.kubernetes.io/instance: {{ .Release.Name }}
name: {{ include "terdut-operator.resourceName" (dict "suffix" "terdutdeadmanswitch-viewer-role" "context" $) }}
rules:
- apiGroups:
- terdut.ryuvia.com
resources:
- terdutdeadmanswitches
verbs:
- get
- list
- watch
- apiGroups:
- terdut.ryuvia.com
resources:
- terdutdeadmanswitches/status
verbs:
- get
{{- end }}
@@ -1,31 +0,0 @@
{{- if .Values.rbac.helpers.enabled }}
apiVersion: rbac.authorization.k8s.io/v1
{{- if .Values.rbac.namespaced }}
kind: Role
{{- else }}
kind: ClusterRole
{{- end }}
metadata:
{{- if .Values.rbac.namespaced }}
namespace: {{ .Release.Namespace }}
{{- end }}
labels:
app.kubernetes.io/managed-by: {{ .Release.Service }}
app.kubernetes.io/name: {{ include "terdut-operator.name" . }}
helm.sh/chart: {{ .Chart.Name }}-{{ .Chart.Version | replace "+" "_" }}
app.kubernetes.io/instance: {{ .Release.Name }}
name: {{ include "terdut-operator.resourceName" (dict "suffix" "terdutescalationrule-admin-role" "context" $) }}
rules:
- apiGroups:
- terdut.ryuvia.com
resources:
- terdutescalationrules
verbs:
- '*'
- apiGroups:
- terdut.ryuvia.com
resources:
- terdutescalationrules/status
verbs:
- get
{{- end }}
@@ -1,37 +0,0 @@
{{- if .Values.rbac.helpers.enabled }}
apiVersion: rbac.authorization.k8s.io/v1
{{- if .Values.rbac.namespaced }}
kind: Role
{{- else }}
kind: ClusterRole
{{- end }}
metadata:
{{- if .Values.rbac.namespaced }}
namespace: {{ .Release.Namespace }}
{{- end }}
labels:
app.kubernetes.io/managed-by: {{ .Release.Service }}
app.kubernetes.io/name: {{ include "terdut-operator.name" . }}
helm.sh/chart: {{ .Chart.Name }}-{{ .Chart.Version | replace "+" "_" }}
app.kubernetes.io/instance: {{ .Release.Name }}
name: {{ include "terdut-operator.resourceName" (dict "suffix" "terdutescalationrule-editor-role" "context" $) }}
rules:
- apiGroups:
- terdut.ryuvia.com
resources:
- terdutescalationrules
verbs:
- create
- delete
- get
- list
- patch
- update
- watch
- apiGroups:
- terdut.ryuvia.com
resources:
- terdutescalationrules/status
verbs:
- get
{{- end }}
@@ -1,33 +0,0 @@
{{- if .Values.rbac.helpers.enabled }}
apiVersion: rbac.authorization.k8s.io/v1
{{- if .Values.rbac.namespaced }}
kind: Role
{{- else }}
kind: ClusterRole
{{- end }}
metadata:
{{- if .Values.rbac.namespaced }}
namespace: {{ .Release.Namespace }}
{{- end }}
labels:
app.kubernetes.io/managed-by: {{ .Release.Service }}
app.kubernetes.io/name: {{ include "terdut-operator.name" . }}
helm.sh/chart: {{ .Chart.Name }}-{{ .Chart.Version | replace "+" "_" }}
app.kubernetes.io/instance: {{ .Release.Name }}
name: {{ include "terdut-operator.resourceName" (dict "suffix" "terdutescalationrule-viewer-role" "context" $) }}
rules:
- apiGroups:
- terdut.ryuvia.com
resources:
- terdutescalationrules
verbs:
- get
- list
- watch
- apiGroups:
- terdut.ryuvia.com
resources:
- terdutescalationrules/status
verbs:
- get
{{- end }}
@@ -34,10 +34,6 @@ spec:
sweeper:
{{- toYaml . | nindent 4 }}
{{- end }}
{{- with .Values.terdutServer.deadman }}
deadman:
{{- toYaml . | nindent 4 }}
{{- end }}
{{- with .Values.terdutServer.notify }}
notify:
{{- toYaml . | nindent 4 }}
+5 -5
View File
@@ -233,7 +233,11 @@ terdutServer:
## Required when terdutServer.enabled.
# tag: ""
replicas: 1
## Safe above 1 since terdut-server v0.36.0 (image.tag above must be that or
## newer): the sweeper, notifier and migration runner are each behind a
## Postgres advisory lock, and incident creation resolves its own insert
## conflict, matching this CRD's own spec.replicas default.
replicas: 2
networking:
## Required when terdutServer.enabled -- terdut-server's own public
@@ -262,10 +266,6 @@ terdutServer:
# sweeper:
# staleAfter: 6h
# archiveAfter: 168h
# deadman:
# matchers: "alertname=Watchdog"
# timeout: 15m
# severity: critical
# notify: {}
# oidc: {}
+6 -36
View File
@@ -166,53 +166,23 @@ func main() {
os.Exit(1)
}
// POD_NAMESPACE is the operator's own namespace, via the Deployment's
// downward API (config/manager/manager.yaml) — every credentials Secret
// TerdutServerReconciler reads or writes lives here, never in a
// TerdutServer's own namespace (DESIGN.md §6). Falling back to "default"
// keeps `go run` usable for local development against a real cluster;
// production always sets it.
operatorNamespace := os.Getenv("POD_NAMESPACE")
if operatorNamespace == "" {
operatorNamespace = "default"
}
if err := (&controller.TerdutServerReconciler{
Client: mgr.GetClient(),
Scheme: mgr.GetScheme(),
OperatorNamespace: operatorNamespace,
Client: mgr.GetClient(),
Scheme: mgr.GetScheme(),
}).SetupWithManager(mgr); err != nil {
setupLog.Error(err, "Failed to create controller", "controller", "terdutserver")
os.Exit(1)
}
if err := (&controller.TerdutTeamReconciler{
Client: mgr.GetClient(),
Scheme: mgr.GetScheme(),
OperatorNamespace: operatorNamespace,
Client: mgr.GetClient(),
Scheme: mgr.GetScheme(),
}).SetupWithManager(mgr); err != nil {
setupLog.Error(err, "Failed to create controller", "controller", "terdutteam")
os.Exit(1)
}
if err := (&controller.TerdutEscalationRuleReconciler{
Client: mgr.GetClient(),
Scheme: mgr.GetScheme(),
OperatorNamespace: operatorNamespace,
}).SetupWithManager(mgr); err != nil {
setupLog.Error(err, "Failed to create controller", "controller", "terdutescalationrule")
os.Exit(1)
}
if err := (&controller.TerdutDeadmanSwitchReconciler{
Client: mgr.GetClient(),
Scheme: mgr.GetScheme(),
OperatorNamespace: operatorNamespace,
}).SetupWithManager(mgr); err != nil {
setupLog.Error(err, "Failed to create controller", "controller", "terdutdeadmanswitch")
os.Exit(1)
}
if err := (&controller.TerdutAlertSourceReconciler{
Client: mgr.GetClient(),
Scheme: mgr.GetScheme(),
OperatorNamespace: operatorNamespace,
Client: mgr.GetClient(),
Scheme: mgr.GetScheme(),
}).SetupWithManager(mgr); err != nil {
setupLog.Error(err, "Failed to create controller", "controller", "terdutalertsource")
os.Exit(1)
@@ -76,9 +76,8 @@ spec:
type: string
teamRef:
description: |-
TerdutTeamRef names the TerdutTeam this resource belongs to. Always
same-namespace as the CR itself (DESIGN.md §1: only TerdutTeam.spec.serverRef
crosses namespaces in v1) -- no namespace field, unlike TerdutServerRef.
TerdutTeamRef names the TerdutTeam a TerdutAlertSource belongs to. Always
same-namespace as the CR itself.
properties:
name:
minLength: 1
@@ -1,180 +0,0 @@
---
apiVersion: apiextensions.k8s.io/v1
kind: CustomResourceDefinition
metadata:
annotations:
controller-gen.kubebuilder.io/version: v0.22.0
name: terdutdeadmanswitches.terdut.ryuvia.com
spec:
group: terdut.ryuvia.com
names:
kind: TerdutDeadmanSwitch
listKind: TerdutDeadmanSwitchList
plural: terdutdeadmanswitches
singular: terdutdeadmanswitch
scope: Namespaced
versions:
- additionalPrinterColumns:
- jsonPath: .spec.teamRef.name
name: Team
type: string
- jsonPath: .status.switchID
name: SwitchID
type: integer
- jsonPath: .status.conditions[?(@.type=="Ready")].status
name: Ready
type: string
- jsonPath: .status.conditions[?(@.type=="Ready")].reason
name: Reason
type: string
name: v1alpha1
schema:
openAPIV3Schema:
description: TerdutDeadmanSwitch is the Schema for the terdutdeadmanswitches
API
properties:
apiVersion:
description: |-
APIVersion defines the versioned schema of this representation of an object.
Servers should convert recognized schemas to the latest internal value, and
may reject unrecognized values.
More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#resources
type: string
kind:
description: |-
Kind is a string value representing the REST resource this object represents.
Servers may infer this from the endpoint the client submits requests to.
Cannot be updated.
In CamelCase.
More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds
type: string
metadata:
type: object
spec:
description: spec defines the desired state of TerdutDeadmanSwitch
properties:
matcher:
description: |-
matcher names the alerts this switch watches, e.g.
"alertname=Watchdog,cluster=prod". One matcher per switch -- add
another TerdutDeadmanSwitch instead of separating with ";"
(terdut-server's own restriction, mirrored here so a bad spec is
rejected at apply time).
minLength: 1
type: string
x-kubernetes-validations:
- message: 'one matcher per switch: add another TerdutDeadmanSwitch
instead of separating with ;'
rule: '!self.contains('';'')'
name:
description: |-
name is optional, same as the API: left empty, terdut-server derives
it from matcher's own canonical form, and that's what the
idempotent-create lookup matches against too.
type: string
severity:
default: critical
enum:
- critical
- error
- warning
- info
type: string
teamRef:
description: |-
TerdutTeamRef names the TerdutTeam this resource belongs to. Always
same-namespace as the CR itself (DESIGN.md §1: only TerdutTeam.spec.serverRef
crosses namespaces in v1) -- no namespace field, unlike TerdutServerRef.
properties:
name:
minLength: 1
type: string
required:
- name
type: object
timeout:
description: timeout is a Go duration string, e.g. "15m".
minLength: 1
type: string
required:
- matcher
- teamRef
- timeout
type: object
status:
description: status defines the observed state of TerdutDeadmanSwitch
properties:
conditions:
items:
description: Condition contains details for one aspect of the current
state of this API Resource.
properties:
lastTransitionTime:
description: |-
lastTransitionTime is the last time the condition transitioned from one status to another.
This should be when the underlying condition changed. If that is not known, then using the time when the API field changed is acceptable.
format: date-time
type: string
message:
description: |-
message is a human readable message indicating details about the transition.
This may be an empty string.
maxLength: 32768
type: string
observedGeneration:
description: |-
observedGeneration represents the .metadata.generation that the condition was set based upon.
For instance, if .metadata.generation is currently 12, but the .status.conditions[x].observedGeneration is 9, the condition is out of date
with respect to the current state of the instance.
format: int64
minimum: 0
type: integer
reason:
description: |-
reason contains a programmatic identifier indicating the reason for the condition's last transition.
Producers of specific condition types may define expected values and meanings for this field,
and whether the values are considered a guaranteed API.
The value should be a CamelCase string.
This field may not be empty.
maxLength: 1024
minLength: 1
pattern: ^[A-Za-z]([A-Za-z0-9_,:]*[A-Za-z0-9_])?$
type: string
status:
description: status of the condition, one of True, False, Unknown.
enum:
- "True"
- "False"
- Unknown
type: string
type:
description: type of condition in CamelCase or in foo.example.com/CamelCase.
maxLength: 316
pattern: ^([a-z0-9]([-a-z0-9]*[a-z0-9])?(\.[a-z0-9]([-a-z0-9]*[a-z0-9])?)*/)?(([A-Za-z0-9][-A-Za-z0-9_.]*)?[A-Za-z0-9])$
type: string
required:
- lastTransitionTime
- message
- reason
- status
- type
type: object
type: array
x-kubernetes-list-map-keys:
- type
x-kubernetes-list-type: map
observedGeneration:
format: int64
type: integer
switchID:
description: switchID is the server-side id.
format: int64
type: integer
type: object
required:
- spec
type: object
served: true
storage: true
subresources:
status: {}
@@ -1,191 +0,0 @@
---
apiVersion: apiextensions.k8s.io/v1
kind: CustomResourceDefinition
metadata:
annotations:
controller-gen.kubebuilder.io/version: v0.22.0
name: terdutescalationrules.terdut.ryuvia.com
spec:
group: terdut.ryuvia.com
names:
kind: TerdutEscalationRule
listKind: TerdutEscalationRuleList
plural: terdutescalationrules
singular: terdutescalationrule
scope: Namespaced
versions:
- additionalPrinterColumns:
- jsonPath: .spec.teamRef.name
name: Team
type: string
- jsonPath: .status.conditions[?(@.type=="Ready")].status
name: Ready
type: string
- jsonPath: .status.conditions[?(@.type=="Ready")].reason
name: Reason
type: string
name: v1alpha1
schema:
openAPIV3Schema:
description: TerdutEscalationRule is the Schema for the terdutescalationrules
API
properties:
apiVersion:
description: |-
APIVersion defines the versioned schema of this representation of an object.
Servers should convert recognized schemas to the latest internal value, and
may reject unrecognized values.
More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#resources
type: string
kind:
description: |-
Kind is a string value representing the REST resource this object represents.
Servers may infer this from the endpoint the client submits requests to.
Cannot be updated.
In CamelCase.
More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds
type: string
metadata:
type: object
spec:
description: spec defines the desired state of TerdutEscalationRule
properties:
fallbackTopic:
type: string
levels:
items:
description: |-
EscalationLevel is one rung of the ladder: how long to wait, and who to
page if nobody's acknowledged by then.
properties:
targets:
items:
description: |-
EscalationTarget is one page within a level. username is required iff
kind is "user" (terdut-server's own validation, internal/api/escalation.go's
handleSetEscalation -- mirrored here as a CEL rule so a bad spec is
rejected at apply time, not discovered on the next failed PUT).
properties:
kind:
description: EscalationTargetKind is who one rung of the
ladder pages.
enum:
- oncall
- user
type: string
username:
type: string
required:
- kind
type: object
x-kubernetes-validations:
- message: username is required when kind is user
rule: self.kind != 'user' || has(self.username)
- message: username must not be set when kind is oncall
rule: self.kind != 'oncall' || !has(self.username)
minItems: 1
type: array
timeout:
description: timeout is a Go duration string, e.g. "5m".
minLength: 1
type: string
required:
- targets
- timeout
type: object
minItems: 1
type: array
repeatCount:
format: int64
maximum: 10
minimum: 0
type: integer
teamRef:
description: |-
TerdutTeamRef names the TerdutTeam this resource belongs to. Always
same-namespace as the CR itself (DESIGN.md §1: only TerdutTeam.spec.serverRef
crosses namespaces in v1) -- no namespace field, unlike TerdutServerRef.
properties:
name:
minLength: 1
type: string
required:
- name
type: object
required:
- levels
- teamRef
type: object
status:
description: status defines the observed state of TerdutEscalationRule
properties:
conditions:
items:
description: Condition contains details for one aspect of the current
state of this API Resource.
properties:
lastTransitionTime:
description: |-
lastTransitionTime is the last time the condition transitioned from one status to another.
This should be when the underlying condition changed. If that is not known, then using the time when the API field changed is acceptable.
format: date-time
type: string
message:
description: |-
message is a human readable message indicating details about the transition.
This may be an empty string.
maxLength: 32768
type: string
observedGeneration:
description: |-
observedGeneration represents the .metadata.generation that the condition was set based upon.
For instance, if .metadata.generation is currently 12, but the .status.conditions[x].observedGeneration is 9, the condition is out of date
with respect to the current state of the instance.
format: int64
minimum: 0
type: integer
reason:
description: |-
reason contains a programmatic identifier indicating the reason for the condition's last transition.
Producers of specific condition types may define expected values and meanings for this field,
and whether the values are considered a guaranteed API.
The value should be a CamelCase string.
This field may not be empty.
maxLength: 1024
minLength: 1
pattern: ^[A-Za-z]([A-Za-z0-9_,:]*[A-Za-z0-9_])?$
type: string
status:
description: status of the condition, one of True, False, Unknown.
enum:
- "True"
- "False"
- Unknown
type: string
type:
description: type of condition in CamelCase or in foo.example.com/CamelCase.
maxLength: 316
pattern: ^([a-z0-9]([-a-z0-9]*[a-z0-9])?(\.[a-z0-9]([-a-z0-9]*[a-z0-9])?)*/)?(([A-Za-z0-9][-A-Za-z0-9_.]*)?[A-Za-z0-9])$
type: string
required:
- lastTransitionTime
- message
- reason
- status
- type
type: object
type: array
x-kubernetes-list-map-keys:
- type
x-kubernetes-list-type: map
observedGeneration:
format: int64
type: integer
type: object
required:
- spec
type: object
served: true
storage: true
subresources:
status: {}
@@ -175,19 +175,6 @@ spec:
- message: exactly one of dsn or postgresClusterRef must be set
rule: '(has(self.dsn) ? 1 : 0) + (has(self.postgresClusterRef) ?
1 : 0) == 1'
deadman:
description: |-
DeadmanSpec controls dead man's switch alerts. Matchers/Timeout/Severity
map straight to TERDUT_DEADMAN_MATCHERS/TERDUT_DEADMAN_TIMEOUT/
TERDUT_DEADMAN_SEVERITY.
properties:
matchers:
type: string
severity:
type: string
timeout:
type: string
type: object
image:
description: ImageSpec is the terdut-server image to run.
properties:
@@ -259,12 +246,7 @@ spec:
type: object
type: object
oidc:
description: |-
OIDCSpec controls single sign-on. Fields the chart also exposes but
DESIGN.md's spec doesn't (usernameClaim, emailClaim, groupsClaim,
trustEmail) use terdut-server's own defaults
(preferred_username/email/groups/false) rather than being added here
speculatively.
description: OIDCSpec controls single sign-on.
properties:
adminGroup:
type: string
@@ -276,12 +258,11 @@ spec:
type: string
clientSecretRef:
description: |-
SecretKeyRef names one data key inside a Secret. Every use of this type in
TerdutServerSpec resolves in the TerdutServer's own namespace (it's wired
straight into the Deployment's pod spec as a secretKeyRef env source,
which Kubernetes itself only allows same-namespace) -- unlike the
generated credentials Secret (DESIGN.md §6), which always lives in the
operator's own namespace and is never referenced through this type.
SecretKeyRef names one data key inside a Secret in the TerdutServer's own
namespace. Every use of this type is wired into the Deployment's pod spec as
a secretKeyRef env source, which Kubernetes only allows same-namespace --
including status.credentialsSecretRef, the operator key the controller
generates there.
properties:
key:
description: key is the data key inside the Secret holding
@@ -296,8 +277,14 @@ spec:
- key
- name
type: object
emailClaim:
default: email
type: string
enabled:
type: boolean
groupsClaim:
default: groups
type: string
issuer:
type: string
name:
@@ -309,6 +296,17 @@ spec:
sessionMaxAge:
default: 12h
type: string
trustEmail:
description: |-
trustEmail links a sign-in to an existing local user by email even when
the provider does not vouch the address is verified (Authentik reports
email_verified false unless told otherwise).
type: boolean
usernameClaim:
default: preferred_username
description: usernameClaim, emailClaim and groupsClaim name the
ID token claims read.
type: string
type: object
passwordLogin:
default: true
@@ -324,11 +322,11 @@ spec:
affinity:
description: |-
affinity covers node affinity, pod affinity and pod anti-affinity in
one field -- unlike a multi-replica-aware operator, this one never
generates a default anti-affinity itself (replicas above 1 isn't a
supported topology, see TerdutServerSpec.Replicas's own doc comment),
so this is pure user-supplied passthrough, not a toggle-plus-generated-
default.
one field -- even though replicas now defaults to 2 (see
TerdutServerSpec.Replicas's own doc comment), this operator still
never generates a default anti-affinity of its own the way a
multi-replica-aware operator typically would, so this stays pure
user-supplied passthrough, not a toggle-plus-generated-default.
properties:
nodeAffinity:
description: Describes node affinity scheduling rules for
@@ -4301,11 +4299,15 @@ spec:
type: array
type: object
replicas:
default: 1
default: 2
description: |-
replicas. terdut-server is not horizontally-scale-tested; keep this
at its default of 1 unless you've verified otherwise -- the sweeper
and the notifier are unsynchronised singletons.
replicas. Defaults to 2: terdut-server v0.36.0 put the sweeper, the
notifier and the migration runner each behind a Postgres advisory
lock, and gave incident creation its own conflict resolution, so
more than one replica no longer double-pages, races a migration, or
drops a webhook payload. image.tag must be v0.36.0 or newer for
that to hold -- an older terdut-server has none of these guards,
and this field does not check the tag for you.
format: int32
type: integer
sweeper:
@@ -4392,10 +4394,9 @@ spec:
x-kubernetes-list-type: map
credentialsSecretRef:
description: |-
credentialsSecretRef is the generated instance-scoped credential
(DESIGN.md §6) -- pure output, always in the operator's own
namespace, under a fixed data key ("token"). Set only once
Bootstrapped is True.
credentialsSecretRef is the operator key this controller generated for
the server (TERDUT_OPERATOR_KEY): pure output, in the TerdutServer's own
namespace and owned by it, under the data key "token".
properties:
key:
description: key is the data key inside the Secret holding the
@@ -4417,12 +4418,6 @@ spec:
tells "applied" from "seen" (DESIGN.md §7).
format: int64
type: integer
serviceName:
description: |-
serviceName is the Service this controller created for the
Deployment, so other objects can reference it without recomputing the
naming convention.
type: string
type: object
required:
- spec
@@ -52,54 +52,121 @@ spec:
spec:
description: spec defines the desired state of TerdutTeam
properties:
deadmanSwitches:
description: |-
deadmanSwitches are this team's dead man's switches, by name. Switches on
the server that are not listed here are removed: in operator mode this
list is the whole truth.
items:
description: |-
DeadmanSwitchSpec is one dead man's switch: the absence of an alert matching
matcher for longer than timeout opens an incident.
properties:
matcher:
description: |-
matcher names the alerts this switch watches, e.g.
"alertname=Watchdog,cluster=prod". One matcher per switch.
maxLength: 512
minLength: 1
type: string
x-kubernetes-validations:
- message: 'one matcher per switch: add another entry instead
of separating with ;'
rule: '!self.contains('';'')'
name:
description: name identifies the switch within the team.
maxLength: 100
minLength: 1
type: string
severity:
default: critical
enum:
- critical
- error
- warning
- info
type: string
timeout:
description: timeout is a Go duration string, e.g. "15m".
maxLength: 32
pattern: ^([0-9]+(\.[0-9]+)?(ns|us|µs|ms|s|m|h))+$
type: string
required:
- matcher
- name
- timeout
type: object
maxItems: 50
type: array
x-kubernetes-list-map-keys:
- name
x-kubernetes-list-type: map
displayName:
description: |-
displayName is this team's name, both in terdut-server's own data
(POST /api/teams {"name": ...}) and as the identity POST /api/teams
and GET /api/teams?name= correlate on (DESIGN.md §5's idempotent-
create rule, via TEAM-LOOKUP.md).
displayName is this team's name on the server. It can be changed freely:
the team is found by the CR's own identity (<namespace>/<name>, sent as
external_id), not by this name.
minLength: 1
type: string
invite:
escalation:
description: |-
TerdutTeamInvite requests a standing invite link into this team, minted
with the team's own team-scoped credential — requireTeamOwner already
treats that credential as owner-equivalent for every /invites route
(ratified, not a gap, as of terdut-server's SERVICE-ACCOUNTS.md). This is
the real answer to "how does a human ever get a first login on a
password-only, operator-managed install" (terdut-server#23): no signup_mode
flip, no admin token, just a link redeemed the same way anyone else's
invite would be.
escalation is this team's escalation ladder. Omitted, the team has none
(the server's plain reminder behaviour applies).
properties:
enabled:
description: |-
enabled mints (and keeps refreshed ahead of terdut-server's own fixed
7-day TTL) an invite link while true. Flipping it back to false
revokes the current one server-side rather than leaving it to expire
on its own.
type: boolean
maxUses:
default: 1
description: |-
maxUses bounds how many times this link may be redeemed before it
stops working, mirroring terdut-server's own 1-100 range
(POST /api/teams/{teamID}/invites). Defaults to 1: a link meant for
one specific person, not a standing door.
format: int64
maximum: 100
minimum: 1
type: integer
role:
default: member
description: |-
role is what the invite grants: member or owner. Defaults to member —
owner by default would make every invite link a standing
administrative credential for the team, a much bigger blast radius
than "let a human see the queue".
enum:
- member
- owner
fallbackTopic:
type: string
levels:
items:
description: |-
EscalationLevel is one rung of the ladder: how long to wait, and who to page
if nobody has acknowledged by then.
properties:
targets:
items:
description: |-
EscalationTarget is one page within a level. username is required iff kind is
"user".
properties:
kind:
description: EscalationTargetKind is who one rung
of the ladder pages.
enum:
- oncall
- user
type: string
username:
maxLength: 255
type: string
required:
- kind
type: object
x-kubernetes-validations:
- message: username is required when kind is user
rule: self.kind != 'user' || has(self.username)
- message: username must not be set when kind is oncall
rule: self.kind != 'oncall' || !has(self.username)
maxItems: 20
minItems: 1
type: array
timeout:
description: timeout is a Go duration string, e.g. "5m".
maxLength: 32
pattern: ^([0-9]+(\.[0-9]+)?(ns|us|µs|ms|s|m|h))+$
type: string
required:
- targets
- timeout
type: object
maxItems: 10
minItems: 1
type: array
repeatCount:
format: int64
maximum: 10
minimum: 0
type: integer
required:
- levels
type: object
oidc:
description: |-
@@ -190,53 +257,9 @@ spec:
x-kubernetes-list-map-keys:
- type
x-kubernetes-list-type: map
credentialsSecretRef:
description: |-
credentialsSecretRef is this team's own scoped credential
(DESIGN.md §6 point 3) -- pure output, always in the operator's own
namespace, under a fixed data key ("token").
properties:
key:
description: key is the data key inside the Secret holding the
raw value.
minLength: 1
type: string
name:
description: name is the Secret's name.
minLength: 1
type: string
required:
- key
- name
type: object
inviteSecretRef:
description: |-
inviteSecretRef is this team's current invite link, if spec.invite.enabled.
Unlike credentialsSecretRef, this lives in the TerdutTeam's OWN
namespace, not the operator's: an invite is bounded, limited-use, and
meant for this namespace's own human operators to read and hand out,
not a durable high-privilege credential — same shape as
TerdutAlertSource's status.webhookURLSecretRef, not TerdutServer's
cross-namespace credentialsSecretRef. Nil whenever spec.invite.enabled
is false or unset.
properties:
name:
description: name is the Secret's name.
minLength: 1
type: string
required:
- name
type: object
observedGeneration:
format: int64
type: integer
serverEndpoint:
description: |-
serverEndpoint is the resolved TerdutServer's base URL, resolved once
here so no child controller (TerdutEscalationRule, TerdutDeadmanSwitch,
TerdutAlertSource) ever needs its own RBAC on terdutservers just to
find out where to send a request (DESIGN.md §5).
type: string
teamID:
description: |-
teamID is the server-side id -- needed by every child object's
-2
View File
@@ -4,8 +4,6 @@
resources:
- bases/terdut.ryuvia.com_terdutservers.yaml
- bases/terdut.ryuvia.com_terdutteams.yaml
- bases/terdut.ryuvia.com_terdutescalationrules.yaml
- bases/terdut.ryuvia.com_terdutdeadmanswitches.yaml
- bases/terdut.ryuvia.com_terdutalertsources.yaml
# +kubebuilder:scaffold:crdkustomizeresource
-6
View File
@@ -23,14 +23,8 @@ resources:
#- ../webhook
# [CERTMANAGER] To enable cert-manager, uncomment all sections with 'CERTMANAGER'. 'WEBHOOK' components are required.
#- ../certmanager
# [PROMETHEUS] To enable prometheus monitor, uncomment all sections with 'PROMETHEUS'.
#- ../prometheus
# [METRICS] Expose the controller manager metrics service.
- metrics_service.yaml
# [NETWORK POLICY] Control ingress to metrics and webhook ports.
# Allow metrics traffic from pods in namespaces labeled 'metrics: enabled'.
# Allow webhook traffic from all sources.
#- ../network-policy
# Uncomment the patches line if you enable Metrics
patches:
@@ -1,26 +0,0 @@
# Allow metrics traffic from pods in namespaces labeled 'metrics: enabled'.
# Add this label to namespaces whose pods should scrape metrics.
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
labels:
app.kubernetes.io/name: terdut-operator
app.kubernetes.io/managed-by: kustomize
name: allow-metrics-traffic
namespace: system
spec:
podSelector:
matchLabels:
control-plane: controller-manager
app.kubernetes.io/name: terdut-operator
policyTypes:
- Ingress
ingress:
# Allow pods in namespaces labeled 'metrics: enabled' to scrape metrics.
- from:
- namespaceSelector:
matchLabels:
metrics: enabled # Only from namespaces with this label
ports:
- port: 8443
protocol: TCP
-2
View File
@@ -1,2 +0,0 @@
resources:
- allow-metrics-traffic.yaml
-11
View File
@@ -1,11 +0,0 @@
resources:
- monitor.yaml
# [PROMETHEUS-WITH-CERTS] The following patch configures the ServiceMonitor in ../prometheus
# to securely reference certificates created and managed by cert-manager.
# Additionally, ensure that you uncomment the [METRICS WITH CERTMANAGER] patch under config/default/kustomization.yaml
# to mount the "metrics-server-cert" secret in the Manager Deployment.
#patches:
# - path: monitor_tls_patch.yaml
# target:
# kind: ServiceMonitor
-27
View File
@@ -1,27 +0,0 @@
# Prometheus Monitor Service (Metrics)
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
labels:
control-plane: controller-manager
app.kubernetes.io/name: terdut-operator
app.kubernetes.io/managed-by: kustomize
name: controller-manager-metrics-monitor
namespace: system
spec:
endpoints:
- path: /metrics
port: https # Ensure this is the name of the port that exposes HTTPS metrics
scheme: https
bearerTokenFile: /var/run/secrets/kubernetes.io/serviceaccount/token
tlsConfig:
# TODO(user): The option insecureSkipVerify: true is not recommended for production since it disables
# certificate verification, exposing the system to potential man-in-the-middle attacks.
# For production environments, it is recommended to use cert-manager for automatic TLS certificate management.
# To apply this configuration, enable cert-manager and use the patch located at config/prometheus/servicemonitor_tls_patch.yaml,
# which securely references the certificate from the 'metrics-server-cert' secret.
insecureSkipVerify: true
selector:
matchLabels:
control-plane: controller-manager
app.kubernetes.io/name: terdut-operator
-19
View File
@@ -1,19 +0,0 @@
# Patch for Prometheus ServiceMonitor to enable secure TLS configuration
# using certificates managed by cert-manager
- op: replace
path: /spec/endpoints/0/tlsConfig
value:
# SERVICE_NAME and SERVICE_NAMESPACE will be substituted by kustomize
serverName: SERVICE_NAME.SERVICE_NAMESPACE.svc
insecureSkipVerify: false
ca:
secret:
name: metrics-server-cert
key: ca.crt
cert:
secret:
name: metrics-server-cert
key: tls.crt
keySecret:
name: metrics-server-cert
key: tls.key
-6
View File
@@ -25,12 +25,6 @@ resources:
- terdutalertsource_admin_role.yaml
- terdutalertsource_editor_role.yaml
- terdutalertsource_viewer_role.yaml
- terdutdeadmanswitch_admin_role.yaml
- terdutdeadmanswitch_editor_role.yaml
- terdutdeadmanswitch_viewer_role.yaml
- terdutescalationrule_admin_role.yaml
- terdutescalationrule_editor_role.yaml
- terdutescalationrule_viewer_role.yaml
- terdutteam_admin_role.yaml
- terdutteam_editor_role.yaml
- terdutteam_viewer_role.yaml
-6
View File
@@ -68,8 +68,6 @@ rules:
- terdut.ryuvia.com
resources:
- terdutalertsources
- terdutdeadmanswitches
- terdutescalationrules
- terdutservers
- terdutteams
verbs:
@@ -84,8 +82,6 @@ rules:
- terdut.ryuvia.com
resources:
- terdutalertsources/finalizers
- terdutdeadmanswitches/finalizers
- terdutescalationrules/finalizers
- terdutservers/finalizers
- terdutteams/finalizers
verbs:
@@ -94,8 +90,6 @@ rules:
- terdut.ryuvia.com
resources:
- terdutalertsources/status
- terdutdeadmanswitches/status
- terdutescalationrules/status
- terdutservers/status
- terdutteams/status
verbs:
@@ -1,27 +0,0 @@
# This rule is not used by the project terdut-operator itself.
# It is provided to allow the cluster admin to help manage permissions for users.
#
# Grants full permissions ('*') over terdut.ryuvia.com.
# This role is intended for users authorized to modify roles and bindings within the cluster,
# enabling them to delegate specific permissions to other users or groups as needed.
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
labels:
app.kubernetes.io/name: terdut-operator
app.kubernetes.io/managed-by: kustomize
name: terdutdeadmanswitch-admin-role
rules:
- apiGroups:
- terdut.ryuvia.com
resources:
- terdutdeadmanswitches
verbs:
- '*'
- apiGroups:
- terdut.ryuvia.com
resources:
- terdutdeadmanswitches/status
verbs:
- get
@@ -1,33 +0,0 @@
# This rule is not used by the project terdut-operator itself.
# It is provided to allow the cluster admin to help manage permissions for users.
#
# Grants permissions to create, update, and delete resources within the terdut.ryuvia.com.
# This role is intended for users who need to manage these resources
# but should not control RBAC or manage permissions for others.
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
labels:
app.kubernetes.io/name: terdut-operator
app.kubernetes.io/managed-by: kustomize
name: terdutdeadmanswitch-editor-role
rules:
- apiGroups:
- terdut.ryuvia.com
resources:
- terdutdeadmanswitches
verbs:
- create
- delete
- get
- list
- patch
- update
- watch
- apiGroups:
- terdut.ryuvia.com
resources:
- terdutdeadmanswitches/status
verbs:
- get
@@ -1,29 +0,0 @@
# This rule is not used by the project terdut-operator itself.
# It is provided to allow the cluster admin to help manage permissions for users.
#
# Grants read-only access to terdut.ryuvia.com resources.
# This role is intended for users who need visibility into these resources
# without permissions to modify them. It is ideal for monitoring purposes and limited-access viewing.
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
labels:
app.kubernetes.io/name: terdut-operator
app.kubernetes.io/managed-by: kustomize
name: terdutdeadmanswitch-viewer-role
rules:
- apiGroups:
- terdut.ryuvia.com
resources:
- terdutdeadmanswitches
verbs:
- get
- list
- watch
- apiGroups:
- terdut.ryuvia.com
resources:
- terdutdeadmanswitches/status
verbs:
- get
@@ -1,27 +0,0 @@
# This rule is not used by the project terdut-operator itself.
# It is provided to allow the cluster admin to help manage permissions for users.
#
# Grants full permissions ('*') over terdut.ryuvia.com.
# This role is intended for users authorized to modify roles and bindings within the cluster,
# enabling them to delegate specific permissions to other users or groups as needed.
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
labels:
app.kubernetes.io/name: terdut-operator
app.kubernetes.io/managed-by: kustomize
name: terdutescalationrule-admin-role
rules:
- apiGroups:
- terdut.ryuvia.com
resources:
- terdutescalationrules
verbs:
- '*'
- apiGroups:
- terdut.ryuvia.com
resources:
- terdutescalationrules/status
verbs:
- get
@@ -1,33 +0,0 @@
# This rule is not used by the project terdut-operator itself.
# It is provided to allow the cluster admin to help manage permissions for users.
#
# Grants permissions to create, update, and delete resources within the terdut.ryuvia.com.
# This role is intended for users who need to manage these resources
# but should not control RBAC or manage permissions for others.
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
labels:
app.kubernetes.io/name: terdut-operator
app.kubernetes.io/managed-by: kustomize
name: terdutescalationrule-editor-role
rules:
- apiGroups:
- terdut.ryuvia.com
resources:
- terdutescalationrules
verbs:
- create
- delete
- get
- list
- patch
- update
- watch
- apiGroups:
- terdut.ryuvia.com
resources:
- terdutescalationrules/status
verbs:
- get
@@ -1,29 +0,0 @@
# This rule is not used by the project terdut-operator itself.
# It is provided to allow the cluster admin to help manage permissions for users.
#
# Grants read-only access to terdut.ryuvia.com resources.
# This role is intended for users who need visibility into these resources
# without permissions to modify them. It is ideal for monitoring purposes and limited-access viewing.
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
labels:
app.kubernetes.io/name: terdut-operator
app.kubernetes.io/managed-by: kustomize
name: terdutescalationrule-viewer-role
rules:
- apiGroups:
- terdut.ryuvia.com
resources:
- terdutescalationrules
verbs:
- get
- list
- watch
- apiGroups:
- terdut.ryuvia.com
resources:
- terdutescalationrules/status
verbs:
- get
-8
View File
@@ -1,8 +0,0 @@
## Append samples of your project ##
resources:
- terdut_v1alpha1_terdutserver.yaml
- terdut_v1alpha1_terdutteam.yaml
- terdut_v1alpha1_terdutescalationrule.yaml
- terdut_v1alpha1_terdutdeadmanswitch.yaml
- terdut_v1alpha1_terdutalertsource.yaml
# +kubebuilder:scaffold:manifestskustomizesamples
@@ -1,19 +0,0 @@
apiVersion: terdut.ryuvia.com/v1alpha1
kind: TerdutAlertSource
metadata:
labels:
app.kubernetes.io/name: terdut-operator
app.kubernetes.io/managed-by: kustomize
name: terdutalertsource-sample
spec:
teamRef:
name: terdutteam-sample
# kind defaults to "alertmanager" -- the only value terdut-server
# supports today. Changing it after this object exists rotates the
# webhook key (DESIGN.md §5): the old integration is deleted and a new
# one created, which breaks whatever still sends to the old URL.
kind: alertmanager
# name is this source's own display name server-side, distinct from this
# object's own metadata.name above -- renaming it is safe and never
# rotates the key.
name: prod-alertmanager
@@ -1,15 +0,0 @@
apiVersion: terdut.ryuvia.com/v1alpha1
kind: TerdutDeadmanSwitch
metadata:
labels:
app.kubernetes.io/name: terdut-operator
app.kubernetes.io/managed-by: kustomize
name: terdutdeadmanswitch-sample
spec:
teamRef:
name: terdutteam-sample
# name is optional -- left empty, terdut-server derives it from matcher's
# own canonical form (DESIGN.md §4.4).
matcher: "alertname=Watchdog"
timeout: 15m
severity: critical
@@ -1,25 +0,0 @@
apiVersion: terdut.ryuvia.com/v1alpha1
kind: TerdutEscalationRule
metadata:
labels:
app.kubernetes.io/name: terdut-operator
app.kubernetes.io/managed-by: kustomize
name: terdutescalationrule-sample
spec:
# One per team (DESIGN.md §4.3) -- a second TerdutEscalationRule naming
# the same teamRef would simply clobber this one every reconcile, since
# there's no admission-time check for it in v1.
teamRef:
name: terdutteam-sample
repeatCount: 2
fallbackTopic: platform-fallback
levels:
# username is required iff kind is "user", and rejected otherwise --
# enforced at apply time via CEL (api/v1alpha1/terdutescalationrule_types.go).
- timeout: 5m
targets:
- kind: user
username: alice
- timeout: 10m
targets:
- kind: oncall
@@ -1,42 +0,0 @@
apiVersion: terdut.ryuvia.com/v1alpha1
kind: TerdutServer
metadata:
labels:
app.kubernetes.io/name: terdut-operator
app.kubernetes.io/managed-by: kustomize
name: terdutserver-sample
spec:
image:
repository: git.ryuvia.com/niklas/terdut-server
tag: v0.20.0
replicas: 1
networking:
hostname: terdut.example.com
servicePort: 8080
# Bring-your-own DSN (simplest path, no external CRD dependency). For the
# Zalando postgres-operator path instead, use:
# database:
# postgresClusterRef:
# name: terdut-postgres
database:
dsn: "postgres://terdut@terdut-postgres:5432/terdut?sslmode=require"
passwordSecretRef:
name: terdut-postgres-password
key: password
sweeper:
staleAfter: 6h
archiveAfter: 168h
deadman:
matchers: "alertname=Watchdog"
timeout: 15m
severity: critical
passwordLogin: true
# Pod-level customization, all optional -- see PodSpec in
# api/v1alpha1/terdutserver_types.go for the full shape (tolerations,
# affinity, topologySpreadConstraints, securityContext,
# serviceAccountName, extraEnv/extraVolumes, imagePullSecrets,
# disruptionBudget, ...). Example:
# pod:
# resources:
# requests: {cpu: 100m, memory: 128Mi}
# limits: {memory: 256Mi}
@@ -1,20 +0,0 @@
apiVersion: terdut.ryuvia.com/v1alpha1
kind: TerdutTeam
metadata:
labels:
app.kubernetes.io/name: terdut-operator
app.kubernetes.io/managed-by: kustomize
name: terdutteam-sample
spec:
# serverRef.namespace is optional, defaulting to this TerdutTeam's own
# namespace (the common case). Set it only to reference a TerdutServer in
# a different namespace -- which needs that TerdutServer's own
# spec.allowedTeams to admit this namespace (DESIGN.md §4.6), otherwise
# this reports Ready: False, reason: RefNotPermitted.
serverRef:
name: terdutserver-sample
displayName: Platform
# oidc is optional -- omit entirely for a password-login-only install.
# oidc:
# memberGroup: terdut-platform-members
# ownerGroup: terdut-platform-owners
+18 -9
View File
@@ -1,6 +1,6 @@
# The one TerdutServer this whole demo runs against. Everything else in
# this directory (teams, escalation rules, dead man's switches, alert
# sources) references it by name.
# this directory (teams with their escalation and dead man's switches,
# alert sources) references it by name.
#
# The operator never creates any ingress/HTTPRoute for this TerdutServer --
# that's a permanent non-goal (DESIGN.md §1, NetworkingSpec's own doc
@@ -14,13 +14,22 @@ metadata:
spec:
image:
repository: git.ryuvia.com/niklas/terdut-server
# v0.34.0: fixes callerMayManageServiceAccount so an instance-scoped
# service account can adopt/rotate a key on a team-scoped account it
# didn't just create in the same call -- without this, terdutteam-*
# can wedge permanently on exactly the crash-window race this demo
# hit live (niklas/terdut-operator#3).
tag: v0.34.0
replicas: 1
# v0.36.0 is the floor now that replicas below is 2 (this demo pins
# the current release, v0.43.0, so it shows the current web UI too): that
# release put the sweeper, the notifier and the migration runner each
# behind a Postgres advisory lock, and gave incident creation its own
# conflict resolution, which is what makes a second replica safe
# instead of racing the first. (Still carries v0.34.0's fix too --
# callerMayManageServiceAccount, so an instance-scoped service account
# can adopt/rotate a key on a team-scoped account it didn't just create
# in the same call -- without which terdutteam-* can wedge permanently
# on the crash-window race this demo hit live, niklas/terdut-operator#3.)
tag: v0.43.0
# Matches this CRD's own spec.replicas default (v0.4.0) -- stated
# explicitly, like every other field in this file, rather than left to
# the default. RollingUpdate follows automatically; this operator does
# not expose Strategy as a spec field.
replicas: 2
networking:
hostname: terdut-operator-demo.example
servicePort: 8080
+29 -7
View File
@@ -1,7 +1,9 @@
# Two teams (this one and 03-team-payments.yaml) so the demo shows
# per-team isolation -- separate incident lists, separate escalation
# ladders, separate alert sources -- rather than one team standing in for
# everything.
# everything. A team carries its own escalation ladder and dead man's
# switches; alert sources are separate objects (04-alertsource-*.yaml)
# because each owns a webhook Secret.
apiVersion: terdut.ryuvia.com/v1alpha1
kind: TerdutTeam
metadata:
@@ -15,9 +17,29 @@ spec:
name: terdut-operator-demo
displayName: Platform
# No oidc block: this demo is password-login only (01-server.yaml).
# A real invite link, minted via this team's own credential, is how
# run-demo.sh's alice actually gets in -- no signup_mode flip, no admin
# token (see niklas/terdut-server#23's resolution). Role/maxUses left at
# their defaults (member, 1): one link for one person.
invite:
enabled: true
# The whole ladder, replaced as one unit. username is required iff kind is
# "user" (CEL validation at apply time). The server resolves it, so a
# username it does not know yet leaves this team at Reason: UnknownUser
# until that person exists -- run-demo.sh creates alice for exactly that.
escalation:
repeatCount: 2
fallbackTopic: platform-fallback
levels:
- timeout: 5m
targets:
- kind: user
username: alice
- timeout: 10m
targets:
- kind: oncall
# Dead man's switches, by name. fire-alerts.sh's "heartbeat" scenario sends
# a matching alert; stop sending it and terdut-server itself opens an
# incident once `timeout` passes with no heartbeat. A switch on the server
# that is not listed here is removed.
deadmanSwitches:
- name: platform-watchdog
matcher: "alertname=PlatformWatchdog"
timeout: 15m
severity: critical
+12 -6
View File
@@ -6,9 +6,15 @@ spec:
serverRef:
name: terdut-operator-demo
displayName: Payments
# No spec.invite here, deliberately: run-demo.sh joins alice to this team
# through POST /api/teams/{teamID}/members instead (this team's own
# credential, same owner-equivalent reach spec.invite relies on, plus her
# user id resolved via GET /api/users), once she already has an account
# from Platform's invite -- showing both onboarding paths this feature
# unlocks, not just the one.
escalation:
repeatCount: 1
fallbackTopic: payments-fallback
levels:
- timeout: 5m
targets:
- kind: oncall
deadmanSwitches:
- name: payments-watchdog
matcher: "alertname=PaymentsWatchdog"
timeout: 15m
severity: critical
-24
View File
@@ -1,24 +0,0 @@
# One per team is the rule (DESIGN.md §4.3) -- a second TerdutEscalationRule
# naming the same teamRef would just clobber this one on the next reconcile.
apiVersion: terdut.ryuvia.com/v1alpha1
kind: TerdutEscalationRule
metadata:
name: terdutescalationrule-platform
spec:
teamRef:
name: terdutteam-platform
repeatCount: 2
fallbackTopic: platform-fallback
levels:
# username is required iff kind is "user", rejected otherwise -- CEL
# validation at apply time (api/v1alpha1/terdutescalationrule_types.go).
# alice won't exist on a fresh demo install -- see README.md for
# creating a real user if you want this level to mean something, or
# just watch it fall through to oncall after 5m.
- timeout: 5m
targets:
- kind: user
username: alice
- timeout: 10m
targets:
- kind: oncall
-13
View File
@@ -1,13 +0,0 @@
apiVersion: terdut.ryuvia.com/v1alpha1
kind: TerdutEscalationRule
metadata:
name: terdutescalationrule-payments
spec:
teamRef:
name: terdutteam-payments
repeatCount: 1
fallbackTopic: payments-fallback
levels:
- timeout: 5m
targets:
- kind: oncall
-17
View File
@@ -1,17 +0,0 @@
# A per-team dead man's switch (DESIGN.md §4.4) -- a different thing from
# 01-server.yaml's spec.deadman, which this demo leaves unset so this CRD
# is what you're actually seeing reconcile. fire-alerts.sh's "heartbeat"
# scenario sends a matching alert; stop sending it and terdut-server
# itself opens an incident once `timeout` passes with no heartbeat.
apiVersion: terdut.ryuvia.com/v1alpha1
kind: TerdutDeadmanSwitch
metadata:
name: terdutdeadmanswitch-platform
spec:
teamRef:
name: terdutteam-platform
# name omitted -- terdut-server derives one from the matcher's own
# canonical form (DESIGN.md §4.4).
matcher: "alertname=PlatformWatchdog"
timeout: 15m
severity: critical
-10
View File
@@ -1,10 +0,0 @@
apiVersion: terdut.ryuvia.com/v1alpha1
kind: TerdutDeadmanSwitch
metadata:
name: terdutdeadmanswitch-payments
spec:
teamRef:
name: terdutteam-payments
matcher: "alertname=PaymentsWatchdog"
timeout: 15m
severity: critical
+43 -43
View File
@@ -1,9 +1,9 @@
# Demo
Every CRD this operator reconciles, wired into one working install: one
`TerdutServer`, two `TerdutTeam`s (Platform and Payments) each with their
own `TerdutEscalationRule`, `TerdutDeadmanSwitch` and `TerdutAlertSource`,
plus a script that fires synthetic Alertmanager webhooks at it so you can
`TerdutServer`, two `TerdutTeam`s (Platform and Payments), each carrying its
own escalation ladder and dead man's switches, and a `TerdutAlertSource` per
team, plus a script that fires synthetic Alertmanager webhooks at it so you can
watch real incidents appear, escalate and resolve.
This is a demo kit, not a reference deployment: `00-postgres.yaml` runs
@@ -13,7 +13,7 @@ directory. Throw the whole namespace away when you're done.
**Want this fully automated instead of walking through it by hand?**
`./run-demo.sh` does everything below itself, against a fresh (or
already-set-up) `kind` cluster — creates the cluster, installs the
operator, applies every CR here, signs `alice` in for real, and fires a
operator, applies every CR here, creates `alice` as the first user, and fires a
few alerts. `./run-demo.sh --help` for the knobs, `./run-demo.sh
--teardown` to tear it back down. The rest of this file is the manual
walkthrough it automates.
@@ -44,16 +44,17 @@ kubectl create namespace terdut-operator-demo
kubectl apply -n terdut-operator-demo -k .
```
Listed and applied in dependency order (server → team → everything that
`teamRef`s it), but you don't have to preserve that order yourself:
every controller here re-queues and waits rather than failing when a ref
isn't resolvable yet (`kubectl describe` shows `Reason: TeamRefNotFound` /
`WaitingForTeam` while that settles).
Listed and applied in dependency order (server → team → alert source), but
you don't have to preserve that order yourself: every controller here
re-queues and waits rather than failing when a ref isn't resolvable yet
(`kubectl describe` shows `Reason: WaitingForServer` / `WaitingForTeam` while
that settles). Platform stays at `Reason: UnknownUser` until `alice` exists:
its escalation ladder names her.
Watch it converge:
```sh
kubectl get terdutservers,terdutteams,terdutescalationrules,terdutdeadmanswitches,terdutalertsources \
kubectl get terdutservers,terdutteams,terdutalertsources \
-n terdut-operator-demo
```
@@ -80,36 +81,21 @@ exposing it for real (Gateway API, Istio, or a plain `Ingress`) instead.
### First login
The operator's own bootstrap (DESIGN.md §6) creates the first user through
`/api/bootstrap` and immediately mints itself a service-account token from
it, then discards the bootstrap user's own key — nobody ever signs in as
that account, and `signup_mode` stays `invite_only` by default. **Don't try
to flip it via the operator's own token**: that token is a service account,
and `/api/admin/settings` is deliberately human-only on terdut-server
(`niklas/terdut-server#23` has the full reasoning — widening that gate was
the wrong fix).
The real path in: `02-team-platform.yaml` turns on `spec.invite`, so
Platform's own `TerdutTeam` mints a real invite link with its own
already-working team-scoped credential (the same reach that lets it manage
its own escalation policy, dead man's switches and integrations — owner-
equivalent, confirmed in terdut-server's `SERVICE-ACCOUNTS.md`). Invite
redemption bypasses `signup_mode` entirely, so this needs no admin
credential at all:
The operator authenticates with a key of its own (the `TerdutServer`'s
`<name>-operator-key` Secret, handed to the server as `TERDUT_OPERATOR_KEY`) and
never creates a user. A person gets in the way anyone does on a fresh
terdut-server: `/api/bootstrap` creates the first user, an administrator,
while no user exists yet:
```sh
secretname=$(kubectl -n terdut-operator-demo get terdutteam terdutteam-platform \
-o jsonpath='{.status.inviteSecretRef.name}')
url=$(kubectl -n terdut-operator-demo get secret "$secretname" -o jsonpath='{.data.url}' | base64 -d)
echo "$url" # open this, or POST /api/signup with {"invite": "<the token after invite=>", ...}
curl -sS -X POST http://localhost:8080/api/bootstrap -H 'Content-Type: application/json' \
-d '{"username":"alice","email":"alice@example.com","password":"a-long-demo-password"}'
```
`04-escalation-platform.yaml` names a user `alice` at its first escalation
level — sign up as `alice` if you want that level to mean something rather
than falling through to on-call after 5 minutes. `run-demo.sh` does exactly
this automatically (and also joins `alice` to Payments, which deliberately
has no `spec.invite` of its own — see that file's comment for the second
onboarding path this demonstrates).
The response carries an API key, shown once. An administrator can manage any
team, so add `alice` to both teams with it (`POST /api/teams/{teamID}/members`;
`kubectl get terdutteam -o jsonpath='{.status.teamID}'` gives the ids), or just
use the UI's team pages. `run-demo.sh` does exactly this for you.
## Fire some alerts
@@ -122,8 +108,8 @@ export NAMESPACE=terdut-operator-demo
./fire-alerts.sh platform disk-full
./fire-alerts.sh payments pod-crash
# Watch it open an incident, escalate per 04/05-escalation-*.yaml's
# levels, and show up on the Platform/Payments team's own incident list.
# Watch it open an incident, escalate per the escalation ladders in
# 02/03-team-*.yaml, and show up on the Platform/Payments team's own incident list.
./fire-alerts.sh platform high-cpu resolve
```
@@ -133,9 +119,23 @@ Each `(team, scenario)` pair is one stable fingerprint, so firing the same
one twice updates the same alert (a real re-fire) and `resolve` closes
exactly that one.
Set `CLUSTER` to send the alert as if it came from one of several clusters:
```sh
CLUSTER=prod-eu ./fire-alerts.sh platform high-cpu
CLUSTER=prod-us ./fire-alerts.sh platform high-cpu # a second incident, not a join
CLUSTER=prod-eu ./fire-alerts.sh platform high-cpu resolve
```
It stands in for a Prometheus external label plus `cluster` in Alertmanager's
`group_by` (terdut-server's README, "Several clusters, one team"): the web UI
then shows the cluster chip on each incident and a cluster filter in the
queue. `CLUSTER` is part of the fingerprint, so resolve with the same value you
fired with. `./run-demo.sh` fires its alerts across `prod-eu` and `prod-us`.
### Dead man's switches
`06-deadman-platform.yaml` / `07-deadman-payments.yaml` expect a heartbeat
The `deadmanSwitches` in `02-team-platform.yaml` / `03-team-payments.yaml` expect a heartbeat
alert on a 15-minute timeout:
```sh
@@ -153,7 +153,7 @@ the switch is watching for silence, not for a signal.
kubectl delete namespace terdut-operator-demo
```
The operator's own finalizers clean up everything cross-namespace
(credentials Secrets in the operator's namespace, server-side team/rule/
integration rows) before this namespace's objects actually disappear —
give it a few seconds past the `kubectl delete` returning.
The operator's finalizers delete each team on the server (with its
escalation, switches and integrations) before this namespace's objects
actually disappear — give it a few seconds past the `kubectl delete`
returning. A team with open incidents is not deleted until they are resolved.
+21 -9
View File
@@ -16,7 +16,15 @@
# is the one part of that URL still usable here.
#
# Usage:
# ./fire-alerts.sh <platform|payments> <high-cpu|disk-full|pod-crash|heartbeat> [resolve]
# [CLUSTER=prod-eu] ./fire-alerts.sh <platform|payments> <high-cpu|disk-full|pod-crash|heartbeat> [resolve]
#
# CLUSTER stands in for a Prometheus externalLabel plus `cluster` in
# Alertmanager's group_by (terdut-server's README, "Several clusters, one
# team"): it is put on the alert's labels and on groupLabels, so the incident
# carries it and the web UI shows the cluster chip and the queue's cluster
# filter. It is also part of the group key and the fingerprint, which is what
# keeps the same alert in two clusters from joining one incident. Unset, the
# alert is sent exactly as before.
#
# Prerequisites: kubectl context pointed at the demo namespace, jq, curl,
# and (in another terminal) a running:
@@ -24,6 +32,7 @@
set -euo pipefail
NAMESPACE="${NAMESPACE:-}"
CLUSTER="${CLUSTER:-}"
BASE_URL="${BASE_URL:-http://localhost:8080}"
usage() {
@@ -35,7 +44,7 @@ scenarios:
disk-full critical -- disk usage above 95%
pod-crash error -- a pod crash-looping
heartbeat critical -- the team's dead man's switch heartbeat
(matches the matcher in 06/07-deadman-*.yaml -- send this
(matches the matcher in the deadmanSwitches in 02/03-team-*.yaml -- send this
repeatedly to keep the switch alive, or stop sending it and
watch terdut-server open an incident on its own once
`timeout` passes with no heartbeat. "resolve" is not a valid
@@ -45,6 +54,8 @@ scenarios:
env vars:
NAMESPACE kubectl -n for reading the webhook Secret (required)
BASE_URL where the port-forwarded terdut-server is (default http://localhost:8080)
CLUSTER optional cluster name, e.g. prod-eu: sent as a `cluster` label and
group label, so the UI shows where the incident came from
EOF
exit 1
}
@@ -63,7 +74,7 @@ case "$scenario" in
disk-full) alertname=TerdutDemoDiskFull severity=critical summary="Disk usage above 95% on /data" ;;
pod-crash) alertname=TerdutDemoPodCrashLooping severity=error summary="Pod web-7f8b9 is crash-looping (5 restarts in 10m)" ;;
heartbeat)
# Must match 06-deadman-platform.yaml / 07-deadman-payments.yaml's own
# Must match 02-team-platform.yaml / 03-team-payments.yaml's own
# matcher exactly -- that's what makes this a heartbeat rather than a
# third ordinary alert.
case "$team" in
@@ -84,13 +95,13 @@ esac
secret_name="terdutalertsource-${team}-terdut-webhook"
key="$(kubectl -n "$NAMESPACE" get secret "$secret_name" -o jsonpath='{.data.key}' | base64 -d)"
[ -n "$key" ] || { echo "fire-alerts.sh: empty key read from Secret $secret_name -- has 08/09-alertsource-*.yaml reconciled yet?" >&2; exit 1; }
[ -n "$key" ] || { echo "fire-alerts.sh: empty key read from Secret $secret_name -- has 04/05-alertsource-*.yaml reconciled yet?" >&2; exit 1; }
# Stable per (team, scenario) so a resolve targets the same alert a fire
# created: terdut-server correlates on (team_id, fingerprint), not on
# anything else in the payload. Real Alertmanager computes this from the
# alert's label set; a fixed string plays the same role here.
fingerprint="demo-${team}-${scenario}"
fingerprint="demo-${team}-${scenario}${CLUSTER:+-$CLUSTER}"
now="$(date -u +%Y-%m-%dT%H:%M:%SZ)"
if [ "$status" = firing ]; then
@@ -101,7 +112,8 @@ fi
payload="$(jq -n \
--arg status "$status" \
--arg groupKey "demo:${team}:${scenario}" \
--arg groupKey "demo:${team}:${scenario}${CLUSTER:+:$CLUSTER}" \
--arg cluster "$CLUSTER" \
--arg alertname "$alertname" \
--arg team "$team" \
--arg severity "$severity" \
@@ -113,10 +125,10 @@ payload="$(jq -n \
version: "4",
status: $status,
groupKey: $groupKey,
groupLabels: { alertname: $alertname, team: $team },
groupLabels: ({ alertname: $alertname, team: $team } + (if $cluster != "" then { cluster: $cluster } else {} end)),
alerts: [{
status: $status,
labels: { alertname: $alertname, severity: $severity, team: $team, instance: "demo" },
labels: ({ alertname: $alertname, severity: $severity, team: $team, instance: "demo" } + (if $cluster != "" then { cluster: $cluster } else {} end)),
annotations: { summary: $summary },
startsAt: $startsAt,
endsAt: $endsAt,
@@ -126,7 +138,7 @@ payload="$(jq -n \
}')"
url="${BASE_URL}/api/integrations/${key}/alertmanager"
echo "POST $url (team=$team scenario=$scenario status=$status)" >&2
echo "POST $url (team=$team scenario=$scenario status=$status${CLUSTER:+ cluster=$CLUSTER})" >&2
code="$(curl -sS -o /tmp/fire-alerts-response.json -w '%{http_code}' \
-X POST "$url" -H 'Content-Type: application/json' -d "$payload")"
echo "-> HTTP $code" >&2
+2 -6
View File
@@ -9,9 +9,5 @@ resources:
- 01-server.yaml
- 02-team-platform.yaml
- 03-team-payments.yaml
- 04-escalation-platform.yaml
- 05-escalation-payments.yaml
- 06-deadman-platform.yaml
- 07-deadman-payments.yaml
- 08-alertsource-platform.yaml
- 09-alertsource-payments.yaml
- 04-alertsource-platform.yaml
- 05-alertsource-payments.yaml
+78 -94
View File
@@ -103,30 +103,35 @@ install_operator() {
apply_demo() {
log "creating namespace $NAMESPACE"
kubectl create namespace "$NAMESPACE" --dry-run=client -o yaml | kubectl apply -f - >/dev/null
log "applying demo CRs (00-09) into $NAMESPACE"
log "applying demo CRs (00-05) into $NAMESPACE"
kubectl apply -n "$NAMESPACE" -k "$SCRIPT_DIR" >/dev/null
}
wait_for_ready() {
local objects=(
"terdutserver/terdut-operator-demo"
"terdutteam/terdutteam-platform"
"terdutteam/terdutteam-payments"
"terdutescalationrule/terdutescalationrule-platform"
"terdutescalationrule/terdutescalationrule-payments"
"terdutdeadmanswitch/terdutdeadmanswitch-platform"
"terdutdeadmanswitch/terdutdeadmanswitch-payments"
"terdutalertsource/terdutalertsource-platform"
"terdutalertsource/terdutalertsource-payments"
)
wait_for_objects() {
local obj
for obj in "${objects[@]}"; do
for obj in "$@"; do
log "waiting for $obj to become Ready"
kubectl wait --for=condition=Ready --timeout "$WAIT_TIMEOUT" -n "$NAMESPACE" "$obj" >/dev/null \
|| die "timed out waiting for $obj -- try: kubectl describe -n $NAMESPACE $obj"
done
}
# Only the server: the teams cannot reach Ready until alice exists (their
# escalation ladders name her, and terdut-server resolves usernames when the
# ladder is applied), and alice is created below, through the server itself.
wait_for_server_ready() {
wait_for_objects "terdutserver/terdut-operator-demo"
}
# Everything that was waiting on alice to exist.
wait_for_remaining_ready() {
wait_for_objects \
"terdutteam/terdutteam-platform" \
"terdutteam/terdutteam-payments" \
"terdutalertsource/terdutalertsource-platform" \
"terdutalertsource/terdutalertsource-payments"
}
start_port_forward() {
# A stale pidfile from an earlier run would otherwise collide with us on
# $LOCAL_PORT -- if that pid is still alive, stop it first.
@@ -158,86 +163,65 @@ start_port_forward() {
log "port-forward ready (pid $STARTED_PF_PID, log $PF_LOGFILE)"
}
# Redeems Platform's own invite link -- minted by its TerdutTeam
# (02-team-platform.yaml's spec.invite.enabled, reconciled through that
# team's own already-working team-scoped credential, which requireTeamOwner
# already treats as owner-equivalent for /invites -- ratified, not a
# workaround, in terdut-server's SERVICE-ACCOUNTS.md) and redeemed through
# the ordinary signup endpoint. Invite redemption bypasses signup_mode
# entirely (terdut-server's internal/api/signup.go), so this needs no admin
# credential, no signup_mode flip, and no direct Postgres access at all --
# unlike an earlier version of this script, before terdut-operator grew
# this feature (see niklas/terdut-server#23).
redeem_platform_invite() {
log "reading Platform's invite link"
local secret_name invite_url invite_token tries=0
until secret_name="$(kubectl -n "$NAMESPACE" get terdutteam terdutteam-platform \
-o jsonpath='{.status.inviteSecretRef.name}' 2>/dev/null)" && [ -n "$secret_name" ]; do
tries=$((tries + 1))
[ "$tries" -lt 30 ] || die "terdutteam-platform never reported status.inviteSecretRef -- check spec.invite.enabled and kubectl describe it"
sleep 1
# Creates alice as the install's first user and administrator through
# /api/bootstrap, which is open until a first user exists -- the operator
# authenticates with its own seeded key and never uses it, so this is how a
# person gets in. An administrator may manage any team, so alice's own key is
# enough to add her to both; the teams' own status.teamID is set as soon as
# the operator has created each team, even while they still wait on her.
create_alice_and_join_teams() {
# A re-run: she exists already (and was added to the teams the first time).
local login_code
login_code="$(curl -sS -o /dev/null -w '%{http_code}' \
-X POST "${BASE_URL}/api/login" -H 'Content-Type: application/json' \
-d "$(jq -n --arg u "$ALICE_USERNAME" --arg p "$DEMO_PASSWORD" '{username: $u, password: $p}')")"
if [ "$login_code" = "200" ]; then
log "account ${ALICE_USERNAME} already exists and can sign in, skipping creation (re-run detected)"
return 0
fi
log "creating ${ALICE_USERNAME} as the first user (administrator)"
local resp_file code alice_key
resp_file="$(mktemp)"
code="$(curl -sS -o "$resp_file" -w '%{http_code}' \
-X POST "${BASE_URL}/api/bootstrap" -H 'Content-Type: application/json' \
-d "$(jq -n --arg u "$ALICE_USERNAME" --arg e "$ALICE_EMAIL" --arg p "$DEMO_PASSWORD" \
'{username: $u, email: $e, password: $p}')")"
[ "$code" = "201" ] || die "bootstrap failed (HTTP $code): $(cat "$resp_file")"
alice_key="$(jq -r '.api_key.key' "$resp_file")"
rm -f "$resp_file"
local team team_id
for team in terdutteam-platform terdutteam-payments; do
team_id=""
local tries=0
until team_id="$(kubectl -n "$NAMESPACE" get terdutteam "$team" -o jsonpath='{.status.teamID}' 2>/dev/null)" \
&& [ -n "$team_id" ]; do
tries=$((tries + 1))
[ "$tries" -lt 60 ] || die "$team never reported status.teamID -- kubectl describe it"
sleep 1
done
local alice_id
alice_id="$(curl -sS -H "Authorization: Bearer $alice_key" "${BASE_URL}/api/users" \
| jq -r --arg u "$ALICE_USERNAME" '.[] | select(.username == $u) | .id')"
code="$(curl -sS -o /dev/null -w '%{http_code}' \
-X POST "${BASE_URL}/api/teams/${team_id}/members" \
-H "Authorization: Bearer $alice_key" -H 'Content-Type: application/json' \
-d "$(jq -n --argjson id "$alice_id" '{user_id: $id, role: "owner"}')")"
[ "$code" = "204" ] || die "adding ${ALICE_USERNAME} to ${team} failed (HTTP $code)"
log "added ${ALICE_USERNAME} to ${team}"
done
invite_url="$(kubectl -n "$NAMESPACE" get secret "$secret_name" -o jsonpath='{.data.url}' | base64 -d)"
invite_token="${invite_url##*invite=}"
[ -n "$invite_token" ] || die "could not parse an invite token out of $invite_url"
log "signing up ${ALICE_USERNAME} via Platform's invite"
local body resp_file code
body="$(jq -n \
--arg u "$ALICE_USERNAME" --arg e "$ALICE_EMAIL" \
--arg p "$DEMO_PASSWORD" --arg i "$invite_token" \
'{username: $u, email: $e, password: $p, invite: $i}')"
resp_file="$(mktemp)"
code="$(curl -sS -o "$resp_file" -w '%{http_code}' \
-X POST "${BASE_URL}/api/signup" -H 'Content-Type: application/json' -d "$body")"
case "$code" in
201) log "created local account ${ALICE_USERNAME} (joined Platform)" ;;
409) log "account ${ALICE_USERNAME} already exists, skipping (re-run detected)" ;;
*) die "signup failed (HTTP $code): $(cat "$resp_file")" ;;
esac
rm -f "$resp_file"
}
# redeem_platform_invite's signup already used up alice's one signup -- a
# second POST /api/signup would just 409 on the taken username, it doesn't
# join an existing account to another team. Payments is joined through the
# ordinary team-scoped member endpoint instead, using that team's own
# already-working team-scoped credential (owner-equivalent, same reach the
# invite route above relies on) and alice's user id resolved via
# GET /api/users -- the same lookup tdclient.GetUserByUsername does
# operator-side. The endpoint upserts on (team_id, user_id), so this is
# naturally idempotent across re-runs with no separate conflict handling
# needed.
join_payments_team() {
log "adding ${ALICE_USERNAME} to Payments"
local payments_id payments_secret payments_key alice_id resp_file code
payments_id="$(kubectl -n "$NAMESPACE" get terdutteam terdutteam-payments -o jsonpath='{.status.teamID}')"
[ -n "$payments_id" ] || die "could not read status.teamID off terdutteam-payments"
payments_secret="$(kubectl -n "$NAMESPACE" get terdutteam terdutteam-payments \
-o jsonpath='{.status.credentialsSecretRef.name}')"
[ -n "$payments_secret" ] || die "terdutteam-payments has no status.credentialsSecretRef yet"
payments_key="$(kubectl -n "$OPERATOR_NAMESPACE" get secret "$payments_secret" -o jsonpath='{.data.token}' | base64 -d)"
alice_id="$(curl -sS -H "Authorization: Bearer $payments_key" "${BASE_URL}/api/users" \
| jq -r --arg u "$ALICE_USERNAME" '.[] | select(.username == $u) | .id')"
[ -n "$alice_id" ] || die "could not resolve ${ALICE_USERNAME}'s user id via GET /api/users"
resp_file="$(mktemp)"
code="$(curl -sS -o "$resp_file" -w '%{http_code}' \
-X POST "${BASE_URL}/api/teams/${payments_id}/members" \
-H "Authorization: Bearer $payments_key" -H 'Content-Type: application/json' \
-d "$(jq -n --argjson id "$alice_id" '{user_id: $id, role: "member"}')")"
[ "$code" = "204" ] || die "adding ${ALICE_USERNAME} to Payments failed (HTTP $code): $(cat "$resp_file")"
log "added ${ALICE_USERNAME} to Payments"
rm -f "$resp_file"
}
fire_demo_alerts() {
log "firing representative demo alerts"
NAMESPACE="$NAMESPACE" BASE_URL="$BASE_URL" "$SCRIPT_DIR/fire-alerts.sh" platform high-cpu
NAMESPACE="$NAMESPACE" BASE_URL="$BASE_URL" "$SCRIPT_DIR/fire-alerts.sh" platform disk-full
NAMESPACE="$NAMESPACE" BASE_URL="$BASE_URL" "$SCRIPT_DIR/fire-alerts.sh" payments pod-crash
# Two clusters, so the queue shows the cluster chip and offers its filter.
# high-cpu fires in both: the same alert in two clusters is two incidents.
local fire="$SCRIPT_DIR/fire-alerts.sh"
CLUSTER=prod-eu NAMESPACE="$NAMESPACE" BASE_URL="$BASE_URL" "$fire" platform high-cpu
CLUSTER=prod-us NAMESPACE="$NAMESPACE" BASE_URL="$BASE_URL" "$fire" platform high-cpu
CLUSTER=prod-us NAMESPACE="$NAMESPACE" BASE_URL="$BASE_URL" "$fire" platform disk-full
CLUSTER=prod-eu NAMESPACE="$NAMESPACE" BASE_URL="$BASE_URL" "$fire" payments pod-crash
}
print_summary() {
@@ -254,8 +238,8 @@ terdut demo is up.
Fire more alerts:
export NAMESPACE=${NAMESPACE} BASE_URL=${BASE_URL}
./fire-alerts.sh platform high-cpu
./fire-alerts.sh platform high-cpu resolve
CLUSTER=prod-eu ./fire-alerts.sh platform high-cpu
CLUSTER=prod-eu ./fire-alerts.sh platform high-cpu resolve
./fire-alerts.sh payments heartbeat # send repeatedly (e.g. every minute)
# to keep a dead man's switch alive;
# stop sending it and, 15 minutes
@@ -319,10 +303,10 @@ main() {
ensure_kind_cluster
install_operator
apply_demo
wait_for_ready
wait_for_server_ready
start_port_forward
redeem_platform_invite
join_payments_team
create_alice_and_join_teams
wait_for_remaining_ready
fire_demo_alerts
print_summary
+20 -10
View File
@@ -5,14 +5,15 @@ import (
"fmt"
apierrors "k8s.io/apimachinery/pkg/api/errors"
"k8s.io/apimachinery/pkg/api/meta"
"sigs.k8s.io/controller-runtime/pkg/client"
terdutv1alpha1 "git.ryuvia.com/niklas/terdut-operator/api/v1alpha1"
"git.ryuvia.com/niklas/terdut-operator/internal/tdclient"
)
// childError carries a condition reason/message, the same role teamError
// and databaseError play for their own controllers: an expected,
// childError carries a condition reason/message, the same role teamError and
// databaseError play for their own controllers: an expected,
// requeue-and-retry outcome, not a reconcile failure.
type childError struct {
reason string
@@ -21,12 +22,13 @@ type childError struct {
func (e *childError) Error() string { return e.message }
// resolveTeamAndClient implements DESIGN.md §5's "every child resolves its
// own teamRef -> TerdutTeam.status, never chains up to TerdutServer"
// rule -- shared by TerdutEscalationRule and TerdutDeadmanSwitch, which
// both need exactly this and nothing else to call terdut-server's API.
// resolveTeamAndClient finds the TerdutTeam a child (TerdutAlertSource) names,
// requires it to be Ready, and returns a client for its server authenticated
// with that server's operator key. The server is read from the team's own
// serverRef: consent was already checked when the team went Ready, and does not
// gate a resource the team owns.
func resolveTeamAndClient(
ctx context.Context, c client.Client, operatorNamespace, namespace string,
ctx context.Context, c client.Client, namespace string,
ref terdutv1alpha1.TerdutTeamRef, newClient func(string) *tdclient.Client,
) (*terdutv1alpha1.TerdutTeam, *tdclient.Client, *childError) {
var team terdutv1alpha1.TerdutTeam
@@ -39,16 +41,24 @@ func resolveTeamAndClient(
}
return nil, nil, &childError{reason: terdutv1alpha1.ReasonTeamRefNotFound, message: err.Error()}
}
if team.Status.CredentialsSecretRef == nil || team.Status.TeamID == 0 {
if team.Status.TeamID == 0 || !meta.IsStatusConditionTrue(team.Status.Conditions, terdutv1alpha1.ConditionReady) {
return nil, nil, &childError{
reason: terdutv1alpha1.ReasonWaitingForTeam,
message: fmt.Sprintf("TerdutTeam %q is not Ready yet", ref.Name),
}
}
teamKey, err := readOperatorSecret(ctx, c, operatorNamespace, team.Status.CredentialsSecretRef)
srvNS := team.Spec.ServerRef.Namespace
if srvNS == "" {
srvNS = team.Namespace
}
var srv terdutv1alpha1.TerdutServer
if err := c.Get(ctx, client.ObjectKey{Namespace: srvNS, Name: team.Spec.ServerRef.Name}, &srv); err != nil {
return nil, nil, &childError{reason: terdutv1alpha1.ReasonWaitingForTeam, message: err.Error()}
}
key, err := readOperatorKey(ctx, c, &srv)
if err != nil {
return nil, nil, &childError{reason: terdutv1alpha1.ReasonWaitingForTeam, message: err.Error()}
}
return &team, newClient(team.Status.ServerEndpoint).WithToken(teamKey), nil
return &team, newClient(serviceURL(&srv)).WithToken(key), nil
}
+375
View File
@@ -0,0 +1,375 @@
package controller
import (
"encoding/json"
"fmt"
"net/http"
"net/http/httptest"
"strconv"
"strings"
"sync"
"git.ryuvia.com/niklas/terdut-operator/internal/tdclient"
)
// errJSONKey is the key every error body the fake writes uses, and
// deadmanSwitchesPath the literal shared by its dead man's switch routes --
// goconst would otherwise flag the repeats.
const (
errJSONKey = "error"
deadmanSwitchesPath = "/deadman/switches"
)
// fakeTerdutServer reproduces the parts of terdut-server's API the controllers
// call, with the same status codes and idempotency rules (POST /api/teams
// keyed by external_id, unique team names, unique switch names per team), so a
// controller test exercises the real contract and not a canned reply.
type fakeTerdutServer struct {
mu sync.Mutex
// expectKey, when set, makes every request without that bearer token a 401.
expectKey string
nextTeamID int64
teams map[string]int64 // name -> id
teamNames map[int64]string // id -> current name
teamExt map[string]int64 // external_id -> id
teamOIDC map[int64][2]string
teamDelete map[int64]bool // id -> true once DELETEd
// unknownUsers are usernames the escalation PUT answers 400 "unknown user" for.
unknownUsers map[string]bool
escalation map[int64]tdclient.SetEscalationRequest
nextSwitchID int64
switches map[int64]map[int64]tdclient.DeadmanSwitch // teamID -> switchID -> switch
switchDelete map[int64]bool
nextIntegrationID int64
integrations map[int64]map[int64]tdclient.Integration
integrationDelete map[int64]bool
}
func newFakeTerdutServer() (*fakeTerdutServer, *httptest.Server) {
f := &fakeTerdutServer{
teams: map[string]int64{},
teamNames: map[int64]string{},
teamExt: map[string]int64{},
teamOIDC: map[int64][2]string{},
teamDelete: map[int64]bool{},
unknownUsers: map[string]bool{},
escalation: map[int64]tdclient.SetEscalationRequest{},
switches: map[int64]map[int64]tdclient.DeadmanSwitch{},
switchDelete: map[int64]bool{},
integrations: map[int64]map[int64]tdclient.Integration{},
integrationDelete: map[int64]bool{},
}
return f, httptest.NewServer(f)
}
func (f *fakeTerdutServer) ServeHTTP(w http.ResponseWriter, r *http.Request) {
f.mu.Lock()
defer f.mu.Unlock()
if f.expectKey != "" && r.Header.Get("Authorization") != "Bearer "+f.expectKey {
writeJSON(w, http.StatusUnauthorized, map[string]string{errJSONKey: "invalid or expired API key"})
return
}
switch {
case r.URL.Path == "/api/version":
writeJSON(w, http.StatusOK, map[string]string{"version": "test"})
case r.URL.Path == "/api/teams" && r.Method == http.MethodPost:
var req struct {
Name string `json:"name"`
ExternalID string `json:"external_id"`
}
_ = json.NewDecoder(r.Body).Decode(&req)
if id, ok := f.teamExt[req.ExternalID]; ok && req.ExternalID != "" {
writeJSON(w, http.StatusOK, tdclient.Team{ID: id, Name: f.teamNames[id]})
return
}
if _, exists := f.teams[req.Name]; exists {
writeJSON(w, http.StatusConflict, map[string]string{errJSONKey: "a team with that name already exists"})
return
}
f.nextTeamID++
id := f.nextTeamID
f.teams[req.Name] = id
f.teamNames[id] = req.Name
if req.ExternalID != "" {
f.teamExt[req.ExternalID] = id
}
writeJSON(w, http.StatusCreated, tdclient.Team{ID: id, Name: req.Name})
default:
if id, rest, ok := parseTeamSubPath(r.URL.Path); ok {
f.handleTeamSubPath(w, r, id, rest)
return
}
w.WriteHeader(http.StatusNotFound)
}
}
// handleTeamSubPath answers everything under /api/teams/{id}. rest is whatever
// parseTeamSubPath found after "/api/teams/{id}" -- "" for the bare resource.
func (f *fakeTerdutServer) handleTeamSubPath(w http.ResponseWriter, r *http.Request, id int64, rest string) {
switch {
case rest == "" && r.Method == http.MethodPut:
var req struct {
Name string `json:"name"`
}
_ = json.NewDecoder(r.Body).Decode(&req)
oldName, exists := f.teamNames[id]
if !exists {
w.WriteHeader(http.StatusNotFound)
return
}
if other, taken := f.teams[req.Name]; taken && other != id {
writeJSON(w, http.StatusConflict, map[string]string{errJSONKey: "a team with that name already exists"})
return
}
delete(f.teams, oldName)
f.teamNames[id] = req.Name
f.teams[req.Name] = id
w.WriteHeader(http.StatusNoContent)
case rest == "" && r.Method == http.MethodDelete:
name, exists := f.teamNames[id]
if !exists {
w.WriteHeader(http.StatusNotFound)
return
}
delete(f.teams, name)
delete(f.teamNames, id)
f.teamDelete[id] = true
w.WriteHeader(http.StatusNoContent)
case rest == "/oidc-groups" && r.Method == http.MethodPut:
var req struct {
MemberGroup string `json:"member_group"`
OwnerGroup string `json:"owner_group"`
}
_ = json.NewDecoder(r.Body).Decode(&req)
if _, exists := f.teamNames[id]; !exists {
w.WriteHeader(http.StatusNotFound)
return
}
f.teamOIDC[id] = [2]string{req.MemberGroup, req.OwnerGroup}
w.WriteHeader(http.StatusNoContent)
case rest == "/escalation" && r.Method == http.MethodPut:
var req tdclient.SetEscalationRequest
_ = json.NewDecoder(r.Body).Decode(&req)
for _, l := range req.Levels {
for _, t := range l.Targets {
if t.Username != "" && f.unknownUsers[t.Username] {
writeJSON(w, http.StatusBadRequest, map[string]string{errJSONKey: fmt.Sprintf("unknown user %q", t.Username)})
return
}
}
}
f.escalation[id] = req
w.WriteHeader(http.StatusNoContent)
case rest == deadmanSwitchesPath || strings.HasPrefix(rest, deadmanSwitchesPath+"/"):
f.handleDeadmanSubPath(w, r, id, rest)
case rest == "/integrations" || strings.HasPrefix(rest, "/integrations/"):
f.handleIntegrationSubPath(w, r, id, rest)
default:
w.WriteHeader(http.StatusNotFound)
}
}
// handleDeadmanSubPath answers GET/POST /api/teams/{id}/deadman/switches and
// PUT/DELETE .../deadman/switches/{switchID} -- split out of
// handleTeamSubPath for the same gocyclo reason as handleIntegrationSubPath.
func (f *fakeTerdutServer) handleDeadmanSubPath(w http.ResponseWriter, r *http.Request, id int64, rest string) {
switch {
case rest == deadmanSwitchesPath && r.Method == http.MethodGet:
existing := f.switches[id]
out := make([]tdclient.DeadmanSwitch, 0, len(existing))
for _, s := range existing {
out = append(out, s)
}
writeJSON(w, http.StatusOK, out)
case rest == deadmanSwitchesPath && r.Method == http.MethodPost:
var req deadmanSwitchFakeRequest
_ = json.NewDecoder(r.Body).Decode(&req)
name := req.Name
for _, s := range f.switches[id] {
if s.Name == name {
writeJSON(w, http.StatusConflict, map[string]string{errJSONKey: "a switch with that name already exists in this team"})
return
}
}
f.nextSwitchID++
switchID := f.nextSwitchID
sw := tdclient.DeadmanSwitch{
ID: switchID, Name: name, Matcher: req.Matcher,
TimeoutSeconds: req.TimeoutSeconds, Severity: req.Severity,
}
if f.switches[id] == nil {
f.switches[id] = map[int64]tdclient.DeadmanSwitch{}
}
f.switches[id][switchID] = sw
writeJSON(w, http.StatusCreated, sw)
case strings.HasPrefix(rest, "/deadman/switches/") && r.Method == http.MethodPut:
switchID, ok := parseTrailingID(rest, "/deadman/switches/")
if !ok {
w.WriteHeader(http.StatusNotFound)
return
}
if _, exists := f.switches[id][switchID]; !exists {
w.WriteHeader(http.StatusNotFound)
return
}
var req deadmanSwitchFakeRequest
_ = json.NewDecoder(r.Body).Decode(&req)
name := req.Name
if name == "" {
name = f.switches[id][switchID].Name
}
f.switches[id][switchID] = tdclient.DeadmanSwitch{
ID: switchID, Name: name, Matcher: req.Matcher,
TimeoutSeconds: req.TimeoutSeconds, Severity: req.Severity,
}
w.WriteHeader(http.StatusNoContent)
case strings.HasPrefix(rest, "/deadman/switches/") && r.Method == http.MethodDelete:
switchID, ok := parseTrailingID(rest, "/deadman/switches/")
if !ok {
w.WriteHeader(http.StatusNotFound)
return
}
if _, exists := f.switches[id][switchID]; !exists {
w.WriteHeader(http.StatusNotFound)
return
}
delete(f.switches[id], switchID)
f.switchDelete[switchID] = true
w.WriteHeader(http.StatusNoContent)
default:
w.WriteHeader(http.StatusNotFound)
}
}
// handleIntegrationSubPath answers POST /api/teams/{id}/integrations,
// PATCH .../integrations/{integrationID} and DELETE .../integrations/{integrationID}
// -- split out of handleTeamSubPath so that switch's own cyclomatic
// complexity stays under golangci-lint's gocyclo threshold.
func (f *fakeTerdutServer) handleIntegrationSubPath(w http.ResponseWriter, r *http.Request, id int64, rest string) {
switch {
case rest == "/integrations" && r.Method == http.MethodPost:
var req struct {
Name string `json:"name"`
Kind string `json:"kind"`
}
_ = json.NewDecoder(r.Body).Decode(&req)
f.nextIntegrationID++
integID := f.nextIntegrationID
integ := tdclient.Integration{
ID: integID, TeamID: id, Kind: req.Kind, Name: req.Name,
// Key/URL are only ever in *this* response -- never again,
// matching terdut-server's own one-time-show semantics
// (DESIGN.md §4.5) -- so what's stored for later GET/PATCH
// calls in this fake deliberately omits them too.
Key: fmt.Sprintf("webhook-key-%d", integID),
URL: fmt.Sprintf("https://terdut.example.invalid/api/integrations/webhook-key-%d/%s", integID, req.Kind),
}
if f.integrations[id] == nil {
f.integrations[id] = map[int64]tdclient.Integration{}
}
f.integrations[id][integID] = tdclient.Integration{ID: integID, TeamID: id, Kind: req.Kind, Name: req.Name}
writeJSON(w, http.StatusCreated, integ)
case strings.HasPrefix(rest, "/integrations/") && r.Method == http.MethodPatch:
integID, ok := parseTrailingID(rest, "/integrations/")
if !ok {
w.WriteHeader(http.StatusNotFound)
return
}
existing, exists := f.integrations[id][integID]
if !exists {
w.WriteHeader(http.StatusNotFound)
return
}
var req struct {
Name string `json:"name"`
}
_ = json.NewDecoder(r.Body).Decode(&req)
existing.Name = req.Name
f.integrations[id][integID] = existing
w.WriteHeader(http.StatusNoContent)
case strings.HasPrefix(rest, "/integrations/") && r.Method == http.MethodDelete:
integID, ok := parseTrailingID(rest, "/integrations/")
if !ok {
w.WriteHeader(http.StatusNotFound)
return
}
if _, exists := f.integrations[id][integID]; !exists {
w.WriteHeader(http.StatusNotFound)
return
}
delete(f.integrations[id], integID)
f.integrationDelete[integID] = true
w.WriteHeader(http.StatusNoContent)
default:
w.WriteHeader(http.StatusNotFound)
}
}
// deadmanSwitchFakeRequest mirrors tdclient's own (unexported)
// deadmanSwitchRequest -- the fake needs its own copy to decode the same
// wire shape without reaching across package boundaries for an internal type.
type deadmanSwitchFakeRequest struct {
Name string `json:"name,omitempty"`
Matcher string `json:"matcher"`
TimeoutSeconds int64 `json:"timeout_seconds"`
Severity string `json:"severity"`
}
// parseTeamSubPath splits "/api/teams/{id}" from anything after it --
// "" for an exact match, "/oidc-groups", "/escalation", "/deadman/switches"
// or "/deadman/switches/{switchID}" otherwise. Doesn't itself validate the
// suffix; handleTeamSubPath's own switch does that.
func parseTeamSubPath(path string) (id int64, rest string, ok bool) {
const prefix = "/api/teams/"
if !strings.HasPrefix(path, prefix) {
return 0, "", false
}
trimmed := path[len(prefix):]
parts := strings.SplitN(trimmed, "/", 2)
parsedID, err := strconv.ParseInt(parts[0], 10, 64)
if err != nil {
return 0, "", false
}
if len(parts) == 1 {
return parsedID, "", true
}
return parsedID, "/" + parts[1], true
}
// parseTrailingID parses the numeric id after prefix within rest, e.g.
// parseTrailingID("/deadman/switches/7", "/deadman/switches/") -> 7, true.
func parseTrailingID(rest, prefix string) (id int64, ok bool) {
parsedID, err := strconv.ParseInt(strings.TrimPrefix(rest, prefix), 10, 64)
if err != nil {
return 0, false
}
return parsedID, true
}
func writeJSON(w http.ResponseWriter, status int, v any) {
w.Header().Set("Content-Type", "application/json")
w.WriteHeader(status)
_ = json.NewEncoder(w).Encode(v)
}
-49
View File
@@ -1,49 +0,0 @@
package controller
import (
"context"
corev1 "k8s.io/api/core/v1"
apierrors "k8s.io/apimachinery/pkg/api/errors"
metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
"sigs.k8s.io/controller-runtime/pkg/client"
terdutv1alpha1 "git.ryuvia.com/niklas/terdut-operator/api/v1alpha1"
)
// credentialsSecretDataKey is the fixed data key every generated
// credentials Secret this controller writes uses — "whatever a human chose"
// only applied to the bring-your-own input an earlier design draft had and
// removed (DESIGN.md §6); every Secret any controller in this repo
// generates itself uses this one key.
const credentialsSecretDataKey = "token"
// writeOperatorSecret creates or replaces a Secret in namespace (always the
// operator's own, DESIGN.md §6) holding one raw value under
// credentialsSecretDataKey. Shared by every controller that generates a
// credential there -- TerdutServer's instance-scoped key and the bootstrap
// admin-key checkpoint, TerdutTeam's team-scoped key.
func writeOperatorSecret(ctx context.Context, c client.Client, namespace, name, rawValue string) error {
secret := &corev1.Secret{
ObjectMeta: metav1.ObjectMeta{Name: name, Namespace: namespace},
Data: map[string][]byte{credentialsSecretDataKey: []byte(rawValue)},
}
if err := c.Create(ctx, secret); err != nil {
if apierrors.IsAlreadyExists(err) {
return c.Update(ctx, secret)
}
return err
}
return nil
}
// readOperatorSecret reads one generated credential back, by the
// SecretKeyRef a controller's own status stores (always resolved in the
// operator's own namespace, never the referencing CR's -- DESIGN.md §6).
func readOperatorSecret(ctx context.Context, c client.Client, namespace string, ref *terdutv1alpha1.SecretKeyRef) (string, error) {
var secret corev1.Secret
if err := c.Get(ctx, client.ObjectKey{Namespace: namespace, Name: ref.Name}, &secret); err != nil {
return "", err
}
return string(secret.Data[ref.Key]), nil
}
@@ -2,7 +2,9 @@ package controller
import (
"context"
"errors"
"fmt"
"net/http"
"strconv"
"time"
@@ -39,9 +41,8 @@ type TerdutAlertSourceReconciler struct {
client.Client
Scheme *runtime.Scheme
OperatorNamespace string
Recorder recorder.EventRecorder
NewClient func(endpoint string) *tdclient.Client
Recorder recorder.EventRecorder
NewClient func(endpoint string) *tdclient.Client
}
// +kubebuilder:rbac:groups=terdut.ryuvia.com,resources=terdutalertsources,verbs=get;list;watch;create;update;patch;delete
@@ -79,7 +80,7 @@ func (r *TerdutAlertSourceReconciler) Reconcile(ctx context.Context, req ctrl.Re
return ctrl.Result{}, nil
}
team, tc, resolveErr := resolveTeamAndClient(ctx, r.Client, r.OperatorNamespace, as.Namespace, as.Spec.TeamRef, newClient)
team, tc, resolveErr := resolveTeamAndClient(ctx, r.Client, as.Namespace, as.Spec.TeamRef, newClient)
if resolveErr != nil {
return r.setNotReady(ctx, &as, resolveErr.reason, resolveErr.message, waitInterval)
}
@@ -111,7 +112,27 @@ func (r *TerdutAlertSourceReconciler) Reconcile(ctx context.Context, req ctrl.Re
return ctrl.Result{}, err
}
} else if err := tc.RenameIntegration(ctx, teamID, as.Status.IntegrationID, as.Spec.Name); err != nil {
return ctrl.Result{}, fmt.Errorf("PATCH /api/teams/%d/integrations/%d: %w", teamID, as.Status.IntegrationID, err)
se, ok := errors.AsType[*tdclient.StatusError](err)
if !ok || se.Code != http.StatusNotFound {
return ctrl.Result{}, fmt.Errorf("PATCH /api/teams/%d/integrations/%d: %w", teamID, as.Status.IntegrationID, err)
}
// Deleted on the server behind our back, so the webhook URL
// already stopped working: recreate it (a new URL, in the same
// Secret) rather than fail forever. reconcileCreate trusts the
// Secret's existence as "already created", so drop it first.
lostID := as.Status.IntegrationID
if err := r.Delete(ctx, &secret); err != nil && !apierrors.IsNotFound(err) {
return ctrl.Result{}, fmt.Errorf("deleting stale webhook Secret %s/%s: %w", as.Namespace, secretName, err)
}
as.Status.IntegrationID = 0
if err := r.reconcileCreate(ctx, &as, tc, teamID); err != nil {
return ctrl.Result{}, err
}
if r.Recorder != nil {
r.Recorder.Eventf(&as, nil, corev1.EventTypeWarning, "IntegrationRecreated", "IntegrationRecreated",
"integration %d no longer exists on the server; created %d with a new webhook URL (Secret %s)",
lostID, as.Status.IntegrationID, secretName)
}
}
}
@@ -279,7 +300,7 @@ func (r *TerdutAlertSourceReconciler) reconcileDelete(
if as.Status.IntegrationID != 0 {
if team, tc, resolveErr := resolveTeamAndClient(
ctx, r.Client, r.OperatorNamespace, as.Namespace, as.Spec.TeamRef, newClient,
ctx, r.Client, as.Namespace, as.Spec.TeamRef, newClient,
); resolveErr == nil {
if err := tc.DeleteIntegration(ctx, team.Status.TeamID, as.Status.IntegrationID); err != nil {
if r.Recorder != nil {
@@ -34,14 +34,13 @@ var _ = Describe("TerdutAlertSource Controller", func() {
fake, fakeSrv = newFakeTerdutServer()
DeferCleanup(fakeSrv.Close)
srv = bootstrapReadyTerdutServer(ctx, uniqueName("asserver"), fakeSrv.URL)
srv = readyTerdutServer(ctx, uniqueName("asserver"))
team = readyTerdutTeam(ctx, operatorNamespace, uniqueName("asteam"), srv, fakeSrv.URL)
reconciler = &TerdutAlertSourceReconciler{
Client: k8sClient,
Scheme: k8sClient.Scheme(),
OperatorNamespace: operatorNamespace,
NewClient: func(string) *tdclient.Client { return tdclient.New(fakeSrv.URL) },
Client: k8sClient,
Scheme: k8sClient.Scheme(),
NewClient: func(string) *tdclient.Client { return tdclient.New(fakeSrv.URL) },
}
srcName = uniqueName("alertsource")
srcKey = types.NamespacedName{Name: srcName, Namespace: operatorNamespace}
@@ -1,199 +0,0 @@
package controller
import (
"context"
"fmt"
"time"
corev1 "k8s.io/api/core/v1"
apierrors "k8s.io/apimachinery/pkg/api/errors"
"k8s.io/apimachinery/pkg/api/meta"
metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
"k8s.io/apimachinery/pkg/runtime"
ctrl "sigs.k8s.io/controller-runtime"
"sigs.k8s.io/controller-runtime/pkg/client"
"sigs.k8s.io/controller-runtime/pkg/controller/controllerutil"
logf "sigs.k8s.io/controller-runtime/pkg/log"
"sigs.k8s.io/controller-runtime/pkg/recorder"
terdutv1alpha1 "git.ryuvia.com/niklas/terdut-operator/api/v1alpha1"
"git.ryuvia.com/niklas/terdut-operator/internal/tdclient"
)
const deadmanFinalizerName = "terdut.ryuvia.com/terdutdeadmanswitch"
// TerdutDeadmanSwitchReconciler reconciles a TerdutDeadmanSwitch object.
type TerdutDeadmanSwitchReconciler struct {
client.Client
Scheme *runtime.Scheme
OperatorNamespace string
Recorder recorder.EventRecorder
NewClient func(endpoint string) *tdclient.Client
}
// +kubebuilder:rbac:groups=terdut.ryuvia.com,resources=terdutdeadmanswitches,verbs=get;list;watch;create;update;patch;delete
// +kubebuilder:rbac:groups=terdut.ryuvia.com,resources=terdutdeadmanswitches/status,verbs=get;update;patch
// +kubebuilder:rbac:groups=terdut.ryuvia.com,resources=terdutdeadmanswitches/finalizers,verbs=update
// +kubebuilder:rbac:groups=terdut.ryuvia.com,resources=terdutteams,verbs=get;list;watch
// +kubebuilder:rbac:groups="",resources=secrets,verbs=get;list;watch
// +kubebuilder:rbac:groups=events.k8s.io,resources=events,verbs=create;patch
func (r *TerdutDeadmanSwitchReconciler) Reconcile(ctx context.Context, req ctrl.Request) (ctrl.Result, error) {
log := logf.FromContext(ctx)
var sw terdutv1alpha1.TerdutDeadmanSwitch
if err := r.Get(ctx, req.NamespacedName, &sw); err != nil {
if apierrors.IsNotFound(err) {
return ctrl.Result{}, nil
}
return ctrl.Result{}, err
}
newClient := r.NewClient
if newClient == nil {
newClient = tdclient.New
}
if !sw.DeletionTimestamp.IsZero() {
return r.reconcileDeadmanDelete(ctx, &sw, newClient)
}
if !controllerutil.ContainsFinalizer(&sw, deadmanFinalizerName) {
controllerutil.AddFinalizer(&sw, deadmanFinalizerName)
if err := r.Update(ctx, &sw); err != nil {
return ctrl.Result{}, err
}
return ctrl.Result{}, nil
}
team, tc, resolveErr := resolveTeamAndClient(ctx, r.Client, r.OperatorNamespace, sw.Namespace, sw.Spec.TeamRef, newClient)
if resolveErr != nil {
return r.setDeadmanNotReady(ctx, &sw, resolveErr.reason, resolveErr.message, waitInterval)
}
timeout, err := time.ParseDuration(sw.Spec.Timeout)
if err != nil {
return ctrl.Result{}, fmt.Errorf("spec.timeout %q: %w", sw.Spec.Timeout, err)
}
severity := sw.Spec.Severity
if severity == "" {
severity = "critical"
}
if sw.Status.SwitchID == 0 {
if err := r.createOrAdoptDeadmanSwitch(ctx, &sw, tc, team.Status.TeamID, timeout, severity); err != nil {
return ctrl.Result{}, err
}
} else if err := tc.UpdateDeadmanSwitch(ctx, team.Status.TeamID, sw.Status.SwitchID, sw.Spec.Name, sw.Spec.Matcher, int64(timeout.Seconds()), severity); err != nil {
return ctrl.Result{}, fmt.Errorf("PUT /api/teams/%d/deadman/switches/%d: %w", team.Status.TeamID, sw.Status.SwitchID, err)
}
meta.SetStatusCondition(&sw.Status.Conditions, metav1.Condition{
Type: terdutv1alpha1.ConditionReady,
Status: metav1.ConditionTrue,
Reason: terdutv1alpha1.ReasonChildAdopted,
Message: fmt.Sprintf("switch %d applied on team %d", sw.Status.SwitchID, team.Status.TeamID),
})
sw.Status.ObservedGeneration = sw.Generation
if err := r.Status().Update(ctx, &sw); err != nil {
return ctrl.Result{}, err
}
if r.Recorder != nil {
r.Recorder.Eventf(&sw, nil, corev1.EventTypeNormal, terdutv1alpha1.ReasonChildAdopted, terdutv1alpha1.ReasonChildAdopted,
"dead man's switch applied")
}
log.Info("TerdutDeadmanSwitch applied", "name", sw.Name, "switchID", sw.Status.SwitchID)
return ctrl.Result{RequeueAfter: resyncInterval}, nil
}
// createOrAdoptDeadmanSwitch implements this resource's own idempotent-
// create shape (DESIGN.md §4.4, §5): there's no unique-name constraint
// server-side to 409 on, so this lists first and matches by name (the
// server's own derived name, when spec.name is empty) rather than adopting
// after a conflict the API would never actually raise.
func (r *TerdutDeadmanSwitchReconciler) createOrAdoptDeadmanSwitch(
ctx context.Context, sw *terdutv1alpha1.TerdutDeadmanSwitch, tc *tdclient.Client,
teamID int64, timeout time.Duration, severity string,
) error {
existing, err := tc.ListDeadmanSwitches(ctx, teamID)
if err != nil {
return fmt.Errorf("GET /api/teams/%d/deadman/switches: %w", teamID, err)
}
if sw.Spec.Name != "" {
for _, s := range existing {
if s.Name == sw.Spec.Name {
sw.Status.SwitchID = s.ID
return nil
}
}
}
created, err := tc.CreateDeadmanSwitch(ctx, teamID, sw.Spec.Name, sw.Spec.Matcher, int64(timeout.Seconds()), severity)
if err != nil {
return fmt.Errorf("POST /api/teams/%d/deadman/switches: %w", teamID, err)
}
sw.Status.SwitchID = created.ID
return nil
}
func (r *TerdutDeadmanSwitchReconciler) setDeadmanNotReady(
ctx context.Context, sw *terdutv1alpha1.TerdutDeadmanSwitch, reason, message string, d time.Duration,
) (ctrl.Result, error) {
meta.SetStatusCondition(&sw.Status.Conditions, metav1.Condition{
Type: terdutv1alpha1.ConditionReady,
Status: metav1.ConditionFalse,
Reason: reason,
Message: message,
})
sw.Status.ObservedGeneration = sw.Generation
if err := r.Status().Update(ctx, sw); err != nil {
return ctrl.Result{}, err
}
if r.Recorder != nil {
r.Recorder.Eventf(sw, nil, corev1.EventTypeWarning, reason, reason, message)
}
return ctrl.Result{RequeueAfter: d}, nil
}
// reconcileDeadmanDelete calls the real DELETE this resource actually has
// (unlike TerdutEscalationRule) if the team is still resolvable and a
// switch was ever created, then removes the finalizer unconditionally.
func (r *TerdutDeadmanSwitchReconciler) reconcileDeadmanDelete(
ctx context.Context, sw *terdutv1alpha1.TerdutDeadmanSwitch, newClient func(string) *tdclient.Client,
) (ctrl.Result, error) {
if !controllerutil.ContainsFinalizer(sw, deadmanFinalizerName) {
return ctrl.Result{}, nil
}
if sw.Status.SwitchID != 0 {
if team, tc, resolveErr := resolveTeamAndClient(
ctx, r.Client, r.OperatorNamespace, sw.Namespace, sw.Spec.TeamRef, newClient,
); resolveErr == nil {
if err := tc.DeleteDeadmanSwitch(ctx, team.Status.TeamID, sw.Status.SwitchID); err != nil {
if r.Recorder != nil {
r.Recorder.Eventf(sw, nil, corev1.EventTypeWarning, "DeleteFailed", "DeleteFailed", err.Error())
}
return ctrl.Result{}, err
}
}
}
controllerutil.RemoveFinalizer(sw, deadmanFinalizerName)
return ctrl.Result{}, r.Update(ctx, sw)
}
// SetupWithManager sets up the controller with the Manager.
func (r *TerdutDeadmanSwitchReconciler) SetupWithManager(mgr ctrl.Manager) error {
if r.NewClient == nil {
r.NewClient = tdclient.New
}
if r.Recorder == nil {
r.Recorder = mgr.GetEventRecorder("terdutdeadmanswitch-controller")
}
return ctrl.NewControllerManagedBy(mgr).
For(&terdutv1alpha1.TerdutDeadmanSwitch{}).
Named("terdutdeadmanswitch").
Complete(r)
}
@@ -1,213 +0,0 @@
package controller
import (
"context"
"net/http/httptest"
. "github.com/onsi/ginkgo/v2"
. "github.com/onsi/gomega"
"k8s.io/apimachinery/pkg/api/meta"
metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
"k8s.io/apimachinery/pkg/types"
"sigs.k8s.io/controller-runtime/pkg/reconcile"
terdutv1alpha1 "git.ryuvia.com/niklas/terdut-operator/api/v1alpha1"
"git.ryuvia.com/niklas/terdut-operator/internal/tdclient"
)
var _ = Describe("TerdutDeadmanSwitch Controller", func() {
const operatorNamespace = "default"
var (
reconciler *TerdutDeadmanSwitchReconciler
fake *fakeTerdutServer
fakeSrv *httptest.Server
srv *terdutv1alpha1.TerdutServer
team *terdutv1alpha1.TerdutTeam
swName string
swKey types.NamespacedName
)
BeforeEach(func(ctx SpecContext) {
fake, fakeSrv = newFakeTerdutServer()
DeferCleanup(fakeSrv.Close)
srv = bootstrapReadyTerdutServer(ctx, uniqueName("dmserver"), fakeSrv.URL)
team = readyTerdutTeam(ctx, operatorNamespace, uniqueName("dmteam"), srv, fakeSrv.URL)
reconciler = &TerdutDeadmanSwitchReconciler{
Client: k8sClient,
Scheme: k8sClient.Scheme(),
OperatorNamespace: operatorNamespace,
NewClient: func(string) *tdclient.Client { return tdclient.New(fakeSrv.URL) },
}
swName = uniqueName("switch")
swKey = types.NamespacedName{Name: swName, Namespace: operatorNamespace}
})
AfterEach(func(ctx SpecContext) {
sw := &terdutv1alpha1.TerdutDeadmanSwitch{}
if err := k8sClient.Get(ctx, swKey, sw); err == nil {
sw.Finalizers = nil
_ = k8sClient.Update(ctx, sw)
_ = k8sClient.Delete(ctx, sw)
}
teamKey := types.NamespacedName{Name: team.Name, Namespace: operatorNamespace}
if err := k8sClient.Get(ctx, teamKey, team); err == nil {
team.Finalizers = nil
_ = k8sClient.Update(ctx, team)
_ = k8sClient.Delete(ctx, team)
}
srvKey := types.NamespacedName{Name: srv.Name, Namespace: operatorNamespace}
if err := k8sClient.Get(ctx, srvKey, srv); err == nil {
srv.Finalizers = nil
_ = k8sClient.Update(ctx, srv)
_ = k8sClient.Delete(ctx, srv)
}
})
createSwitch := func(ctx context.Context, teamRef terdutv1alpha1.TerdutTeamRef, name, matcher, timeout string) {
sw := &terdutv1alpha1.TerdutDeadmanSwitch{
ObjectMeta: metav1.ObjectMeta{Name: swName, Namespace: operatorNamespace},
Spec: terdutv1alpha1.TerdutDeadmanSwitchSpec{
TeamRef: teamRef,
Name: name,
Matcher: matcher,
Timeout: timeout,
},
}
Expect(k8sClient.Create(ctx, sw)).To(Succeed())
}
reconcileOnce := func(ctx context.Context) {
_, err := reconciler.Reconcile(ctx, reconcile.Request{NamespacedName: swKey})
Expect(err).NotTo(HaveOccurred())
}
readyCondition := func(ctx context.Context) metav1.Condition {
sw := &terdutv1alpha1.TerdutDeadmanSwitch{}
Expect(k8sClient.Get(ctx, swKey, sw)).To(Succeed())
c := meta.FindStatusCondition(sw.Status.Conditions, terdutv1alpha1.ConditionReady)
Expect(c).NotTo(BeNil())
return *c
}
sameTeamRef := func() terdutv1alpha1.TerdutTeamRef {
return terdutv1alpha1.TerdutTeamRef{Name: team.Name}
}
Describe("the happy path", func() {
It("creates the switch server-side", func(ctx SpecContext) {
createSwitch(ctx, sameTeamRef(), "heartbeat", "alertname=Watchdog", "15m")
reconcileOnce(ctx) // finalizer
reconcileOnce(ctx) // create
cond := readyCondition(ctx)
Expect(cond.Status).To(Equal(metav1.ConditionTrue))
Expect(cond.Reason).To(Equal(terdutv1alpha1.ReasonChildAdopted))
sw := &terdutv1alpha1.TerdutDeadmanSwitch{}
Expect(k8sClient.Get(ctx, swKey, sw)).To(Succeed())
Expect(sw.Status.SwitchID).NotTo(BeZero())
created, ok := fake.switches[team.Status.TeamID][sw.Status.SwitchID]
Expect(ok).To(BeTrue())
Expect(created.Matcher).To(Equal("alertname=Watchdog"))
Expect(created.TimeoutSeconds).To(Equal(int64(900)))
Expect(created.Severity).To(Equal("critical")) // kubebuilder default
})
})
Describe("list-and-match-by-name adoption", func() {
It("adopts an already-created switch instead of creating a duplicate", func(ctx SpecContext) {
// Simulates a prior, interrupted reconcile that got as far as
// POSTing the switch -- no 409 signal exists for this resource
// (DESIGN.md §4.4), so the recovery path is GET-list-and-match,
// not adopt-on-409.
fake.nextSwitchID = 1
fake.switches[team.Status.TeamID] = map[int64]tdclient.DeadmanSwitch{
1: {ID: 1, Name: "heartbeat", Matcher: "alertname=Watchdog", TimeoutSeconds: 900, Severity: "critical"},
}
createSwitch(ctx, sameTeamRef(), "heartbeat", "alertname=Watchdog", "15m")
reconcileOnce(ctx) // finalizer
reconcileOnce(ctx)
sw := &terdutv1alpha1.TerdutDeadmanSwitch{}
Expect(k8sClient.Get(ctx, swKey, sw)).To(Succeed())
Expect(sw.Status.SwitchID).To(Equal(int64(1)))
Expect(fake.switches[team.Status.TeamID]).To(HaveLen(1), "should not have created a second switch")
})
})
Describe("update-in-place on spec drift", func() {
It("PUTs the new spec rather than creating a second switch", func(ctx SpecContext) {
createSwitch(ctx, sameTeamRef(), "heartbeat", "alertname=Watchdog", "15m")
reconcileOnce(ctx) // finalizer
reconcileOnce(ctx)
Expect(readyCondition(ctx).Status).To(Equal(metav1.ConditionTrue))
sw := &terdutv1alpha1.TerdutDeadmanSwitch{}
Expect(k8sClient.Get(ctx, swKey, sw)).To(Succeed())
switchID := sw.Status.SwitchID
sw.Spec.Timeout = "30m"
Expect(k8sClient.Update(ctx, sw)).To(Succeed())
reconcileOnce(ctx)
Expect(fake.switches[team.Status.TeamID]).To(HaveLen(1), "update-in-place, not a second switch")
Expect(fake.switches[team.Status.TeamID][switchID].TimeoutSeconds).To(Equal(int64(1800)))
})
})
Describe("waiting on the referenced TerdutTeam", func() {
It("reports TeamRefNotFound when the TerdutTeam doesn't exist", func(ctx SpecContext) {
createSwitch(ctx, terdutv1alpha1.TerdutTeamRef{Name: testRefNotFoundName}, "heartbeat", "alertname=Watchdog", "15m")
reconcileOnce(ctx) // finalizer
reconcileOnce(ctx)
Expect(readyCondition(ctx).Reason).To(Equal(terdutv1alpha1.ReasonTeamRefNotFound))
})
It("reports WaitingForTeam when the TerdutTeam exists but isn't Ready yet", func(ctx SpecContext) {
unreadyName := uniqueName("dmteam-unready")
unready := &terdutv1alpha1.TerdutTeam{
ObjectMeta: metav1.ObjectMeta{Name: unreadyName, Namespace: operatorNamespace},
Spec: terdutv1alpha1.TerdutTeamSpec{
ServerRef: terdutv1alpha1.TerdutServerRef{Name: srv.Name},
DisplayName: testUnreadyDisplayName,
},
}
Expect(k8sClient.Create(ctx, unready)).To(Succeed())
DeferCleanup(func() { _ = k8sClient.Delete(ctx, unready) })
createSwitch(ctx, terdutv1alpha1.TerdutTeamRef{Name: unreadyName}, "heartbeat", "alertname=Watchdog", "15m")
reconcileOnce(ctx) // finalizer
reconcileOnce(ctx)
Expect(readyCondition(ctx).Reason).To(Equal(terdutv1alpha1.ReasonWaitingForTeam))
})
})
Describe("deletion", func() {
It("deletes the switch server-side (the real DELETE this resource has) and removes the finalizer", func(ctx SpecContext) {
createSwitch(ctx, sameTeamRef(), "heartbeat", "alertname=Watchdog", "15m")
reconcileOnce(ctx)
reconcileOnce(ctx)
Expect(readyCondition(ctx).Status).To(Equal(metav1.ConditionTrue))
sw := &terdutv1alpha1.TerdutDeadmanSwitch{}
Expect(k8sClient.Get(ctx, swKey, sw)).To(Succeed())
switchID := sw.Status.SwitchID
Expect(k8sClient.Delete(ctx, sw)).To(Succeed())
reconcileOnce(ctx) // runs the finalizer
Expect(fake.switchDelete[switchID]).To(BeTrue())
err := k8sClient.Get(ctx, swKey, sw)
Expect(err).To(HaveOccurred(), "the TerdutDeadmanSwitch itself should be gone once the finalizer clears")
})
})
})
@@ -1,217 +0,0 @@
package controller
import (
"context"
"fmt"
"time"
corev1 "k8s.io/api/core/v1"
apierrors "k8s.io/apimachinery/pkg/api/errors"
"k8s.io/apimachinery/pkg/api/meta"
metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
"k8s.io/apimachinery/pkg/runtime"
ctrl "sigs.k8s.io/controller-runtime"
"sigs.k8s.io/controller-runtime/pkg/client"
"sigs.k8s.io/controller-runtime/pkg/controller/controllerutil"
logf "sigs.k8s.io/controller-runtime/pkg/log"
"sigs.k8s.io/controller-runtime/pkg/recorder"
terdutv1alpha1 "git.ryuvia.com/niklas/terdut-operator/api/v1alpha1"
"git.ryuvia.com/niklas/terdut-operator/internal/tdclient"
)
// escalationFinalizerName exists for the one undo PUT /api/teams/{teamID}/escalation
// supports, on delete: an empty policy (DESIGN.md §5's general finalizer
// rule expects a server-side counterpart to be undone, but this resource's
// API is GET/PUT-only, with no DELETE at all -- a zeroed PUT is the closest
// equivalent terdut-server itself recognizes as "no policy"
// (escalationPolicy.configured(), confirmed against source: "a policy row
// with no levels is the same as no policy").
const escalationFinalizerName = "terdut.ryuvia.com/terdutescalationrule"
// TerdutEscalationRuleReconciler reconciles a TerdutEscalationRule object.
type TerdutEscalationRuleReconciler struct {
client.Client
Scheme *runtime.Scheme
OperatorNamespace string
Recorder recorder.EventRecorder
NewClient func(endpoint string) *tdclient.Client
}
// +kubebuilder:rbac:groups=terdut.ryuvia.com,resources=terdutescalationrules,verbs=get;list;watch;create;update;patch;delete
// +kubebuilder:rbac:groups=terdut.ryuvia.com,resources=terdutescalationrules/status,verbs=get;update;patch
// +kubebuilder:rbac:groups=terdut.ryuvia.com,resources=terdutescalationrules/finalizers,verbs=update
// +kubebuilder:rbac:groups=terdut.ryuvia.com,resources=terdutteams,verbs=get;list;watch
// +kubebuilder:rbac:groups="",resources=secrets,verbs=get;list;watch
// +kubebuilder:rbac:groups=events.k8s.io,resources=events,verbs=create;patch
func (r *TerdutEscalationRuleReconciler) Reconcile(ctx context.Context, req ctrl.Request) (ctrl.Result, error) {
log := logf.FromContext(ctx)
var rule terdutv1alpha1.TerdutEscalationRule
if err := r.Get(ctx, req.NamespacedName, &rule); err != nil {
if apierrors.IsNotFound(err) {
return ctrl.Result{}, nil
}
return ctrl.Result{}, err
}
newClient := r.NewClient
if newClient == nil {
newClient = tdclient.New
}
if !rule.DeletionTimestamp.IsZero() {
return r.reconcileEscalationDelete(ctx, &rule, newClient)
}
if !controllerutil.ContainsFinalizer(&rule, escalationFinalizerName) {
controllerutil.AddFinalizer(&rule, escalationFinalizerName)
if err := r.Update(ctx, &rule); err != nil {
return ctrl.Result{}, err
}
return ctrl.Result{}, nil
}
team, tc, resolveErr := resolveTeamAndClient(ctx, r.Client, r.OperatorNamespace, rule.Namespace, rule.Spec.TeamRef, newClient)
if resolveErr != nil {
return r.setEscalationNotReady(ctx, &rule, resolveErr.reason, resolveErr.message, waitInterval)
}
body, unknownUser, err := buildEscalationRequest(ctx, tc, rule.Spec)
if err != nil {
return ctrl.Result{}, err
}
if unknownUser != "" {
return r.setEscalationNotReady(ctx, &rule, terdutv1alpha1.ReasonUnknownUser,
fmt.Sprintf("username %q does not resolve to any user", unknownUser), waitInterval)
}
if err := tc.SetEscalation(ctx, team.Status.TeamID, body); err != nil {
return ctrl.Result{}, fmt.Errorf("PUT /api/teams/%d/escalation: %w", team.Status.TeamID, err)
}
meta.SetStatusCondition(&rule.Status.Conditions, metav1.Condition{
Type: terdutv1alpha1.ConditionReady, // "Ready" -- same name, shared across every CRD (DESIGN.md §7)
Status: metav1.ConditionTrue,
Reason: terdutv1alpha1.ReasonChildAdopted,
Message: fmt.Sprintf("escalation policy applied to team %d", team.Status.TeamID),
})
rule.Status.ObservedGeneration = rule.Generation
if err := r.Status().Update(ctx, &rule); err != nil {
return ctrl.Result{}, err
}
if r.Recorder != nil {
r.Recorder.Eventf(&rule, nil, corev1.EventTypeNormal, terdutv1alpha1.ReasonChildAdopted, terdutv1alpha1.ReasonChildAdopted,
"escalation policy applied")
}
log.Info("TerdutEscalationRule applied", "name", rule.Name, "teamID", team.Status.TeamID)
return ctrl.Result{RequeueAfter: resyncInterval}, nil
}
// buildEscalationRequest resolves every "user" target's username to a
// user_id (DESIGN.md §4.3) and translates spec into the wire shape
// SetEscalation sends. Returns the first unresolvable username, if any,
// distinct from a plain error: that's an expected, reportable condition
// (ReasonUnknownUser), not a reconcile failure.
func buildEscalationRequest(
ctx context.Context, tc *tdclient.Client, spec terdutv1alpha1.TerdutEscalationRuleSpec,
) (tdclient.SetEscalationRequest, string, error) {
levels := make([]tdclient.EscalationLevelRequest, len(spec.Levels))
for i, lvl := range spec.Levels {
timeout, err := time.ParseDuration(lvl.Timeout)
if err != nil {
return tdclient.SetEscalationRequest{}, "", fmt.Errorf("spec.levels[%d].timeout %q: %w", i, lvl.Timeout, err)
}
targets := make([]tdclient.EscalationTargetRequest, len(lvl.Targets))
for j, t := range lvl.Targets {
if t.Kind == terdutv1alpha1.EscalationTargetOncall {
targets[j] = tdclient.EscalationTargetRequest{Kind: string(terdutv1alpha1.EscalationTargetOncall)}
continue
}
user, err := tc.GetUserByUsername(ctx, t.Username)
if err != nil {
return tdclient.SetEscalationRequest{}, "", fmt.Errorf("GET /api/users (resolving %q): %w", t.Username, err)
}
if user == nil {
return tdclient.SetEscalationRequest{}, t.Username, nil
}
targets[j] = tdclient.EscalationTargetRequest{Kind: string(terdutv1alpha1.EscalationTargetUser), UserID: &user.ID}
}
levels[i] = tdclient.EscalationLevelRequest{
Position: int64(i + 1),
TimeoutSeconds: int64(timeout.Seconds()),
Targets: targets,
}
}
return tdclient.SetEscalationRequest{
RepeatCount: spec.RepeatCount,
FallbackTopic: spec.FallbackTopic,
Levels: levels,
}, "", nil
}
func (r *TerdutEscalationRuleReconciler) setEscalationNotReady(
ctx context.Context, rule *terdutv1alpha1.TerdutEscalationRule, reason, message string, d time.Duration,
) (ctrl.Result, error) {
meta.SetStatusCondition(&rule.Status.Conditions, metav1.Condition{
Type: terdutv1alpha1.ConditionReady,
Status: metav1.ConditionFalse,
Reason: reason,
Message: message,
})
rule.Status.ObservedGeneration = rule.Generation
if err := r.Status().Update(ctx, rule); err != nil {
return ctrl.Result{}, err
}
if r.Recorder != nil {
r.Recorder.Eventf(rule, nil, corev1.EventTypeWarning, reason, reason, message)
}
return ctrl.Result{RequeueAfter: d}, nil
}
// reconcileEscalationDelete PUTs an empty policy (this resource's only
// available "undo", per this file's own const comment) if the team is
// still resolvable, then removes the finalizer unconditionally -- same
// "the parent's probably going away too" reasoning TerdutTeam's own delete
// path uses for a TerdutServer that's gone.
func (r *TerdutEscalationRuleReconciler) reconcileEscalationDelete(
ctx context.Context, rule *terdutv1alpha1.TerdutEscalationRule, newClient func(string) *tdclient.Client,
) (ctrl.Result, error) {
if !controllerutil.ContainsFinalizer(rule, escalationFinalizerName) {
return ctrl.Result{}, nil
}
if team, tc, resolveErr := resolveTeamAndClient(
ctx, r.Client, r.OperatorNamespace, rule.Namespace, rule.Spec.TeamRef, newClient,
); resolveErr == nil {
if err := tc.SetEscalation(ctx, team.Status.TeamID, tdclient.SetEscalationRequest{Levels: []tdclient.EscalationLevelRequest{}}); err != nil {
if r.Recorder != nil {
r.Recorder.Eventf(rule, nil, corev1.EventTypeWarning, "DeleteFailed", "DeleteFailed", err.Error())
}
return ctrl.Result{}, err
}
}
controllerutil.RemoveFinalizer(rule, escalationFinalizerName)
return ctrl.Result{}, r.Update(ctx, rule)
}
// SetupWithManager sets up the controller with the Manager.
func (r *TerdutEscalationRuleReconciler) SetupWithManager(mgr ctrl.Manager) error {
if r.NewClient == nil {
r.NewClient = tdclient.New
}
if r.Recorder == nil {
r.Recorder = mgr.GetEventRecorder("terdutescalationrule-controller")
}
return ctrl.NewControllerManagedBy(mgr).
For(&terdutv1alpha1.TerdutEscalationRule{}).
Named("terdutescalationrule").
Complete(r)
}
@@ -1,206 +0,0 @@
package controller
import (
"context"
"net/http/httptest"
. "github.com/onsi/ginkgo/v2"
. "github.com/onsi/gomega"
"k8s.io/apimachinery/pkg/api/meta"
metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
"k8s.io/apimachinery/pkg/types"
"sigs.k8s.io/controller-runtime/pkg/reconcile"
terdutv1alpha1 "git.ryuvia.com/niklas/terdut-operator/api/v1alpha1"
"git.ryuvia.com/niklas/terdut-operator/internal/tdclient"
)
var _ = Describe("TerdutEscalationRule Controller", func() {
const operatorNamespace = "default"
var (
reconciler *TerdutEscalationRuleReconciler
fake *fakeTerdutServer
fakeSrv *httptest.Server
srv *terdutv1alpha1.TerdutServer
team *terdutv1alpha1.TerdutTeam
ruleName string
ruleKey types.NamespacedName
)
BeforeEach(func(ctx SpecContext) {
fake, fakeSrv = newFakeTerdutServer()
DeferCleanup(fakeSrv.Close)
srv = bootstrapReadyTerdutServer(ctx, uniqueName("erserver"), fakeSrv.URL)
team = readyTerdutTeam(ctx, operatorNamespace, uniqueName("erteam"), srv, fakeSrv.URL)
reconciler = &TerdutEscalationRuleReconciler{
Client: k8sClient,
Scheme: k8sClient.Scheme(),
OperatorNamespace: operatorNamespace,
NewClient: func(string) *tdclient.Client { return tdclient.New(fakeSrv.URL) },
}
ruleName = uniqueName("escalation")
ruleKey = types.NamespacedName{Name: ruleName, Namespace: operatorNamespace}
})
AfterEach(func(ctx SpecContext) {
rule := &terdutv1alpha1.TerdutEscalationRule{}
if err := k8sClient.Get(ctx, ruleKey, rule); err == nil {
rule.Finalizers = nil
_ = k8sClient.Update(ctx, rule)
_ = k8sClient.Delete(ctx, rule)
}
// team/srv carry their own finalizers from readyTerdutTeam/
// bootstrapReadyTerdutServer -- clear them directly the same way
// terdutteam_controller_test.go's own AfterEach does, rather than
// relying on either reconciler to ever run again here.
teamKey := types.NamespacedName{Name: team.Name, Namespace: operatorNamespace}
if err := k8sClient.Get(ctx, teamKey, team); err == nil {
team.Finalizers = nil
_ = k8sClient.Update(ctx, team)
_ = k8sClient.Delete(ctx, team)
}
srvKey := types.NamespacedName{Name: srv.Name, Namespace: operatorNamespace}
if err := k8sClient.Get(ctx, srvKey, srv); err == nil {
srv.Finalizers = nil
_ = k8sClient.Update(ctx, srv)
_ = k8sClient.Delete(ctx, srv)
}
})
createRule := func(ctx context.Context, teamRef terdutv1alpha1.TerdutTeamRef, levels []terdutv1alpha1.EscalationLevel) {
rule := &terdutv1alpha1.TerdutEscalationRule{
ObjectMeta: metav1.ObjectMeta{Name: ruleName, Namespace: operatorNamespace},
Spec: terdutv1alpha1.TerdutEscalationRuleSpec{
TeamRef: teamRef,
Levels: levels,
},
}
Expect(k8sClient.Create(ctx, rule)).To(Succeed())
}
reconcileOnce := func(ctx context.Context) {
_, err := reconciler.Reconcile(ctx, reconcile.Request{NamespacedName: ruleKey})
Expect(err).NotTo(HaveOccurred())
}
readyCondition := func(ctx context.Context) metav1.Condition {
rule := &terdutv1alpha1.TerdutEscalationRule{}
Expect(k8sClient.Get(ctx, ruleKey, rule)).To(Succeed())
c := meta.FindStatusCondition(rule.Status.Conditions, terdutv1alpha1.ConditionReady)
Expect(c).NotTo(BeNil())
return *c
}
sameTeamRef := func() terdutv1alpha1.TerdutTeamRef {
return terdutv1alpha1.TerdutTeamRef{Name: team.Name}
}
Describe("the happy path", func() {
It("resolves usernames and applies the escalation policy", func(ctx SpecContext) {
fake.seedUser("alice")
createRule(ctx, sameTeamRef(), []terdutv1alpha1.EscalationLevel{
{Timeout: "5m", Targets: []terdutv1alpha1.EscalationTarget{
{Kind: terdutv1alpha1.EscalationTargetUser, Username: "alice"},
}},
{Timeout: "10m", Targets: []terdutv1alpha1.EscalationTarget{
{Kind: terdutv1alpha1.EscalationTargetOncall},
}},
})
reconcileOnce(ctx) // finalizer
reconcileOnce(ctx) // resolve + apply
cond := readyCondition(ctx)
Expect(cond.Status).To(Equal(metav1.ConditionTrue))
Expect(cond.Reason).To(Equal(terdutv1alpha1.ReasonChildAdopted))
applied, ok := fake.escalation[team.Status.TeamID]
Expect(ok).To(BeTrue())
Expect(applied.Levels).To(HaveLen(2))
Expect(applied.Levels[0].Targets[0].Kind).To(Equal("user"))
Expect(applied.Levels[0].Targets[0].UserID).NotTo(BeNil())
Expect(*applied.Levels[0].Targets[0].UserID).To(Equal(fake.users["alice"]))
Expect(applied.Levels[1].Targets[0].Kind).To(Equal("oncall"))
Expect(applied.Levels[1].Targets[0].UserID).To(BeNil())
})
})
Describe("an unresolvable username", func() {
It("reports UnknownUser and never calls PUT /api/teams/{id}/escalation", func(ctx SpecContext) {
createRule(ctx, sameTeamRef(), []terdutv1alpha1.EscalationLevel{
{Timeout: "5m", Targets: []terdutv1alpha1.EscalationTarget{
{Kind: terdutv1alpha1.EscalationTargetUser, Username: "ghost"},
}},
})
reconcileOnce(ctx) // finalizer
reconcileOnce(ctx)
cond := readyCondition(ctx)
Expect(cond.Status).To(Equal(metav1.ConditionFalse))
Expect(cond.Reason).To(Equal(terdutv1alpha1.ReasonUnknownUser))
_, applied := fake.escalation[team.Status.TeamID]
Expect(applied).To(BeFalse())
})
})
Describe("waiting on the referenced TerdutTeam", func() {
It("reports TeamRefNotFound when the TerdutTeam doesn't exist", func(ctx SpecContext) {
createRule(ctx, terdutv1alpha1.TerdutTeamRef{Name: testRefNotFoundName}, []terdutv1alpha1.EscalationLevel{
{Timeout: "5m", Targets: []terdutv1alpha1.EscalationTarget{{Kind: terdutv1alpha1.EscalationTargetOncall}}},
})
reconcileOnce(ctx) // finalizer
reconcileOnce(ctx)
Expect(readyCondition(ctx).Reason).To(Equal(terdutv1alpha1.ReasonTeamRefNotFound))
})
It("reports WaitingForTeam when the TerdutTeam exists but isn't Ready yet", func(ctx SpecContext) {
unreadyName := uniqueName("erteam-unready")
unready := &terdutv1alpha1.TerdutTeam{
ObjectMeta: metav1.ObjectMeta{Name: unreadyName, Namespace: operatorNamespace},
Spec: terdutv1alpha1.TerdutTeamSpec{
ServerRef: terdutv1alpha1.TerdutServerRef{Name: srv.Name},
DisplayName: testUnreadyDisplayName,
},
}
Expect(k8sClient.Create(ctx, unready)).To(Succeed())
DeferCleanup(func() { _ = k8sClient.Delete(ctx, unready) })
createRule(ctx, terdutv1alpha1.TerdutTeamRef{Name: unreadyName}, []terdutv1alpha1.EscalationLevel{
{Timeout: "5m", Targets: []terdutv1alpha1.EscalationTarget{{Kind: terdutv1alpha1.EscalationTargetOncall}}},
})
reconcileOnce(ctx) // finalizer
reconcileOnce(ctx)
Expect(readyCondition(ctx).Reason).To(Equal(terdutv1alpha1.ReasonWaitingForTeam))
})
})
Describe("deletion", func() {
It("PUTs an empty policy (this resource's only available undo) and removes the finalizer", func(ctx SpecContext) {
fake.seedUser("alice")
createRule(ctx, sameTeamRef(), []terdutv1alpha1.EscalationLevel{
{Timeout: "5m", Targets: []terdutv1alpha1.EscalationTarget{
{Kind: terdutv1alpha1.EscalationTargetUser, Username: "alice"},
}},
})
reconcileOnce(ctx)
reconcileOnce(ctx)
Expect(readyCondition(ctx).Status).To(Equal(metav1.ConditionTrue))
Expect(fake.escalation[team.Status.TeamID].Levels).To(HaveLen(1))
rule := &terdutv1alpha1.TerdutEscalationRule{}
Expect(k8sClient.Get(ctx, ruleKey, rule)).To(Succeed())
Expect(k8sClient.Delete(ctx, rule)).To(Succeed())
reconcileOnce(ctx) // runs the finalizer
Expect(fake.escalation[team.Status.TeamID].Levels).To(BeEmpty())
err := k8sClient.Get(ctx, ruleKey, rule)
Expect(err).To(HaveOccurred(), "the TerdutEscalationRule itself should be gone once the finalizer clears")
})
})
})
@@ -1,131 +0,0 @@
package controller
import (
"context"
"errors"
"fmt"
"net/http"
corev1 "k8s.io/api/core/v1"
apierrors "k8s.io/apimachinery/pkg/api/errors"
metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
"sigs.k8s.io/controller-runtime/pkg/client"
terdutv1alpha1 "git.ryuvia.com/niklas/terdut-operator/api/v1alpha1"
"git.ryuvia.com/niklas/terdut-operator/internal/tdclient"
)
// bootstrapStateLostError is DESIGN.md §6's one genuinely pathological
// case: a checkpointed admin credential was used and then lost before the
// lasting credential it was for could be persisted. Distinct from a plain
// error so Reconcile can route it to a Ready: False condition (the
// documented recovery is delete-and-recreate, not an automatic retry) rather
// than treating it as a transient reconcile failure.
type bootstrapStateLostError struct{ detail string }
func (e *bootstrapStateLostError) Error() string {
return fmt.Sprintf(
"server reports already bootstrapped, but neither status.credentialsSecretRef nor a "+
"checkpointed admin credential exist here: %s. This TerdutServer cannot recover a "+
"credential on its own; delete and recreate it", e.detail)
}
// reconcileBootstrap implements DESIGN.md §6 point 1's self-registration
// flow, checkpointed against the two real crash windows in it rather than
// leaving them as theoretical gaps. Only called once
// srv.Status.CredentialsSecretRef is nil and the Deployment has a ready
// replica.
func (r *TerdutServerReconciler) reconcileBootstrap(ctx context.Context, srv *terdutv1alpha1.TerdutServer) error {
adminKey, err := r.getOrCreateCheckpointedAdminKey(ctx, srv)
if err != nil {
return err
}
bc := r.NewClient(serviceURL(srv)).WithToken(adminKey)
instanceKey, err := r.getOrMintInstanceServiceAccountKey(ctx, bc)
if err != nil {
return err
}
credsName := credentialsSecretName(srv)
if err := writeOperatorSecret(ctx, r.Client, r.OperatorNamespace, credsName, instanceKey); err != nil {
return err
}
srv.Status.CredentialsSecretRef = &terdutv1alpha1.SecretKeyRef{Name: credsName, Key: credentialsSecretDataKey}
// Best-effort: the checkpoint has done its job. Leaving it behind on a
// delete failure here isn't a correctness problem (the next reconcile
// finds status.CredentialsSecretRef already set and never looks at the
// checkpoint again) — it would just be an unused Secret sitting around,
// cleaned up for real by the finalizer on delete.
checkpoint := &corev1.Secret{ObjectMeta: metav1.ObjectMeta{Name: checkpointSecretName(srv), Namespace: r.OperatorNamespace}}
_ = r.Delete(ctx, checkpoint)
return nil
}
// getOrCreateCheckpointedAdminKey returns a usable admin key: from the
// checkpoint Secret if an earlier, interrupted attempt already got one, or
// freshly from /api/bootstrap, immediately checkpointed before it's used
// for anything else.
func (r *TerdutServerReconciler) getOrCreateCheckpointedAdminKey(ctx context.Context, srv *terdutv1alpha1.TerdutServer) (string, error) {
checkpointName := checkpointSecretName(srv)
var checkpoint corev1.Secret
err := r.Get(ctx, client.ObjectKey{Namespace: r.OperatorNamespace, Name: checkpointName}, &checkpoint)
switch {
case err == nil:
return string(checkpoint.Data[credentialsSecretDataKey]), nil
case !apierrors.IsNotFound(err):
return "", err
}
bc := r.NewClient(serviceURL(srv))
result, err := bc.Bootstrap(ctx, bootstrapUsername, bootstrapEmail)
if err != nil {
if statusErr, ok := errors.AsType[*tdclient.StatusError](err); ok && statusErr.Code == http.StatusForbidden {
// §1: this operator is the only thing that ever bootstraps a
// server it created, so a 403 here (no checkpoint, no
// status.credentialsSecretRef) means a prior reconcile already
// won this exact race and its checkpoint was lost afterward --
// the one case §6 doesn't try to paper over.
return "", &bootstrapStateLostError{detail: "/api/bootstrap returned 403"}
}
return "", fmt.Errorf("POST /api/bootstrap: %w", err)
}
if err := writeOperatorSecret(ctx, r.Client, r.OperatorNamespace, checkpointName, result.APIKey.Key); err != nil {
return "", fmt.Errorf("checkpointing admin key: %w", err)
}
return result.APIKey.Key, nil
}
// getOrMintInstanceServiceAccountKey mints the operator's own instance-
// scoped service account, or, if an earlier interrupted attempt already
// created it (409), adopts it and mints a fresh key rather than treating
// the conflict as an error (DESIGN.md §6 point 1, §5's general
// adopt-on-conflict rule).
func (r *TerdutServerReconciler) getOrMintInstanceServiceAccountKey(ctx context.Context, bc *tdclient.Client) (string, error) {
result, err := bc.CreateInstanceServiceAccount(ctx, serviceAccountName)
if err == nil {
return result.Key.Key, nil
}
statusErr, ok := errors.AsType[*tdclient.StatusError](err)
if !ok || statusErr.Code != http.StatusConflict {
return "", fmt.Errorf("POST /api/service-accounts: %w", err)
}
sa, err := bc.GetServiceAccountByName(ctx, serviceAccountName)
if err != nil {
return "", fmt.Errorf("GET /api/service-accounts?name=%s (adopting after 409): %w", serviceAccountName, err)
}
if sa == nil {
return "", fmt.Errorf("POST /api/service-accounts 409'd for %q but GET found nothing", serviceAccountName)
}
key, err := bc.CreateServiceAccountKey(ctx, sa.ID, "initial")
if err != nil {
return "", fmt.Errorf("POST /api/service-accounts/%d/keys (adopting after 409): %w", sa.ID, err)
}
return key.Key, nil
}
+10 -110
View File
@@ -2,7 +2,7 @@ package controller
import (
"context"
"errors"
"fmt"
"time"
@@ -16,12 +16,10 @@ import (
"k8s.io/apimachinery/pkg/runtime/schema"
ctrl "sigs.k8s.io/controller-runtime"
"sigs.k8s.io/controller-runtime/pkg/client"
"sigs.k8s.io/controller-runtime/pkg/controller/controllerutil"
logf "sigs.k8s.io/controller-runtime/pkg/log"
"sigs.k8s.io/controller-runtime/pkg/recorder"
terdutv1alpha1 "git.ryuvia.com/niklas/terdut-operator/api/v1alpha1"
"git.ryuvia.com/niklas/terdut-operator/internal/tdclient"
)
// resyncInterval is the periodic requeue on a successful reconcile (DESIGN.md
@@ -39,27 +37,6 @@ const resyncInterval = 5 * time.Minute
// than "someone edited something out of band."
const waitInterval = 15 * time.Second
// finalizerName cleans up the credentials Secret(s) this controller
// generates in the operator's own namespace on delete — the Deployment and
// Service are owned (OwnerReference, DESIGN.md §7) and need no finalizer of
// their own.
const finalizerName = "terdut.ryuvia.com/terdutserver"
// serviceAccountName is the name the operator registers itself under
// server-side (DESIGN.md §6) — a fixed, repo-wide constant, not a spec
// field: it names the automation, not anything about this one TerdutServer.
const serviceAccountName = "terdut-operator"
// bootstrapUsername/bootstrapEmail found the one human-shaped user every
// fresh install needs (terdut-server's handleBootstrap requires both).
// Nobody signs in as this user afterward — its only purpose is minting the
// admin key the controller immediately trades for a real service-account
// key — so these are fixed, not spec fields.
const (
bootstrapUsername = "terdut-operator-bootstrap"
bootstrapEmail = "bootstrap@terdut-operator.local"
)
// postgresqlGVK is the Zalando postgres-operator's CR (DESIGN.md §8).
// Resolved via unstructured rather than vendoring Zalando's own client, to
// keep this operator's dependency on it to "an optional CRD read" rather
@@ -75,21 +52,9 @@ type TerdutServerReconciler struct {
client.Client
Scheme *runtime.Scheme
// OperatorNamespace is where every credentials Secret this controller
// generates lives (DESIGN.md §6) — never the TerdutServer's own
// namespace. Set from the POD_NAMESPACE downward-API env var in
// production (cmd/main.go); tests set it directly.
OperatorNamespace string
// Recorder emits the Kubernetes Events DESIGN.md §12 asks for on every
// externally-visible outcome.
Recorder recorder.EventRecorder
// NewClient builds the terdut-server API client for a given endpoint. A
// field, not a direct tdclient.New call, so tests can substitute an
// httptest.Server's client without a real network round trip. Defaults
// to tdclient.New via SetupWithManager.
NewClient func(endpoint string) *tdclient.Client
}
// +kubebuilder:rbac:groups=terdut.ryuvia.com,resources=terdutservers,verbs=get;list;watch;create;update;patch;delete
@@ -113,26 +78,18 @@ func (r *TerdutServerReconciler) Reconcile(ctx context.Context, req ctrl.Request
return ctrl.Result{}, err
}
if !srv.DeletionTimestamp.IsZero() {
return r.reconcileDelete(ctx, &srv)
}
if !controllerutil.ContainsFinalizer(&srv, finalizerName) {
controllerutil.AddFinalizer(&srv, finalizerName)
if err := r.Update(ctx, &srv); err != nil {
return ctrl.Result{}, err
}
// The Update above re-triggers a reconcile via the watch; nothing
// further to do on this pass.
return ctrl.Result{}, nil
}
dbEnv, dbErr := r.resolveDatabaseEnv(ctx, &srv)
if dbErr != nil {
return r.setNotReady(ctx, &srv, dbErr.reason, dbErr.message, waitInterval)
}
deploy, err := r.reconcileDeployment(ctx, &srv, dbEnv)
operatorKey, err := r.reconcileOperatorKey(ctx, &srv)
if err != nil {
return ctrl.Result{}, err
}
srv.Status.CredentialsSecretRef = operatorKeyRef(&srv)
deploy, err := r.reconcileDeployment(ctx, &srv, dbEnv, operatorKeyHash(operatorKey))
if err != nil {
return ctrl.Result{}, err
}
@@ -143,13 +100,6 @@ func (r *TerdutServerReconciler) Reconcile(ctx context.Context, req ctrl.Request
return ctrl.Result{}, err
}
meta.SetStatusCondition(&srv.Status.Conditions, metav1.Condition{
Type: terdutv1alpha1.ConditionDatabaseReady,
Status: metav1.ConditionTrue,
Reason: terdutv1alpha1.ReasonAdopted,
Message: "database resolved",
})
if deploy.Status.ReadyReplicas < 1 {
return r.setNotReady(ctx, &srv,
terdutv1alpha1.ReasonWaitingForDeployment,
@@ -157,28 +107,12 @@ func (r *TerdutServerReconciler) Reconcile(ctx context.Context, req ctrl.Request
waitInterval)
}
if srv.Status.CredentialsSecretRef == nil {
if err := r.reconcileBootstrap(ctx, &srv); err != nil {
if pending, ok := errors.AsType[*bootstrapStateLostError](err); ok {
return r.setNotReady(ctx, &srv, terdutv1alpha1.ReasonBootstrapStateLost, pending.Error(), waitInterval)
}
return ctrl.Result{}, err
}
}
meta.SetStatusCondition(&srv.Status.Conditions, metav1.Condition{
Type: terdutv1alpha1.ConditionBootstrapped,
Status: metav1.ConditionTrue,
Reason: terdutv1alpha1.ReasonAdopted,
Message: fmt.Sprintf("credentials in Secret %q", srv.Status.CredentialsSecretRef.Name),
})
meta.SetStatusCondition(&srv.Status.Conditions, metav1.Condition{
Type: terdutv1alpha1.ConditionReady,
Status: metav1.ConditionTrue,
Reason: terdutv1alpha1.ReasonAdopted,
Message: "deployment ready, database resolved, credentials bootstrapped",
Message: "deployment ready, database resolved, operator key seeded",
})
srv.Status.ServiceName = srv.Name
srv.Status.ObservedGeneration = srv.Generation
if err := r.Status().Update(ctx, &srv); err != nil {
return ctrl.Result{}, err
@@ -215,35 +149,8 @@ func (r *TerdutServerReconciler) setNotReady(
return ctrl.Result{RequeueAfter: d}, nil
}
// reconcileDelete cleans up the credentials Secret(s) this controller
// generated in the operator's own namespace. The Deployment and Service are
// owned (OwnerReference, DESIGN.md §7) and need no attention here — normal
// GC handles them. There is no server-side "delete this install" call to
// make: bootstrap created a user and a service account, and terdut-server's
// API has no way to delete either (only to revoke individual keys), so
// there is nothing meaningful to undo there either.
func (r *TerdutServerReconciler) reconcileDelete(ctx context.Context, srv *terdutv1alpha1.TerdutServer) (ctrl.Result, error) {
if !controllerutil.ContainsFinalizer(srv, finalizerName) {
return ctrl.Result{}, nil
}
for _, name := range []string{checkpointSecretName(srv), credentialsSecretName(srv)} {
sec := &corev1.Secret{ObjectMeta: metav1.ObjectMeta{Name: name, Namespace: r.OperatorNamespace}}
if err := r.Delete(ctx, sec); err != nil && !apierrors.IsNotFound(err) {
return ctrl.Result{}, err
}
}
controllerutil.RemoveFinalizer(srv, finalizerName)
if err := r.Update(ctx, srv); err != nil {
return ctrl.Result{}, err
}
return ctrl.Result{}, nil
}
// SetupWithManager sets up the controller with the Manager.
func (r *TerdutServerReconciler) SetupWithManager(mgr ctrl.Manager) error {
if r.NewClient == nil {
r.NewClient = tdclient.New
}
if r.Recorder == nil {
r.Recorder = mgr.GetEventRecorder("terdutserver-controller")
}
@@ -252,20 +159,13 @@ func (r *TerdutServerReconciler) SetupWithManager(mgr ctrl.Manager) error {
Owns(&appsv1.Deployment{}).
Owns(&corev1.Service{}).
Owns(&policyv1.PodDisruptionBudget{}).
Owns(&corev1.Secret{}).
Named("terdutserver").
Complete(r)
}
// --- naming ---
func checkpointSecretName(srv *terdutv1alpha1.TerdutServer) string {
return fmt.Sprintf("%s.%s-bootstrap-admin", srv.Namespace, srv.Name)
}
func credentialsSecretName(srv *terdutv1alpha1.TerdutServer) string {
return fmt.Sprintf("%s.%s-instance-credentials", srv.Namespace, srv.Name)
}
func serviceURL(srv *terdutv1alpha1.TerdutServer) string {
port := srv.Spec.Networking.ServicePort
if port == 0 {
@@ -2,14 +2,7 @@ package controller
import (
"context"
"encoding/json"
"fmt"
"net/http"
"net/http/httptest"
"strconv"
"strings"
"sync"
"time"
. "github.com/onsi/ginkgo/v2"
. "github.com/onsi/gomega"
@@ -27,514 +20,8 @@ import (
"sigs.k8s.io/controller-runtime/pkg/reconcile"
terdutv1alpha1 "git.ryuvia.com/niklas/terdut-operator/api/v1alpha1"
"git.ryuvia.com/niklas/terdut-operator/internal/tdclient"
)
// fakeVersionString/errJSONKey are shared by every response fakeTerdutServer
// writes -- goconst would otherwise flag "test" and "error" as repeated
// literals across its handlers.
const (
fakeVersionString = "test"
errJSONKey = "error"
// deadmanSwitchesPath is the literal path (not Printf'd like the others
// below) shared by the exact-match collection route and the dispatcher
// that routes into it -- goconst flags three occurrences of the same
// string, so this is that string, once.
deadmanSwitchesPath = "/deadman/switches"
)
// fakeTerdutServer reproduces the exact stateful semantics of
// /api/bootstrap, /api/service-accounts and /api/version that the
// bootstrap flow depends on (DESIGN.md §6, §11: "terdut-server's REST API
// is faked with a small httptest.Server per controller test ... matching
// the real handlers' request/response shapes"), including the 403-after-
// first-success and 409-on-name-conflict behavior the adopt-on-conflict
// recovery path exists for.
type fakeTerdutServer struct {
mu sync.Mutex
bootstrapped bool
bootstrap403 bool // force every /api/bootstrap call to 403, even the first
nextID int64
accounts map[string]int64 // name -> id
keyMints map[int64]int // id -> number of keys minted so far
nextTeamID int64
teams map[string]int64 // name -> id
teamNames map[int64]string // id -> current name (renames update this)
teamOIDC map[int64][2]string
teamDelete map[int64]bool // id -> true once DELETEd, for 404-on-redelete
// users backs GET /api/users for TerdutEscalationRule's username
// resolution (DESIGN.md §4.3) -- a fixed, pre-seeded directory, since
// nothing in this controller's own flow ever creates a user.
users map[string]int64 // username -> id
// escalation backs PUT /api/teams/{id}/escalation -- an upsert
// server-side (confirmed against source), so this is just "the last
// body PUT for this team", keyed by teamID, with no separate create
// step to model.
escalation map[int64]tdclient.SetEscalationRequest
// switches/nextSwitchID/switchDelete back the dead man's switch
// endpoints -- no unique-name constraint server-side (DESIGN.md §4.4),
// so switches is keyed by id, not name, same as the real API's own
// GET-list-and-match-by-name idempotent-create shape requires.
nextSwitchID int64
switches map[int64]map[int64]tdclient.DeadmanSwitch // teamID -> switchID -> switch
switchDelete map[int64]bool // switchID -> true once DELETEd, for 404-on-redelete
// integrations/nextIntegrationID/integrationDelete back the alert
// source endpoints -- no unique-name constraint server-side either
// (DESIGN.md §4.5), same shape as switches, keyed by id.
nextIntegrationID int64
integrations map[int64]map[int64]tdclient.Integration // teamID -> integrationID -> integration
integrationDelete map[int64]bool // integrationID -> true once DELETEd, for 404-on-redelete
// invites/nextInviteID/inviteDelete back TerdutTeam's own invite-minting
// feature -- no unique constraint on an invite server-side either (every
// POST mints a brand new row, confirmed against source), same keyed-by-id
// shape as switches/integrations.
nextInviteID int64
invites map[int64]map[int64]tdclient.Invite // teamID -> inviteID -> invite
inviteDelete map[int64]bool // inviteID -> true once DELETEd, for 404-on-redelete
}
func newFakeTerdutServer() (*fakeTerdutServer, *httptest.Server) {
f := &fakeTerdutServer{
accounts: map[string]int64{},
keyMints: map[int64]int{},
teams: map[string]int64{},
teamNames: map[int64]string{},
teamOIDC: map[int64][2]string{},
teamDelete: map[int64]bool{},
users: map[string]int64{},
escalation: map[int64]tdclient.SetEscalationRequest{},
switches: map[int64]map[int64]tdclient.DeadmanSwitch{},
switchDelete: map[int64]bool{},
integrations: map[int64]map[int64]tdclient.Integration{},
integrationDelete: map[int64]bool{},
invites: map[int64]map[int64]tdclient.Invite{},
inviteDelete: map[int64]bool{},
}
return f, httptest.NewServer(f)
}
// seedUser registers a username the fake GET /api/users will return --
// called from test setup, before the controller under test ever runs.
// Callers read the assigned id back from f.users themselves, so this has
// nothing left to return.
func (f *fakeTerdutServer) seedUser(username string) {
f.mu.Lock()
defer f.mu.Unlock()
f.nextID++
f.users[username] = f.nextID
}
func (f *fakeTerdutServer) ServeHTTP(w http.ResponseWriter, r *http.Request) {
f.mu.Lock()
defer f.mu.Unlock()
switch {
case r.URL.Path == "/api/version":
writeJSON(w, http.StatusOK, map[string]string{"version": fakeVersionString})
case r.URL.Path == "/api/bootstrap" && r.Method == http.MethodPost:
if f.bootstrapped || f.bootstrap403 {
writeJSON(w, http.StatusForbidden, map[string]string{errJSONKey: "bootstrap already completed"})
return
}
f.bootstrapped = true
writeJSON(w, http.StatusCreated, tdclient.BootstrapResult{
APIKey: tdclient.APIKey{ID: 1, Name: "bootstrap", Key: "admin-key-raw"},
})
case r.URL.Path == "/api/service-accounts" && r.Method == http.MethodPost:
var req struct {
Name string `json:"name"`
Scope string `json:"scope"`
}
_ = json.NewDecoder(r.Body).Decode(&req)
if _, exists := f.accounts[req.Name]; exists {
writeJSON(w, http.StatusConflict, map[string]string{errJSONKey: "a service account with that name already exists"})
return
}
f.nextID++
id := f.nextID
f.accounts[req.Name] = id
f.keyMints[id] = 1
writeJSON(w, http.StatusCreated, tdclient.CreateServiceAccountResult{
ServiceAccount: tdclient.ServiceAccount{ID: id, Name: req.Name, Scope: req.Scope},
Key: tdclient.APIKey{ID: 1, Name: "initial", Key: fmt.Sprintf("instance-key-%d-initial", id)},
})
case r.URL.Path == "/api/service-accounts" && r.Method == http.MethodGet:
name := r.URL.Query().Get("name")
id, exists := f.accounts[name]
if !exists {
writeJSON(w, http.StatusOK, []tdclient.ServiceAccount{})
return
}
writeJSON(w, http.StatusOK, []tdclient.ServiceAccount{{ID: id, Name: name, Scope: "instance"}})
case r.URL.Path == "/api/teams" && r.Method == http.MethodPost:
var req struct {
Name string `json:"name"`
}
_ = json.NewDecoder(r.Body).Decode(&req)
if _, exists := f.teams[req.Name]; exists {
writeJSON(w, http.StatusConflict, map[string]string{errJSONKey: "a team with that name already exists"})
return
}
f.nextTeamID++
id := f.nextTeamID
f.teams[req.Name] = id
f.teamNames[id] = req.Name
writeJSON(w, http.StatusCreated, tdclient.Team{ID: id, Name: req.Name})
case r.URL.Path == "/api/teams" && r.Method == http.MethodGet:
// The controller never calls this without ?name= (TEAM-LOOKUP.md's
// own lookup shape) -- the fake only needs to answer that form.
name := r.URL.Query().Get("name")
id, exists := f.teams[name]
if !exists {
writeJSON(w, http.StatusOK, []tdclient.Team{})
return
}
writeJSON(w, http.StatusOK, []tdclient.Team{{ID: id, Name: f.teamNames[id]}})
case r.URL.Path == "/api/users" && r.Method == http.MethodGet:
// No query filter -- GetUserByUsername fetches the whole list and
// matches client-side (confirmed against source: no server-side
// filter either), so the fake does the same.
users := make([]tdclient.User, 0, len(f.users))
for name, id := range f.users {
users = append(users, tdclient.User{ID: id, Username: name})
}
writeJSON(w, http.StatusOK, users)
default:
if id, name, ok := parseKeysPath(r.URL.Path); ok && r.Method == http.MethodPost {
f.keyMints[id]++
writeJSON(w, http.StatusCreated, tdclient.APIKey{
ID: int64(f.keyMints[id]), Name: name,
Key: fmt.Sprintf("instance-key-%d-mint%d", id, f.keyMints[id]),
})
return
}
if id, rest, ok := parseTeamSubPath(r.URL.Path); ok {
f.handleTeamSubPath(w, r, id, rest)
return
}
w.WriteHeader(http.StatusNotFound)
}
}
// handleTeamSubPath answers everything under /api/teams/{id}: PUT (rename),
// DELETE, PUT .../oidc-groups, PUT .../escalation, and the dead man's
// switch collection/item endpoints. rest is whatever parseTeamSubPath found
// after "/api/teams/{id}" -- "" for the bare resource.
func (f *fakeTerdutServer) handleTeamSubPath(w http.ResponseWriter, r *http.Request, id int64, rest string) {
switch {
case rest == "" && r.Method == http.MethodPut:
var req struct {
Name string `json:"name"`
}
_ = json.NewDecoder(r.Body).Decode(&req)
oldName, exists := f.teamNames[id]
if !exists {
w.WriteHeader(http.StatusNotFound)
return
}
delete(f.teams, oldName)
f.teamNames[id] = req.Name
f.teams[req.Name] = id
w.WriteHeader(http.StatusNoContent)
case rest == "" && r.Method == http.MethodDelete:
name, exists := f.teamNames[id]
if !exists {
w.WriteHeader(http.StatusNotFound)
return
}
delete(f.teams, name)
delete(f.teamNames, id)
f.teamDelete[id] = true
w.WriteHeader(http.StatusNoContent)
case rest == "/oidc-groups" && r.Method == http.MethodPut:
var req struct {
MemberGroup string `json:"member_group"`
OwnerGroup string `json:"owner_group"`
}
_ = json.NewDecoder(r.Body).Decode(&req)
if _, exists := f.teamNames[id]; !exists {
w.WriteHeader(http.StatusNotFound)
return
}
f.teamOIDC[id] = [2]string{req.MemberGroup, req.OwnerGroup}
w.WriteHeader(http.StatusNoContent)
case rest == "/escalation" && r.Method == http.MethodPut:
var req tdclient.SetEscalationRequest
_ = json.NewDecoder(r.Body).Decode(&req)
f.escalation[id] = req
w.WriteHeader(http.StatusNoContent)
case rest == deadmanSwitchesPath || strings.HasPrefix(rest, deadmanSwitchesPath+"/"):
f.handleDeadmanSubPath(w, r, id, rest)
case rest == "/integrations" || strings.HasPrefix(rest, "/integrations/"):
f.handleIntegrationSubPath(w, r, id, rest)
case rest == "/invites" || strings.HasPrefix(rest, "/invites/"):
f.handleInviteSubPath(w, r, id, rest)
default:
w.WriteHeader(http.StatusNotFound)
}
}
// handleDeadmanSubPath answers GET/POST /api/teams/{id}/deadman/switches and
// PUT/DELETE .../deadman/switches/{switchID} -- split out of
// handleTeamSubPath for the same gocyclo reason as handleIntegrationSubPath.
func (f *fakeTerdutServer) handleDeadmanSubPath(w http.ResponseWriter, r *http.Request, id int64, rest string) {
switch {
case rest == deadmanSwitchesPath && r.Method == http.MethodGet:
existing := f.switches[id]
out := make([]tdclient.DeadmanSwitch, 0, len(existing))
for _, s := range existing {
out = append(out, s)
}
writeJSON(w, http.StatusOK, out)
case rest == deadmanSwitchesPath && r.Method == http.MethodPost:
var req deadmanSwitchFakeRequest
_ = json.NewDecoder(r.Body).Decode(&req)
f.nextSwitchID++
switchID := f.nextSwitchID
name := req.Name
if name == "" {
name = "derived-" + req.Matcher
}
sw := tdclient.DeadmanSwitch{
ID: switchID, Name: name, Matcher: req.Matcher,
TimeoutSeconds: req.TimeoutSeconds, Severity: req.Severity,
}
if f.switches[id] == nil {
f.switches[id] = map[int64]tdclient.DeadmanSwitch{}
}
f.switches[id][switchID] = sw
writeJSON(w, http.StatusCreated, sw)
case strings.HasPrefix(rest, "/deadman/switches/") && r.Method == http.MethodPut:
switchID, ok := parseTrailingID(rest, "/deadman/switches/")
if !ok {
w.WriteHeader(http.StatusNotFound)
return
}
if _, exists := f.switches[id][switchID]; !exists {
w.WriteHeader(http.StatusNotFound)
return
}
var req deadmanSwitchFakeRequest
_ = json.NewDecoder(r.Body).Decode(&req)
name := req.Name
if name == "" {
name = f.switches[id][switchID].Name
}
f.switches[id][switchID] = tdclient.DeadmanSwitch{
ID: switchID, Name: name, Matcher: req.Matcher,
TimeoutSeconds: req.TimeoutSeconds, Severity: req.Severity,
}
w.WriteHeader(http.StatusNoContent)
case strings.HasPrefix(rest, "/deadman/switches/") && r.Method == http.MethodDelete:
switchID, ok := parseTrailingID(rest, "/deadman/switches/")
if !ok {
w.WriteHeader(http.StatusNotFound)
return
}
if _, exists := f.switches[id][switchID]; !exists {
w.WriteHeader(http.StatusNotFound)
return
}
delete(f.switches[id], switchID)
f.switchDelete[switchID] = true
w.WriteHeader(http.StatusNoContent)
default:
w.WriteHeader(http.StatusNotFound)
}
}
// handleInviteSubPath answers POST /api/teams/{id}/invites and
// DELETE .../invites/{inviteID} -- split out for the same gocyclo reason as
// handleIntegrationSubPath.
func (f *fakeTerdutServer) handleInviteSubPath(w http.ResponseWriter, r *http.Request, id int64, rest string) {
switch {
case rest == "/invites" && r.Method == http.MethodPost:
var req struct {
Role string `json:"role"`
MaxUses int64 `json:"max_uses"`
}
_ = json.NewDecoder(r.Body).Decode(&req)
f.nextInviteID++
inviteID := f.nextInviteID
inv := tdclient.Invite{
ID: inviteID, TeamID: id, Role: req.Role, MaxUses: req.MaxUses,
ExpiresAt: time.Now().Add(7 * 24 * time.Hour),
URL: fmt.Sprintf("https://terdut.example.invalid/signup?invite=invite-token-%d", inviteID),
}
if f.invites[id] == nil {
f.invites[id] = map[int64]tdclient.Invite{}
}
f.invites[id][inviteID] = inv
writeJSON(w, http.StatusCreated, inv)
case strings.HasPrefix(rest, "/invites/") && r.Method == http.MethodDelete:
inviteID, ok := parseTrailingID(rest, "/invites/")
if !ok {
w.WriteHeader(http.StatusNotFound)
return
}
if _, exists := f.invites[id][inviteID]; !exists {
w.WriteHeader(http.StatusNotFound)
return
}
delete(f.invites[id], inviteID)
f.inviteDelete[inviteID] = true
w.WriteHeader(http.StatusNoContent)
default:
w.WriteHeader(http.StatusNotFound)
}
}
// handleIntegrationSubPath answers POST /api/teams/{id}/integrations,
// PATCH .../integrations/{integrationID} and DELETE .../integrations/{integrationID}
// -- split out of handleTeamSubPath so that switch's own cyclomatic
// complexity stays under golangci-lint's gocyclo threshold.
func (f *fakeTerdutServer) handleIntegrationSubPath(w http.ResponseWriter, r *http.Request, id int64, rest string) {
switch {
case rest == "/integrations" && r.Method == http.MethodPost:
var req struct {
Name string `json:"name"`
Kind string `json:"kind"`
}
_ = json.NewDecoder(r.Body).Decode(&req)
f.nextIntegrationID++
integID := f.nextIntegrationID
integ := tdclient.Integration{
ID: integID, TeamID: id, Kind: req.Kind, Name: req.Name,
// Key/URL are only ever in *this* response -- never again,
// matching terdut-server's own one-time-show semantics
// (DESIGN.md §4.5) -- so what's stored for later GET/PATCH
// calls in this fake deliberately omits them too.
Key: fmt.Sprintf("webhook-key-%d", integID),
URL: fmt.Sprintf("https://terdut.example.invalid/api/integrations/webhook-key-%d/%s", integID, req.Kind),
}
if f.integrations[id] == nil {
f.integrations[id] = map[int64]tdclient.Integration{}
}
f.integrations[id][integID] = tdclient.Integration{ID: integID, TeamID: id, Kind: req.Kind, Name: req.Name}
writeJSON(w, http.StatusCreated, integ)
case strings.HasPrefix(rest, "/integrations/") && r.Method == http.MethodPatch:
integID, ok := parseTrailingID(rest, "/integrations/")
if !ok {
w.WriteHeader(http.StatusNotFound)
return
}
existing, exists := f.integrations[id][integID]
if !exists {
w.WriteHeader(http.StatusNotFound)
return
}
var req struct {
Name string `json:"name"`
}
_ = json.NewDecoder(r.Body).Decode(&req)
existing.Name = req.Name
f.integrations[id][integID] = existing
w.WriteHeader(http.StatusNoContent)
case strings.HasPrefix(rest, "/integrations/") && r.Method == http.MethodDelete:
integID, ok := parseTrailingID(rest, "/integrations/")
if !ok {
w.WriteHeader(http.StatusNotFound)
return
}
if _, exists := f.integrations[id][integID]; !exists {
w.WriteHeader(http.StatusNotFound)
return
}
delete(f.integrations[id], integID)
f.integrationDelete[integID] = true
w.WriteHeader(http.StatusNoContent)
default:
w.WriteHeader(http.StatusNotFound)
}
}
// deadmanSwitchFakeRequest mirrors tdclient's own (unexported)
// deadmanSwitchRequest -- the fake needs its own copy to decode the same
// wire shape without reaching across package boundaries for an internal type.
type deadmanSwitchFakeRequest struct {
Name string `json:"name,omitempty"`
Matcher string `json:"matcher"`
TimeoutSeconds int64 `json:"timeout_seconds"`
Severity string `json:"severity"`
}
// parseTeamSubPath splits "/api/teams/{id}" from anything after it --
// "" for an exact match, "/oidc-groups", "/escalation", "/deadman/switches"
// or "/deadman/switches/{switchID}" otherwise. Doesn't itself validate the
// suffix; handleTeamSubPath's own switch does that.
func parseTeamSubPath(path string) (id int64, rest string, ok bool) {
const prefix = "/api/teams/"
if !strings.HasPrefix(path, prefix) {
return 0, "", false
}
trimmed := path[len(prefix):]
parts := strings.SplitN(trimmed, "/", 2)
parsedID, err := strconv.ParseInt(parts[0], 10, 64)
if err != nil {
return 0, "", false
}
if len(parts) == 1 {
return parsedID, "", true
}
return parsedID, "/" + parts[1], true
}
// parseTrailingID parses the numeric id after prefix within rest, e.g.
// parseTrailingID("/deadman/switches/7", "/deadman/switches/") -> 7, true.
func parseTrailingID(rest, prefix string) (id int64, ok bool) {
parsedID, err := strconv.ParseInt(strings.TrimPrefix(rest, prefix), 10, 64)
if err != nil {
return 0, false
}
return parsedID, true
}
func parseKeysPath(path string) (id int64, mintName string, ok bool) {
var parsedID int64
n, err := fmt.Sscanf(path, "/api/service-accounts/%d/keys", &parsedID)
if err != nil || n != 1 {
return 0, "", false
}
return parsedID, "minted", true
}
func writeJSON(w http.ResponseWriter, status int, v any) {
w.Header().Set("Content-Type", "application/json")
w.WriteHeader(status)
_ = json.NewEncoder(w).Encode(v)
}
var _ = Describe("TerdutServer Controller", func() {
const operatorNamespace = "default"
@@ -546,9 +33,8 @@ var _ = Describe("TerdutServer Controller", func() {
BeforeEach(func() {
reconciler = &TerdutServerReconciler{
Client: k8sClient,
Scheme: k8sClient.Scheme(),
OperatorNamespace: operatorNamespace,
Client: k8sClient,
Scheme: k8sClient.Scheme(),
}
name = fmt.Sprintf("test-server-%d-%d", GinkgoRandomSeed(), GinkgoParallelProcess())
objKey = types.NamespacedName{Name: name, Namespace: operatorNamespace}
@@ -561,9 +47,7 @@ var _ = Describe("TerdutServer Controller", func() {
_ = k8sClient.Update(ctx, srv)
_ = k8sClient.Delete(ctx, srv)
}
for _, n := range []string{checkpointSecretNameFor(name), credentialsSecretNameFor(name)} {
_ = k8sClient.Delete(ctx, &corev1.Secret{ObjectMeta: metav1.ObjectMeta{Name: n, Namespace: operatorNamespace}})
}
_ = k8sClient.Delete(ctx, &corev1.Secret{ObjectMeta: metav1.ObjectMeta{Name: name + "-operator-key", Namespace: operatorNamespace}})
})
dsnSpec := func() terdutv1alpha1.TerdutServerSpec {
@@ -600,127 +84,117 @@ var _ = Describe("TerdutServer Controller", func() {
srv := &terdutv1alpha1.TerdutServer{}
Expect(k8sClient.Get(ctx, objKey, srv)).To(Succeed())
c := meta.FindStatusCondition(srv.Status.Conditions, terdutv1alpha1.ConditionReady)
Expect(c).NotTo(BeNil(), "Ready condition should always be set after a reconcile past the finalizer-add pass")
Expect(c).NotTo(BeNil(), "Ready condition should always be set after a reconcile")
return *c
}
Describe("bring-your-own DSN path", func() {
It("reaches Ready through the full lifecycle: finalizer, Deployment/Service, wait, bootstrap", func(ctx SpecContext) {
fake, fakeSrv := newFakeTerdutServer()
DeferCleanup(fakeSrv.Close)
reconciler.NewClient = func(string) *tdclient.Client { return tdclient.New(fakeSrv.URL) }
envOf := func(ctx context.Context) map[string]corev1.EnvVar {
var deploy appsv1.Deployment
Expect(k8sClient.Get(ctx, objKey, &deploy)).To(Succeed())
out := map[string]corev1.EnvVar{}
for _, e := range deploy.Spec.Template.Spec.Containers[0].Env {
out[e.Name] = e
}
return out
}
Describe("bring-your-own DSN path", func() {
It("creates Deployment, Service and operator key, then reaches Ready once a replica is up", func(ctx SpecContext) {
createServer(ctx, dsnSpec())
// Pass 1: adds the finalizer and returns early.
reconcileOnce(ctx)
srv := &terdutv1alpha1.TerdutServer{}
Expect(k8sClient.Get(ctx, objKey, srv)).To(Succeed())
Expect(srv.Finalizers).To(ContainElement(finalizerName))
// Pass 2: creates Deployment + Service, waits for readiness.
reconcileOnce(ctx)
Expect(readyCondition(ctx).Reason).To(Equal(terdutv1alpha1.ReasonWaitingForDeployment))
var deploy appsv1.Deployment
Expect(k8sClient.Get(ctx, objKey, &deploy)).To(Succeed())
Expect(deploy.Spec.Template.Spec.Containers[0].Image).To(Equal("example.invalid/terdut-server:test"))
envNames := map[string]string{}
for _, e := range deploy.Spec.Template.Spec.Containers[0].Env {
envNames[e.Name] = e.Value
}
Expect(envNames).To(HaveKeyWithValue("TERDUT_DB_DSN", testDSN))
Expect(envNames).To(HaveKeyWithValue("TERDUT_OPERATOR_MODE", "true"))
env := envOf(ctx)
Expect(env["TERDUT_DB_DSN"].Value).To(Equal(testDSN))
Expect(env["TERDUT_OPERATOR_MODE"].Value).To(Equal("true"))
Expect(env).NotTo(HaveKey("TERDUT_DEADMAN_MATCHERS"), "dead man's switches are per team now")
var svc corev1.Service
Expect(k8sClient.Get(ctx, objKey, &svc)).To(Succeed())
Expect(svc.Spec.Ports[0].Port).To(Equal(int32(8080)))
// Simulate the Deployment becoming ready (envtest has no
// kubelet/deployment-controller to do this for real).
markDeploymentReady(ctx)
// Pass 3: bootstraps for real against the fake server.
reconcileOnce(ctx)
Expect(fake.bootstrapped).To(BeTrue())
Expect(k8sClient.Get(ctx, objKey, srv)).To(Succeed())
ready := meta.FindStatusCondition(srv.Status.Conditions, terdutv1alpha1.ConditionReady)
Expect(ready.Status).To(Equal(metav1.ConditionTrue))
Expect(ready.Reason).To(Equal(terdutv1alpha1.ReasonAdopted))
Expect(srv.Status.CredentialsSecretRef).NotTo(BeNil())
Expect(srv.Status.CredentialsSecretRef.Key).To(Equal(credentialsSecretDataKey))
var credsSecret corev1.Secret
Expect(k8sClient.Get(ctx, types.NamespacedName{
Name: srv.Status.CredentialsSecretRef.Name, Namespace: operatorNamespace,
}, &credsSecret)).To(Succeed())
Expect(string(credsSecret.Data[credentialsSecretDataKey])).To(Equal("instance-key-1-initial"))
// The checkpoint is cleaned up once the lasting credential is
// written (DESIGN.md §6 point 2).
var checkpoint corev1.Secret
err := k8sClient.Get(ctx, types.NamespacedName{Name: checkpointSecretName(srv), Namespace: operatorNamespace}, &checkpoint)
Expect(err).To(HaveOccurred())
})
})
Describe("the adopt-on-409 recovery path", func() {
It("mints a fresh key instead of erroring when the service account already exists", func(ctx SpecContext) {
fake, fakeSrv := newFakeTerdutServer()
DeferCleanup(fakeSrv.Close)
fake.bootstrapped = true // an earlier attempt already bootstrapped...
fake.nextID = 1
fake.accounts[serviceAccountName] = 1 // ...and already created the service account.
fake.keyMints[1] = 1
reconciler.NewClient = func(string) *tdclient.Client { return tdclient.New(fakeSrv.URL) }
createServer(ctx, dsnSpec())
reconcileOnce(ctx) // finalizer
reconcileOnce(ctx) // Deployment/Service, WaitingForDeployment
markDeploymentReady(ctx)
// The earlier attempt's checkpoint survived (that's how this
// reconcile can authenticate at all to recover).
Expect(writeOperatorSecret(ctx, k8sClient, operatorNamespace, checkpointSecretName(&terdutv1alpha1.TerdutServer{
ObjectMeta: metav1.ObjectMeta{Name: name, Namespace: operatorNamespace},
}), "admin-key-raw")).To(Succeed())
reconcileOnce(ctx)
srv := &terdutv1alpha1.TerdutServer{}
Expect(k8sClient.Get(ctx, objKey, srv)).To(Succeed())
Expect(meta.FindStatusCondition(srv.Status.Conditions, terdutv1alpha1.ConditionReady).Status).To(Equal(metav1.ConditionTrue))
var credsSecret corev1.Secret
Expect(k8sClient.Get(ctx, types.NamespacedName{
Name: srv.Status.CredentialsSecretRef.Name, Namespace: operatorNamespace,
}, &credsSecret)).To(Succeed())
// Minted fresh, not the (never-seen-by-this-reconcile) "initial"
// key from the account's original creation.
Expect(string(credsSecret.Data[credentialsSecretDataKey])).To(Equal("instance-key-1-mint2"))
ready := meta.FindStatusCondition(srv.Status.Conditions, terdutv1alpha1.ConditionReady)
Expect(ready.Status).To(Equal(metav1.ConditionTrue))
Expect(ready.Reason).To(Equal(terdutv1alpha1.ReasonAdopted))
Expect(srv.Status.CredentialsSecretRef).To(Equal(&terdutv1alpha1.SecretKeyRef{
Name: name + "-operator-key", Key: operatorKeyDataKey,
}))
})
})
Describe("the BootstrapStateLost path", func() {
It("fails closed when the server reports already-bootstrapped with no checkpoint to recover from", func(ctx SpecContext) {
fake, fakeSrv := newFakeTerdutServer()
_ = fake
DeferCleanup(fakeSrv.Close)
fake.bootstrap403 = true
reconciler.NewClient = func(string) *tdclient.Client { return tdclient.New(fakeSrv.URL) }
Describe("the operator key", func() {
keySecret := func(ctx context.Context) corev1.Secret {
var s corev1.Secret
Expect(k8sClient.Get(ctx, types.NamespacedName{Name: name + "-operator-key", Namespace: operatorNamespace}, &s)).To(Succeed())
return s
}
It("is generated once, owned by the TerdutServer, and wired into the pod", func(ctx SpecContext) {
createServer(ctx, dsnSpec())
reconcileOnce(ctx) // finalizer
reconcileOnce(ctx) // Deployment/Service
markDeploymentReady(ctx)
reconcileOnce(ctx)
cond := readyCondition(ctx)
Expect(cond.Status).To(Equal(metav1.ConditionFalse))
Expect(cond.Reason).To(Equal(terdutv1alpha1.ReasonBootstrapStateLost))
s := keySecret(ctx)
value := string(s.Data[operatorKeyDataKey])
Expect(len(value)).To(BeNumerically(">=", 32), "terdut-server refuses a shorter TERDUT_OPERATOR_KEY")
Expect(value).To(HavePrefix(operatorKeyPrefix))
Expect(s.OwnerReferences).To(ContainElement(HaveField("Name", name)))
env := envOf(ctx)
Expect(env["TERDUT_OPERATOR_KEY"].ValueFrom).NotTo(BeNil())
Expect(env["TERDUT_OPERATOR_KEY"].ValueFrom.SecretKeyRef.Name).To(Equal(name + "-operator-key"))
Expect(env["TERDUT_OPERATOR_KEY"].Value).To(BeEmpty(), "the key itself must never appear in the pod spec")
var deploy appsv1.Deployment
Expect(k8sClient.Get(ctx, objKey, &deploy)).To(Succeed())
Expect(deploy.Spec.Template.Annotations).To(HaveKeyWithValue(operatorKeyHashAnnotation, operatorKeyHash(value)))
// A later reconcile keeps the same key: the running server was seeded with it.
reconcileOnce(ctx)
Expect(string(keySecret(ctx).Data[operatorKeyDataKey])).To(Equal(value))
})
It("rolls the Deployment when the Secret is replaced", func(ctx SpecContext) {
createServer(ctx, dsnSpec())
reconcileOnce(ctx)
first := string(keySecret(ctx).Data[operatorKeyDataKey])
s := keySecret(ctx)
Expect(k8sClient.Delete(ctx, &s)).To(Succeed())
reconcileOnce(ctx)
second := string(keySecret(ctx).Data[operatorKeyDataKey])
Expect(second).NotTo(Equal(first))
var deploy appsv1.Deployment
Expect(k8sClient.Get(ctx, objKey, &deploy)).To(Succeed())
Expect(deploy.Spec.Template.Annotations).To(HaveKeyWithValue(operatorKeyHashAnnotation, operatorKeyHash(second)))
})
})
Describe("spec.oidc claims", func() {
It("passes the claim names and trustEmail through", func(ctx SpecContext) {
spec := dsnSpec()
spec.OIDC = terdutv1alpha1.OIDCSpec{
Enabled: true, Issuer: "https://sso.example.invalid", ClientID: "terdut",
UsernameClaim: "upn", EmailClaim: "mail", GroupsClaim: "roles", TrustEmail: true,
SessionMaxAge: "12h",
}
createServer(ctx, spec)
reconcileOnce(ctx)
env := envOf(ctx)
Expect(env["TERDUT_OIDC_USERNAME_CLAIM"].Value).To(Equal("upn"))
Expect(env["TERDUT_OIDC_EMAIL_CLAIM"].Value).To(Equal("mail"))
Expect(env["TERDUT_OIDC_GROUPS_CLAIM"].Value).To(Equal("roles"))
Expect(env["TERDUT_OIDC_TRUST_EMAIL"].Value).To(Equal("true"))
})
})
@@ -735,7 +209,6 @@ var _ = Describe("TerdutServer Controller", func() {
It("waits with reason PostgresClusterNotFound when the postgresql CR doesn't exist yet", func(ctx SpecContext) {
createServer(ctx, zalandoSpec("missing-cluster"))
reconcileOnce(ctx) // finalizer
reconcileOnce(ctx)
Expect(readyCondition(ctx).Reason).To(Equal(terdutv1alpha1.ReasonPostgresClusterNotFound))
@@ -759,13 +232,7 @@ var _ = Describe("TerdutServer Controller", func() {
Expect(k8sClient.Create(ctx, zalandoSecret)).To(Succeed())
DeferCleanup(func() { _ = k8sClient.Delete(ctx, zalandoSecret) })
fake, fakeSrv := newFakeTerdutServer()
_ = fake
DeferCleanup(fakeSrv.Close)
reconciler.NewClient = func(string) *tdclient.Client { return tdclient.New(fakeSrv.URL) }
createServer(ctx, zalandoSpec(clusterName))
reconcileOnce(ctx) // finalizer
reconcileOnce(ctx) // Deployment/Service created with resolved DB env
var deploy appsv1.Deployment
@@ -787,11 +254,6 @@ var _ = Describe("TerdutServer Controller", func() {
Describe("spec.pod", func() {
It("wires pod-level customization onto the right spot on the Deployment", func(ctx SpecContext) {
fake, fakeSrv := newFakeTerdutServer()
_ = fake
DeferCleanup(fakeSrv.Close)
reconciler.NewClient = func(string) *tdclient.Client { return tdclient.New(fakeSrv.URL) }
spec := dsnSpec()
qty := resource.MustParse("250m")
spec.Pod = terdutv1alpha1.PodSpec{
@@ -803,8 +265,7 @@ var _ = Describe("TerdutServer Controller", func() {
ExtraVolumeMounts: []corev1.VolumeMount{{Name: "extra-ca", MountPath: "/etc/extra-ca"}},
}
createServer(ctx, spec)
reconcileOnce(ctx) // finalizer
reconcileOnce(ctx) // Deployment/Service/PDB
reconcileOnce(ctx)
var deploy appsv1.Deployment
Expect(k8sClient.Get(ctx, objKey, &deploy)).To(Succeed())
@@ -828,17 +289,11 @@ var _ = Describe("TerdutServer Controller", func() {
Describe("spec.pod.disruptionBudget", func() {
It("creates an owned PodDisruptionBudget when set, and deletes it once cleared", func(ctx SpecContext) {
fake, fakeSrv := newFakeTerdutServer()
_ = fake
DeferCleanup(fakeSrv.Close)
reconciler.NewClient = func(string) *tdclient.Client { return tdclient.New(fakeSrv.URL) }
spec := dsnSpec()
minAvail := intstr.FromInt32(1)
spec.Pod.DisruptionBudget = &terdutv1alpha1.PodDisruptionBudgetSpec{MinAvailable: &minAvail}
createServer(ctx, spec)
reconcileOnce(ctx) // finalizer
reconcileOnce(ctx) // Deployment/Service/PDB
reconcileOnce(ctx)
var pdb policyv1.PodDisruptionBudget
Expect(k8sClient.Get(ctx, objKey, &pdb)).To(Succeed())
@@ -874,35 +329,6 @@ var _ = Describe("TerdutServer Controller", func() {
})
})
Describe("deletion", func() {
It("removes the credentials and checkpoint Secrets and the finalizer", func(ctx SpecContext) {
fake, fakeSrv := newFakeTerdutServer()
_ = fake
DeferCleanup(fakeSrv.Close)
reconciler.NewClient = func(string) *tdclient.Client { return tdclient.New(fakeSrv.URL) }
createServer(ctx, dsnSpec())
reconcileOnce(ctx)
reconcileOnce(ctx)
markDeploymentReady(ctx)
reconcileOnce(ctx)
srv := &terdutv1alpha1.TerdutServer{}
Expect(k8sClient.Get(ctx, objKey, srv)).To(Succeed())
credsName := srv.Status.CredentialsSecretRef.Name
Expect(k8sClient.Delete(ctx, srv)).To(Succeed())
reconcileOnce(ctx) // runs the finalizer
err := k8sClient.Get(ctx, objKey, srv)
Expect(err).To(HaveOccurred(), "the TerdutServer itself should be gone once the finalizer clears")
var leftover corev1.Secret
err = k8sClient.Get(ctx, types.NamespacedName{Name: credsName, Namespace: operatorNamespace}, &leftover)
Expect(err).To(HaveOccurred(), "the credentials Secret should have been cleaned up by the finalizer")
})
})
When("the TerdutServer object no longer exists", func() {
It("returns no error (deleted between enqueue and reconcile)", func(ctx SpecContext) {
_, err := reconciler.Reconcile(ctx, reconcile.Request{
@@ -912,13 +338,3 @@ var _ = Describe("TerdutServer Controller", func() {
})
})
})
// checkpointSecretNameFor/credentialsSecretNameFor let AfterEach clean up
// without needing a live TerdutServer object (it may already be gone by
// then in the deletion test).
func checkpointSecretNameFor(name string) string {
return fmt.Sprintf("default.%s-bootstrap-admin", name)
}
func credentialsSecretNameFor(name string) string {
return fmt.Sprintf("default.%s-instance-credentials", name)
}
@@ -0,0 +1,101 @@
package controller
import (
"context"
"crypto/rand"
"crypto/sha256"
"encoding/hex"
"fmt"
corev1 "k8s.io/api/core/v1"
apierrors "k8s.io/apimachinery/pkg/api/errors"
metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
"sigs.k8s.io/controller-runtime/pkg/client"
"sigs.k8s.io/controller-runtime/pkg/controller/controllerutil"
terdutv1alpha1 "git.ryuvia.com/niklas/terdut-operator/api/v1alpha1"
)
// operatorKeyDataKey is the data key of the operator key Secret.
const operatorKeyDataKey = "token"
// operatorKeyPrefix marks the key as a service-account credential in logs, the
// way terdut-server's own generated ones are.
const operatorKeyPrefix = "tdsa_"
// operatorKeySecretName is the Secret holding a TerdutServer's operator key. It
// lives beside the Deployment because the pod mounts it (TERDUT_OPERATOR_KEY)
// and a pod can only reference Secrets of its own namespace; it is owned by the
// TerdutServer, so deleting that deletes the key and a recreated one starts
// with a fresh one the server re-seeds on its next start.
func operatorKeySecretName(srv *terdutv1alpha1.TerdutServer) string {
return srv.Name + "-operator-key"
}
// reconcileOperatorKey returns this server's operator key, generating the Secret
// on first use. An existing Secret is never overwritten: the running server was
// seeded with that value, and replacing it would lock the operator out until
// every pod restarted.
func (r *TerdutServerReconciler) reconcileOperatorKey(ctx context.Context, srv *terdutv1alpha1.TerdutServer) (string, error) {
var secret corev1.Secret
key := client.ObjectKey{Namespace: srv.Namespace, Name: operatorKeySecretName(srv)}
err := r.Get(ctx, key, &secret)
if err == nil {
if v := string(secret.Data[operatorKeyDataKey]); v != "" {
return v, nil
}
// Present but empty or foreign: treat as ours to fill rather than
// failing every reconcile on it.
} else if !apierrors.IsNotFound(err) {
return "", err
}
raw := make([]byte, 24)
if _, err := rand.Read(raw); err != nil {
return "", fmt.Errorf("generating operator key: %w", err)
}
value := operatorKeyPrefix + hex.EncodeToString(raw)
secret = corev1.Secret{
ObjectMeta: metav1.ObjectMeta{Name: key.Name, Namespace: key.Namespace},
Data: map[string][]byte{operatorKeyDataKey: []byte(value)},
}
if err := controllerutil.SetControllerReference(srv, &secret, r.Scheme); err != nil {
return "", err
}
if err := r.Create(ctx, &secret); err != nil {
if apierrors.IsAlreadyExists(err) {
// Lost a race with a stale cache read: the next reconcile finds it.
return "", fmt.Errorf("operator key Secret %s appeared during creation; retrying", key.Name)
}
return "", err
}
return value, nil
}
// operatorKeyHash is a short digest of the key, stamped on the pod template so
// a replaced Secret rolls the Deployment: the server only reads the env var at
// start.
func operatorKeyHash(value string) string {
sum := sha256.Sum256([]byte(value))
return hex.EncodeToString(sum[:8])
}
// operatorKeyRef is the reference status.credentialsSecretRef reports.
func operatorKeyRef(srv *terdutv1alpha1.TerdutServer) *terdutv1alpha1.SecretKeyRef {
return &terdutv1alpha1.SecretKeyRef{Name: operatorKeySecretName(srv), Key: operatorKeyDataKey}
}
// readOperatorKey reads a TerdutServer's key from its own namespace, for the
// controllers that call the server's API.
func readOperatorKey(ctx context.Context, c client.Client, srv *terdutv1alpha1.TerdutServer) (string, error) {
var secret corev1.Secret
if err := c.Get(ctx, client.ObjectKey{Namespace: srv.Namespace, Name: operatorKeySecretName(srv)}, &secret); err != nil {
return "", err
}
v := string(secret.Data[operatorKeyDataKey])
if v == "" {
return "", fmt.Errorf("operator key Secret %s/%s is empty", srv.Namespace, secret.Name)
}
return v, nil
}
+38 -29
View File
@@ -3,6 +3,7 @@ package controller
import (
"context"
"fmt"
"maps"
"strings"
appsv1 "k8s.io/api/apps/v1"
@@ -31,32 +32,37 @@ func labelsFor(srv *terdutv1alpha1.TerdutServer) map[string]string {
// object (not just the desired one) so the caller can check
// status.readyReplicas.
func (r *TerdutServerReconciler) reconcileDeployment(
ctx context.Context, srv *terdutv1alpha1.TerdutServer, dbEnv []corev1.EnvVar,
ctx context.Context, srv *terdutv1alpha1.TerdutServer, dbEnv []corev1.EnvVar, keyHash string,
) (*appsv1.Deployment, error) {
deploy := &appsv1.Deployment{ObjectMeta: metav1.ObjectMeta{Name: srv.Name, Namespace: srv.Namespace}}
_, err := controllerutil.CreateOrUpdate(ctx, r.Client, deploy, func() error {
replicas := srv.Spec.Replicas
if replicas == 0 {
replicas = 1
// Only reachable for a TerdutServer stored before the
// +kubebuilder:default=2 marker existed -- the API server's own
// CRD defaulting fills this in for anything created or updated
// through it, so a fresh zero value here means a pre-existing
// object that predates the default, not a deliberate "none"
// (there is no way to request zero replicas).
replicas = 2
}
labels := labelsFor(srv)
deploy.Spec.Replicas = &replicas
deploy.Spec.Selector = &metav1.LabelSelector{MatchLabels: labels}
// Recreate, not RollingUpdate: the sweeper and the notifier are
// unsynchronised singletons inside terdut-server, and two replicas
// overlapping during a rollout would both page for the same
// incident (matches the chart's own deployment.yaml comment).
deploy.Spec.Strategy = appsv1.DeploymentStrategy{Type: appsv1.RecreateDeploymentStrategyType}
// RollingUpdate, not Recreate: terdut-server v0.36.0 put the sweeper,
// the notifier and the migration runner each behind a Postgres
// advisory lock, and gave incident creation its own conflict
// resolution, so two replicas overlapping during a rollout no longer
// double-page, race a migration, or drop a webhook payload (matches
// the chart's own deployment.yaml comment). No explicit
// maxUnavailable/maxSurge: left at the 25%/25% default, which rounds
// to 0/1 at the default replicas: 2 -- already zero-downtime.
deploy.Spec.Strategy = appsv1.DeploymentStrategy{Type: appsv1.RollingUpdateDeploymentStrategyType}
pod := srv.Spec.Pod
deploy.Spec.Template = corev1.PodTemplateSpec{
// pod.Annotations is assigned directly, not merged -- nothing
// else sets pod-template annotations today. If a future change
// needs the controller to set one of its own (e.g. a
// Prometheus-scrape annotation), this needs to become a real
// map merge with a stated precedence rather than silently
// clobbering one side.
ObjectMeta: metav1.ObjectMeta{Labels: labels, Annotations: pod.Annotations},
// The controller's own annotation wins over a same-named user one.
ObjectMeta: metav1.ObjectMeta{Labels: labels, Annotations: podAnnotations(pod.Annotations, keyHash)},
Spec: corev1.PodSpec{
EnableServiceLinks: new(false),
NodeSelector: pod.NodeSelector,
@@ -117,6 +123,17 @@ func (r *TerdutServerReconciler) reconcileService(ctx context.Context, srv *terd
return err
}
// operatorKeyHashAnnotation records which operator key the pods were started
// with, so replacing the key's Secret rolls them.
const operatorKeyHashAnnotation = "terdut.ryuvia.com/operator-key-hash"
func podAnnotations(user map[string]string, keyHash string) map[string]string {
out := make(map[string]string, len(user)+1)
maps.Copy(out, user)
out[operatorKeyHashAnnotation] = keyHash
return out
}
func servicePort(srv *terdutv1alpha1.TerdutServer) int32 {
if srv.Spec.Networking.ServicePort == 0 {
return 8080
@@ -184,13 +201,6 @@ func buildEnv(srv *terdutv1alpha1.TerdutServer, dbEnv []corev1.EnvVar) []corev1.
corev1.EnvVar{Name: "TERDUT_ARCHIVE_AFTER", Value: sweeper.ArchiveAfter},
)
deadman := srv.Spec.Deadman
env = append(env,
corev1.EnvVar{Name: "TERDUT_DEADMAN_MATCHERS", Value: deadman.Matchers},
corev1.EnvVar{Name: "TERDUT_DEADMAN_TIMEOUT", Value: deadman.Timeout},
corev1.EnvVar{Name: "TERDUT_DEADMAN_SEVERITY", Value: deadman.Severity},
)
notify := srv.Spec.Notify
if notify.NtfyURL != "" {
env = append(env,
@@ -220,6 +230,9 @@ func buildEnv(srv *terdutv1alpha1.TerdutServer, dbEnv []corev1.EnvVar) []corev1.
// here on" -- which is unconditionally true for anything this
// operator creates).
corev1.EnvVar{Name: "TERDUT_OPERATOR_MODE", Value: "true"},
// The credential this operator calls the API with; the server
// creates or re-keys its instance-scoped account from it at start.
corev1.EnvVar{Name: "TERDUT_OPERATOR_KEY", ValueFrom: secretEnvSource(operatorKeyRef(srv))},
)
if oidc := srv.Spec.OIDC; oidc.Enabled {
@@ -228,14 +241,10 @@ func buildEnv(srv *terdutv1alpha1.TerdutServer, dbEnv []corev1.EnvVar) []corev1.
corev1.EnvVar{Name: "TERDUT_OIDC_CLIENT_ID", Value: oidc.ClientID},
corev1.EnvVar{Name: "TERDUT_OIDC_NAME", Value: oidc.Name},
corev1.EnvVar{Name: "TERDUT_OIDC_SCOPES", Value: oidc.Scopes},
// terdut-server's own defaults for the claims/trust-email knobs
// the chart exposes but DESIGN.md's spec doesn't (§4.1's doc
// comment on OIDCSpec) -- not configurable here, not an
// oversight.
corev1.EnvVar{Name: "TERDUT_OIDC_USERNAME_CLAIM", Value: "preferred_username"},
corev1.EnvVar{Name: "TERDUT_OIDC_EMAIL_CLAIM", Value: "email"},
corev1.EnvVar{Name: "TERDUT_OIDC_GROUPS_CLAIM", Value: "groups"},
corev1.EnvVar{Name: "TERDUT_OIDC_TRUST_EMAIL", Value: "false"},
corev1.EnvVar{Name: "TERDUT_OIDC_USERNAME_CLAIM", Value: oidc.UsernameClaim},
corev1.EnvVar{Name: "TERDUT_OIDC_EMAIL_CLAIM", Value: oidc.EmailClaim},
corev1.EnvVar{Name: "TERDUT_OIDC_GROUPS_CLAIM", Value: oidc.GroupsClaim},
corev1.EnvVar{Name: "TERDUT_OIDC_TRUST_EMAIL", Value: boolString(oidc.TrustEmail)},
corev1.EnvVar{Name: "TERDUT_OIDC_SESSION_MAX_AGE", Value: oidc.SessionMaxAge},
)
if oidc.ClientSecretRef != nil {
@@ -1,84 +0,0 @@
package controller
import (
"context"
"errors"
"fmt"
"net/http"
terdutv1alpha1 "git.ryuvia.com/niklas/terdut-operator/api/v1alpha1"
"git.ryuvia.com/niklas/terdut-operator/internal/tdclient"
)
// createOrAdoptTeam calls POST /api/teams with the TerdutServer's
// instance-scoped credential, or, if an earlier interrupted attempt
// already created this name (409), adopts it via GET /api/teams?name=
// (TEAM-LOOKUP.md) rather than treating the conflict as an error --
// DESIGN.md §5's general adopt-on-conflict rule, the same shape
// TerdutServer's own bootstrap flow uses for minting its instance account.
func (r *TerdutTeamReconciler) createOrAdoptTeam(ctx context.Context, team *terdutv1alpha1.TerdutTeam, instanceClient *tdclient.Client) error {
created, err := instanceClient.CreateTeam(ctx, team.Spec.DisplayName)
if err == nil {
team.Status.TeamID = created.ID
return nil
}
statusErr, ok := errors.AsType[*tdclient.StatusError](err)
if !ok || statusErr.Code != http.StatusConflict {
return fmt.Errorf("POST /api/teams: %w", err)
}
found, err := instanceClient.GetTeamByName(ctx, team.Spec.DisplayName)
if err != nil {
return fmt.Errorf("GET /api/teams?name=%s (adopting after 409): %w", team.Spec.DisplayName, err)
}
if found == nil {
// Genuinely pathological, not just a narrow crash window: the name
// was taken a moment ago and isn't now. Surfaced as a plain error
// (standard requeue-with-backoff) rather than a dedicated
// condition -- there's no documented recovery to point at that
// differs from "try again".
return fmt.Errorf("POST /api/teams 409'd for %q but GET found nothing", team.Spec.DisplayName)
}
team.Status.TeamID = found.ID
return nil
}
// mintTeamCredential mints this team's own team-scoped service account,
// using the TerdutServer's instance-scoped credential (DESIGN.md §6 point
// 3: an instance-scoped caller may do this against any team). Adopts via
// GET+mint-new-key on a 409, the same pattern TerdutServer's own bootstrap
// flow uses.
func (r *TerdutTeamReconciler) mintTeamCredential(ctx context.Context, team *terdutv1alpha1.TerdutTeam, instanceClient *tdclient.Client) error {
saName := teamServiceAccountName(team)
result, err := instanceClient.CreateTeamServiceAccount(ctx, saName, team.Status.TeamID)
var key string
if err == nil {
key = result.Key.Key
} else {
statusErr, ok := errors.AsType[*tdclient.StatusError](err)
if !ok || statusErr.Code != http.StatusConflict {
return fmt.Errorf("POST /api/service-accounts (team scope): %w", err)
}
sa, err := instanceClient.GetServiceAccountByName(ctx, saName)
if err != nil {
return fmt.Errorf("GET /api/service-accounts?name=%s (adopting after 409): %w", saName, err)
}
if sa == nil {
return fmt.Errorf("POST /api/service-accounts 409'd for %q but GET found nothing", saName)
}
minted, err := instanceClient.CreateServiceAccountKey(ctx, sa.ID, "initial")
if err != nil {
return fmt.Errorf("POST /api/service-accounts/%d/keys (adopting after 409): %w", sa.ID, err)
}
key = minted.Key
}
credsName := teamCredentialsSecretName(team)
if err := writeOperatorSecret(ctx, r.Client, r.OperatorNamespace, credsName, key); err != nil {
return err
}
team.Status.CredentialsSecretRef = &terdutv1alpha1.SecretKeyRef{Name: credsName, Key: credentialsSecretDataKey}
return nil
}
+127
View File
@@ -0,0 +1,127 @@
package controller
import (
"context"
"errors"
"fmt"
"net/http"
"strings"
"time"
terdutv1alpha1 "git.ryuvia.com/niklas/terdut-operator/api/v1alpha1"
"git.ryuvia.com/niklas/terdut-operator/internal/tdclient"
)
// reconcileEscalation applies spec.escalation as the team's whole ladder, or
// clears it when the field is absent (an empty ladder is the server's "none").
// The second return is an expected, reportable condition (a bad duration, an
// unknown user), the third a failure to retry.
func (r *TerdutTeamReconciler) reconcileEscalation(
ctx context.Context, tc *tdclient.Client, team *terdutv1alpha1.TerdutTeam,
) (*teamError, error) {
body, err := buildEscalationRequest(team.Spec.Escalation)
if err != nil {
return &teamError{reason: terdutv1alpha1.ReasonInvalidSpec, message: err.Error()}, nil
}
if err := tc.SetEscalation(ctx, team.Status.TeamID, body); err != nil {
// The server resolves usernames; one it does not know is a state to
// wait out (the person may be created later), not a failure.
if se, ok := errors.AsType[*tdclient.StatusError](err); ok &&
se.Code == http.StatusBadRequest && strings.HasPrefix(se.Message, "unknown user") {
return &teamError{reason: terdutv1alpha1.ReasonUnknownUser, message: se.Message}, nil
}
return nil, fmt.Errorf("PUT /api/teams/%d/escalation: %w", team.Status.TeamID, err)
}
return nil, nil
}
func buildEscalationRequest(spec *terdutv1alpha1.EscalationSpec) (tdclient.SetEscalationRequest, error) {
if spec == nil {
return tdclient.SetEscalationRequest{Levels: []tdclient.EscalationLevelRequest{}}, nil
}
levels := make([]tdclient.EscalationLevelRequest, len(spec.Levels))
for i, lvl := range spec.Levels {
timeout, err := time.ParseDuration(lvl.Timeout)
if err != nil || timeout <= 0 {
return tdclient.SetEscalationRequest{}, fmt.Errorf("spec.escalation.levels[%d].timeout %q is not a positive duration", i, lvl.Timeout)
}
targets := make([]tdclient.EscalationTargetRequest, len(lvl.Targets))
for j, t := range lvl.Targets {
targets[j] = tdclient.EscalationTargetRequest{Kind: string(t.Kind), Username: t.Username}
}
levels[i] = tdclient.EscalationLevelRequest{
Position: int64(i + 1),
TimeoutSeconds: int64(timeout.Seconds()),
Targets: targets,
}
}
return tdclient.SetEscalationRequest{
RepeatCount: spec.RepeatCount,
FallbackTopic: spec.FallbackTopic,
Levels: levels,
}, nil
}
// reconcileDeadmanSwitches makes the team's switches on the server exactly
// spec.deadmanSwitches, matched by name (unique per team server-side): create
// what is missing, update what differs, delete what is not listed. In operator
// mode nobody else can add one, so anything extra is leftover to remove.
func (r *TerdutTeamReconciler) reconcileDeadmanSwitches(
ctx context.Context, tc *tdclient.Client, team *terdutv1alpha1.TerdutTeam,
) (*teamError, error) {
type want struct {
spec terdutv1alpha1.DeadmanSwitchSpec
seconds int64
}
desired := make(map[string]want, len(team.Spec.DeadmanSwitches))
for _, sw := range team.Spec.DeadmanSwitches {
timeout, err := time.ParseDuration(sw.Timeout)
if err != nil || timeout <= 0 {
return &teamError{
reason: terdutv1alpha1.ReasonInvalidSpec,
message: fmt.Sprintf("spec.deadmanSwitches[%q].timeout %q is not a positive duration", sw.Name, sw.Timeout),
}, nil
}
desired[sw.Name] = want{spec: sw, seconds: int64(timeout.Seconds())}
}
existing, err := tc.ListDeadmanSwitches(ctx, team.Status.TeamID)
if err != nil {
return nil, fmt.Errorf("GET /api/teams/%d/deadman/switches: %w", team.Status.TeamID, err)
}
seen := make(map[string]bool, len(existing))
for _, have := range existing {
w, keep := desired[have.Name]
if !keep {
if err := tc.DeleteDeadmanSwitch(ctx, team.Status.TeamID, have.ID); err != nil {
return nil, fmt.Errorf("DELETE deadman switch %q: %w", have.Name, err)
}
continue
}
seen[have.Name] = true
severity := severityOrDefault(w.spec.Severity)
if have.Matcher == w.spec.Matcher && have.TimeoutSeconds == w.seconds && have.Severity == severity {
continue
}
if err := tc.UpdateDeadmanSwitch(ctx, team.Status.TeamID, have.ID, w.spec.Name, w.spec.Matcher, w.seconds, severity); err != nil {
return nil, fmt.Errorf("PUT deadman switch %q: %w", have.Name, err)
}
}
for _, sw := range team.Spec.DeadmanSwitches {
if seen[sw.Name] {
continue
}
w := desired[sw.Name]
if _, err := tc.CreateDeadmanSwitch(ctx, team.Status.TeamID, sw.Name, sw.Matcher, w.seconds, severityOrDefault(sw.Severity)); err != nil {
return nil, fmt.Errorf("POST deadman switch %q: %w", sw.Name, err)
}
}
return nil, nil
}
func severityOrDefault(s string) string {
if s == "" {
return "critical"
}
return s
}
+130 -119
View File
@@ -5,7 +5,6 @@ import (
"errors"
"fmt"
"net/http"
"time"
corev1 "k8s.io/api/core/v1"
apierrors "k8s.io/apimachinery/pkg/api/errors"
@@ -15,33 +14,25 @@ import (
ctrl "sigs.k8s.io/controller-runtime"
"sigs.k8s.io/controller-runtime/pkg/client"
"sigs.k8s.io/controller-runtime/pkg/controller/controllerutil"
"sigs.k8s.io/controller-runtime/pkg/handler"
logf "sigs.k8s.io/controller-runtime/pkg/log"
"sigs.k8s.io/controller-runtime/pkg/reconcile"
"sigs.k8s.io/controller-runtime/pkg/recorder"
terdutv1alpha1 "git.ryuvia.com/niklas/terdut-operator/api/v1alpha1"
"git.ryuvia.com/niklas/terdut-operator/internal/tdclient"
)
// teamFinalizerName cleans up the team-scoped credentials Secret this
// controller generates in the operator's own namespace on delete.
// teamFinalizerName deletes the team on the server before the CR goes.
const teamFinalizerName = "terdut.ryuvia.com/terdutteam"
// TerdutTeamReconciler reconciles a TerdutTeam object.
//
// Every child resolves its own teamRef/serverRef independently and never
// chains up through another controller (DESIGN.md §5) — this one talks
// directly to the TerdutServer it references and to terdut-server's API,
// never to TerdutServerReconciler.
// TerdutTeamReconciler reconciles a TerdutTeam: the team itself plus its
// escalation ladder and dead man's switches, all through the referenced
// TerdutServer's operator key.
type TerdutTeamReconciler struct {
client.Client
Scheme *runtime.Scheme
// OperatorNamespace is where every credentials Secret this controller
// reads (the referenced TerdutServer's) or writes (this team's own)
// lives (DESIGN.md §6) — never a TerdutTeam's or TerdutServer's own
// namespace.
OperatorNamespace string
Recorder recorder.EventRecorder
NewClient func(endpoint string) *tdclient.Client
}
@@ -50,10 +41,16 @@ type TerdutTeamReconciler struct {
// +kubebuilder:rbac:groups=terdut.ryuvia.com,resources=terdutteams/status,verbs=get;update;patch
// +kubebuilder:rbac:groups=terdut.ryuvia.com,resources=terdutteams/finalizers,verbs=update
// +kubebuilder:rbac:groups=terdut.ryuvia.com,resources=terdutservers,verbs=get;list;watch
// +kubebuilder:rbac:groups="",resources=secrets,verbs=get;list;watch;create;update;patch;delete
// +kubebuilder:rbac:groups="",resources=secrets,verbs=get;list;watch
// +kubebuilder:rbac:groups="",resources=namespaces,verbs=get;list;watch
// +kubebuilder:rbac:groups=events.k8s.io,resources=events,verbs=create;patch
// teamExternalID is the identity this CR gives its team on the server, so the
// team is found by who owns it and not by its (editable, global) display name.
func teamExternalID(team *terdutv1alpha1.TerdutTeam) string {
return team.Namespace + "/" + team.Name
}
func (r *TerdutTeamReconciler) Reconcile(ctx context.Context, req ctrl.Request) (ctrl.Result, error) {
log := logf.FromContext(ctx)
@@ -79,67 +76,56 @@ func (r *TerdutTeamReconciler) Reconcile(ctx context.Context, req ctrl.Request)
srv, resolveErr := r.resolveServerRef(ctx, &team)
if resolveErr != nil {
return r.setTeamNotReady(ctx, &team, resolveErr.reason, resolveErr.message, waitInterval)
return r.setTeamNotReady(ctx, &team, resolveErr.reason, resolveErr.message)
}
if srv.Status.CredentialsSecretRef == nil {
if !meta.IsStatusConditionTrue(srv.Status.Conditions, terdutv1alpha1.ConditionReady) {
return r.setTeamNotReady(ctx, &team, terdutv1alpha1.ReasonWaitingForServer,
fmt.Sprintf("TerdutServer %q is not Bootstrapped yet", srv.Name), waitInterval)
fmt.Sprintf("TerdutServer %q is not Ready yet", srv.Name))
}
newClient := r.NewClient
if newClient == nil {
newClient = tdclient.New
}
instanceKey, err := readOperatorSecret(ctx, r.Client, r.OperatorNamespace, srv.Status.CredentialsSecretRef)
tc, err := r.serverClient(ctx, srv)
if err != nil {
return ctrl.Result{}, err
}
endpoint := serviceURL(srv)
team.Status.ServerEndpoint = endpoint
instanceClient := newClient(endpoint).WithToken(instanceKey)
if team.Status.TeamID == 0 {
if err := r.createOrAdoptTeam(ctx, &team, instanceClient); err != nil {
return ctrl.Result{}, err
}
}
if team.Status.CredentialsSecretRef == nil {
if err := r.mintTeamCredential(ctx, &team, instanceClient); err != nil {
return ctrl.Result{}, err
}
}
teamKey, err := readOperatorSecret(ctx, r.Client, r.OperatorNamespace, team.Status.CredentialsSecretRef)
// Idempotent on the external id: the same team comes back whether this is
// the first reconcile, a retry after a lost status write, or a repair after
// somebody deleted the team behind our back.
created, err := tc.CreateTeam(ctx, team.Spec.DisplayName, teamExternalID(&team))
if err != nil {
return ctrl.Result{}, err
if se, ok := errors.AsType[*tdclient.StatusError](err); ok && se.Code == http.StatusConflict {
return r.setTeamNotReady(ctx, &team, terdutv1alpha1.ReasonTeamNameTaken, nameTakenMessage(&team))
}
return ctrl.Result{}, fmt.Errorf("POST /api/teams: %w", err)
}
teamClient := newClient(endpoint).WithToken(teamKey)
team.Status.TeamID = created.ID
// Owner-gated on terdut-server, so this always runs with the
// team-scoped credential just minted above, never the instance-scoped
// one used to create the team (internal/api/teams.go's requireTeamOwner
// has no branch for an instance-scoped service account, confirmed
// against source). Applied unconditionally rather than diffed against a
// stored "last-applied" value: both calls are idempotent PUTs of the
// whole resource, the same "cheap because it's small" reasoning §5
// already applies to the escalation policy's whole-policy PUT.
if err := teamClient.RenameTeam(ctx, team.Status.TeamID, team.Spec.DisplayName); err != nil {
return ctrl.Result{}, fmt.Errorf("PUT /api/teams/%d: %w", team.Status.TeamID, err)
if err := tc.RenameTeam(ctx, created.ID, team.Spec.DisplayName); err != nil {
if se, ok := errors.AsType[*tdclient.StatusError](err); ok && se.Code == http.StatusConflict {
return r.setTeamNotReady(ctx, &team, terdutv1alpha1.ReasonTeamNameTaken, nameTakenMessage(&team))
}
return ctrl.Result{}, fmt.Errorf("PUT /api/teams/%d: %w", created.ID, err)
}
if err := teamClient.SetTeamOIDCGroups(ctx, team.Status.TeamID, team.Spec.OIDC.MemberGroup, team.Spec.OIDC.OwnerGroup); err != nil {
return ctrl.Result{}, fmt.Errorf("PUT /api/teams/%d/oidc-groups: %w", team.Status.TeamID, err)
if err := tc.SetTeamOIDCGroups(ctx, created.ID, team.Spec.OIDC.MemberGroup, team.Spec.OIDC.OwnerGroup); err != nil {
return ctrl.Result{}, fmt.Errorf("PUT /api/teams/%d/oidc-groups: %w", created.ID, err)
}
if err := r.reconcileInvite(ctx, &team, teamClient); err != nil {
if condErr, err := r.reconcileEscalation(ctx, tc, &team); err != nil {
return ctrl.Result{}, err
} else if condErr != nil {
return r.setTeamNotReady(ctx, &team, condErr.reason, condErr.message)
}
if condErr, err := r.reconcileDeadmanSwitches(ctx, tc, &team); err != nil {
return ctrl.Result{}, err
} else if condErr != nil {
return r.setTeamNotReady(ctx, &team, condErr.reason, condErr.message)
}
meta.SetStatusCondition(&team.Status.Conditions, metav1.Condition{
Type: terdutv1alpha1.ConditionReady,
Status: metav1.ConditionTrue,
Reason: terdutv1alpha1.ReasonTeamAdopted,
Message: fmt.Sprintf("team %d ready, credentials in Secret %q", team.Status.TeamID, team.Status.CredentialsSecretRef.Name),
Message: fmt.Sprintf("team %d applied", team.Status.TeamID),
})
team.Status.ObservedGeneration = team.Generation
if err := r.Status().Update(ctx, &team); err != nil {
@@ -147,13 +133,31 @@ func (r *TerdutTeamReconciler) Reconcile(ctx context.Context, req ctrl.Request)
}
if r.Recorder != nil {
r.Recorder.Eventf(&team, nil, corev1.EventTypeNormal, terdutv1alpha1.ReasonTeamAdopted, terdutv1alpha1.ReasonTeamAdopted,
"team ready")
"team applied")
}
log.Info("TerdutTeam ready", "name", team.Name, "teamID", team.Status.TeamID)
return ctrl.Result{RequeueAfter: resyncInterval}, nil
}
func nameTakenMessage(team *terdutv1alpha1.TerdutTeam) string {
return fmt.Sprintf("the server already has a different team named %q; choose another spec.displayName, "+
"or remove that team", team.Spec.DisplayName)
}
// serverClient builds a client for srv, authenticated with its operator key.
func (r *TerdutTeamReconciler) serverClient(ctx context.Context, srv *terdutv1alpha1.TerdutServer) (*tdclient.Client, error) {
key, err := readOperatorKey(ctx, r.Client, srv)
if err != nil {
return nil, err
}
newClient := r.NewClient
if newClient == nil {
newClient = tdclient.New
}
return newClient(serviceURL(srv)).WithToken(key), nil
}
// teamError carries a condition reason/message, the same role databaseError
// plays for TerdutServer: an expected, requeue-and-retry outcome, not a
// reconcile failure.
@@ -168,6 +172,21 @@ func (e *teamError) Error() string { return e.message }
// consent check (DESIGN.md §4.6) when serverRef.namespace differs from this
// TerdutTeam's own.
func (r *TerdutTeamReconciler) resolveServerRef(ctx context.Context, team *terdutv1alpha1.TerdutTeam) (*terdutv1alpha1.TerdutServer, *teamError) {
srv, ns, getErr := r.getServerRef(ctx, team)
if getErr != nil {
return nil, getErr
}
if ns == team.Namespace {
return srv, nil
}
return r.checkConsent(ctx, team, srv, ns)
}
// getServerRef fetches the referenced TerdutServer and the namespace it was
// looked up in, without the consent check. Deleting a team uses it directly:
// consent gates what the operator will start acting on, not whether it may
// clean up after itself once the server owner has narrowed allowedTeams.
func (r *TerdutTeamReconciler) getServerRef(ctx context.Context, team *terdutv1alpha1.TerdutTeam) (*terdutv1alpha1.TerdutServer, string, *teamError) {
ns := team.Spec.ServerRef.Namespace
if ns == "" {
ns = team.Namespace
@@ -176,18 +195,19 @@ func (r *TerdutTeamReconciler) resolveServerRef(ctx context.Context, team *terdu
var srv terdutv1alpha1.TerdutServer
if err := r.Get(ctx, client.ObjectKey{Namespace: ns, Name: team.Spec.ServerRef.Name}, &srv); err != nil {
if apierrors.IsNotFound(err) {
return nil, &teamError{
return nil, ns, &teamError{
reason: terdutv1alpha1.ReasonServerRefNotFound,
message: fmt.Sprintf("TerdutServer %q not found in namespace %q", team.Spec.ServerRef.Name, ns),
}
}
return nil, &teamError{reason: terdutv1alpha1.ReasonServerRefNotFound, message: err.Error()}
}
if ns == team.Namespace {
return &srv, nil
return nil, ns, &teamError{reason: terdutv1alpha1.ReasonServerRefNotFound, message: err.Error()}
}
return &srv, ns, nil
}
// checkConsent applies DESIGN.md §4.6's allowedTeams gate to a cross-namespace
// reference.
func (r *TerdutTeamReconciler) checkConsent(ctx context.Context, team *terdutv1alpha1.TerdutTeam, srv *terdutv1alpha1.TerdutServer, ns string) (*terdutv1alpha1.TerdutServer, *teamError) {
var ownNamespace corev1.Namespace
if err := r.Get(ctx, client.ObjectKey{Name: team.Namespace}, &ownNamespace); err != nil {
return nil, &teamError{
@@ -206,11 +226,11 @@ func (r *TerdutTeamReconciler) resolveServerRef(ctx context.Context, team *terdu
ns, team.Spec.ServerRef.Name, team.Namespace),
}
}
return &srv, nil
return srv, nil
}
func (r *TerdutTeamReconciler) setTeamNotReady(
ctx context.Context, team *terdutv1alpha1.TerdutTeam, reason, message string, d time.Duration,
ctx context.Context, team *terdutv1alpha1.TerdutTeam, reason, message string,
) (ctrl.Result, error) {
meta.SetStatusCondition(&team.Status.Conditions, metav1.Condition{
Type: terdutv1alpha1.ConditionReady,
@@ -225,42 +245,18 @@ func (r *TerdutTeamReconciler) setTeamNotReady(
if r.Recorder != nil {
r.Recorder.Eventf(team, nil, corev1.EventTypeWarning, reason, reason, message)
}
return ctrl.Result{RequeueAfter: d}, nil
return ctrl.Result{RequeueAfter: waitInterval}, nil
}
// teamCredentialsSecretName/teamServiceAccountName follow the same
// <namespace>.<name>-suffix convention TerdutServer's secrets use
// (DESIGN.md §6 point 3), keyed on the TerdutTeam CR's own identity rather
// than its (mutable) displayName -- a service account's own name is
// globally unique across the whole install (internal/db/migrations/
// 014_service_accounts.sql's UNIQUE constraint, confirmed against source),
// so this has to be collision-safe the same way the credentials Secret
// names already are.
func teamCredentialsSecretName(team *terdutv1alpha1.TerdutTeam) string {
return fmt.Sprintf("%s.%s-team-credentials", team.Namespace, team.Name)
}
func teamServiceAccountName(team *terdutv1alpha1.TerdutTeam) string {
return fmt.Sprintf("terdut-team.%s.%s", team.Namespace, team.Name)
}
// reconcileTeamDelete cleans up the team-scoped credentials Secret and, if
// one was ever minted, deletes the team server-side first. If
// status.teamID or status.credentialsSecretRef was never set (the CR was
// deleted before reconciliation ever got that far), there is deliberately
// no attempt to clean up server-side: terdut-server's DELETE /api/teams/
// {teamID} is owner-gated (requireTeamOwner), and an instance-scoped
// credential -- the only one this controller would otherwise hold -- does
// not satisfy that check (confirmed against source, same finding as
// TEAM-LOOKUP.md's). A team created but never fully reconciled to Ready is
// left orphaned server-side for a human with real owner/admin access to
// clean up -- a known, documented limitation, not a silent gap.
// reconcileTeamDelete removes the team on the server, then the finalizer. The
// server refuses while the team has open incidents (409), which surfaces as a
// retried error: tidying up must not be how a live page disappears.
func (r *TerdutTeamReconciler) reconcileTeamDelete(ctx context.Context, team *terdutv1alpha1.TerdutTeam) (ctrl.Result, error) {
if !controllerutil.ContainsFinalizer(team, teamFinalizerName) {
return ctrl.Result{}, nil
}
if team.Status.TeamID != 0 && team.Status.CredentialsSecretRef != nil {
if team.Status.TeamID != 0 {
if err := r.deleteTeamServerSide(ctx, team); err != nil {
if r.Recorder != nil {
r.Recorder.Eventf(team, nil, corev1.EventTypeWarning, "DeleteFailed", "DeleteFailed", err.Error())
@@ -269,40 +265,33 @@ func (r *TerdutTeamReconciler) reconcileTeamDelete(ctx context.Context, team *te
}
}
secret := &corev1.Secret{ObjectMeta: metav1.ObjectMeta{Name: teamCredentialsSecretName(team), Namespace: r.OperatorNamespace}}
if err := r.Delete(ctx, secret); err != nil && !apierrors.IsNotFound(err) {
return ctrl.Result{}, err
}
controllerutil.RemoveFinalizer(team, teamFinalizerName)
return ctrl.Result{}, r.Update(ctx, team)
}
func (r *TerdutTeamReconciler) deleteTeamServerSide(ctx context.Context, team *terdutv1alpha1.TerdutTeam) error {
srv, resolveErr := r.resolveServerRef(ctx, team)
if resolveErr != nil {
// The TerdutServer (or the namespace consent for it) is gone too --
// most likely the whole install is being torn down together.
// Nothing to delete against; proceed rather than block forever on
// a parent that no longer exists.
return nil
}
teamKey, err := readOperatorSecret(ctx, r.Client, r.OperatorNamespace, team.Status.CredentialsSecretRef)
if err != nil {
// Consent is deliberately not checked here: it gates what the operator
// starts acting on, not whether it may clean up after itself once the
// server owner has narrowed allowedTeams. Only a TerdutServer that is
// really gone lets the delete proceed (there is nothing to delete against);
// any other failure is retried rather than orphaning the team.
srv, ns, getErr := r.getServerRef(ctx, team)
if getErr != nil {
var probe terdutv1alpha1.TerdutServer
err := r.Get(ctx, client.ObjectKey{Namespace: ns, Name: team.Spec.ServerRef.Name}, &probe)
if apierrors.IsNotFound(err) {
return nil
}
return fmt.Errorf("reading TerdutServer %q/%q: %w", ns, team.Spec.ServerRef.Name, getErr)
}
tc, err := r.serverClient(ctx, srv)
if err != nil {
if apierrors.IsNotFound(err) {
return nil // the key Secret went with its TerdutServer
}
return err
}
newClient := r.NewClient
if newClient == nil {
newClient = tdclient.New
}
err = newClient(serviceURL(srv)).WithToken(teamKey).DeleteTeam(ctx, team.Status.TeamID)
if statusErr, ok := errors.AsType[*tdclient.StatusError](err); ok && statusErr.Code == http.StatusNotFound {
return nil
}
return err
return tc.DeleteTeam(ctx, team.Status.TeamID)
}
// SetupWithManager sets up the controller with the Manager.
@@ -315,6 +304,28 @@ func (r *TerdutTeamReconciler) SetupWithManager(mgr ctrl.Manager) error {
}
return ctrl.NewControllerManagedBy(mgr).
For(&terdutv1alpha1.TerdutTeam{}).
Watches(&terdutv1alpha1.TerdutServer{}, handler.EnqueueRequestsFromMapFunc(r.teamsForServer)).
Named("terdutteam").
Complete(r)
}
// teamsForServer re-reconciles every TerdutTeam that references a TerdutServer
// when it changes (becomes Ready, say), instead of polling for it.
func (r *TerdutTeamReconciler) teamsForServer(ctx context.Context, obj client.Object) []reconcile.Request {
var teams terdutv1alpha1.TerdutTeamList
if err := r.List(ctx, &teams); err != nil {
return nil
}
var out []reconcile.Request
for i := range teams.Items {
t := &teams.Items[i]
ns := t.Spec.ServerRef.Namespace
if ns == "" {
ns = t.Namespace
}
if t.Spec.ServerRef.Name == obj.GetName() && ns == obj.GetNamespace() {
out = append(out, reconcile.Request{NamespacedName: client.ObjectKeyFromObject(t)})
}
}
return out
}
+294 -143
View File
@@ -2,9 +2,7 @@ package controller
import (
"context"
"fmt"
"net/http/httptest"
"time"
. "github.com/onsi/ginkgo/v2"
. "github.com/onsi/gomega"
@@ -18,6 +16,12 @@ import (
"git.ryuvia.com/niklas/terdut-operator/internal/tdclient"
)
// platformName is the display name most specs give their team.
const platformName = "platform"
// allNamespaces is allowedTeams.namespaces.from admitting every namespace.
const allNamespaces = "All"
var _ = Describe("TerdutTeam Controller", func() {
const operatorNamespace = "default"
@@ -34,13 +38,12 @@ var _ = Describe("TerdutTeam Controller", func() {
fake, fakeSrv = newFakeTerdutServer()
DeferCleanup(fakeSrv.Close)
srv = bootstrapReadyTerdutServer(ctx, uniqueName("ttserver"), fakeSrv.URL)
srv = readyTerdutServer(ctx, uniqueName("ttserver"))
reconciler = &TerdutTeamReconciler{
Client: k8sClient,
Scheme: k8sClient.Scheme(),
OperatorNamespace: operatorNamespace,
NewClient: func(string) *tdclient.Client { return tdclient.New(fakeSrv.URL) },
Client: k8sClient,
Scheme: k8sClient.Scheme(),
NewClient: func(string) *tdclient.Client { return tdclient.New(fakeSrv.URL) },
}
teamName = uniqueName("team")
teamKey = types.NamespacedName{Name: teamName, Namespace: operatorNamespace}
@@ -53,13 +56,7 @@ var _ = Describe("TerdutTeam Controller", func() {
_ = k8sClient.Update(ctx, team)
_ = k8sClient.Delete(ctx, team)
}
_ = k8sClient.Delete(ctx, &corev1.Secret{ObjectMeta: metav1.ObjectMeta{
Name: fmt.Sprintf("%s.%s-team-credentials", operatorNamespace, teamName), Namespace: operatorNamespace,
}})
// The TerdutServer bootstrapReadyTerdutServer created in BeforeEach
// carries its own finalizer; clear it directly the same way, rather
// than relying on TerdutServerReconciler to ever run again here.
srvKey := types.NamespacedName{Name: srv.Name, Namespace: operatorNamespace}
if err := k8sClient.Get(ctx, srvKey, srv); err == nil {
srv.Finalizers = nil
@@ -67,7 +64,7 @@ var _ = Describe("TerdutTeam Controller", func() {
_ = k8sClient.Delete(ctx, srv)
}
_ = k8sClient.Delete(ctx, &corev1.Secret{ObjectMeta: metav1.ObjectMeta{
Name: fmt.Sprintf("%s.%s-instance-credentials", operatorNamespace, srv.Name), Namespace: operatorNamespace,
Name: srv.Name + "-operator-key", Namespace: operatorNamespace,
}})
})
@@ -96,39 +93,56 @@ var _ = Describe("TerdutTeam Controller", func() {
return terdutv1alpha1.TerdutServerRef{Name: srv.Name}
}
getTeam := func(ctx context.Context) *terdutv1alpha1.TerdutTeam {
team := &terdutv1alpha1.TerdutTeam{}
Expect(k8sClient.Get(ctx, teamKey, team)).To(Succeed())
return team
}
// reconcileToReady drives a freshly created team through the finalizer pass
// and one apply pass.
reconcileToReady := func(ctx context.Context) {
reconcileOnce(ctx) // finalizer
reconcileOnce(ctx) // apply
}
Describe("the happy path", func() {
It("creates the team, mints its credential, and applies rename/oidc-groups", func(ctx SpecContext) {
createTeam(ctx, "platform", sameNSRef())
reconcileOnce(ctx) // finalizer
reconcileOnce(ctx) // full create+mint+apply
It("creates the team under its CR identity with the operator key, and applies oidc groups", func(ctx SpecContext) {
// The fake only answers a request carrying the server's operator key.
var keySecret corev1.Secret
Expect(k8sClient.Get(ctx, types.NamespacedName{Name: srv.Name + "-operator-key", Namespace: operatorNamespace}, &keySecret)).To(Succeed())
fake.expectKey = string(keySecret.Data[operatorKeyDataKey])
team := &terdutv1alpha1.TerdutTeam{
ObjectMeta: metav1.ObjectMeta{Name: teamName, Namespace: operatorNamespace},
Spec: terdutv1alpha1.TerdutTeamSpec{
ServerRef: sameNSRef(), DisplayName: platformName,
OIDC: terdutv1alpha1.TerdutTeamOIDC{MemberGroup: "members", OwnerGroup: "owners"},
},
}
Expect(k8sClient.Create(ctx, team)).To(Succeed())
reconcileToReady(ctx)
cond := readyCondition(ctx)
Expect(cond.Status).To(Equal(metav1.ConditionTrue))
Expect(cond.Reason).To(Equal(terdutv1alpha1.ReasonTeamAdopted))
team := &terdutv1alpha1.TerdutTeam{}
Expect(k8sClient.Get(ctx, teamKey, team)).To(Succeed())
Expect(team.Status.TeamID).To(Equal(fake.teams["platform"]))
Expect(team.Status.CredentialsSecretRef).NotTo(BeNil())
var credsSecret corev1.Secret
Expect(k8sClient.Get(ctx, types.NamespacedName{
Name: team.Status.CredentialsSecretRef.Name, Namespace: operatorNamespace,
}, &credsSecret)).To(Succeed())
Expect(credsSecret.Data[credentialsSecretDataKey]).NotTo(BeEmpty())
team = getTeam(ctx)
Expect(team.Status.TeamID).To(Equal(fake.teams[platformName]))
Expect(fake.teamExt).To(HaveKeyWithValue(operatorNamespace+"/"+teamName, team.Status.TeamID))
Expect(fake.teamOIDC[team.Status.TeamID]).To(Equal([2]string{"members", "owners"}))
})
})
Describe("waiting on the referenced TerdutServer", func() {
It("reports ServerRefNotFound when the TerdutServer doesn't exist", func(ctx SpecContext) {
createTeam(ctx, "orphan", terdutv1alpha1.TerdutServerRef{Name: testRefNotFoundName})
reconcileOnce(ctx) // finalizer
reconcileOnce(ctx)
reconcileToReady(ctx)
Expect(readyCondition(ctx).Reason).To(Equal(terdutv1alpha1.ReasonServerRefNotFound))
})
It("reports WaitingForServer when the TerdutServer exists but isn't Bootstrapped yet", func(ctx SpecContext) {
It("reports WaitingForServer when the TerdutServer exists but isn't Ready yet", func(ctx SpecContext) {
unreadyName := uniqueName("ttserver-unready")
unready := &terdutv1alpha1.TerdutServer{
ObjectMeta: metav1.ObjectMeta{Name: unreadyName, Namespace: operatorNamespace},
@@ -142,8 +156,7 @@ var _ = Describe("TerdutTeam Controller", func() {
DeferCleanup(func() { _ = k8sClient.Delete(ctx, unready) })
createTeam(ctx, "waiting", terdutv1alpha1.TerdutServerRef{Name: unreadyName})
reconcileOnce(ctx) // finalizer
reconcileOnce(ctx)
reconcileToReady(ctx)
Expect(readyCondition(ctx).Reason).To(Equal(terdutv1alpha1.ReasonWaitingForServer))
})
@@ -182,7 +195,7 @@ var _ = Describe("TerdutTeam Controller", func() {
It("permits when the TerdutServer's allowedTeams.namespaces.from is All", func(ctx SpecContext) {
Expect(k8sClient.Get(ctx, types.NamespacedName{Name: srv.Name, Namespace: operatorNamespace}, srv)).To(Succeed())
srv.Spec.AllowedTeams.Namespaces.From = "All"
srv.Spec.AllowedTeams.Namespaces.From = allNamespaces
Expect(k8sClient.Update(ctx, srv)).To(Succeed())
team := &terdutv1alpha1.TerdutTeam{
@@ -209,154 +222,292 @@ var _ = Describe("TerdutTeam Controller", func() {
})
})
Describe("adopt-on-409 recovery", func() {
It("adopts an already-created team instead of erroring", func(ctx SpecContext) {
Describe("team identity", func() {
It("refuses a name that belongs to a team this CR did not create", func(ctx SpecContext) {
fake.nextTeamID = 1
fake.teams["platform"] = 1
fake.teamNames[1] = "platform"
fake.teams[platformName] = 1
fake.teamNames[1] = platformName
createTeam(ctx, "platform", sameNSRef())
reconcileOnce(ctx) // finalizer
reconcileOnce(ctx)
createTeam(ctx, platformName, sameNSRef())
reconcileToReady(ctx)
team := &terdutv1alpha1.TerdutTeam{}
Expect(k8sClient.Get(ctx, teamKey, team)).To(Succeed())
Expect(team.Status.TeamID).To(Equal(int64(1)))
Expect(readyCondition(ctx).Status).To(Equal(metav1.ConditionTrue))
cond := readyCondition(ctx)
Expect(cond.Status).To(Equal(metav1.ConditionFalse))
Expect(cond.Reason).To(Equal(terdutv1alpha1.ReasonTeamNameTaken))
Expect(getTeam(ctx).Status.TeamID).To(BeZero())
Expect(fake.teamOIDC).To(BeEmpty(), "nothing is configured on a team that is not ours")
})
It("adopts an already-created team-scoped service account instead of erroring", func(ctx SpecContext) {
createTeam(ctx, "platform", sameNSRef())
reconcileOnce(ctx) // finalizer
It("keeps the same team when spec.displayName changes", func(ctx SpecContext) {
createTeam(ctx, platformName, sameNSRef())
reconcileToReady(ctx)
id := getTeam(ctx).Status.TeamID
// Pre-seed the service account the mint step is about to try to
// create, simulating an attempt that got this far before being
// interrupted.
fake.nextID = 1
saName := fmt.Sprintf("terdut-team.%s.%s", operatorNamespace, teamName)
fake.accounts[saName] = 1
fake.keyMints[1] = 1
team := getTeam(ctx)
team.Spec.DisplayName = "platform-renamed"
Expect(k8sClient.Update(ctx, team)).To(Succeed())
reconcileOnce(ctx)
Expect(getTeam(ctx).Status.TeamID).To(Equal(id))
Expect(fake.teamNames[id]).To(Equal("platform-renamed"))
Expect(fake.nextTeamID).To(Equal(id), "renamed in place, not recreated")
})
It("reports TeamNameTaken when a rename collides with another team", func(ctx SpecContext) {
createTeam(ctx, platformName, sameNSRef())
reconcileToReady(ctx)
fake.nextTeamID++
fake.teams["other"] = fake.nextTeamID
fake.teamNames[fake.nextTeamID] = "other"
team := getTeam(ctx)
team.Spec.DisplayName = "other"
Expect(k8sClient.Update(ctx, team)).To(Succeed())
reconcileOnce(ctx)
Expect(readyCondition(ctx).Reason).To(Equal(terdutv1alpha1.ReasonTeamNameTaken))
})
It("finds its team again when the status was lost", func(ctx SpecContext) {
createTeam(ctx, platformName, sameNSRef())
reconcileToReady(ctx)
id := getTeam(ctx).Status.TeamID
team := getTeam(ctx)
team.Status.TeamID = 0
Expect(k8sClient.Status().Update(ctx, team)).To(Succeed())
reconcileOnce(ctx)
Expect(getTeam(ctx).Status.TeamID).To(Equal(id))
Expect(fake.nextTeamID).To(Equal(id), "no second team was created")
})
It("recreates a team deleted behind its back", func(ctx SpecContext) {
createTeam(ctx, platformName, sameNSRef())
reconcileToReady(ctx)
id := getTeam(ctx).Status.TeamID
delete(fake.teams, platformName)
delete(fake.teamNames, id)
delete(fake.teamExt, operatorNamespace+"/"+teamName)
reconcileOnce(ctx)
Expect(getTeam(ctx).Status.TeamID).NotTo(Equal(id))
Expect(readyCondition(ctx).Status).To(Equal(metav1.ConditionTrue))
team := &terdutv1alpha1.TerdutTeam{}
Expect(k8sClient.Get(ctx, teamKey, team)).To(Succeed())
var credsSecret corev1.Secret
Expect(k8sClient.Get(ctx, types.NamespacedName{
Name: team.Status.CredentialsSecretRef.Name, Namespace: operatorNamespace,
}, &credsSecret)).To(Succeed())
// Minted fresh (mint2), not the pre-seeded account's original
// (never-issued-to-this-reconcile) key.
Expect(string(credsSecret.Data[credentialsSecretDataKey])).To(Equal("instance-key-1-mint2"))
})
})
Describe("spec.invite", func() {
It("mints a link into the TerdutTeam's own namespace, not the operator's", func(ctx SpecContext) {
createTeam(ctx, "platform", sameNSRef())
reconcileOnce(ctx) // finalizer
reconcileOnce(ctx) // create+mint+apply
Describe("spec.escalation", func() {
ladder := func() *terdutv1alpha1.EscalationSpec {
return &terdutv1alpha1.EscalationSpec{
RepeatCount: 2, FallbackTopic: "oncall",
Levels: []terdutv1alpha1.EscalationLevel{
{Timeout: "5m", Targets: []terdutv1alpha1.EscalationTarget{{Kind: terdutv1alpha1.EscalationTargetOncall}}},
{Timeout: "15m", Targets: []terdutv1alpha1.EscalationTarget{{Kind: terdutv1alpha1.EscalationTargetUser, Username: "alice"}}},
},
}
}
createWithLadder := func(ctx context.Context, e *terdutv1alpha1.EscalationSpec) {
Expect(k8sClient.Create(ctx, &terdutv1alpha1.TerdutTeam{
ObjectMeta: metav1.ObjectMeta{Name: teamName, Namespace: operatorNamespace},
Spec: terdutv1alpha1.TerdutTeamSpec{ServerRef: sameNSRef(), DisplayName: platformName, Escalation: e},
})).To(Succeed())
}
team := &terdutv1alpha1.TerdutTeam{}
Expect(k8sClient.Get(ctx, teamKey, team)).To(Succeed())
team.Spec.Invite.Enabled = true
Expect(k8sClient.Update(ctx, team)).To(Succeed())
reconcileOnce(ctx)
It("PUTs the ladder with usernames for the server to resolve", func(ctx SpecContext) {
createWithLadder(ctx, ladder())
reconcileToReady(ctx)
Expect(k8sClient.Get(ctx, teamKey, team)).To(Succeed())
Expect(team.Status.InviteSecretRef).NotTo(BeNil())
var secret corev1.Secret
Expect(k8sClient.Get(ctx, types.NamespacedName{
Name: team.Status.InviteSecretRef.Name, Namespace: team.Namespace,
}, &secret)).To(Succeed())
Expect(string(secret.Data[inviteSecretURLKey])).To(ContainSubstring("invite="))
Expect(string(secret.Data[inviteSecretInviteIDKey])).To(Equal("1"))
inv := fake.invites[team.Status.TeamID][1]
Expect(inv.Role).To(Equal("member"), "default role")
Expect(inv.MaxUses).To(Equal(int64(1)), "default max uses")
got := fake.escalation[getTeam(ctx).Status.TeamID]
Expect(got.RepeatCount).To(Equal(int64(2)))
Expect(got.FallbackTopic).To(Equal("oncall"))
Expect(got.Levels).To(HaveLen(2))
Expect(got.Levels[0].TimeoutSeconds).To(Equal(int64(300)))
Expect(got.Levels[1].Targets[0]).To(Equal(tdclient.EscalationTargetRequest{Kind: "user", Username: "alice"}))
})
It("refreshes a link that's within a day of terdut-server's 7-day TTL", func(ctx SpecContext) {
createTeam(ctx, "platform", sameNSRef())
It("waits with UnknownUser while the server does not know a username", func(ctx SpecContext) {
fake.unknownUsers["alice"] = true
createWithLadder(ctx, ladder())
reconcileToReady(ctx)
cond := readyCondition(ctx)
Expect(cond.Reason).To(Equal(terdutv1alpha1.ReasonUnknownUser))
Expect(cond.Message).To(ContainSubstring("alice"))
delete(fake.unknownUsers, "alice")
reconcileOnce(ctx)
reconcileOnce(ctx)
team := &terdutv1alpha1.TerdutTeam{}
Expect(k8sClient.Get(ctx, teamKey, team)).To(Succeed())
team.Spec.Invite.Enabled = true
Expect(k8sClient.Update(ctx, team)).To(Succeed())
reconcileOnce(ctx)
Expect(k8sClient.Get(ctx, teamKey, team)).To(Succeed())
firstSecretName := team.Status.InviteSecretRef.Name
var secret corev1.Secret
Expect(k8sClient.Get(ctx, types.NamespacedName{Name: firstSecretName, Namespace: team.Namespace}, &secret)).To(Succeed())
// Simulate the stored link being within the refresh window of
// expiry, the way it genuinely would be six days from now,
// without the test waiting six days.
secret.Data[inviteSecretExpiresAtKey] = []byte(time.Now().Add(12 * time.Hour).Format(time.RFC3339)) // inside inviteRefreshWindow
Expect(k8sClient.Update(ctx, &secret)).To(Succeed())
reconcileOnce(ctx)
Expect(fake.inviteDelete[1]).To(BeTrue(), "the stale invite should have been revoked")
Expect(fake.invites[team.Status.TeamID]).To(HaveKey(int64(2)), "a replacement should have been minted")
Expect(readyCondition(ctx).Status).To(Equal(metav1.ConditionTrue))
})
It("revokes the invite when spec.invite.enabled flips back to false", func(ctx SpecContext) {
createTeam(ctx, "platform", sameNSRef())
reconcileOnce(ctx)
reconcileOnce(ctx)
It("clears the ladder when spec.escalation is removed", func(ctx SpecContext) {
createWithLadder(ctx, ladder())
reconcileToReady(ctx)
team := &terdutv1alpha1.TerdutTeam{}
Expect(k8sClient.Get(ctx, teamKey, team)).To(Succeed())
team.Spec.Invite.Enabled = true
team := getTeam(ctx)
team.Spec.Escalation = nil
Expect(k8sClient.Update(ctx, team)).To(Succeed())
reconcileOnce(ctx)
Expect(k8sClient.Get(ctx, teamKey, team)).To(Succeed())
secretName := team.Status.InviteSecretRef.Name
team.Spec.Invite.Enabled = false
Expect(fake.escalation[team.Status.TeamID].Levels).To(BeEmpty())
})
It("reports InvalidSpec for a duration that does not parse", func(ctx SpecContext) {
bad := ladder()
bad.Levels[0].Timeout = "5 minutes"
// The CRD pattern rejects it at admission when the API server enforces it;
// envtest does, so write it past the pattern by creating a valid one first.
bad.Levels[0].Timeout = "5m"
createWithLadder(ctx, bad)
reconcileToReady(ctx)
team := getTeam(ctx)
team.Spec.Escalation.Levels[0].Timeout = "0s"
Expect(k8sClient.Update(ctx, team)).To(Succeed())
reconcileOnce(ctx)
Expect(fake.inviteDelete[1]).To(BeTrue())
Expect(k8sClient.Get(ctx, teamKey, team)).To(Succeed())
Expect(team.Status.InviteSecretRef).To(BeNil())
Expect(readyCondition(ctx).Reason).To(Equal(terdutv1alpha1.ReasonInvalidSpec))
})
})
var secret corev1.Secret
err := k8sClient.Get(ctx, types.NamespacedName{Name: secretName, Namespace: team.Namespace}, &secret)
Expect(err).To(HaveOccurred(), "the invite Secret should have been deleted")
Describe("spec.deadmanSwitches", func() {
sw := func(name, matcher, timeout string) terdutv1alpha1.DeadmanSwitchSpec {
return terdutv1alpha1.DeadmanSwitchSpec{Name: name, Matcher: matcher, Timeout: timeout, Severity: "critical"}
}
setSwitches := func(ctx context.Context, switches ...terdutv1alpha1.DeadmanSwitchSpec) {
team := getTeam(ctx)
team.Spec.DeadmanSwitches = switches
Expect(k8sClient.Update(ctx, team)).To(Succeed())
reconcileOnce(ctx)
}
names := func(teamID int64) []string {
out := make([]string, 0, len(fake.switches[teamID]))
for _, s := range fake.switches[teamID] {
out = append(out, s.Name)
}
return out
}
It("creates, updates and prunes switches by name", func(ctx SpecContext) {
createTeam(ctx, platformName, sameNSRef())
reconcileToReady(ctx)
id := getTeam(ctx).Status.TeamID
setSwitches(ctx, sw("watchdog", "alertname=Watchdog", "15m"), sw("edge", "alertname=Edge", "5m"))
Expect(names(id)).To(ConsistOf("watchdog", "edge"))
// Same name, new timeout: updated in place, same id.
var watchdogID int64
for _, s := range fake.switches[id] {
if s.Name == "watchdog" {
watchdogID = s.ID
}
}
setSwitches(ctx, sw("watchdog", "alertname=Watchdog", "30m"))
Expect(names(id)).To(ConsistOf("watchdog"))
Expect(fake.switches[id][watchdogID].TimeoutSeconds).To(Equal(int64(1800)))
Expect(readyCondition(ctx).Status).To(Equal(metav1.ConditionTrue))
})
It("removes a switch somebody added on the server", func(ctx SpecContext) {
createTeam(ctx, platformName, sameNSRef())
reconcileToReady(ctx)
id := getTeam(ctx).Status.TeamID
fake.nextSwitchID++
fake.switches[id] = map[int64]tdclient.DeadmanSwitch{
fake.nextSwitchID: {ID: fake.nextSwitchID, Name: "stray", Matcher: "alertname=X", TimeoutSeconds: 60, Severity: "critical"},
}
reconcileOnce(ctx)
Expect(fake.switches[id]).To(BeEmpty())
})
It("recreates a switch deleted on the server", func(ctx SpecContext) {
createTeam(ctx, platformName, sameNSRef())
reconcileToReady(ctx)
id := getTeam(ctx).Status.TeamID
setSwitches(ctx, sw("watchdog", "alertname=Watchdog", "15m"))
fake.switches[id] = map[int64]tdclient.DeadmanSwitch{}
reconcileOnce(ctx)
Expect(names(id)).To(ConsistOf("watchdog"))
})
It("reports InvalidSpec for a timeout that does not parse", func(ctx SpecContext) {
createTeam(ctx, platformName, sameNSRef())
reconcileToReady(ctx)
setSwitches(ctx, sw("watchdog", "alertname=Watchdog", "0s"))
Expect(readyCondition(ctx).Reason).To(Equal(terdutv1alpha1.ReasonInvalidSpec))
})
})
Describe("deletion", func() {
It("deletes the team server-side and removes the credentials Secret", func(ctx SpecContext) {
It("deletes the team on the server and clears the finalizer", func(ctx SpecContext) {
createTeam(ctx, "to-delete", sameNSRef())
reconcileOnce(ctx)
reconcileOnce(ctx)
Expect(readyCondition(ctx).Status).To(Equal(metav1.ConditionTrue))
reconcileToReady(ctx)
teamID := getTeam(ctx).Status.TeamID
team := &terdutv1alpha1.TerdutTeam{}
Expect(k8sClient.Get(ctx, teamKey, team)).To(Succeed())
teamID := team.Status.TeamID
credsName := team.Status.CredentialsSecretRef.Name
Expect(k8sClient.Delete(ctx, team)).To(Succeed())
Expect(k8sClient.Delete(ctx, getTeam(ctx))).To(Succeed())
reconcileOnce(ctx) // runs the finalizer
Expect(fake.teamDelete[teamID]).To(BeTrue())
Expect(k8sClient.Get(ctx, teamKey, &terdutv1alpha1.TerdutTeam{})).NotTo(Succeed(),
"the TerdutTeam should be gone once the finalizer clears")
})
err := k8sClient.Get(ctx, teamKey, team)
Expect(err).To(HaveOccurred(), "the TerdutTeam itself should be gone once the finalizer clears")
It("still deletes the team after the server's allowedTeams consent is revoked", func(ctx SpecContext) {
otherNS := uniqueName("ns")
Expect(k8sClient.Create(ctx, &corev1.Namespace{ObjectMeta: metav1.ObjectMeta{Name: otherNS}})).To(Succeed())
DeferCleanup(func() { _ = k8sClient.Delete(ctx, &corev1.Namespace{ObjectMeta: metav1.ObjectMeta{Name: otherNS}}) })
var leftover corev1.Secret
err = k8sClient.Get(ctx, types.NamespacedName{Name: credsName, Namespace: operatorNamespace}, &leftover)
Expect(err).To(HaveOccurred(), "the team-credentials Secret should have been cleaned up")
Expect(k8sClient.Get(ctx, types.NamespacedName{Name: srv.Name, Namespace: operatorNamespace}, srv)).To(Succeed())
srv.Spec.AllowedTeams.Namespaces.From = allNamespaces
Expect(k8sClient.Update(ctx, srv)).To(Succeed())
crossKey := types.NamespacedName{Name: teamName, Namespace: otherNS}
Expect(k8sClient.Create(ctx, &terdutv1alpha1.TerdutTeam{
ObjectMeta: metav1.ObjectMeta{Name: teamName, Namespace: otherNS},
Spec: terdutv1alpha1.TerdutTeamSpec{
ServerRef: terdutv1alpha1.TerdutServerRef{Name: srv.Name, Namespace: operatorNamespace},
DisplayName: "revoked",
},
})).To(Succeed())
for range 2 {
_, err := reconciler.Reconcile(ctx, reconcile.Request{NamespacedName: crossKey})
Expect(err).NotTo(HaveOccurred())
}
cross := &terdutv1alpha1.TerdutTeam{}
Expect(k8sClient.Get(ctx, crossKey, cross)).To(Succeed())
teamID := cross.Status.TeamID
Expect(teamID).NotTo(BeZero())
Expect(k8sClient.Get(ctx, types.NamespacedName{Name: srv.Name, Namespace: operatorNamespace}, srv)).To(Succeed())
srv.Spec.AllowedTeams.Namespaces.From = "None"
Expect(k8sClient.Update(ctx, srv)).To(Succeed())
Expect(k8sClient.Delete(ctx, cross)).To(Succeed())
_, err := reconciler.Reconcile(ctx, reconcile.Request{NamespacedName: crossKey})
Expect(err).NotTo(HaveOccurred())
Expect(fake.teamDelete[teamID]).To(BeTrue(), "revoking consent must not orphan the team")
})
It("lets the CR go when its TerdutServer is already gone", func(ctx SpecContext) {
createTeam(ctx, "to-delete", sameNSRef())
reconcileToReady(ctx)
serverKey := types.NamespacedName{Name: srv.Name, Namespace: operatorNamespace}
Expect(k8sClient.Get(ctx, serverKey, srv)).To(Succeed())
srv.Finalizers = nil
Expect(k8sClient.Update(ctx, srv)).To(Succeed())
Expect(k8sClient.Delete(ctx, srv)).To(Succeed())
Expect(k8sClient.Delete(ctx, getTeam(ctx))).To(Succeed())
reconcileOnce(ctx)
Expect(k8sClient.Get(ctx, teamKey, &terdutv1alpha1.TerdutTeam{})).NotTo(Succeed())
})
})
})
-149
View File
@@ -1,149 +0,0 @@
package controller
import (
"context"
"fmt"
"strconv"
"time"
corev1 "k8s.io/api/core/v1"
apierrors "k8s.io/apimachinery/pkg/api/errors"
metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
"sigs.k8s.io/controller-runtime/pkg/client"
"sigs.k8s.io/controller-runtime/pkg/controller/controllerutil"
terdutv1alpha1 "git.ryuvia.com/niklas/terdut-operator/api/v1alpha1"
"git.ryuvia.com/niklas/terdut-operator/internal/tdclient"
)
// Data keys inside the generated invite Secret, following the same naming
// shape as TerdutAlertSource's webhookSecret*Key constants.
const (
inviteSecretURLKey = "url"
inviteSecretInviteIDKey = "inviteID"
inviteSecretExpiresAtKey = "expiresAt"
)
// inviteRefreshWindow is how far ahead of expiry this controller mints a
// replacement link, so a human reading status.inviteSecretRef never finds a
// dead link mid-use. terdut-server's invite TTL is a fixed, unconfigurable
// 7 days (internal/api/signup.go's inviteTTL) -- refreshing a full day
// ahead of that leaves comfortable margin against this controller's own
// 5-minute resync interval ever being delayed.
const inviteRefreshWindow = 24 * time.Hour
func inviteSecretName(team *terdutv1alpha1.TerdutTeam) string {
return team.Name + "-terdut-invite"
}
// reconcileInvite applies spec.invite against teamClient -- this team's own
// team-scoped credential, already owner-equivalent for every /invites route
// (terdut-server's SERVICE-ACCOUNTS.md, ratified not accidental). Mints,
// refreshes ahead of expiry, or revokes, entirely independent of this
// team's own Ready condition: an invite is a convenience for onboarding a
// human, never something anything else in this reconcile waits on.
func (r *TerdutTeamReconciler) reconcileInvite(ctx context.Context, team *terdutv1alpha1.TerdutTeam, teamClient *tdclient.Client) error {
if !team.Spec.Invite.Enabled {
return r.revokeInvite(ctx, team, teamClient)
}
secretName := inviteSecretName(team)
var secret corev1.Secret
err := r.Get(ctx, client.ObjectKey{Namespace: team.Namespace, Name: secretName}, &secret)
switch {
case err == nil:
expiresAt, parseErr := time.Parse(time.RFC3339, string(secret.Data[inviteSecretExpiresAtKey]))
if parseErr == nil && time.Until(expiresAt) > inviteRefreshWindow {
team.Status.InviteSecretRef = &terdutv1alpha1.LocalSecretRef{Name: secretName}
return nil // still fresh, nothing to do this reconcile
}
// Expired, about to expire, or unreadable: mint a replacement.
// Revoke the old row by id first (best-effort) so a leaked old link
// stops working immediately rather than lingering unrevoked until
// its own TTL -- failure here is not fatal, since the replacement
// below is what actually matters.
if oldID, idErr := strconv.ParseInt(string(secret.Data[inviteSecretInviteIDKey]), 10, 64); idErr == nil {
_ = teamClient.RevokeInvite(ctx, team.Status.TeamID, oldID)
}
return r.mintInvite(ctx, team, teamClient, secretName)
case apierrors.IsNotFound(err):
// Low stakes, unlike TerdutAlertSource's webhook URL: nothing
// external holds a durable dependency on one specific invite link
// staying stable the way an Alertmanager config depends on a
// webhook URL -- it's read once by one human and handed out. So
// this silently re-mints rather than failing closed the way
// TerdutAlertSource's ReasonWebhookSecretLost does for its Secret.
return r.mintInvite(ctx, team, teamClient, secretName)
default:
return err
}
}
func (r *TerdutTeamReconciler) mintInvite(
ctx context.Context, team *terdutv1alpha1.TerdutTeam, teamClient *tdclient.Client, secretName string,
) error {
role := team.Spec.Invite.Role
if role == "" {
role = "member"
}
maxUses := team.Spec.Invite.MaxUses
if maxUses == 0 {
maxUses = 1
}
inv, err := teamClient.CreateInvite(ctx, team.Status.TeamID, role, maxUses)
if err != nil {
return fmt.Errorf("POST /api/teams/%d/invites: %w", team.Status.TeamID, err)
}
secret := &corev1.Secret{ObjectMeta: metav1.ObjectMeta{Name: secretName, Namespace: team.Namespace}}
if _, err := controllerutil.CreateOrUpdate(ctx, r.Client, secret, func() error {
secret.Data = map[string][]byte{
inviteSecretURLKey: []byte(inv.URL),
inviteSecretInviteIDKey: []byte(strconv.FormatInt(inv.ID, 10)),
inviteSecretExpiresAtKey: []byte(inv.ExpiresAt.Format(time.RFC3339)),
}
return controllerutil.SetControllerReference(team, secret, r.Scheme)
}); err != nil {
return err
}
team.Status.InviteSecretRef = &terdutv1alpha1.LocalSecretRef{Name: secretName}
if r.Recorder != nil {
r.Recorder.Eventf(team, nil, corev1.EventTypeNormal, terdutv1alpha1.ReasonInviteMinted, terdutv1alpha1.ReasonInviteMinted,
"invite link minted into Secret %q", secretName)
}
return nil
}
// revokeInvite tears down spec.invite's Secret and server-side row when
// spec.invite.enabled is false (or was never set). Same-namespace and
// OwnerReference'd, so deleting the TerdutTeam itself already garbage-
// collects this Secret -- this path exists for the narrower case of
// flipping enabled back to false on an otherwise-live TerdutTeam.
func (r *TerdutTeamReconciler) revokeInvite(ctx context.Context, team *terdutv1alpha1.TerdutTeam, teamClient *tdclient.Client) error {
if team.Status.InviteSecretRef == nil {
return nil
}
name := team.Status.InviteSecretRef.Name
var secret corev1.Secret
err := r.Get(ctx, client.ObjectKey{Namespace: team.Namespace, Name: name}, &secret)
switch {
case err == nil:
if id, idErr := strconv.ParseInt(string(secret.Data[inviteSecretInviteIDKey]), 10, 64); idErr == nil {
if err := teamClient.RevokeInvite(ctx, team.Status.TeamID, id); err != nil {
return fmt.Errorf("DELETE /api/teams/%d/invites/%d: %w", team.Status.TeamID, id, err)
}
}
if err := r.Delete(ctx, &secret); err != nil && !apierrors.IsNotFound(err) {
return err
}
case !apierrors.IsNotFound(err):
return err
}
team.Status.InviteSecretRef = nil
if r.Recorder != nil {
r.Recorder.Eventf(team, nil, corev1.EventTypeNormal, terdutv1alpha1.ReasonInviteRevoked, terdutv1alpha1.ReasonInviteRevoked,
"invite link revoked")
}
return nil
}
+18 -27
View File
@@ -8,6 +8,7 @@ import (
. "github.com/onsi/ginkgo/v2"
. "github.com/onsi/gomega"
appsv1 "k8s.io/api/apps/v1"
"k8s.io/apimachinery/pkg/api/meta"
metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
"k8s.io/apimachinery/pkg/types"
"sigs.k8s.io/controller-runtime/pkg/reconcile"
@@ -47,21 +48,15 @@ const (
testOperatorNamespace = "default"
)
// bootstrapReadyTerdutServer creates a TerdutServer with a bring-your-own
// DSN and drives it to Ready against fakeURL, the same three-pass sequence
// terdutserver_controller_test.go's own happy-path test exercises directly
// -- shared here so TerdutTeam's tests (which need a real, Ready
// TerdutServer to resolve against) don't duplicate it.
func bootstrapReadyTerdutServer(ctx context.Context, name, fakeURL string) *terdutv1alpha1.TerdutServer {
// readyTerdutServer creates a TerdutServer with a bring-your-own DSN and drives
// it to Ready the way the happy-path test does -- shared so the tests of
// everything that references a server (TerdutTeam, TerdutAlertSource) start
// from a real, Ready one.
func readyTerdutServer(ctx context.Context, name string) *terdutv1alpha1.TerdutServer {
GinkgoHelper()
namespace := testOperatorNamespace
reconciler := &TerdutServerReconciler{
Client: k8sClient,
Scheme: k8sClient.Scheme(),
OperatorNamespace: namespace,
NewClient: func(string) *tdclient.Client { return tdclient.New(fakeURL) },
}
reconciler := &TerdutServerReconciler{Client: k8sClient, Scheme: k8sClient.Scheme()}
objKey := types.NamespacedName{Name: name, Namespace: namespace}
srv := &terdutv1alpha1.TerdutServer{
@@ -74,9 +69,7 @@ func bootstrapReadyTerdutServer(ctx context.Context, name, fakeURL string) *terd
}
Expect(k8sClient.Create(ctx, srv)).To(Succeed())
_, err := reconciler.Reconcile(ctx, reconcile.Request{NamespacedName: objKey}) // finalizer
Expect(err).NotTo(HaveOccurred())
_, err = reconciler.Reconcile(ctx, reconcile.Request{NamespacedName: objKey}) // Deployment/Service
_, err := reconciler.Reconcile(ctx, reconcile.Request{NamespacedName: objKey}) // Deployment, Service, key
Expect(err).NotTo(HaveOccurred())
var deploy appsv1.Deployment
@@ -85,27 +78,25 @@ func bootstrapReadyTerdutServer(ctx context.Context, name, fakeURL string) *terd
deploy.Status.Replicas = 1
Expect(k8sClient.Status().Update(ctx, &deploy)).To(Succeed())
_, err = reconciler.Reconcile(ctx, reconcile.Request{NamespacedName: objKey}) // bootstrap
_, err = reconciler.Reconcile(ctx, reconcile.Request{NamespacedName: objKey}) // Ready
Expect(err).NotTo(HaveOccurred())
Expect(k8sClient.Get(ctx, objKey, srv)).To(Succeed())
Expect(srv.Status.CredentialsSecretRef).NotTo(BeNil(), "test setup: TerdutServer %s/%s did not reach Bootstrapped", namespace, name)
Expect(meta.IsStatusConditionTrue(srv.Status.Conditions, terdutv1alpha1.ConditionReady)).To(BeTrue(),
"test setup: TerdutServer %s/%s did not reach Ready", namespace, name)
return srv
}
// readyTerdutTeam creates a TerdutTeam under srv (an already-Ready
// TerdutServer, e.g. from bootstrapReadyTerdutServer) and drives it to
// Ready against fakeURL -- shared by TerdutEscalationRule's and
// TerdutDeadmanSwitch's own tests, which both just need a resolvable
// teamRef (DESIGN.md §5), not TerdutTeam's own behavior.
// TerdutServer, from readyTerdutServer) and drives it to Ready against the fake
// at fakeURL.
func readyTerdutTeam(ctx context.Context, namespace, name string, srv *terdutv1alpha1.TerdutServer, fakeURL string) *terdutv1alpha1.TerdutTeam {
GinkgoHelper()
reconciler := &TerdutTeamReconciler{
Client: k8sClient,
Scheme: k8sClient.Scheme(),
OperatorNamespace: namespace,
NewClient: func(string) *tdclient.Client { return tdclient.New(fakeURL) },
Client: k8sClient,
Scheme: k8sClient.Scheme(),
NewClient: func(string) *tdclient.Client { return tdclient.New(fakeURL) },
}
objKey := types.NamespacedName{Name: name, Namespace: namespace}
@@ -120,11 +111,11 @@ func readyTerdutTeam(ctx context.Context, namespace, name string, srv *terdutv1a
_, err := reconciler.Reconcile(ctx, reconcile.Request{NamespacedName: objKey}) // finalizer
Expect(err).NotTo(HaveOccurred())
_, err = reconciler.Reconcile(ctx, reconcile.Request{NamespacedName: objKey}) // create+mint+apply
_, err = reconciler.Reconcile(ctx, reconcile.Request{NamespacedName: objKey}) // apply
Expect(err).NotTo(HaveOccurred())
Expect(k8sClient.Get(ctx, objKey, team)).To(Succeed())
Expect(team.Status.CredentialsSecretRef).NotTo(BeNil(), "test setup: TerdutTeam %s/%s did not reach Ready", namespace, name)
Expect(team.Status.TeamID).NotTo(BeZero(), "test setup: TerdutTeam %s/%s did not reach Ready", namespace, name)
return team
}
+28 -277
View File
@@ -10,6 +10,7 @@ package tdclient
import (
"context"
"encoding/json"
"errors"
"fmt"
"net/http"
"strings"
@@ -57,6 +58,16 @@ type StatusError struct {
Message string
}
// ignoreNotFound turns a 404 into success. Every DELETE here is idempotent: the
// thing already being gone is the outcome the caller wanted, and a finalizer
// that failed on it could never be removed.
func ignoreNotFound(err error) error {
if se, ok := errors.AsType[*StatusError](err); ok && se.Code == http.StatusNotFound {
return nil
}
return err
}
func (e *StatusError) Error() string {
if e.Message != "" {
return fmt.Sprintf("server returned %d: %s", e.Code, e.Message)
@@ -116,187 +127,23 @@ func (c *Client) do(req *http.Request, out any) error {
return nil
}
// Version calls GET /api/version — unauthenticated, per terdut-server's own
// router.go comment ("a client deciding whether it can talk to this server —
// terdut-tui, terdut-operator — needs to ask before it holds a credential
// for it"). Used here purely as a reachability probe: a bad endpoint fails
// here, clearly, rather than on whatever the controller tries first.
func (c *Client) Version(ctx context.Context) (string, error) {
req, err := c.newRequest(ctx, http.MethodGet, "/api/version", nil)
if err != nil {
return "", err
}
var v struct {
Version string `json:"version"`
}
if err := c.do(req, &v); err != nil {
return "", err
}
return v.Version, nil
}
// APIKey is the raw key a bootstrap or service-account-key mint hands back —
// the one moment its value exists outside the request that generated it.
// Mirrors terdut-server's models.APIKey/models.ServiceAccountKey shape
// (internal/models in that repo) for the fields this client actually reads.
type APIKey struct {
ID int64 `json:"id"`
Name string `json:"name"`
Key string `json:"key"`
CreatedAt string `json:"created_at"`
}
// BootstrapResult is /api/bootstrap's 201 response body.
type BootstrapResult struct {
User struct {
ID int64 `json:"id"`
Username string `json:"username"`
Email string `json:"email"`
} `json:"user"`
APIKey APIKey `json:"api_key"`
}
// Bootstrap calls POST /api/bootstrap — unauthenticated, single-shot per
// install (internal/api/users.go's handleBootstrap in terdut-server:
// gated on SELECT COUNT(*) FROM users). Returns the raw admin key directly;
// DESIGN.md §6 has the controller use it for exactly one further call
// (CreateServiceAccount) and discard it, never storing it as the lasting
// credential.
//
// A StatusError with Code 403 means this install already has a user —
// per this operator's design (DESIGN.md §1), that only happens if this
// exact TerdutServer's own controller already won this race on an earlier,
// interrupted reconcile; see the checkpoint-Secret handling in the
// controller, not a retry loop here.
func (c *Client) Bootstrap(ctx context.Context, username, email string) (*BootstrapResult, error) {
req, err := c.newRequest(ctx, http.MethodPost, "/api/bootstrap", map[string]string{
"username": username,
"email": email,
})
if err != nil {
return nil, err
}
var result BootstrapResult
if err := c.do(req, &result); err != nil {
return nil, err
}
return &result, nil
}
// ServiceAccount mirrors terdut-server's models.ServiceAccount
// (internal/models/service_account.go), minus fields this client never
// reads.
type ServiceAccount struct {
ID int64 `json:"id"`
Name string `json:"name"`
Scope string `json:"scope"`
}
// CreateServiceAccountResult is POST /api/service-accounts' 201 response.
type CreateServiceAccountResult struct {
ServiceAccount ServiceAccount `json:"service_account"`
Key APIKey `json:"key"`
}
// CreateInstanceServiceAccount calls POST /api/service-accounts with
// scope "instance", authenticated with c's current token (the raw admin key
// from Bootstrap, for the operator's own first-ever call). A StatusError
// with Code 409 means a prior, interrupted attempt already created this
// name — DESIGN.md §6's adopt-rather-than-error rule: the caller should
// fall back to GetServiceAccountByName + CreateServiceAccountKey, not treat
// this as a hard failure.
func (c *Client) CreateInstanceServiceAccount(ctx context.Context, name string) (*CreateServiceAccountResult, error) {
req, err := c.newRequest(ctx, http.MethodPost, "/api/service-accounts", map[string]string{
fieldName: name,
"scope": "instance",
})
if err != nil {
return nil, err
}
var result CreateServiceAccountResult
if err := c.do(req, &result); err != nil {
return nil, err
}
return &result, nil
}
// CreateTeamServiceAccount calls POST /api/service-accounts with scope
// "team" for teamID, authenticated with c's current token -- the
// TerdutServer's instance-scoped credential, per DESIGN.md §6 point 3: an
// instance-scoped caller may mint a team-scoped account against any team
// (internal/api/service_accounts.go's handleCreateServiceAccount,
// confirmed against source), which is what lets TerdutTeam's own
// controller do this without ever touching a human credential. Same
// adopt-on-409 contract as CreateInstanceServiceAccount.
func (c *Client) CreateTeamServiceAccount(ctx context.Context, name string, teamID int64) (*CreateServiceAccountResult, error) {
req, err := c.newRequest(ctx, http.MethodPost, "/api/service-accounts", map[string]any{
fieldName: name,
"scope": "team",
"team_id": teamID,
})
if err != nil {
return nil, err
}
var result CreateServiceAccountResult
if err := c.do(req, &result); err != nil {
return nil, err
}
return &result, nil
}
// GetServiceAccountByName calls GET /api/service-accounts?name=... --
// authenticated (terdut-server's AuthMiddleware hard-rejects any
// unauthenticated request before this endpoint's own, more permissive
// internal check ever runs; DESIGN.md §6). Returns nil, nil if nothing
// matches, not an error -- the server's own distinction between "found
// nothing" and "the call failed".
func (c *Client) GetServiceAccountByName(ctx context.Context, name string) (*ServiceAccount, error) {
req, err := c.newRequest(ctx, http.MethodGet, "/api/service-accounts?name="+name, nil)
if err != nil {
return nil, err
}
var accounts []ServiceAccount
if err := c.do(req, &accounts); err != nil {
return nil, err
}
if len(accounts) == 0 {
return nil, nil
}
return &accounts[0], nil
}
// CreateServiceAccountKey calls POST /api/service-accounts/{id}/keys to
// mint an additional key on an existing account -- rotation (DESIGN.md §6
// point 6), and the adopt-on-409 recovery path in point 1.
func (c *Client) CreateServiceAccountKey(ctx context.Context, serviceAccountID int64, name string) (*APIKey, error) {
req, err := c.newRequest(ctx, http.MethodPost,
fmt.Sprintf("/api/service-accounts/%d/keys", serviceAccountID),
map[string]string{fieldName: name})
if err != nil {
return nil, err
}
var key APIKey
if err := c.do(req, &key); err != nil {
return nil, err
}
return &key, nil
}
// Team mirrors terdut-server's models.Team (internal/models/team.go), minus
// Role/Source, which are only ever populated for a human caller's own
// membership and never apply to a service account's view of a team.
type Team struct {
ID int64 `json:"id"`
Name string `json:"name"`
ID int64 `json:"id"`
Name string `json:"name"`
ExternalID *string `json:"external_id,omitempty"`
}
// CreateTeam calls POST /api/teams, authenticated with c's current token --
// the TerdutServer's instance-scoped credential (DESIGN.md §6 point 3). A
// StatusError with Code 409 means a prior, interrupted attempt already
// created this name — GetTeamByName (TEAM-LOOKUP.md) is the adopt-rather-
// than-error recovery, the same contract CreateInstanceServiceAccount has.
func (c *Client) CreateTeam(ctx context.Context, name string) (*Team, error) {
req, err := c.newRequest(ctx, http.MethodPost, "/api/teams", map[string]string{fieldName: name})
// CreateTeam calls POST /api/teams with an external_id, which makes it
// idempotent: a team that already carries that id is returned (200) instead of
// created (201), so a client that crashed before recording the id finds its own
// team again. A StatusError with Code 409 means the name belongs to a different
// team. Needs an instance-scoped credential.
func (c *Client) CreateTeam(ctx context.Context, name, externalID string) (*Team, error) {
req, err := c.newRequest(ctx, http.MethodPost, "/api/teams",
map[string]string{fieldName: name, "external_id": externalID})
if err != nil {
return nil, err
}
@@ -307,24 +154,6 @@ func (c *Client) CreateTeam(ctx context.Context, name string) (*Team, error) {
return &team, nil
}
// GetTeamByName calls GET /api/teams?name=... (TEAM-LOOKUP.md) — open to
// any authenticated caller, not just the team's own members. Returns nil,
// nil on no match, not an error.
func (c *Client) GetTeamByName(ctx context.Context, name string) (*Team, error) {
req, err := c.newRequest(ctx, http.MethodGet, "/api/teams?name="+name, nil)
if err != nil {
return nil, err
}
var teams []Team
if err := c.do(req, &teams); err != nil {
return nil, err
}
if len(teams) == 0 {
return nil, nil
}
return &teams[0], nil
}
// RenameTeam calls PUT /api/teams/{teamID} — owner-gated server-side
// (requireTeamOwner), so c must hold this team's own team-scoped
// credential, not the instance-scoped one CreateTeam used.
@@ -346,7 +175,7 @@ func (c *Client) DeleteTeam(ctx context.Context, teamID int64) error {
if err != nil {
return err
}
return c.do(req, nil)
return ignoreNotFound(c.do(req, nil))
}
// SetTeamOIDCGroups calls PUT /api/teams/{teamID}/oidc-groups — owner-gated,
@@ -362,46 +191,15 @@ func (c *Client) SetTeamOIDCGroups(ctx context.Context, teamID int64, memberGrou
return c.do(req, nil)
}
// User mirrors terdut-server's models.User, minus fields this client never
// reads.
type User struct {
ID int64 `json:"id"`
Username string `json:"username"`
}
// GetUserByUsername calls GET /api/users and finds the one matching exactly
// -- confirmed open to any authenticated caller, not gated by team
// membership or admin (internal/api/router.go's own comment: "readable by
// anyone signed in"), so the team-scoped credential a TerdutEscalationRule's
// controller already holds is enough. There is no server-side filter, so
// this always fetches the whole list; terdut-server's own query has no
// pagination either (confirmed against source), so this matches what the
// server itself considers an acceptable cost. Returns nil, nil on no match.
func (c *Client) GetUserByUsername(ctx context.Context, username string) (*User, error) {
req, err := c.newRequest(ctx, http.MethodGet, "/api/users", nil)
if err != nil {
return nil, err
}
var users []User
if err := c.do(req, &users); err != nil {
return nil, err
}
for _, u := range users {
if u.Username == username {
return &u, nil
}
}
return nil, nil
}
// EscalationTargetRequest/EscalationLevelRequest/SetEscalationRequest mirror
// terdut-server's escalationTargetJSON/escalationLevelJSON/escalationJSON
// (internal/api/escalation.go) -- the PUT body, not the richer GET response
// (escalationView), which this client never needs to decode since the
// controller always computes its own desired state fresh from spec.
type EscalationTargetRequest struct {
Kind string `json:"kind"`
UserID *int64 `json:"user_id,omitempty"`
Kind string `json:"kind"`
// Username names the person for a "user" target; the server resolves it.
Username string `json:"username,omitempty"`
}
type EscalationLevelRequest struct {
@@ -498,7 +296,7 @@ func (c *Client) DeleteDeadmanSwitch(ctx context.Context, teamID, switchID int64
if err != nil {
return err
}
return c.do(req, nil)
return ignoreNotFound(c.do(req, nil))
}
// Integration mirrors terdut-server's models.Integration, minus
@@ -564,52 +362,5 @@ func (c *Client) DeleteIntegration(ctx context.Context, teamID, integrationID in
if err != nil {
return err
}
return c.do(req, nil)
}
// Invite is a standing link into a team (POST /api/teams/{teamID}/invites'
// own response shape). URL carries the raw token exactly once, at creation
// -- terdut-server never shows it again (same one-time-shown shape as an
// integration's webhook key) -- so a caller that needs it later has to have
// kept this response, not re-fetched it.
type Invite struct {
ID int64 `json:"id"`
TeamID int64 `json:"team_id"`
Role string `json:"role"`
ExpiresAt time.Time `json:"expires_at"`
MaxUses int64 `json:"max_uses"`
URL string `json:"url,omitempty"`
}
// CreateInvite calls POST /api/teams/{teamID}/invites -- owner-gated
// (requireTeamOwner), so c must hold this team's own team-scoped
// credential, which already satisfies that check via its synthetic owner
// membership (terdut-server's SERVICE-ACCOUNTS.md). No conflict handling
// needed: unlike a team or a service account, an invite has no unique name
// to collide on -- every call mints a brand new row.
func (c *Client) CreateInvite(ctx context.Context, teamID int64, role string, maxUses int64) (*Invite, error) {
req, err := c.newRequest(ctx, http.MethodPost, fmt.Sprintf("/api/teams/%d/invites", teamID),
map[string]any{"role": role, "max_uses": maxUses})
if err != nil {
return nil, err
}
var inv Invite
if err := c.do(req, &inv); err != nil {
return nil, err
}
return &inv, nil
}
// RevokeInvite calls DELETE /api/teams/{teamID}/invites/{inviteID} -- same
// credential requirement as CreateInvite. A 404 (already revoked, or never
// existed) is the caller's to treat as success if it wants to, the same way
// DeleteTeam's own 404 handling works -- this method itself just reports
// whatever terdut-server said.
func (c *Client) RevokeInvite(ctx context.Context, teamID, inviteID int64) error {
req, err := c.newRequest(ctx, http.MethodDelete,
fmt.Sprintf("/api/teams/%d/invites/%d", teamID, inviteID), nil)
if err != nil {
return err
}
return c.do(req, nil)
return ignoreNotFound(c.do(req, nil))
}
+30 -69
View File
@@ -7,85 +7,46 @@ import (
"testing"
)
func TestVersion(t *testing.T) {
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
if r.URL.Path != "/api/version" {
t.Errorf("unexpected path %q", r.URL.Path)
}
// Unauthenticated per terdut-server's own router.go comment: no
// Authorization header should be required, and none is sent here.
w.Header().Set("Content-Type", "application/json")
w.Write([]byte(`{"version":"v0.20.0"}`)) //nolint:errcheck
}))
defer srv.Close()
got, err := New(srv.URL).Version(context.Background())
if err != nil {
t.Fatalf("Version() error = %v", err)
}
if got != "v0.20.0" {
t.Errorf("Version() = %q, want %q", got, "v0.20.0")
}
}
func TestVersionUnreachable(t *testing.T) {
// A closed server: connection refused, the same shape a bad
// spec.endpoint produces against a real cluster.
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {}))
srv.Close()
if _, err := New(srv.URL).Version(context.Background()); err == nil {
t.Fatal("Version() error = nil, want a connection error")
}
}
func TestVersionErrorStatus(t *testing.T) {
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
w.WriteHeader(http.StatusInternalServerError)
w.Write([]byte(`{"error":"internal error"}`)) //nolint:errcheck
}))
defer srv.Close()
_, err := New(srv.URL).Version(context.Background())
if err == nil {
t.Fatal("Version() error = nil, want a StatusError")
}
var statusErr *StatusError
if !asStatusError(err, &statusErr) {
t.Fatalf("Version() error = %v (%T), want *StatusError", err, err)
}
if statusErr.Code != http.StatusInternalServerError {
t.Errorf("StatusError.Code = %d, want %d", statusErr.Code, http.StatusInternalServerError)
}
if statusErr.Message != "internal error" {
t.Errorf("StatusError.Message = %q, want %q", statusErr.Message, "internal error")
}
}
// asStatusError is errors.As without importing errors twice in a tiny test
// file — kept local since no other test here needs it.
func asStatusError(err error, target **StatusError) bool {
se, ok := err.(*StatusError)
if !ok {
return false
}
*target = se
return true
}
func TestWithTokenSetsAuthorizationHeader(t *testing.T) {
var gotAuth string
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
gotAuth = r.Header.Get("Authorization")
w.Write([]byte(`{"version":"v0.20.0"}`)) //nolint:errcheck
w.WriteHeader(http.StatusNoContent)
}))
defer srv.Close()
c := New(srv.URL).WithToken("tdsa_abc123")
if _, err := c.Version(context.Background()); err != nil {
t.Fatalf("Version() error = %v", err)
if err := c.DeleteTeam(context.Background(), 1); err != nil {
t.Fatalf("DeleteTeam() error = %v", err)
}
if want := "Bearer tdsa_abc123"; gotAuth != want {
t.Errorf("Authorization header = %q, want %q", gotAuth, want)
}
}
// Every delete is idempotent: the thing already being gone is success, any
// other failure is not.
func TestDeletesTreatNotFoundAsSuccess(t *testing.T) {
status := http.StatusNotFound
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
w.WriteHeader(status)
}))
defer srv.Close()
c := New(srv.URL)
ctx := context.Background()
for name, call := range map[string]func() error{
"team": func() error { return c.DeleteTeam(ctx, 1) },
"switch": func() error { return c.DeleteDeadmanSwitch(ctx, 1, 2) },
"integration": func() error { return c.DeleteIntegration(ctx, 1, 2) },
} {
status = http.StatusNotFound
if err := call(); err != nil {
t.Errorf("%s: 404 should be success, got %v", name, err)
}
status = http.StatusConflict
if err := call(); err == nil {
t.Errorf("%s: 409 should be an error", name)
}
}
}
-103
View File
@@ -1,103 +0,0 @@
//go:build e2e
// +build e2e
package e2e
import (
"fmt"
"os"
"os/exec"
"testing"
. "github.com/onsi/ginkgo/v2"
. "github.com/onsi/gomega"
"git.ryuvia.com/niklas/terdut-operator/test/utils"
)
var (
// managerImage is the manager image to be built and loaded for testing.
managerImage = "example.com/terdut-operator:v0.0.1"
// shouldCleanupCertManager tracks whether CertManager was installed by this suite.
shouldCleanupCertManager = false
)
// TestE2E runs the e2e test suite to validate the solution in an isolated environment.
// The default setup requires Kind and CertManager.
//
// To enable kubectl kuberc (use custom kubectl configurations), set: KUBECTL_KUBERC=true
// By default, kuberc is disabled to ensure consistent test behavior across different environments.
// To skip CertManager installation, set: CERT_MANAGER_INSTALL_SKIP=true
func TestE2E(t *testing.T) {
RegisterFailHandler(Fail)
_, _ = fmt.Fprintf(GinkgoWriter, "Starting terdut-operator e2e test suite\n")
RunSpecs(t, "e2e suite")
}
var _ = BeforeSuite(func() {
By("building the manager image")
cmd := exec.Command("make", "docker-build", fmt.Sprintf("IMG=%s", managerImage))
_, err := utils.Run(cmd)
ExpectWithOffset(1, err).NotTo(HaveOccurred(), "Failed to build the manager image")
// TODO(user): If you want to change the e2e test vendor from Kind,
// ensure the image is built and available, then remove the following block.
By("loading the manager image on Kind")
err = utils.LoadImageToKindClusterWithName(managerImage)
ExpectWithOffset(1, err).NotTo(HaveOccurred(), "Failed to load the manager image into Kind")
configureKubectlKubeRC()
setupCertManager()
})
var _ = AfterSuite(func() {
teardownCertManager()
})
// Disable kubectl kuberc by default for test isolation.
// This prevents local kubectl configurations from affecting test behavior.
// To enable kuberc, set: KUBECTL_KUBERC=true
func configureKubectlKubeRC() {
if os.Getenv("KUBECTL_KUBERC") != "true" {
By("disabling kubectl kuberc for test isolation")
err := os.Setenv("KUBECTL_KUBERC", "false")
ExpectWithOffset(1, err).NotTo(HaveOccurred(), "Failed to disable kubectl kuberc")
_, _ = fmt.Fprintf(GinkgoWriter,
"kubectl kuberc disabled for consistent test behavior (override with KUBECTL_KUBERC=true)\n")
} else {
_, _ = fmt.Fprintf(GinkgoWriter, "kubectl kuberc enabled (KUBECTL_KUBERC=true)\n")
}
}
// setupCertManager installs CertManager if needed for webhook tests.
// Skips installation if CERT_MANAGER_INSTALL_SKIP=true or if already present.
func setupCertManager() {
if os.Getenv("CERT_MANAGER_INSTALL_SKIP") == "true" {
_, _ = fmt.Fprintf(GinkgoWriter, "Skipping CertManager installation (CERT_MANAGER_INSTALL_SKIP=true)\n")
return
}
By("checking if CertManager is already installed")
if utils.IsCertManagerCRDsInstalled() {
_, _ = fmt.Fprintf(GinkgoWriter, "CertManager is already installed. Skipping installation.\n")
return
}
// Mark for cleanup before installation to handle interruptions and partial installs.
shouldCleanupCertManager = true
By("installing CertManager")
Expect(utils.InstallCertManager()).To(Succeed(), "Failed to install CertManager")
}
// teardownCertManager uninstalls CertManager if it was installed by setupCertManager.
// This ensures we only remove what we installed.
func teardownCertManager() {
if !shouldCleanupCertManager {
_, _ = fmt.Fprintf(GinkgoWriter, "Skipping CertManager cleanup (not installed by this suite)\n")
return
}
By("uninstalling CertManager")
utils.UninstallCertManager()
}
-323
View File
@@ -1,323 +0,0 @@
//go:build e2e
// +build e2e
package e2e
import (
"encoding/json"
"fmt"
"os"
"os/exec"
"path/filepath"
"time"
. "github.com/onsi/ginkgo/v2"
. "github.com/onsi/gomega"
"git.ryuvia.com/niklas/terdut-operator/test/utils"
)
// namespace where the project is deployed in
const namespace = "terdut-operator-system"
// serviceAccountName created for the project
const serviceAccountName = "terdut-operator-controller-manager"
// metricsServiceName is the name of the metrics service of the project
const metricsServiceName = "terdut-operator-controller-manager-metrics-service"
// metricsRoleBindingName is the name of the RBAC that will be created to allow get the metrics data
const metricsRoleBindingName = "terdut-operator-metrics-binding"
var _ = Describe("Manager", Ordered, func() {
var controllerPodName string
// Before running the tests, set up the environment by creating the namespace,
// enforce the restricted security policy to the namespace, installing CRDs,
// and deploying the controller.
BeforeAll(func() {
By("creating manager namespace")
cmd := exec.Command("kubectl", "create", "ns", namespace)
_, err := utils.Run(cmd)
Expect(err).NotTo(HaveOccurred(), "Failed to create namespace")
By("labeling the namespace to enforce the restricted security policy")
cmd = exec.Command("kubectl", "label", "--overwrite", "ns", namespace,
"pod-security.kubernetes.io/enforce=restricted")
_, err = utils.Run(cmd)
Expect(err).NotTo(HaveOccurred(), "Failed to label namespace with restricted policy")
By("installing CRDs")
cmd = exec.Command("make", "install")
_, err = utils.Run(cmd)
Expect(err).NotTo(HaveOccurred(), "Failed to install CRDs")
By("deploying the controller-manager")
cmd = exec.Command("make", "deploy", fmt.Sprintf("IMG=%s", managerImage))
_, err = utils.Run(cmd)
Expect(err).NotTo(HaveOccurred(), "Failed to deploy the controller-manager")
})
// After all tests have been executed, clean up by undeploying the controller, uninstalling CRDs,
// and deleting the namespace.
AfterAll(func() {
By("cleaning up the curl pod for metrics")
cmd := exec.Command("kubectl", "delete", "pod", "curl-metrics", "-n", namespace)
_, _ = utils.Run(cmd)
By("undeploying the controller-manager")
cmd = exec.Command("make", "undeploy")
_, _ = utils.Run(cmd)
By("uninstalling CRDs")
cmd = exec.Command("make", "uninstall")
_, _ = utils.Run(cmd)
By("removing manager namespace")
cmd = exec.Command("kubectl", "delete", "ns", namespace)
_, _ = utils.Run(cmd)
})
// After each test, check for failures and collect logs, events,
// and pod descriptions for debugging.
AfterEach(func() {
specReport := CurrentSpecReport()
if specReport.Failed() {
By("Fetching controller manager pod logs")
cmd := exec.Command("kubectl", "logs", controllerPodName, "-n", namespace)
controllerLogs, err := utils.Run(cmd)
if err == nil {
_, _ = fmt.Fprintf(GinkgoWriter, "Controller logs:\n %s", controllerLogs)
} else {
_, _ = fmt.Fprintf(GinkgoWriter, "Failed to get Controller logs: %s", err)
}
By("Fetching Kubernetes events")
cmd = exec.Command("kubectl", "get", "events", "-n", namespace, "--sort-by=.lastTimestamp")
eventsOutput, err := utils.Run(cmd)
if err == nil {
_, _ = fmt.Fprintf(GinkgoWriter, "Kubernetes events:\n%s", eventsOutput)
} else {
_, _ = fmt.Fprintf(GinkgoWriter, "Failed to get Kubernetes events: %s", err)
}
By("Fetching curl-metrics logs")
cmd = exec.Command("kubectl", "logs", "curl-metrics", "-n", namespace)
metricsOutput, err := utils.Run(cmd)
if err == nil {
_, _ = fmt.Fprintf(GinkgoWriter, "Metrics logs:\n %s", metricsOutput)
} else {
_, _ = fmt.Fprintf(GinkgoWriter, "Failed to get curl-metrics logs: %s", err)
}
By("Fetching controller manager pod description")
cmd = exec.Command("kubectl", "describe", "pod", controllerPodName, "-n", namespace)
podDescription, err := utils.Run(cmd)
if err == nil {
fmt.Println("Pod description:\n", podDescription)
} else {
fmt.Println("Failed to describe controller pod")
}
}
})
SetDefaultEventuallyTimeout(2 * time.Minute)
SetDefaultEventuallyPollingInterval(time.Second)
Context("Manager", func() {
It("should run successfully", func() {
By("validating that the controller-manager pod is running as expected")
verifyControllerUp := func(g Gomega) {
By("getting the name of the controller-manager pod")
cmd := exec.Command("kubectl", "get",
"pods", "-l", "control-plane=controller-manager",
"-o", "go-template={{ range .items }}"+
"{{ if not .metadata.deletionTimestamp }}"+
"{{ .metadata.name }}"+
"{{ \"\\n\" }}{{ end }}{{ end }}",
"-n", namespace,
)
podOutput, err := utils.Run(cmd)
g.Expect(err).NotTo(HaveOccurred(), "Failed to retrieve controller-manager pod information")
podNames := utils.GetNonEmptyLines(podOutput)
g.Expect(podNames).To(HaveLen(1), "expected 1 controller pod running")
controllerPodName = podNames[0]
g.Expect(controllerPodName).To(ContainSubstring("controller-manager"))
By("validating the pod's status")
cmd = exec.Command("kubectl", "get",
"pods", controllerPodName, "-o", "jsonpath={.status.phase}",
"-n", namespace,
)
output, err := utils.Run(cmd)
g.Expect(err).NotTo(HaveOccurred())
g.Expect(output).To(Equal("Running"), "Incorrect controller-manager pod status")
}
Eventually(verifyControllerUp).Should(Succeed())
})
It("should ensure the metrics endpoint is serving metrics", func() {
By("creating a ClusterRoleBinding for the service account to allow access to metrics")
cmd := exec.Command("kubectl", "create", "clusterrolebinding", metricsRoleBindingName,
"--clusterrole=terdut-operator-metrics-reader",
fmt.Sprintf("--serviceaccount=%s:%s", namespace, serviceAccountName),
)
_, err := utils.Run(cmd)
Expect(err).NotTo(HaveOccurred(), "Failed to create ClusterRoleBinding")
By("validating that the metrics service is available")
cmd = exec.Command("kubectl", "get", "service", metricsServiceName, "-n", namespace)
_, err = utils.Run(cmd)
Expect(err).NotTo(HaveOccurred(), "Metrics service should exist")
By("getting the service account token")
token, err := serviceAccountToken()
Expect(err).NotTo(HaveOccurred())
Expect(token).NotTo(BeEmpty())
By("ensuring the controller pod is ready")
verifyControllerPodReady := func(g Gomega) {
cmd := exec.Command("kubectl", "get", "pod", controllerPodName, "-n", namespace,
"-o", "jsonpath={.status.conditions[?(@.type=='Ready')].status}")
output, err := utils.Run(cmd)
g.Expect(err).NotTo(HaveOccurred())
g.Expect(output).To(Equal("True"), "Controller pod not ready")
}
Eventually(verifyControllerPodReady, 3*time.Minute, time.Second).Should(Succeed())
By("verifying that the controller manager is serving the metrics server")
verifyMetricsServerStarted := func(g Gomega) {
cmd := exec.Command("kubectl", "logs", controllerPodName, "-n", namespace)
output, err := utils.Run(cmd)
g.Expect(err).NotTo(HaveOccurred())
g.Expect(output).To(ContainSubstring("Serving metrics server"),
"Metrics server not yet started")
}
Eventually(verifyMetricsServerStarted, 3*time.Minute, time.Second).Should(Succeed())
// +kubebuilder:scaffold:e2e-metrics-webhooks-readiness
By("creating the curl-metrics pod to access the metrics endpoint")
cmd = exec.Command("kubectl", "run", "curl-metrics", "--restart=Never",
"--namespace", namespace,
"--image=curlimages/curl:latest",
"--overrides",
fmt.Sprintf(`{
"spec": {
"containers": [{
"name": "curl",
"image": "curlimages/curl:latest",
"command": ["/bin/sh", "-c"],
"args": [
"for i in $(seq 1 30); do curl -v -k -H 'Authorization: Bearer %s' https://%s.%s.svc.cluster.local:8443/metrics && exit 0 || sleep 2; done; exit 1"
],
"securityContext": {
"readOnlyRootFilesystem": true,
"allowPrivilegeEscalation": false,
"capabilities": {
"drop": ["ALL"]
},
"runAsNonRoot": true,
"runAsUser": 1000,
"seccompProfile": {
"type": "RuntimeDefault"
}
}
}],
"serviceAccountName": "%s"
}
}`, token, metricsServiceName, namespace, serviceAccountName))
_, err = utils.Run(cmd)
Expect(err).NotTo(HaveOccurred(), "Failed to create curl-metrics pod")
By("waiting for the curl-metrics pod to complete.")
verifyCurlUp := func(g Gomega) {
cmd := exec.Command("kubectl", "get", "pods", "curl-metrics",
"-o", "jsonpath={.status.phase}",
"-n", namespace)
output, err := utils.Run(cmd)
g.Expect(err).NotTo(HaveOccurred())
g.Expect(output).To(Equal("Succeeded"), "curl pod in wrong status")
}
Eventually(verifyCurlUp, 5*time.Minute).Should(Succeed())
By("getting the metrics by checking curl-metrics logs")
verifyMetricsAvailable := func(g Gomega) {
metricsOutput, err := getMetricsOutput()
g.Expect(err).NotTo(HaveOccurred(), "Failed to retrieve logs from curl pod")
g.Expect(metricsOutput).NotTo(BeEmpty())
g.Expect(metricsOutput).To(ContainSubstring("< HTTP/1.1 200 OK"))
}
Eventually(verifyMetricsAvailable, 2*time.Minute).Should(Succeed())
})
// +kubebuilder:scaffold:e2e-webhooks-checks
// TODO: Customize the e2e test suite with scenarios specific to your project.
// Consider applying sample/CR(s) and check their status and/or verifying
// the reconciliation by using the metrics, i.e.:
// metricsOutput, err := getMetricsOutput()
// Expect(err).NotTo(HaveOccurred(), "Failed to retrieve logs from curl pod")
// Expect(metricsOutput).To(ContainSubstring(
// fmt.Sprintf(`controller_runtime_reconcile_total{controller="%s",result="success"} 1`,
// strings.ToLower(<Kind>),
// ))
})
})
// serviceAccountToken returns a token for the specified service account in the given namespace.
// It uses the Kubernetes TokenRequest API to generate a token by directly sending a request
// and parsing the resulting token from the API response.
func serviceAccountToken() (string, error) {
const tokenRequestRawString = `{
"apiVersion": "authentication.k8s.io/v1",
"kind": "TokenRequest"
}`
By("creating temporary file to store the token request")
secretName := fmt.Sprintf("%s-token-request", serviceAccountName)
tokenRequestFile := filepath.Join("/tmp", secretName)
err := os.WriteFile(tokenRequestFile, []byte(tokenRequestRawString), os.FileMode(0o644))
if err != nil {
return "", err
}
var out string
verifyTokenCreation := func(g Gomega) {
By("executing kubectl command to create the token")
cmd := exec.Command("kubectl", "create", "--raw", fmt.Sprintf(
"/api/v1/namespaces/%s/serviceaccounts/%s/token",
namespace,
serviceAccountName,
), "-f", tokenRequestFile)
output, err := cmd.CombinedOutput()
g.Expect(err).NotTo(HaveOccurred())
By("parsing the JSON output to extract the token")
var token tokenRequest
err = json.Unmarshal(output, &token)
g.Expect(err).NotTo(HaveOccurred())
out = token.Status.Token
}
Eventually(verifyTokenCreation).Should(Succeed())
return out, err
}
// getMetricsOutput retrieves and returns the logs from the curl pod used to access the metrics endpoint.
func getMetricsOutput() (string, error) {
By("getting the curl-metrics logs")
cmd := exec.Command("kubectl", "logs", "curl-metrics", "-n", namespace)
return utils.Run(cmd)
}
// tokenRequest is a simplified representation of the Kubernetes TokenRequest API response,
// containing only the token field that we need to extract.
type tokenRequest struct {
Status struct {
Token string `json:"token"`
} `json:"status"`
}
-210
View File
@@ -1,210 +0,0 @@
package utils
import (
"bufio"
"bytes"
"fmt"
"os"
"os/exec"
"strings"
. "github.com/onsi/ginkgo/v2" // nolint:revive,staticcheck
)
const (
certmanagerVersion = "v1.21.1"
certmanagerURLTmpl = "https://github.com/cert-manager/cert-manager/releases/download/%s/cert-manager.yaml"
defaultKindBinary = "kind"
defaultKindCluster = "kind"
)
func warnError(err error) {
_, _ = fmt.Fprintf(GinkgoWriter, "warning: %v\n", err)
}
// Run executes the provided command within this context
func Run(cmd *exec.Cmd) (string, error) {
dir, _ := GetProjectDir()
cmd.Dir = dir
if err := os.Chdir(cmd.Dir); err != nil {
_, _ = fmt.Fprintf(GinkgoWriter, "chdir dir: %q\n", err)
}
cmd.Env = append(os.Environ(), "GO111MODULE=on")
command := strings.Join(cmd.Args, " ")
_, _ = fmt.Fprintf(GinkgoWriter, "running: %q\n", command)
output, err := cmd.CombinedOutput()
if err != nil {
return string(output), fmt.Errorf("%q failed with error %q: %w", command, string(output), err)
}
return string(output), nil
}
// UninstallCertManager uninstalls the cert manager
func UninstallCertManager() {
url := fmt.Sprintf(certmanagerURLTmpl, certmanagerVersion)
cmd := exec.Command("kubectl", "delete", "-f", url)
if _, err := Run(cmd); err != nil {
warnError(err)
}
// Delete leftover leases in kube-system (not cleaned by default)
kubeSystemLeases := []string{
"cert-manager-cainjector-leader-election",
"cert-manager-controller",
}
for _, lease := range kubeSystemLeases {
cmd = exec.Command("kubectl", "delete", "lease", lease,
"-n", "kube-system", "--ignore-not-found", "--force", "--grace-period=0")
if _, err := Run(cmd); err != nil {
warnError(err)
}
}
}
// InstallCertManager installs the cert manager bundle.
func InstallCertManager() error {
url := fmt.Sprintf(certmanagerURLTmpl, certmanagerVersion)
cmd := exec.Command("kubectl", "apply", "-f", url)
if _, err := Run(cmd); err != nil {
return err
}
// Wait for cert-manager-webhook to be ready, which can take time if cert-manager
// was re-installed after uninstalling on a cluster.
cmd = exec.Command("kubectl", "wait", "deployment.apps/cert-manager-webhook",
"--for", "condition=Available",
"--namespace", "cert-manager",
"--timeout", "5m",
)
_, err := Run(cmd)
return err
}
// IsCertManagerCRDsInstalled checks if any Cert Manager CRDs are installed
// by verifying the existence of key CRDs related to Cert Manager.
func IsCertManagerCRDsInstalled() bool {
// List of common Cert Manager CRDs
certManagerCRDs := []string{
"certificates.cert-manager.io",
"issuers.cert-manager.io",
"clusterissuers.cert-manager.io",
"certificaterequests.cert-manager.io",
"orders.acme.cert-manager.io",
"challenges.acme.cert-manager.io",
}
// Execute the kubectl command to get all CRDs
cmd := exec.Command("kubectl", "get", "crds")
output, err := Run(cmd)
if err != nil {
return false
}
// Check if any of the Cert Manager CRDs are present
crdList := GetNonEmptyLines(output)
for _, crd := range certManagerCRDs {
for _, line := range crdList {
if strings.Contains(line, crd) {
return true
}
}
}
return false
}
// LoadImageToKindClusterWithName loads a local docker image to the kind cluster
func LoadImageToKindClusterWithName(name string) error {
cluster := defaultKindCluster
if v, ok := os.LookupEnv("KIND_CLUSTER"); ok {
cluster = v
}
kindOptions := []string{"load", "docker-image", name, "--name", cluster}
kindBinary := defaultKindBinary
if v, ok := os.LookupEnv("KIND"); ok {
kindBinary = v
}
cmd := exec.Command(kindBinary, kindOptions...)
_, err := Run(cmd)
return err
}
// GetNonEmptyLines converts given command output string into individual objects
// according to line breakers, and ignores the empty elements in it.
func GetNonEmptyLines(output string) []string {
var res []string
elements := strings.SplitSeq(output, "\n")
for element := range elements {
if element != "" {
res = append(res, element)
}
}
return res
}
// GetProjectDir will return the directory where the project is
func GetProjectDir() (string, error) {
wd, err := os.Getwd()
if err != nil {
return wd, fmt.Errorf("failed to get current working directory: %w", err)
}
wd = strings.ReplaceAll(wd, "/test/e2e", "")
return wd, nil
}
// UncommentCode searches for target in the file and remove the comment prefix
// of the target content. The target content may span multiple lines.
func UncommentCode(filename, target, prefix string) error {
// false positive
// nolint:gosec
content, err := os.ReadFile(filename)
if err != nil {
return fmt.Errorf("failed to read file %q: %w", filename, err)
}
strContent := string(content)
idx := strings.Index(strContent, target)
if idx < 0 {
return fmt.Errorf("unable to find the code %q to be uncommented", target)
}
out := new(bytes.Buffer)
_, err = out.Write(content[:idx])
if err != nil {
return fmt.Errorf("failed to write to output: %w", err)
}
scanner := bufio.NewScanner(bytes.NewBufferString(target))
if !scanner.Scan() {
return nil
}
for {
if _, err = out.WriteString(strings.TrimPrefix(scanner.Text(), prefix)); err != nil {
return fmt.Errorf("failed to write to output: %w", err)
}
// Avoid writing a newline in case the previous line was the last in target.
if !scanner.Scan() {
break
}
if _, err = out.WriteString("\n"); err != nil {
return fmt.Errorf("failed to write to output: %w", err)
}
}
if _, err = out.Write(content[idx+len(target):]); err != nil {
return fmt.Errorf("failed to write to output: %w", err)
}
// false positive
// nolint:gosec
if err = os.WriteFile(filename, out.Bytes(), 0644); err != nil {
return fmt.Errorf("failed to write file %q: %w", filename, err)
}
return nil
}