The README was 1,240 lines of reference material and still described a SQLite quick start. It is now a short tour (highlights, screenshots of the web UI, an accurate quick start against Postgres), and each topic has its own page under docs/ with an index: deployment, configuration, Alertmanager, incidents, notifications, escalation, dead man's switches, single sign-on, web UI, API and development. SERVICE-ACCOUNTS.md is rewritten from a proposal into a reference, and TEAM-LOOKUP.md is gone with the endpoint it described. The "Upgrading to ..." sections for an unreleased product are dropped. Claude-Session: https://claude.ai/code/session_016mBLURvJoMuUEr9cB2RpUN
2.1 KiB
Alertmanager configuration
Pointing Alertmanager at the server. Back to the README and the documentation index.
Alerts arrive on a team's integration key, which says both that the sender may post and which team the alerts belong to. Mint one as an owner of the team:
curl -X POST https://terdut.example.com/api/teams/1/integrations \
-H "Authorization: Bearer $TERDUT_API_KEY" \
-H 'Content-Type: application/json' \
-d '{"name":"prod alertmanager"}'
The response carries the key and the full URL once; only a SHA-256 hash is
stored. Put it in your alertmanager.yml:
receivers:
- name: terdut
webhook_configs:
- url: http://terdut-server:8080/api/integrations/<key>/alertmanager
send_resolved: true
route:
receiver: terdut
The whole URL is a credential, so treat it like one. Alertmanager 0.26 and
later can read it from a file with url_file: instead, which keeps it out of
your configuration repository:
- url_file: /etc/alertmanager/secrets/terdut-webhook-url/url
send_resolved: true
The webhook endpoint requires no authentication.
If you use the dead man's switch — and the default configuration does — give the heartbeat a route of its own, because the deadline is only as tight as the interval feeding it:
route:
receiver: terdut
repeat_interval: 4h
routes:
- matchers: [ 'alertname = "Watchdog"' ]
receiver: terdut
group_wait: 0s
group_interval: 1m
repeat_interval: 1m
That delivers a heartbeat every 2 minutes, not every minute. Alertmanager only reconsiders a
group every group_interval, and at exactly one elapsed interval repeat_interval has not quite
passed, so the send slips to the next tick — equal values give 2×. Two minutes against the 15 minute
default is seven heartbeats per window, which is the point; use group_interval: 30s if you want
the numbers to mean what they say.
kube-prometheus-stack users get the Watchdog alert (expr: vector(1)) for free; it just needs
routing to terdut rather than to null.