Compare commits
10 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| 4c85e7646c | |||
| 05f82220a6 | |||
| 5227eb0d5f | |||
| 74359c72ab | |||
| 43f69272f0 | |||
| 2de5c8412d | |||
| a4fbd60441 | |||
| 1377d9005b | |||
| e3ad19c110 | |||
| a8ee742533 |
@@ -35,6 +35,27 @@ jobs:
|
||||
- go-mod-cache:/go/pkg/mod
|
||||
- go-build-cache:/root/.cache/go-build
|
||||
- gobin-cache:/go/bin
|
||||
|
||||
# The same database ci.yaml's test job gets, for the same reason: `make test` needs a
|
||||
# real Postgres and fails without TERDUT_TEST_DSN rather than skipping. This job is
|
||||
# the gate every publishing job below hangs off, so it has to be able to run the
|
||||
# suite -- v0.11.0 was tagged with the service here missing and published nothing.
|
||||
services:
|
||||
postgres:
|
||||
image: postgres:17-alpine
|
||||
env:
|
||||
POSTGRES_USER: terdut
|
||||
POSTGRES_PASSWORD: terdut
|
||||
POSTGRES_DB: terdut_test
|
||||
options: >-
|
||||
--health-cmd "pg_isready -U terdut -d terdut_test"
|
||||
--health-interval 5s
|
||||
--health-timeout 5s
|
||||
--health-retries 12
|
||||
|
||||
env:
|
||||
TERDUT_TEST_DSN: postgres://terdut:terdut@postgres:5432/terdut_test?sslmode=disable
|
||||
|
||||
steps:
|
||||
- name: Checkout
|
||||
env:
|
||||
|
||||
@@ -161,7 +161,7 @@ somewhere to exec. The sidecar, the PVC and the `backupSidecar` values are all g
|
||||
| `TERDUT_DB_DSN` | — | **Required.** Postgres connection string, e.g. `postgres://terdut:secret@localhost:5432/terdut?sslmode=require` |
|
||||
| `TERDUT_ARCHIVE_AFTER` | `168h` (7d) | How long a resolved alert or incident stays in the default list before being auto-archived |
|
||||
| `TERDUT_STALE_AFTER` | `6h` | How long a firing alert may go without a refreshing webhook before it is treated as resolved — **must exceed your Alertmanager `repeat_interval`** |
|
||||
| `TERDUT_DEADMAN_MATCHERS` | `alertname=Watchdog` | Which alerts are [dead man's switches](#dead-mans-switch). `;` separates matchers, `,` the label conditions within one, `=` is exact equality. Every matcher must name an `alertname` |
|
||||
| `TERDUT_DEADMAN_MATCHERS` | `alertname=Watchdog` | The **default** matchers a team starts with — switches are per team now, and this seeds teams that have no configuration of their own. `;` separates matchers, `,` the label conditions within one, `=` is exact equality. Every matcher must name an `alertname` |
|
||||
| `TERDUT_DEADMAN_TIMEOUT` | `15m` | How long a heartbeat may go unheard before its switch is declared dead — **must be shorter than the `repeat_interval` of the route carrying it**. `0` disables dead man's switch handling |
|
||||
| `TERDUT_DEADMAN_SEVERITY` | `critical` | Severity a dead man's switch incident opens at |
|
||||
| `TERDUT_NTFY_URL` | — | ntfy server to publish push notifications to. Empty disables notifications entirely |
|
||||
@@ -378,11 +378,23 @@ kube-prometheus-stack already ships the alert for this. `Watchdog` is
|
||||
nothing unless something downstream notices it stop. That is what
|
||||
`TERDUT_DEADMAN_MATCHERS` defaults to.
|
||||
|
||||
**Switches belong to a team**, which decides which of its own alerts are
|
||||
heartbeats and how long a silence has to last. An owner sets them through
|
||||
`PUT /api/teams/{teamID}/deadman`; a missed heartbeat opens an incident in the
|
||||
team whose integration received it.
|
||||
|
||||
The environment variables are the starting point, not the setting: at startup
|
||||
every team **without** a configuration of its own is given one from them, and an
|
||||
owner's later edit is never overwritten by a redeploy. A team created after
|
||||
that starts watching nothing until its owner says otherwise — inheriting an
|
||||
install-wide heartbeat would page a new team about a source it has never heard
|
||||
of.
|
||||
|
||||
A matcher is a set of exact label conditions, one of which must be the
|
||||
`alertname`:
|
||||
`alertname`, in the same format the environment variable uses:
|
||||
|
||||
```
|
||||
TERDUT_DEADMAN_MATCHERS="alertname=Watchdog,cluster=prod; alertname=EdgeHeartbeat"
|
||||
alertname=Watchdog,cluster=prod; alertname=EdgeHeartbeat
|
||||
```
|
||||
|
||||
**The unit of monitoring is the fingerprint, not the alert name.** Two clusters
|
||||
@@ -445,6 +457,30 @@ Authorization: Bearer <api-key>
|
||||
or the web UI's session cookie. A request that carries an `Authorization` header
|
||||
is judged on that header alone.
|
||||
|
||||
Two kinds of user exist. An **administrator** manages accounts: creating and
|
||||
deleting users, setting anybody's password, minting keys for anybody, and
|
||||
granting the flag itself. Everybody else works incidents — acknowledging,
|
||||
assigning, snoozing, resolving, noting — and manages their own account and
|
||||
nobody else's. An API key carries exactly the rights of the user it belongs to.
|
||||
|
||||
The first user, from `/api/bootstrap`, is an administrator. Users created
|
||||
afterwards are not, until an administrator says so. An install always keeps at
|
||||
least one: the last administrator can be neither deleted nor demoted, and
|
||||
nobody can delete or demote themselves.
|
||||
|
||||
Endpoints that require the flag answer `403` with
|
||||
`{"error":"administrator access required"}`.
|
||||
|
||||
**Teams** are the unit of tenancy, and are a separate axis from the administrator
|
||||
flag. A team owns its incidents, alerts, schedule and integrations, and a user
|
||||
sees exactly the teams they belong to — an administrator is not implicitly in
|
||||
every team, because administration is about accounts, not about reading other
|
||||
people's incidents. Within a team an **owner** configures it (schedule,
|
||||
integrations, membership) and a **member** works its incidents.
|
||||
|
||||
Anything belonging to a team you are not in answers `404`, not `403`: whether an
|
||||
incident exists is itself something only its team should learn.
|
||||
|
||||
| Method | Path | Description |
|
||||
|---|---|---|
|
||||
| `POST` | `/api/login` | `{"username","password"}` → sets the session cookie, returns `{user, has_password}`. `429` after too many failures |
|
||||
@@ -455,20 +491,52 @@ is judged on that header alone.
|
||||
|
||||
| Method | Path | Description |
|
||||
|---|---|---|
|
||||
| `POST` | `/api/bootstrap` | Create first user + API key `{"username","email","password"?}` (only works on empty DB) |
|
||||
| `GET` | `/api/users` | List users |
|
||||
| `POST` | `/api/users` | Create user `{"username","email"}` |
|
||||
| `DELETE` | `/api/users/{id}` | Delete user (cascades to keys) |
|
||||
| `PUT` | `/api/users/{id}/notify` | Set push notification target `{"ntfy_topic"}` — empty string clears it |
|
||||
| `PUT` | `/api/users/{id}/password` | Set web UI password `{"password","current_password"}`. `current_password` is required only when changing your own existing password. Ends the user's other sessions |
|
||||
| `POST` | `/api/users/{id}/api-keys` | Issue API key `{"name"}` — key shown once |
|
||||
| `DELETE` | `/api/users/{id}/api-keys/{keyID}` | Revoke API key |
|
||||
**admin** marks an endpoint that requires the administrator flag; **self or
|
||||
admin** marks one you may use on your own account and an administrator may use
|
||||
on anybody's.
|
||||
|
||||
| Method | Path | Who | Description |
|
||||
|---|---|---|---|
|
||||
| `POST` | `/api/bootstrap` | — | Create first user + API key `{"username","email","password"?}` (only works on empty DB). The user is an administrator |
|
||||
| `GET` | `/api/users` | any | List users. Open to everybody: the queue's assignment control and the schedule both have to name people |
|
||||
| `POST` | `/api/users` | **admin** | Create user `{"username","email"}`. Not an administrator |
|
||||
| `DELETE` | `/api/users/{id}` | **admin** | Delete user (cascades to keys). `409` for yourself or the last administrator |
|
||||
| `PUT` | `/api/users/{id}/admin` | **admin** | Grant or revoke the administrator flag `{"is_admin"}`. `409` for yourself or the last administrator |
|
||||
| `PUT` | `/api/users/{id}/notify` | self or admin | Set push notification target `{"ntfy_topic"}` — empty string clears it |
|
||||
| `PUT` | `/api/users/{id}/password` | self or admin | Set web UI password `{"password","current_password"}`. `current_password` is required only when changing your own existing password. Ends the user's other sessions |
|
||||
| `POST` | `/api/users/{id}/api-keys` | self or admin | Issue API key `{"name"}` — key shown once |
|
||||
| `DELETE` | `/api/users/{id}/api-keys/{keyID}` | self or admin | Revoke API key |
|
||||
|
||||
### Alert ingestion
|
||||
|
||||
Alerts arrive on a team's integration key. The key is both the credential and the
|
||||
routing: it says that the sender may post, and which team the alerts belong to.
|
||||
Create one with `POST /api/teams/{teamID}/integrations`, which returns the key
|
||||
and the full URL once and stores only a SHA-256 hash.
|
||||
|
||||
| Method | Path | Description |
|
||||
|---|---|---|
|
||||
| `POST` | `/api/alertmanager/webhook` | Alertmanager v4 webhook receiver (no auth) |
|
||||
| `POST` | `/api/integrations/{key}/alertmanager` | Alertmanager v4 webhook receiver for the key's team. `401` for an unknown key |
|
||||
| `POST` | `/api/alertmanager/webhook` | **Deprecated, unauthenticated.** The pre-teams receiver, kept for one release so an upgrade does not stop delivering while the Alertmanager config is edited. Routes everything to the oldest team |
|
||||
|
||||
The deprecated path is why anything that can reach the port can still open an
|
||||
incident. Move senders to a key and it goes away.
|
||||
|
||||
### Teams
|
||||
|
||||
| Method | Path | Who | Description |
|
||||
|---|---|---|---|
|
||||
| `GET` | `/api/teams` | any | The caller's own teams, each with their role |
|
||||
| `POST` | `/api/teams` | any | Create a team `{"name"}`; the creator becomes its first owner |
|
||||
| `DELETE` | `/api/teams/{teamID}` | **owner** | Delete a team and everything under it. `409` while it has open incidents |
|
||||
| `GET` | `/api/teams/{teamID}/members` | member | Who is in the team |
|
||||
| `POST` | `/api/teams/{teamID}/members` | **owner** | Add a member, or change their role `{"user_id","role"}` |
|
||||
| `DELETE` | `/api/teams/{teamID}/members/{userID}` | **owner** | Remove a member. `409` for the last owner |
|
||||
| `GET` | `/api/teams/{teamID}/integrations` | member | List integrations. Never returns keys |
|
||||
| `POST` | `/api/teams/{teamID}/integrations` | **owner** | Mint an integration `{"name","kind"}` — key and URL shown once |
|
||||
| `DELETE` | `/api/teams/{teamID}/integrations/{integrationID}` | **owner** | Revoke an integration |
|
||||
| `GET` | `/api/teams/{teamID}/deadman` | member | The team's [dead man's switch](#dead-mans-switch) configuration `{matchers, timeout_seconds, severity}` |
|
||||
| `PUT` | `/api/teams/{teamID}/deadman` | **owner** | Replace it. `400` when no matcher names an `alertname`, because a switch that silently watches nothing is the failure this feature exists to prevent |
|
||||
|
||||
### Notifications
|
||||
|
||||
@@ -655,13 +723,23 @@ unknown" rather than being rejected.
|
||||
|
||||
| Method | Path | Description |
|
||||
|---|---|---|
|
||||
| `POST` | `/api/schedule` | Assign user to dates `{"user_id", "dates":["YYYY-MM-DD",...], "replace"}` — all-or-nothing |
|
||||
| `GET` | `/api/schedule` | List entries. Filters: `?from=YYYY-MM-DD`, `?to=YYYY-MM-DD` |
|
||||
| `GET` | `/api/schedule/current` | Today's on-call user (UTC), 404 if none |
|
||||
| `DELETE` | `/api/schedule/{id}` | Remove schedule entry |
|
||||
Each team keeps its own rota, so two teams can have two different people on call
|
||||
on the same day. The person taking a shift has to be in the team — paging
|
||||
somebody who cannot open the incident is worse than paging nobody.
|
||||
|
||||
| Method | Path | Who | Description |
|
||||
|---|---|---|---|
|
||||
| `POST` | `/api/teams/{teamID}/schedule` | **owner** | Assign user to dates `{"user_id", "dates":["YYYY-MM-DD",...], "replace"}` — all-or-nothing |
|
||||
| `GET` | `/api/teams/{teamID}/schedule` | member | List entries. Filters: `?from=YYYY-MM-DD`, `?to=YYYY-MM-DD` |
|
||||
| `DELETE` | `/api/teams/{teamID}/schedule/{id}` | **owner** | Remove schedule entry |
|
||||
| `GET` | `/api/schedule/current` | any | Who is on call today (UTC) in **every** team the caller is in — one entry per team, `[]` when nobody anywhere |
|
||||
|
||||
### Statistics
|
||||
|
||||
Every figure counts the caller's own teams only: a report that counted other
|
||||
teams' incidents would leak their volume, and their alert names through the
|
||||
top-alerts list, and would not be a number about the reader's work anyway.
|
||||
|
||||
All stat endpoints accept optional `?from=YYYY-MM-DD` and `?to=YYYY-MM-DD`, and exclude archived rows to match the default list views. Alert stats filter on `received_at`; incident stats filter on `triggered_at`.
|
||||
|
||||
| Method | Path | Description |
|
||||
@@ -678,6 +756,58 @@ averages over incidents that have actually been acknowledged or resolved, and ar
|
||||
|
||||
---
|
||||
|
||||
## Upgrading to teams
|
||||
|
||||
Everything that existed before teams moves into one team called **Default**, and
|
||||
every existing user becomes an owner of it. The upgrade is a no-op for the
|
||||
people using it: the same queue, the same schedule, the same incidents, with a
|
||||
name on them.
|
||||
|
||||
What changes, and will need attention:
|
||||
|
||||
- **Alert ingestion moved.** `POST /api/alertmanager/webhook` still works but is
|
||||
deprecated and unauthenticated, and routes everything to the oldest team. Mint
|
||||
a key with `POST /api/teams/{teamID}/integrations` and point Alertmanager at
|
||||
the URL it returns. The old path goes away in a later release.
|
||||
- **The schedule endpoints moved** under `/api/teams/{teamID}/schedule`, and
|
||||
editing the rota is now an owner's job. `GET /api/schedule/current` stayed
|
||||
where it was but now returns an **array** — one entry per team with somebody
|
||||
on call — instead of a single object or a 404. This is a breaking API change
|
||||
for anything that reads it, terdut-tui included.
|
||||
- **Uniqueness is per team now.** Two teams can legitimately see the same alert
|
||||
fingerprint, the same Alertmanager groupKey, and put somebody on call on the
|
||||
same date.
|
||||
|
||||
**Dead man's switches moved too.** `TERDUT_DEADMAN_MATCHERS`, `_TIMEOUT` and
|
||||
`_SEVERITY` are no longer the setting; they are the default each existing team
|
||||
is seeded with at startup, after which an owner edits them per team through
|
||||
`PUT /api/teams/{teamID}/deadman` and a redeploy never overwrites that.
|
||||
|
||||
Nothing else about an incident changes, and incidents never move between teams:
|
||||
an alert belongs to whichever team's key it arrived on.
|
||||
|
||||
## Upgrading to roles
|
||||
|
||||
Before this release every authenticated caller could create and delete users,
|
||||
set anybody's password and mint anybody's API keys. That is now the
|
||||
administrator flag, and the migration **makes every existing user an
|
||||
administrator** — they already held those powers, so nobody's access changes on
|
||||
upgrade and demotion is a deliberate act afterwards. Promoting only the first
|
||||
user would have silently stripped the rest, and could leave an install whose
|
||||
only administrator is an account nobody has a password for.
|
||||
|
||||
Users created after the upgrade are not administrators. Hand the flag out with:
|
||||
|
||||
```bash
|
||||
curl -X PUT https://terdut.example.com/api/users/7/admin \
|
||||
-H "Authorization: Bearer $TERDUT_API_KEY" \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{"is_admin": true}'
|
||||
```
|
||||
|
||||
Nothing in the API changed shape, so terdut-tui needs no new version — but a
|
||||
non-administrator now gets `403` where a `200` used to come back.
|
||||
|
||||
## Upgrading from SQLite
|
||||
|
||||
Versions up to v0.10.2 stored everything in a SQLite file. From the Postgres release onwards
|
||||
|
||||
@@ -15,5 +15,5 @@ type: application
|
||||
# appVersion and image.tag in values.yaml no longer agree, and that is not an oversight:
|
||||
# image.tag stays "latest", which is what a local install actually pulls. appVersion is
|
||||
# metadata and drives nothing.
|
||||
version: 0.11.0
|
||||
appVersion: "v0.11.0"
|
||||
version: 0.12.0
|
||||
appVersion: "v0.12.0"
|
||||
|
||||
+9
-2
@@ -36,9 +36,16 @@ func main() {
|
||||
RepeatEvery: cfg.NotifyRepeat,
|
||||
}
|
||||
|
||||
// Dead man's switches live per team now. The environment variables are the
|
||||
// defaults a team starts from: every team without a configuration of its
|
||||
// own gets one from them here, and an owner's later edit is never
|
||||
// overwritten by a redeploy.
|
||||
deadman := api.ParseDeadmanConfig(cfg.DeadmanMatchers, cfg.DeadmanTimeout, cfg.DeadmanSeverity)
|
||||
if err := api.SeedDeadmanConfigs(context.Background(), database, deadman); err != nil {
|
||||
log.Fatalf("seed dead man's switch defaults: %v", err)
|
||||
}
|
||||
|
||||
router := api.NewRouter(database, notify, deadman)
|
||||
router := api.NewRouter(database, notify)
|
||||
|
||||
srv := &http.Server{
|
||||
Addr: cfg.Addr,
|
||||
@@ -51,7 +58,7 @@ func main() {
|
||||
ctx, stop := signal.NotifyContext(context.Background(), syscall.SIGINT, syscall.SIGTERM)
|
||||
defer stop()
|
||||
|
||||
go api.StartArchiver(ctx, database, cfg.ArchiveAfter, cfg.StaleAfter, deadman, notify)
|
||||
go api.StartArchiver(ctx, database, cfg.ArchiveAfter, cfg.StaleAfter, notify)
|
||||
go api.StartNotifier(ctx, database, notify)
|
||||
|
||||
go func() {
|
||||
|
||||
@@ -0,0 +1,264 @@
|
||||
package api_test
|
||||
|
||||
import (
|
||||
"bytes"
|
||||
"encoding/json"
|
||||
"io"
|
||||
"net/http"
|
||||
"strconv"
|
||||
"testing"
|
||||
)
|
||||
|
||||
// The bootstrap user is an administrator; everybody it creates afterwards is
|
||||
// not. These tests are about the line between them.
|
||||
|
||||
// id64 spells an id into a path segment.
|
||||
func id64(n int64) string { return strconv.FormatInt(n, 10) }
|
||||
|
||||
// member creates an ordinary user and an API key for it, and returns a caller
|
||||
// that authenticates as them. Minting the key goes through the admin's own
|
||||
// credentials, which is how a real install hands one out.
|
||||
func member(t *testing.T, s *ts, username string) (id int64, call func(method, path string, body any) *http.Response) {
|
||||
t.Helper()
|
||||
|
||||
resp := s.req(t, http.MethodPost, "/api/users",
|
||||
map[string]string{"username": username, "email": username + "@test.com"})
|
||||
if resp.StatusCode != http.StatusCreated {
|
||||
t.Fatalf("create %s: %d", username, resp.StatusCode)
|
||||
}
|
||||
var user struct {
|
||||
ID int64 `json:"id"`
|
||||
IsAdmin bool `json:"is_admin"`
|
||||
}
|
||||
decode(t, resp, &user)
|
||||
if user.IsAdmin {
|
||||
t.Fatalf("a created user must not be an administrator")
|
||||
}
|
||||
|
||||
// Into the default team as a plain member: being in a team is what lets
|
||||
// somebody work its incidents, and is separate from administering accounts.
|
||||
resp = s.req(t, http.MethodPost, "/api/teams/"+defaultTeam+"/members",
|
||||
map[string]any{"user_id": user.ID, "role": "member"})
|
||||
resp.Body.Close()
|
||||
if resp.StatusCode != http.StatusNoContent {
|
||||
t.Fatalf("add %s to the team: %d", username, resp.StatusCode)
|
||||
}
|
||||
|
||||
resp = s.req(t, http.MethodPost, "/api/users/"+id64(user.ID)+"/api-keys",
|
||||
map[string]string{"name": "test"})
|
||||
if resp.StatusCode != http.StatusCreated {
|
||||
t.Fatalf("mint key for %s: %d", username, resp.StatusCode)
|
||||
}
|
||||
var key struct {
|
||||
Key string `json:"key"`
|
||||
}
|
||||
decode(t, resp, &key)
|
||||
|
||||
return user.ID, func(method, path string, body any) *http.Response {
|
||||
t.Helper()
|
||||
var r io.Reader
|
||||
if body != nil {
|
||||
data, _ := json.Marshal(body)
|
||||
r = bytes.NewReader(data)
|
||||
}
|
||||
req, _ := http.NewRequest(method, s.URL+path, r)
|
||||
req.Header.Set("Authorization", "Bearer "+key.Key)
|
||||
if body != nil {
|
||||
req.Header.Set("Content-Type", "application/json")
|
||||
}
|
||||
resp, err := http.DefaultClient.Do(req)
|
||||
if err != nil {
|
||||
t.Fatalf("%s %s: %v", method, path, err)
|
||||
}
|
||||
return resp
|
||||
}
|
||||
}
|
||||
|
||||
// The whole point of the release: a user who is not an administrator cannot
|
||||
// manage other people's accounts. Every one of these was open to any
|
||||
// authenticated caller before.
|
||||
func TestAdmin_MemberIsRefusedAdministration(t *testing.T) {
|
||||
s := newTS(t)
|
||||
memberID, call := member(t, s, "member")
|
||||
|
||||
cases := []struct {
|
||||
name string
|
||||
method string
|
||||
path string
|
||||
body any
|
||||
}{
|
||||
{"create a user", http.MethodPost, "/api/users",
|
||||
map[string]string{"username": "sneaky", "email": "sneaky@test.com"}},
|
||||
{"delete the admin", http.MethodDelete, "/api/users/1", nil},
|
||||
{"grant themselves admin", http.MethodPut, "/api/users/" + id64(memberID) + "/admin",
|
||||
map[string]bool{"is_admin": true}},
|
||||
{"set the admin's password", http.MethodPut, "/api/users/1/password",
|
||||
map[string]string{"password": "hunter2-hunter2"}},
|
||||
{"mint a key for the admin", http.MethodPost, "/api/users/1/api-keys",
|
||||
map[string]string{"name": "borrowed"}},
|
||||
{"retarget the admin's notifications", http.MethodPut, "/api/users/1/notify",
|
||||
map[string]string{"ntfy_topic": "attacker-topic"}},
|
||||
}
|
||||
|
||||
for _, c := range cases {
|
||||
resp := call(c.method, c.path, c.body)
|
||||
resp.Body.Close()
|
||||
if resp.StatusCode != http.StatusForbidden {
|
||||
t.Errorf("%s: expected 403, got %d", c.name, resp.StatusCode)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// Being refused other people's accounts must not cost a user their own.
|
||||
func TestAdmin_MemberKeepsTheirOwnAccount(t *testing.T) {
|
||||
s := newTS(t)
|
||||
memberID, call := member(t, s, "member")
|
||||
self := "/api/users/" + id64(memberID)
|
||||
|
||||
resp := call(http.MethodPut, self+"/notify", map[string]string{"ntfy_topic": "terdut-member"})
|
||||
resp.Body.Close()
|
||||
if resp.StatusCode != http.StatusOK {
|
||||
t.Errorf("own notify target: %d", resp.StatusCode)
|
||||
}
|
||||
|
||||
resp = call(http.MethodPut, self+"/password", map[string]string{"password": "correct-horse-battery"})
|
||||
resp.Body.Close()
|
||||
if resp.StatusCode != http.StatusOK && resp.StatusCode != http.StatusNoContent {
|
||||
t.Errorf("own password: %d", resp.StatusCode)
|
||||
}
|
||||
|
||||
// An API key carries exactly the rights of its owner, so minting your own
|
||||
// is no more than signing in again.
|
||||
resp = call(http.MethodPost, self+"/api-keys", map[string]string{"name": "laptop"})
|
||||
resp.Body.Close()
|
||||
if resp.StatusCode != http.StatusCreated {
|
||||
t.Errorf("own API key: %d", resp.StatusCode)
|
||||
}
|
||||
|
||||
// And the queue still has to be able to name people.
|
||||
resp = call(http.MethodGet, "/api/users", nil)
|
||||
resp.Body.Close()
|
||||
if resp.StatusCode != http.StatusOK {
|
||||
t.Errorf("list users: %d", resp.StatusCode)
|
||||
}
|
||||
}
|
||||
|
||||
// Incident work is everybody's job; none of it is administration.
|
||||
func TestAdmin_MemberCanWorkIncidents(t *testing.T) {
|
||||
s := newTS(t)
|
||||
_, call := member(t, s, "responder")
|
||||
postWebhook(t, s, []map[string]any{
|
||||
amAlert("fp-admin", "DiskFull", "firing", "2026-09-20T10:00:00Z", zeroTime, nil),
|
||||
})
|
||||
|
||||
for _, c := range []struct {
|
||||
name string
|
||||
method string
|
||||
path string
|
||||
}{
|
||||
{"list", http.MethodGet, "/api/incidents"},
|
||||
{"acknowledge", http.MethodPost, "/api/incidents/1/acknowledge"},
|
||||
{"resolve", http.MethodPost, "/api/incidents/1/resolve"},
|
||||
} {
|
||||
resp := call(c.method, c.path, nil)
|
||||
resp.Body.Close()
|
||||
if resp.StatusCode != http.StatusOK {
|
||||
t.Errorf("%s: expected 200, got %d", c.name, resp.StatusCode)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// An install must never be left with nobody who can administer it.
|
||||
func TestAdmin_LastAdministratorIsProtected(t *testing.T) {
|
||||
s := newTS(t)
|
||||
|
||||
resp := s.req(t, http.MethodPut, "/api/users/1/admin", map[string]bool{"is_admin": false})
|
||||
resp.Body.Close()
|
||||
if resp.StatusCode != http.StatusConflict {
|
||||
t.Errorf("self-demotion: expected 409, got %d", resp.StatusCode)
|
||||
}
|
||||
|
||||
resp = s.req(t, http.MethodDelete, "/api/users/1", nil)
|
||||
resp.Body.Close()
|
||||
if resp.StatusCode != http.StatusConflict {
|
||||
t.Errorf("deleting yourself: expected 409, got %d", resp.StatusCode)
|
||||
}
|
||||
|
||||
// With a second administrator the first may stand down, but not while they
|
||||
// are the only one — which is the same rule from the other side.
|
||||
otherID, _ := member(t, s, "second")
|
||||
resp = s.req(t, http.MethodPut, "/api/users/"+id64(otherID)+"/admin", map[string]bool{"is_admin": true})
|
||||
resp.Body.Close()
|
||||
if resp.StatusCode != http.StatusOK {
|
||||
t.Fatalf("granting admin: %d", resp.StatusCode)
|
||||
}
|
||||
|
||||
resp = s.req(t, http.MethodDelete, "/api/users/"+id64(otherID), nil)
|
||||
resp.Body.Close()
|
||||
if resp.StatusCode != http.StatusNoContent {
|
||||
t.Errorf("deleting the second admin: expected 204, got %d", resp.StatusCode)
|
||||
}
|
||||
}
|
||||
|
||||
// A promoted user gets the powers with the flag, and loses them with it.
|
||||
func TestAdmin_GrantAndRevokeChangeWhatIsAllowed(t *testing.T) {
|
||||
s := newTS(t)
|
||||
memberID, call := member(t, s, "promotee")
|
||||
admin := "/api/users/" + id64(memberID) + "/admin"
|
||||
|
||||
resp := call(http.MethodPost, "/api/users", map[string]string{"username": "a", "email": "a@test.com"})
|
||||
resp.Body.Close()
|
||||
if resp.StatusCode != http.StatusForbidden {
|
||||
t.Fatalf("before the grant: %d", resp.StatusCode)
|
||||
}
|
||||
|
||||
resp = s.req(t, http.MethodPut, admin, map[string]bool{"is_admin": true})
|
||||
resp.Body.Close()
|
||||
if resp.StatusCode != http.StatusOK {
|
||||
t.Fatalf("grant: %d", resp.StatusCode)
|
||||
}
|
||||
|
||||
resp = call(http.MethodPost, "/api/users", map[string]string{"username": "b", "email": "b@test.com"})
|
||||
resp.Body.Close()
|
||||
if resp.StatusCode != http.StatusCreated {
|
||||
t.Errorf("after the grant: expected 201, got %d", resp.StatusCode)
|
||||
}
|
||||
|
||||
resp = s.req(t, http.MethodPut, admin, map[string]bool{"is_admin": false})
|
||||
resp.Body.Close()
|
||||
if resp.StatusCode != http.StatusOK {
|
||||
t.Fatalf("revoke: %d", resp.StatusCode)
|
||||
}
|
||||
|
||||
resp = call(http.MethodPost, "/api/users", map[string]string{"username": "c", "email": "c@test.com"})
|
||||
resp.Body.Close()
|
||||
if resp.StatusCode != http.StatusForbidden {
|
||||
t.Errorf("after the revoke: expected 403, got %d", resp.StatusCode)
|
||||
}
|
||||
}
|
||||
|
||||
// The flag has to reach the client, or the web UI cannot decide what to show.
|
||||
func TestAdmin_MeReportsTheFlag(t *testing.T) {
|
||||
s := newTS(t)
|
||||
|
||||
var me struct {
|
||||
User struct {
|
||||
IsAdmin bool `json:"is_admin"`
|
||||
} `json:"user"`
|
||||
}
|
||||
decode(t, s.req(t, http.MethodGet, "/api/me", nil), &me)
|
||||
if !me.User.IsAdmin {
|
||||
t.Error("the bootstrap user should be an administrator")
|
||||
}
|
||||
|
||||
_, call := member(t, s, "plain")
|
||||
var theirs struct {
|
||||
User struct {
|
||||
IsAdmin bool `json:"is_admin"`
|
||||
} `json:"user"`
|
||||
}
|
||||
decode(t, call(http.MethodGet, "/api/me", nil), &theirs)
|
||||
if theirs.User.IsAdmin {
|
||||
t.Error("a created user should not be an administrator")
|
||||
}
|
||||
}
|
||||
@@ -4,9 +4,12 @@ import (
|
||||
"context"
|
||||
"database/sql"
|
||||
"encoding/json"
|
||||
"errors"
|
||||
"log"
|
||||
"net/http"
|
||||
"time"
|
||||
|
||||
"github.com/go-chi/chi/v5"
|
||||
)
|
||||
|
||||
// Values for alerts.resolution_source, recording why an alert left the firing
|
||||
@@ -70,8 +73,49 @@ type ingested struct {
|
||||
deadman bool
|
||||
}
|
||||
|
||||
func handleAlertmanagerWebhook(db *sql.DB, notify NotifyConfig, deadman DeadmanConfig) http.HandlerFunc {
|
||||
// handleIntegrationWebhook receives alerts on a team's own integration key.
|
||||
// The key in the path is both the credential and the routing: it says who may
|
||||
// post, and which team the alerts belong to.
|
||||
func handleIntegrationWebhook(db *sql.DB, notify NotifyConfig) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
teamID, err := teamIDForKey(r.Context(), db, chi.URLParam(r, "key"))
|
||||
if err != nil {
|
||||
if errors.Is(err, errUnknownIntegration) {
|
||||
// 401 and not 404: the path is real, the key is not, and a
|
||||
// sender misconfigured this way should say so in its own logs
|
||||
// rather than believe it is delivering.
|
||||
respond(w, http.StatusUnauthorized, errResp("unknown integration key"))
|
||||
return
|
||||
}
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
receiveWebhook(w, r, db, notify, teamID)
|
||||
}
|
||||
}
|
||||
|
||||
// handleLegacyWebhook is the pre-teams unauthenticated endpoint, kept for one
|
||||
// release so an upgrade does not silently stop delivering while somebody edits
|
||||
// the Alertmanager config. It routes to the oldest team, which on an upgraded
|
||||
// install is the Default team everything was moved into.
|
||||
//
|
||||
// It is deprecated and unauthenticated — anything that can reach the port can
|
||||
// open an incident. Move senders to an integration key and this goes away.
|
||||
func handleLegacyWebhook(db *sql.DB, notify NotifyConfig) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
teamID, err := defaultTeamID(r.Context(), db)
|
||||
if err != nil {
|
||||
log.Printf("legacy webhook: no team to route to: %v", err)
|
||||
w.WriteHeader(http.StatusOK)
|
||||
return
|
||||
}
|
||||
log.Printf("legacy webhook: unauthenticated payload routed to team %d; "+
|
||||
"move this sender to an integration key", teamID)
|
||||
receiveWebhook(w, r, db, notify, teamID)
|
||||
}
|
||||
}
|
||||
|
||||
func receiveWebhook(w http.ResponseWriter, r *http.Request, db *sql.DB, notify NotifyConfig, teamID int64) {
|
||||
var payload amPayload
|
||||
if err := decodeJSON(r, &payload); err != nil {
|
||||
respond(w, http.StatusBadRequest, errResp("invalid payload"))
|
||||
@@ -81,25 +125,32 @@ func handleAlertmanagerWebhook(db *sql.DB, notify NotifyConfig, deadman DeadmanC
|
||||
// Alertmanager retries anything that is not 2xx, and a retry of a payload
|
||||
// we failed to store is more useful than an error it cannot act on — so
|
||||
// failures are logged, not surfaced.
|
||||
if err := ingest(r.Context(), db, notify, deadman, payload); err != nil {
|
||||
log.Printf("webhook ingest (group %q): %v", payload.GroupKey, err)
|
||||
if err := ingest(r.Context(), db, notify, teamID, payload); err != nil {
|
||||
log.Printf("webhook ingest (team %d, group %q): %v", teamID, payload.GroupKey, err)
|
||||
}
|
||||
|
||||
w.WriteHeader(http.StatusOK)
|
||||
}
|
||||
}
|
||||
|
||||
// ingest stores a payload's alerts and reconciles the incident for its group.
|
||||
// The whole payload is one transaction: an incident that opened but whose alerts
|
||||
// failed to link would be a work item nobody could act on.
|
||||
func ingest(ctx context.Context, db *sql.DB, notify NotifyConfig, deadman DeadmanConfig, payload amPayload) error {
|
||||
func ingest(ctx context.Context, db *sql.DB, notify NotifyConfig, teamID int64, payload amPayload) error {
|
||||
tx, err := db.BeginTx(ctx, nil)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
defer tx.Rollback() //nolint:errcheck
|
||||
|
||||
accepted, err := upsertAlerts(ctx, tx, deadman, payload.Alerts)
|
||||
// Which arriving alerts are heartbeats is the team's own answer, read
|
||||
// inside the transaction so an owner editing it mid-payload cannot split
|
||||
// one webhook across two interpretations.
|
||||
deadman, err := deadmanConfigForTeam(ctx, tx, teamID)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
|
||||
accepted, err := upsertAlerts(ctx, tx, deadman, teamID, payload.Alerts)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
@@ -108,7 +159,7 @@ func ingest(ctx context.Context, db *sql.DB, notify NotifyConfig, deadman Deadma
|
||||
// resolution cascade are recomputed once per incident at the end.
|
||||
touched := map[int64]bool{}
|
||||
|
||||
incidentID, err := incidentForGroup(ctx, tx, notify, payload, accepted)
|
||||
incidentID, err := incidentForGroup(ctx, tx, notify, teamID, payload, accepted)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
@@ -156,7 +207,7 @@ func ingest(ctx context.Context, db *sql.DB, notify NotifyConfig, deadman Deadma
|
||||
|
||||
// upsertAlerts stores each alert of a payload and reports what changed. Payloads
|
||||
// the ordering guard rejected are left out entirely.
|
||||
func upsertAlerts(ctx context.Context, tx *sql.Tx, deadman DeadmanConfig, alerts []amAlert) ([]ingested, error) {
|
||||
func upsertAlerts(ctx context.Context, tx *sql.Tx, deadman DeadmanConfig, teamID int64, alerts []amAlert) ([]ingested, error) {
|
||||
now := time.Now().Unix()
|
||||
accepted := make([]ingested, 0, len(alerts))
|
||||
|
||||
@@ -171,7 +222,8 @@ func upsertAlerts(ctx context.Context, tx *sql.Tx, deadman DeadmanConfig, alerts
|
||||
var prevStartsAt int64
|
||||
existed := true
|
||||
switch err := tx.QueryRowContext(ctx,
|
||||
"SELECT status, starts_at FROM alerts WHERE fingerprint = $1", a.Fingerprint,
|
||||
"SELECT status, starts_at FROM alerts WHERE team_id = $1 AND fingerprint = $2",
|
||||
teamID, a.Fingerprint,
|
||||
).Scan(&prevStatus, &prevStartsAt); {
|
||||
case err == sql.ErrNoRows:
|
||||
existed = false
|
||||
@@ -210,10 +262,10 @@ func upsertAlerts(ctx context.Context, tx *sql.Tx, deadman DeadmanConfig, alerts
|
||||
// undone by a stale retry.
|
||||
if _, err := tx.ExecContext(ctx, `
|
||||
INSERT INTO alerts
|
||||
(fingerprint, name, status, labels, annotations, starts_at, ends_at,
|
||||
(team_id, fingerprint, name, status, labels, annotations, starts_at, ends_at,
|
||||
generator_url, received_at, resolution_source)
|
||||
VALUES ($1, $2, $3, $4::jsonb, $5::jsonb, $6, $7, $8, $9, $10)
|
||||
ON CONFLICT (fingerprint) DO UPDATE SET
|
||||
VALUES ($1, $2, $3, $4, $5::jsonb, $6::jsonb, $7, $8, $9, $10, $11)
|
||||
ON CONFLICT (team_id, fingerprint) DO UPDATE SET
|
||||
status = excluded.status,
|
||||
labels = excluded.labels,
|
||||
annotations = excluded.annotations,
|
||||
@@ -233,7 +285,7 @@ func upsertAlerts(ctx context.Context, tx *sql.Tx, deadman DeadmanConfig, alerts
|
||||
OR (excluded.starts_at = alerts.starts_at
|
||||
AND (alerts.resolution_source = '`+resolutionDeadman+`'
|
||||
OR NOT (alerts.status = 'resolved' AND excluded.status = 'firing')))`,
|
||||
a.Fingerprint, name, a.Status,
|
||||
teamID, a.Fingerprint, name, a.Status,
|
||||
string(labelsJSON), string(annotationsJSON),
|
||||
a.StartsAt.Unix(), endsAtUnix,
|
||||
a.GeneratorURL, now, resolutionSource,
|
||||
@@ -245,7 +297,8 @@ func upsertAlerts(ctx context.Context, tx *sql.Tx, deadman DeadmanConfig, alerts
|
||||
var curStatus string
|
||||
var curStartsAt int64
|
||||
if err := tx.QueryRowContext(ctx,
|
||||
"SELECT id, status, starts_at FROM alerts WHERE fingerprint = $1", a.Fingerprint,
|
||||
"SELECT id, status, starts_at FROM alerts WHERE team_id = $1 AND fingerprint = $2",
|
||||
teamID, a.Fingerprint,
|
||||
).Scan(&id, &curStatus, &curStartsAt); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
@@ -284,7 +337,7 @@ func upsertAlerts(ctx context.Context, tx *sql.Tx, deadman DeadmanConfig, alerts
|
||||
// Heartbeats do not count as anything here. A group of nothing but dead man's
|
||||
// switch alerts opens no incident at all, and a mixed group gets an incident for
|
||||
// its real alerts only.
|
||||
func incidentForGroup(ctx context.Context, tx *sql.Tx, notify NotifyConfig, payload amPayload, accepted []ingested) (int64, error) {
|
||||
func incidentForGroup(ctx context.Context, tx *sql.Tx, notify NotifyConfig, teamID int64, payload amPayload, accepted []ingested) (int64, error) {
|
||||
var firstName string
|
||||
anyFiring, anyNew := false, false
|
||||
for _, a := range accepted {
|
||||
@@ -315,7 +368,8 @@ func incidentForGroup(ctx context.Context, tx *sql.Tx, notify NotifyConfig, payl
|
||||
|
||||
var id int64
|
||||
switch err := tx.QueryRowContext(ctx,
|
||||
"SELECT id FROM incidents WHERE group_key = $1 AND resolved_at IS NULL", groupKey,
|
||||
"SELECT id FROM incidents WHERE team_id = $1 AND group_key = $2 AND resolved_at IS NULL",
|
||||
teamID, groupKey,
|
||||
).Scan(&id); {
|
||||
case err == nil:
|
||||
return id, nil
|
||||
@@ -326,7 +380,7 @@ func incidentForGroup(ctx context.Context, tx *sql.Tx, notify NotifyConfig, payl
|
||||
if !anyNew {
|
||||
return 0, nil
|
||||
}
|
||||
return openIncident(ctx, tx, notify, groupKey,
|
||||
return openIncident(ctx, tx, notify, teamID, groupKey,
|
||||
incidentTitle(payload.GroupLabels, firstName), payload.GroupLabels, nil)
|
||||
}
|
||||
|
||||
@@ -338,8 +392,8 @@ func incidentForGroup(ctx context.Context, tx *sql.Tx, notify NotifyConfig, payl
|
||||
// its own. Hence the querier rather than a *sql.Tx. A nil severity leaves the
|
||||
// column for refreshSeverity to fill from the member alerts; the sweeper passes
|
||||
// one because its incidents have no members to derive it from.
|
||||
func openIncident(ctx context.Context, q querier, notify NotifyConfig, groupKey, title string, groupLabels map[string]string, severity *string) (int64, error) {
|
||||
onCall, err := currentOnCall(ctx, q)
|
||||
func openIncident(ctx context.Context, q querier, notify NotifyConfig, teamID int64, groupKey, title string, groupLabels map[string]string, severity *string) (int64, error) {
|
||||
onCall, err := currentOnCall(ctx, q, teamID)
|
||||
if err != nil {
|
||||
return 0, err
|
||||
}
|
||||
@@ -351,10 +405,10 @@ func openIncident(ctx context.Context, q querier, notify NotifyConfig, groupKey,
|
||||
|
||||
var id int64
|
||||
err = q.QueryRowContext(ctx, `
|
||||
INSERT INTO incidents (group_key, title, group_labels, status, severity, triggered_at, assigned_to)
|
||||
VALUES ($1, $2, $3::jsonb, 'triggered', $4, $5, $6)
|
||||
INSERT INTO incidents (team_id, group_key, title, group_labels, status, severity, triggered_at, assigned_to)
|
||||
VALUES ($1, $2, $3, $4::jsonb, 'triggered', $5, $6, $7)
|
||||
RETURNING id`,
|
||||
groupKey, title, string(labelsJSON), severity,
|
||||
teamID, groupKey, title, string(labelsJSON), severity,
|
||||
time.Now().Unix(), onCall).Scan(&id)
|
||||
if err != nil {
|
||||
return 0, err
|
||||
|
||||
+15
-6
@@ -19,7 +19,7 @@ import (
|
||||
// incident_alerts rather than as a column here, because one alert row is reused
|
||||
// across occurrences and belongs to a different incident each time.
|
||||
const alertSelectFrom = `
|
||||
SELECT a.id, a.fingerprint, a.name, a.status,
|
||||
SELECT a.id, a.team_id, t.name, a.fingerprint, a.name, a.status,
|
||||
a.labels, a.annotations,
|
||||
a.starts_at, a.ends_at, a.generator_url, a.received_at,
|
||||
(SELECT ia.incident_id
|
||||
@@ -29,7 +29,8 @@ const alertSelectFrom = `
|
||||
ORDER BY i.triggered_at DESC, i.id DESC
|
||||
LIMIT 1),
|
||||
a.resolution_source, a.archived_at
|
||||
FROM alerts a`
|
||||
FROM alerts a
|
||||
JOIN teams t ON t.id = a.team_id`
|
||||
|
||||
func handleListAlerts(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
@@ -38,6 +39,13 @@ func handleListAlerts(db *sql.DB) http.HandlerFunc {
|
||||
where := []string{}
|
||||
args := &sqlArgs{}
|
||||
|
||||
where = append(where, "a.team_id = ANY("+args.add(callerTeamIDs(r.Context()))+")")
|
||||
if team := q.Get("team_id"); team != "" {
|
||||
if n, err := strconv.ParseInt(team, 10, 64); err == nil {
|
||||
where = append(where, "a.team_id = "+args.add(n))
|
||||
}
|
||||
}
|
||||
|
||||
if status := q.Get("status"); status != "" {
|
||||
where = append(where, "a.status = "+args.add(status))
|
||||
}
|
||||
@@ -107,7 +115,7 @@ func handleGetAlert(db *sql.DB) http.HandlerFunc {
|
||||
respond(w, http.StatusBadRequest, errResp("invalid alert id"))
|
||||
return
|
||||
}
|
||||
a, err := fetchAlert(r.Context(), db, id)
|
||||
a, err := fetchAlert(r.Context(), db, id, callerTeamIDs(r.Context()))
|
||||
if err == sql.ErrNoRows {
|
||||
respond(w, http.StatusNotFound, errResp("alert not found"))
|
||||
return
|
||||
@@ -121,8 +129,9 @@ func handleGetAlert(db *sql.DB) http.HandlerFunc {
|
||||
}
|
||||
|
||||
// fetchAlert loads a single alert by ID using the shared query.
|
||||
func fetchAlert(ctx context.Context, db *sql.DB, id int64) (models.Alert, error) {
|
||||
return scanAlert(db.QueryRowContext(ctx, alertSelectFrom+" WHERE a.id = $1", id))
|
||||
func fetchAlert(ctx context.Context, db *sql.DB, id int64, teamIDs []int64) (models.Alert, error) {
|
||||
return scanAlert(db.QueryRowContext(ctx,
|
||||
alertSelectFrom+" WHERE a.id = $1 AND a.team_id = ANY($2)", id, teamIDs))
|
||||
}
|
||||
|
||||
// scanner is satisfied by both *sql.Row and *sql.Rows.
|
||||
@@ -137,7 +146,7 @@ func scanAlert(s scanner) (models.Alert, error) {
|
||||
var endsAtUnix, archivedAtUnix *int64
|
||||
|
||||
if err := s.Scan(
|
||||
&a.ID, &a.Fingerprint, &a.Name, &a.Status,
|
||||
&a.ID, &a.TeamID, &a.TeamName, &a.Fingerprint, &a.Name, &a.Status,
|
||||
&labelsJSON, &annotationsJSON,
|
||||
&startsAtUnix, &endsAtUnix,
|
||||
&a.GeneratorURL, &receivedAtUnix,
|
||||
|
||||
+68
-21
@@ -9,6 +9,8 @@ import (
|
||||
"io"
|
||||
"net/http"
|
||||
"net/http/httptest"
|
||||
"sort"
|
||||
"strings"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
@@ -37,7 +39,7 @@ func newTS(t *testing.T, notify ...api.NotifyConfig) *ts {
|
||||
return newDeadmanTS(t, api.DeadmanConfig{}, cfg)
|
||||
}
|
||||
|
||||
// newDeadmanTS is newTS with dead man's switch handling configured.
|
||||
// newDeadmanTS is newTS with the default team's dead man's switches configured.
|
||||
func newDeadmanTS(t *testing.T, deadman api.DeadmanConfig, notify ...api.NotifyConfig) *ts {
|
||||
t.Helper()
|
||||
var cfg api.NotifyConfig
|
||||
@@ -46,7 +48,7 @@ func newDeadmanTS(t *testing.T, deadman api.DeadmanConfig, notify ...api.NotifyC
|
||||
}
|
||||
|
||||
database := newTestDB(t)
|
||||
srv := httptest.NewServer(api.NewRouter(database, cfg, deadman))
|
||||
srv := httptest.NewServer(api.NewRouter(database, cfg))
|
||||
t.Cleanup(srv.Close)
|
||||
|
||||
body, _ := json.Marshal(map[string]string{"username": "admin", "email": "admin@test.com"})
|
||||
@@ -62,7 +64,38 @@ func newDeadmanTS(t *testing.T, deadman api.DeadmanConfig, notify ...api.NotifyC
|
||||
json.NewDecoder(resp.Body).Decode(&result)
|
||||
key := result["api_key"].(map[string]any)["key"].(string)
|
||||
|
||||
return &ts{Server: srv, key: key, db: database, notify: cfg, deadman: deadman}
|
||||
s := &ts{Server: srv, key: key, db: database, notify: cfg, deadman: deadman}
|
||||
|
||||
// Dead man's switches belong to a team now, so a test that wants them
|
||||
// configures the default team the way an owner would.
|
||||
if deadman.Timeout > 0 {
|
||||
setTeamDeadman(t, s, deadman)
|
||||
}
|
||||
return s
|
||||
}
|
||||
|
||||
// setTeamDeadman configures the default team's switches over the API, rendering
|
||||
// the matchers back into the string form the endpoint takes.
|
||||
func setTeamDeadman(t *testing.T, s *ts, cfg api.DeadmanConfig) {
|
||||
t.Helper()
|
||||
matchers := make([]string, 0, len(cfg.Matchers))
|
||||
for _, m := range cfg.Matchers {
|
||||
parts := []string{"alertname=" + m.Name}
|
||||
for k, v := range m.Labels {
|
||||
parts = append(parts, k+"="+v)
|
||||
}
|
||||
sort.Strings(parts[1:])
|
||||
matchers = append(matchers, strings.Join(parts, ","))
|
||||
}
|
||||
resp := s.req(t, http.MethodPut, "/api/teams/"+defaultTeam+"/deadman", map[string]any{
|
||||
"matchers": strings.Join(matchers, "; "),
|
||||
"timeout_seconds": int64(cfg.Timeout.Seconds()),
|
||||
"severity": cfg.Severity,
|
||||
})
|
||||
defer resp.Body.Close()
|
||||
if resp.StatusCode != http.StatusOK {
|
||||
t.Fatalf("configure the team's dead man's switches: %d", resp.StatusCode)
|
||||
}
|
||||
}
|
||||
|
||||
// exec runs a statement against the test database.
|
||||
@@ -276,14 +309,14 @@ func TestAlertUpsert_DifferentFingerprintsStored(t *testing.T) {
|
||||
func TestSchedule_ConflictOnSameDate(t *testing.T) {
|
||||
s := newTS(t)
|
||||
|
||||
first := s.req(t, http.MethodPost, "/api/schedule",
|
||||
first := s.req(t, http.MethodPost, "/api/teams/"+defaultTeam+"/schedule",
|
||||
map[string]any{"user_id": 1, "dates": []string{"2026-06-01"}})
|
||||
if first.StatusCode != http.StatusCreated {
|
||||
t.Fatalf("first assignment returned %d", first.StatusCode)
|
||||
}
|
||||
first.Body.Close()
|
||||
|
||||
second := s.req(t, http.MethodPost, "/api/schedule",
|
||||
second := s.req(t, http.MethodPost, "/api/teams/"+defaultTeam+"/schedule",
|
||||
map[string]any{"user_id": 1, "dates": []string{"2026-06-01"}})
|
||||
if second.StatusCode != http.StatusConflict {
|
||||
t.Errorf("expected 409 on duplicate date, got %d", second.StatusCode)
|
||||
@@ -295,11 +328,11 @@ func TestSchedule_MultiDateRollbackOnConflict(t *testing.T) {
|
||||
s := newTS(t)
|
||||
|
||||
// Claim 2026-06-10 first.
|
||||
s.req(t, http.MethodPost, "/api/schedule",
|
||||
s.req(t, http.MethodPost, "/api/teams/"+defaultTeam+"/schedule",
|
||||
map[string]any{"user_id": 1, "dates": []string{"2026-06-10"}}).Body.Close()
|
||||
|
||||
// Try to assign two dates in one request where the second conflicts.
|
||||
resp := s.req(t, http.MethodPost, "/api/schedule",
|
||||
resp := s.req(t, http.MethodPost, "/api/teams/"+defaultTeam+"/schedule",
|
||||
map[string]any{"user_id": 1, "dates": []string{"2026-06-09", "2026-06-10"}})
|
||||
if resp.StatusCode != http.StatusConflict {
|
||||
t.Fatalf("expected 409, got %d", resp.StatusCode)
|
||||
@@ -307,7 +340,7 @@ func TestSchedule_MultiDateRollbackOnConflict(t *testing.T) {
|
||||
resp.Body.Close()
|
||||
|
||||
// 2026-06-09 must NOT have been committed (transaction rolled back).
|
||||
listResp := s.req(t, http.MethodGet, "/api/schedule?from=2026-06-09&to=2026-06-09", nil)
|
||||
listResp := s.req(t, http.MethodGet, "/api/teams/"+defaultTeam+"/schedule?from=2026-06-09&to=2026-06-09", nil)
|
||||
var entries []any
|
||||
decode(t, listResp, &entries)
|
||||
if len(entries) != 0 {
|
||||
@@ -321,21 +354,35 @@ func TestSchedule_MultiDateRollbackOnConflict(t *testing.T) {
|
||||
|
||||
// addUser creates a second person to hand a shift to. The bootstrap user is
|
||||
// admin, id 1.
|
||||
// addUser creates a user and puts them in the default team, because a user who
|
||||
// is in no team can be paged by nobody and take no shift — which is the rule
|
||||
// these tests exercise around, not the one they are testing.
|
||||
func addUser(t *testing.T, s *ts, username string) {
|
||||
t.Helper()
|
||||
resp := s.req(t, http.MethodPost, "/api/users",
|
||||
map[string]any{"username": username, "email": username + "@test.com"})
|
||||
defer resp.Body.Close()
|
||||
if resp.StatusCode != http.StatusCreated {
|
||||
resp.Body.Close()
|
||||
t.Fatalf("create user returned %d", resp.StatusCode)
|
||||
}
|
||||
var user struct {
|
||||
ID int64 `json:"id"`
|
||||
}
|
||||
decode(t, resp, &user)
|
||||
|
||||
member := s.req(t, http.MethodPost, "/api/teams/"+defaultTeam+"/members",
|
||||
map[string]any{"user_id": user.ID, "role": "member"})
|
||||
defer member.Body.Close()
|
||||
if member.StatusCode != http.StatusNoContent {
|
||||
t.Fatalf("add %s to the team returned %d", username, member.StatusCode)
|
||||
}
|
||||
}
|
||||
|
||||
// scheduleHolder reports who is on call for one date, or "" for nobody.
|
||||
func scheduleHolder(t *testing.T, s *ts, date string) string {
|
||||
t.Helper()
|
||||
var entries []map[string]any
|
||||
decode(t, s.req(t, http.MethodGet, "/api/schedule?from="+date+"&to="+date, nil), &entries)
|
||||
decode(t, s.req(t, http.MethodGet, "/api/teams/"+defaultTeam+"/schedule?from="+date+"&to="+date, nil), &entries)
|
||||
if len(entries) == 0 {
|
||||
return ""
|
||||
}
|
||||
@@ -347,10 +394,10 @@ func TestSchedule_ReplaceTakesAnAssignedDate(t *testing.T) {
|
||||
s := newTS(t)
|
||||
addUser(t, s, "alex")
|
||||
|
||||
s.req(t, http.MethodPost, "/api/schedule",
|
||||
s.req(t, http.MethodPost, "/api/teams/"+defaultTeam+"/schedule",
|
||||
map[string]any{"user_id": 1, "dates": []string{"2026-06-01"}}).Body.Close()
|
||||
|
||||
resp := s.req(t, http.MethodPost, "/api/schedule",
|
||||
resp := s.req(t, http.MethodPost, "/api/teams/"+defaultTeam+"/schedule",
|
||||
map[string]any{"user_id": 2, "dates": []string{"2026-06-01"}, "replace": true})
|
||||
if resp.StatusCode != http.StatusCreated {
|
||||
t.Fatalf("expected replace to succeed, got %d", resp.StatusCode)
|
||||
@@ -364,7 +411,7 @@ func TestSchedule_ReplaceTakesAnAssignedDate(t *testing.T) {
|
||||
// One row, not two: two entries for a date would mean two people believing
|
||||
// they are on call for it.
|
||||
var entries []map[string]any
|
||||
decode(t, s.req(t, http.MethodGet, "/api/schedule?from=2026-06-01&to=2026-06-01", nil), &entries)
|
||||
decode(t, s.req(t, http.MethodGet, "/api/teams/"+defaultTeam+"/schedule?from=2026-06-01&to=2026-06-01", nil), &entries)
|
||||
if len(entries) != 1 {
|
||||
t.Errorf("expected exactly one entry for the date, got %d", len(entries))
|
||||
}
|
||||
@@ -376,11 +423,11 @@ func TestSchedule_ReplaceMixedWeek(t *testing.T) {
|
||||
s := newTS(t)
|
||||
addUser(t, s, "alex")
|
||||
|
||||
s.req(t, http.MethodPost, "/api/schedule",
|
||||
s.req(t, http.MethodPost, "/api/teams/"+defaultTeam+"/schedule",
|
||||
map[string]any{"user_id": 1, "dates": []string{"2026-06-02", "2026-06-04"}}).Body.Close()
|
||||
|
||||
week := []string{"2026-06-01", "2026-06-02", "2026-06-03", "2026-06-04", "2026-06-05"}
|
||||
resp := s.req(t, http.MethodPost, "/api/schedule",
|
||||
resp := s.req(t, http.MethodPost, "/api/teams/"+defaultTeam+"/schedule",
|
||||
map[string]any{"user_id": 2, "dates": week, "replace": true})
|
||||
if resp.StatusCode != http.StatusCreated {
|
||||
t.Fatalf("expected the mixed week to succeed, got %d", resp.StatusCode)
|
||||
@@ -399,10 +446,10 @@ func TestSchedule_ReplaceDefaultsOff(t *testing.T) {
|
||||
s := newTS(t)
|
||||
addUser(t, s, "alex")
|
||||
|
||||
s.req(t, http.MethodPost, "/api/schedule",
|
||||
s.req(t, http.MethodPost, "/api/teams/"+defaultTeam+"/schedule",
|
||||
map[string]any{"user_id": 1, "dates": []string{"2026-06-01"}}).Body.Close()
|
||||
|
||||
resp := s.req(t, http.MethodPost, "/api/schedule",
|
||||
resp := s.req(t, http.MethodPost, "/api/teams/"+defaultTeam+"/schedule",
|
||||
map[string]any{"user_id": 2, "dates": []string{"2026-06-01"}})
|
||||
if resp.StatusCode != http.StatusConflict {
|
||||
t.Fatalf("expected 409 without replace, got %d", resp.StatusCode)
|
||||
@@ -420,7 +467,7 @@ func TestSchedule_ReplaceDefaultsOff(t *testing.T) {
|
||||
func TestSchedule_ReplaceCollapsesRepeatedDates(t *testing.T) {
|
||||
s := newTS(t)
|
||||
|
||||
resp := s.req(t, http.MethodPost, "/api/schedule",
|
||||
resp := s.req(t, http.MethodPost, "/api/teams/"+defaultTeam+"/schedule",
|
||||
map[string]any{"user_id": 1, "dates": []string{"2026-06-01", "2026-06-01"}, "replace": true})
|
||||
if resp.StatusCode != http.StatusCreated {
|
||||
t.Fatalf("expected a repeated date to be accepted under replace, got %d", resp.StatusCode)
|
||||
@@ -428,7 +475,7 @@ func TestSchedule_ReplaceCollapsesRepeatedDates(t *testing.T) {
|
||||
resp.Body.Close()
|
||||
|
||||
var entries []map[string]any
|
||||
decode(t, s.req(t, http.MethodGet, "/api/schedule?from=2026-06-01&to=2026-06-01", nil), &entries)
|
||||
decode(t, s.req(t, http.MethodGet, "/api/teams/"+defaultTeam+"/schedule?from=2026-06-01&to=2026-06-01", nil), &entries)
|
||||
if len(entries) != 1 {
|
||||
t.Errorf("expected one entry for the repeated date, got %d", len(entries))
|
||||
}
|
||||
@@ -498,7 +545,7 @@ func TestArchive_AlertListFilter(t *testing.T) {
|
||||
}
|
||||
|
||||
// 2. Let the sweeper archive it: ends_at is already well past archiveAfter.
|
||||
api.Sweep(context.Background(), s.db, time.Hour, 6*time.Hour, s.deadman, s.notify)
|
||||
api.Sweep(context.Background(), s.db, time.Hour, 6*time.Hour, s.notify)
|
||||
|
||||
// 3. Default list excludes it.
|
||||
decode(t, s.req(t, http.MethodGet, "/api/alerts", nil), &alerts)
|
||||
@@ -539,7 +586,7 @@ func postAlert(t *testing.T, s *ts, fingerprint, status, startsAt, endsAt string
|
||||
|
||||
func sweep(t *testing.T, s *ts, staleAfter time.Duration) {
|
||||
t.Helper()
|
||||
api.Sweep(context.Background(), s.db, noArchive, staleAfter, s.deadman, s.notify)
|
||||
api.Sweep(context.Background(), s.db, noArchive, staleAfter, s.notify)
|
||||
}
|
||||
|
||||
// A firing alert Alertmanager stopped refreshing is resolved via the
|
||||
|
||||
@@ -19,15 +19,15 @@ const (
|
||||
|
||||
// StartArchiver runs the alert sweeper until ctx is cancelled, starting with an
|
||||
// immediate pass so a restart reconciles state right away.
|
||||
func StartArchiver(ctx context.Context, db *sql.DB, archiveAfter, staleAfter time.Duration, deadman DeadmanConfig, notify NotifyConfig) {
|
||||
func StartArchiver(ctx context.Context, db *sql.DB, archiveAfter, staleAfter time.Duration, notify NotifyConfig) {
|
||||
ticker := time.NewTicker(sweepInterval)
|
||||
defer ticker.Stop()
|
||||
|
||||
Sweep(ctx, db, archiveAfter, staleAfter, deadman, notify)
|
||||
Sweep(ctx, db, archiveAfter, staleAfter, notify)
|
||||
for {
|
||||
select {
|
||||
case <-ticker.C:
|
||||
Sweep(ctx, db, archiveAfter, staleAfter, deadman, notify)
|
||||
Sweep(ctx, db, archiveAfter, staleAfter, notify)
|
||||
case <-ctx.Done():
|
||||
return
|
||||
}
|
||||
@@ -44,8 +44,8 @@ func StartArchiver(ctx context.Context, db *sql.DB, archiveAfter, staleAfter tim
|
||||
// touch: a heartbeat answers to its own, much tighter, timeout, and the generic
|
||||
// staleness rules would otherwise resolve it as 'expiry' long before that.
|
||||
// Exported so tests can drive a pass without waiting on the ticker.
|
||||
func Sweep(ctx context.Context, db *sql.DB, archiveAfter, staleAfter time.Duration, deadman DeadmanConfig, notify NotifyConfig) {
|
||||
heartbeats := sweepDeadman(ctx, db, deadman, notify)
|
||||
func Sweep(ctx context.Context, db *sql.DB, archiveAfter, staleAfter time.Duration, notify NotifyConfig) {
|
||||
heartbeats := sweepDeadman(ctx, db, notify)
|
||||
expireStale(ctx, db, staleAfter, heartbeats)
|
||||
resolveSettledIncidents(ctx, db)
|
||||
archiveResolved(ctx, db, archiveAfter)
|
||||
|
||||
@@ -266,8 +266,9 @@ func handleMe(db *sql.DB) http.HandlerFunc {
|
||||
//
|
||||
// Changing your own password takes the current one, when there is one, so an
|
||||
// unattended signed-in browser cannot be used to take the account over. Setting
|
||||
// somebody else's is how an admin gives a user their first password, and like
|
||||
// the other user endpoints it is open to any authenticated caller.
|
||||
// somebody else's is how an admin gives a user their first password, and is
|
||||
// restricted to administrators: it hands over an account outright, without
|
||||
// knowing the password it replaces.
|
||||
//
|
||||
// Every other session of the target is ended: a password change is what you
|
||||
// do when you think someone else is signed in.
|
||||
@@ -278,6 +279,9 @@ func handleSetPassword(db *sql.DB) http.HandlerFunc {
|
||||
respond(w, http.StatusBadRequest, errResp("invalid user id"))
|
||||
return
|
||||
}
|
||||
if !requireSelfOrAdmin(w, r, id) {
|
||||
return
|
||||
}
|
||||
var req struct {
|
||||
Password string `json:"password"`
|
||||
CurrentPassword string `json:"current_password"`
|
||||
|
||||
@@ -248,7 +248,7 @@ func TestSession_ExpiredIsRejected(t *testing.T) {
|
||||
if code := status(t, b.do(t, http.MethodGet, "/api/me", nil)); code != http.StatusUnauthorized {
|
||||
t.Errorf("expired session: %d", code)
|
||||
}
|
||||
api.Sweep(t.Context(), s.db, 0, 0, api.DeadmanConfig{}, api.NotifyConfig{})
|
||||
api.Sweep(t.Context(), s.db, 0, 0, api.NotifyConfig{})
|
||||
var n int
|
||||
s.db.QueryRow("SELECT COUNT(*) FROM sessions").Scan(&n)
|
||||
if n != 0 {
|
||||
@@ -301,7 +301,7 @@ func TestSetPassword_EndsOtherSessionsButNotThisOne(t *testing.T) {
|
||||
|
||||
func TestBootstrap_WithPassword(t *testing.T) {
|
||||
database := newTestDB(t)
|
||||
srv := httptest.NewServer(api.NewRouter(database, api.NotifyConfig{}, api.DeadmanConfig{}))
|
||||
srv := httptest.NewServer(api.NewRouter(database, api.NotifyConfig{}))
|
||||
t.Cleanup(srv.Close)
|
||||
|
||||
body := `{"username":"admin","email":"a@test.com","password":"` + adminPassword + `"}`
|
||||
|
||||
+139
-17
@@ -171,6 +171,7 @@ func ParseDeadmanConfig(matchers string, timeout time.Duration, severity string)
|
||||
// deadmanAlert is one switch: the alert row carrying its last heartbeat.
|
||||
type deadmanAlert struct {
|
||||
id int64
|
||||
teamID int64
|
||||
fingerprint string
|
||||
labels map[string]string
|
||||
matcher DeadmanMatcher
|
||||
@@ -188,27 +189,33 @@ func (a deadmanAlert) groupKey() string { return deadmanGroupPrefix + a.fingerpr
|
||||
// It returns the ids of the alerts it owns, because the generic staleness
|
||||
// expiry must leave them alone — staleAfter and ends_at would otherwise resolve
|
||||
// a heartbeat long before its own, much tighter, timeout ever fired.
|
||||
func sweepDeadman(ctx context.Context, db *sql.DB, cfg DeadmanConfig, notify NotifyConfig) map[int64]bool {
|
||||
// Each team is swept against its own configuration: its own matchers, its own
|
||||
// timeout, its own severity. A team watching nothing is skipped entirely, which
|
||||
// is most of them.
|
||||
func sweepDeadman(ctx context.Context, db *sql.DB, notify NotifyConfig) map[int64]bool {
|
||||
owned := map[int64]bool{}
|
||||
if !cfg.enabled() {
|
||||
return owned
|
||||
}
|
||||
|
||||
switches, err := deadmanAlerts(ctx, db, cfg)
|
||||
configs, err := deadmanConfigs(ctx, db)
|
||||
if err != nil {
|
||||
log.Printf("deadman: load switches: %v", err)
|
||||
log.Printf("deadman: load configs: %v", err)
|
||||
return owned
|
||||
}
|
||||
|
||||
now := time.Now()
|
||||
for teamID, cfg := range configs {
|
||||
switches, err := deadmanAlerts(ctx, db, teamID, cfg)
|
||||
if err != nil {
|
||||
log.Printf("deadman: load switches for team %d: %v", teamID, err)
|
||||
continue
|
||||
}
|
||||
cutoff := now.Add(-cfg.Timeout).Unix()
|
||||
|
||||
for _, sw := range switches {
|
||||
owned[sw.id] = true
|
||||
|
||||
// An explicit resolved from Alertmanager is a stronger death signal than
|
||||
// mere absence: the sender is telling us the heartbeat stopped, so there
|
||||
// is nothing left to wait out.
|
||||
// An explicit resolved from Alertmanager is a stronger death signal
|
||||
// than mere absence: the sender is telling us the heartbeat
|
||||
// stopped, so there is nothing left to wait out.
|
||||
if sw.resolved || sw.receivedAt < cutoff {
|
||||
if err := deadmanDied(ctx, db, cfg, notify, sw, now); err != nil {
|
||||
log.Printf("deadman: open incident for %s: %v", sw.matcher.Name, err)
|
||||
@@ -219,6 +226,7 @@ func sweepDeadman(ctx context.Context, db *sql.DB, cfg DeadmanConfig, notify Not
|
||||
log.Printf("deadman: resolve incident for %s: %v", sw.matcher.Name, err)
|
||||
}
|
||||
}
|
||||
}
|
||||
return owned
|
||||
}
|
||||
|
||||
@@ -227,7 +235,7 @@ func sweepDeadman(ctx context.Context, db *sql.DB, cfg DeadmanConfig, notify Not
|
||||
// happens in Go, which keeps one implementation of the rules. The rows are read
|
||||
// in full before the caller writes, so the writes do not run against an open
|
||||
// cursor over the same table.
|
||||
func deadmanAlerts(ctx context.Context, db *sql.DB, cfg DeadmanConfig) ([]deadmanAlert, error) {
|
||||
func deadmanAlerts(ctx context.Context, db *sql.DB, teamID int64, cfg DeadmanConfig) ([]deadmanAlert, error) {
|
||||
names := cfg.names()
|
||||
args := &sqlArgs{}
|
||||
nameList := make([]any, len(names))
|
||||
@@ -236,9 +244,10 @@ func deadmanAlerts(ctx context.Context, db *sql.DB, cfg DeadmanConfig) ([]deadma
|
||||
}
|
||||
|
||||
rows, err := db.QueryContext(ctx, `
|
||||
SELECT id, fingerprint, labels, status, received_at
|
||||
SELECT id, team_id, fingerprint, labels, status, received_at
|
||||
FROM alerts
|
||||
WHERE name IN (`+args.addList(nameList)+`)
|
||||
WHERE team_id = `+args.add(teamID)+`
|
||||
AND name IN (`+args.addList(nameList)+`)
|
||||
AND archived_at IS NULL`, args.all()...)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
@@ -249,7 +258,7 @@ func deadmanAlerts(ctx context.Context, db *sql.DB, cfg DeadmanConfig) ([]deadma
|
||||
for rows.Next() {
|
||||
var a deadmanAlert
|
||||
var labelsJSON, status string
|
||||
if err := rows.Scan(&a.id, &a.fingerprint, &labelsJSON, &status, &a.receivedAt); err != nil {
|
||||
if err := rows.Scan(&a.id, &a.teamID, &a.fingerprint, &labelsJSON, &status, &a.receivedAt); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
json.Unmarshal([]byte(labelsJSON), &a.labels) //nolint:errcheck
|
||||
@@ -280,8 +289,8 @@ func deadmanDied(ctx context.Context, db *sql.DB, cfg DeadmanConfig, notify Noti
|
||||
if err := db.QueryRowContext(ctx, `
|
||||
SELECT COALESCE(MAX(triggered_at), 0),
|
||||
COUNT(*) FILTER (WHERE resolved_at IS NULL)
|
||||
FROM incidents WHERE group_key = $1`,
|
||||
sw.groupKey()).Scan(&lastTriggered, &open); err != nil {
|
||||
FROM incidents WHERE team_id = $1 AND group_key = $2`,
|
||||
sw.teamID, sw.groupKey()).Scan(&lastTriggered, &open); err != nil {
|
||||
return err
|
||||
}
|
||||
if open > 0 || sw.receivedAt <= lastTriggered {
|
||||
@@ -314,7 +323,9 @@ func deadmanDied(ctx context.Context, db *sql.DB, cfg DeadmanConfig, notify Noti
|
||||
sev = &severity
|
||||
}
|
||||
|
||||
incidentID, err := openIncident(ctx, tx, notify, sw.groupKey(),
|
||||
// The incident opens in the team whose integration received the heartbeat:
|
||||
// the switch belongs to whoever is watching that source, not to the install.
|
||||
incidentID, err := openIncident(ctx, tx, notify, sw.teamID, sw.groupKey(),
|
||||
"No heartbeat from "+sw.matcher.String(), sw.labels, sev)
|
||||
if err != nil {
|
||||
return err
|
||||
@@ -343,7 +354,8 @@ func deadmanRecovered(ctx context.Context, db *sql.DB, sw deadmanAlert) error {
|
||||
var incidentID int64
|
||||
switch err := db.QueryRowContext(ctx, `
|
||||
SELECT id FROM incidents
|
||||
WHERE group_key = $1 AND resolved_at IS NULL`, sw.groupKey()).Scan(&incidentID); {
|
||||
WHERE team_id = $1 AND group_key = $2 AND resolved_at IS NULL`,
|
||||
sw.teamID, sw.groupKey()).Scan(&incidentID); {
|
||||
case err == sql.ErrNoRows:
|
||||
return nil
|
||||
case err != nil:
|
||||
@@ -378,3 +390,113 @@ func deadmanRecovered(ctx context.Context, db *sql.DB, sw deadmanAlert) error {
|
||||
log.Printf("deadman: %s is back, resolved incident %d", sw.matcher.String(), incidentID)
|
||||
return nil
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Per-team configuration
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
// deadmanConfigForTeam reads one team's switches. A team with no row, or with
|
||||
// nothing configured, gets a disabled config — which is the right answer rather
|
||||
// than an error: most teams watch no heartbeat at all.
|
||||
func deadmanConfigForTeam(ctx context.Context, q querier, teamID int64) (DeadmanConfig, error) {
|
||||
var matchers, severity string
|
||||
var timeout int64
|
||||
err := q.QueryRowContext(ctx,
|
||||
"SELECT matchers, timeout_seconds, severity FROM deadman_configs WHERE team_id = $1",
|
||||
teamID).Scan(&matchers, &timeout, &severity)
|
||||
if err == sql.ErrNoRows {
|
||||
return DeadmanConfig{}, nil
|
||||
}
|
||||
if err != nil {
|
||||
return DeadmanConfig{}, err
|
||||
}
|
||||
return parseDeadmanQuietly(matchers, time.Duration(timeout)*time.Second, severity), nil
|
||||
}
|
||||
|
||||
// deadmanConfigs reads every team's switches in one query, for the sweeper.
|
||||
func deadmanConfigs(ctx context.Context, db *sql.DB) (map[int64]DeadmanConfig, error) {
|
||||
rows, err := db.QueryContext(ctx,
|
||||
"SELECT team_id, matchers, timeout_seconds, severity FROM deadman_configs")
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
defer rows.Close()
|
||||
|
||||
out := map[int64]DeadmanConfig{}
|
||||
for rows.Next() {
|
||||
var teamID, timeout int64
|
||||
var matchers, severity string
|
||||
if err := rows.Scan(&teamID, &matchers, &timeout, &severity); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
cfg := parseDeadmanQuietly(matchers, time.Duration(timeout)*time.Second, severity)
|
||||
if cfg.enabled() {
|
||||
out[teamID] = cfg
|
||||
}
|
||||
}
|
||||
return out, rows.Err()
|
||||
}
|
||||
|
||||
// SeedDeadmanConfigs gives every team without a row the server's environment
|
||||
// configuration, so the install that upgrades into per-team switches keeps
|
||||
// watching exactly what it was watching before.
|
||||
//
|
||||
// Idempotent, and never overwrites: once a team has a row it owns its own
|
||||
// configuration, and a redeploy must not quietly put the environment's value
|
||||
// back over an owner's edit.
|
||||
//
|
||||
// A team created after startup gets no row and therefore watches nothing until
|
||||
// its owner says otherwise. That is deliberate: inheriting an install-wide
|
||||
// heartbeat would page a new team about a source it has never heard of, and a
|
||||
// switch nobody chose is the kind that gets muted rather than fixed.
|
||||
func SeedDeadmanConfigs(ctx context.Context, db *sql.DB, cfg DeadmanConfig) error {
|
||||
matchers := make([]string, 0, len(cfg.Matchers))
|
||||
for _, m := range cfg.Matchers {
|
||||
parts := []string{"alertname=" + m.Name}
|
||||
for k, v := range m.Labels {
|
||||
parts = append(parts, k+"="+v)
|
||||
}
|
||||
sort.Strings(parts[1:])
|
||||
matchers = append(matchers, strings.Join(parts, ","))
|
||||
}
|
||||
|
||||
_, err := db.ExecContext(ctx, `
|
||||
INSERT INTO deadman_configs (team_id, matchers, timeout_seconds, severity)
|
||||
SELECT id, $1, $2, $3 FROM teams
|
||||
ON CONFLICT (team_id) DO NOTHING`,
|
||||
strings.Join(matchers, "; "), int64(cfg.Timeout.Seconds()), cfg.Severity)
|
||||
return err
|
||||
}
|
||||
|
||||
// parseDeadmanQuietly is ParseDeadmanConfig without the startup logging: a
|
||||
// team's configuration is read on every sweep and every webhook, and logging it
|
||||
// each time would bury everything else.
|
||||
func parseDeadmanQuietly(matchers string, timeout time.Duration, severity string) DeadmanConfig {
|
||||
cfg := DeadmanConfig{Timeout: timeout, Severity: severity}
|
||||
for _, entry := range strings.Split(matchers, ";") {
|
||||
entry = strings.TrimSpace(entry)
|
||||
if entry == "" {
|
||||
continue
|
||||
}
|
||||
m := DeadmanMatcher{Labels: map[string]string{}}
|
||||
malformed := false
|
||||
for _, cond := range strings.Split(entry, ",") {
|
||||
k, v, ok := strings.Cut(cond, "=")
|
||||
k, v = strings.TrimSpace(k), strings.TrimSpace(v)
|
||||
if !ok || k == "" || v == "" {
|
||||
malformed = true
|
||||
break
|
||||
}
|
||||
if k == "alertname" {
|
||||
m.Name = v
|
||||
continue
|
||||
}
|
||||
m.Labels[k] = v
|
||||
}
|
||||
if malformed || m.Name == "" {
|
||||
continue
|
||||
}
|
||||
cfg.Matchers = append(cfg.Matchers, m)
|
||||
}
|
||||
return cfg
|
||||
}
|
||||
|
||||
@@ -2,6 +2,7 @@ package api_test
|
||||
|
||||
import (
|
||||
"net/http"
|
||||
"strings"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
@@ -469,3 +470,123 @@ func TestDeadman_DisabledConfigIsInert(t *testing.T) {
|
||||
t.Errorf("expected the generic sweeper to own the alert, got %v", source)
|
||||
}
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Per-team configuration
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
// Each team decides for itself what a heartbeat is. The same alert is a
|
||||
// heartbeat in one team and an ordinary problem in another.
|
||||
func TestDeadman_ConfigurationIsPerTeam(t *testing.T) {
|
||||
s, _ := deadmanTS(t, deadmanCfg())
|
||||
watched := newTeam(t, s, "watched")
|
||||
unwatched := newTeam(t, s, "unwatched")
|
||||
|
||||
// Only the first team calls Watchdog a heartbeat.
|
||||
resp := s.req(t, http.MethodPut, "/api/teams/"+id64(watched.id)+"/deadman", map[string]any{
|
||||
"matchers": "alertname=Watchdog",
|
||||
"timeout_seconds": 3600,
|
||||
"severity": "critical",
|
||||
})
|
||||
resp.Body.Close()
|
||||
if resp.StatusCode != http.StatusOK {
|
||||
t.Fatalf("configure the watched team: %d", resp.StatusCode)
|
||||
}
|
||||
|
||||
postToIntegration(t, s, watched.key, "fp-watched", "Watchdog")
|
||||
postToIntegration(t, s, unwatched.key, "fp-unwatched", "Watchdog")
|
||||
|
||||
// A heartbeat opens nothing where it is one; an ordinary alert opens an
|
||||
// incident where it is not.
|
||||
if got := len(list(t, watched.call(http.MethodGet, "/api/incidents", nil))); got != 0 {
|
||||
t.Errorf("the watched team's heartbeat opened %d incident(s), want 0", got)
|
||||
}
|
||||
if got := len(list(t, unwatched.call(http.MethodGet, "/api/incidents", nil))); got != 1 {
|
||||
t.Errorf("the unwatched team's Watchdog opened %d incident(s), want 1", got)
|
||||
}
|
||||
|
||||
// Silence pages only the team that is watching.
|
||||
s.exec(t, "UPDATE alerts SET received_at = $1 WHERE fingerprint = $2",
|
||||
time.Now().Add(-2*time.Hour).Unix(), "fp-watched")
|
||||
s.exec(t, "UPDATE alerts SET received_at = $1 WHERE fingerprint = $2",
|
||||
time.Now().Add(-2*time.Hour).Unix(), "fp-unwatched")
|
||||
sweep(t, s, noArchive)
|
||||
|
||||
watchedIncidents := list(t, watched.call(http.MethodGet, "/api/incidents", nil))
|
||||
if len(watchedIncidents) != 1 {
|
||||
t.Fatalf("silence opened %d incident(s) for the watching team, want 1", len(watchedIncidents))
|
||||
}
|
||||
if title := watchedIncidents[0]["title"].(string); title != "No heartbeat from Watchdog" {
|
||||
t.Errorf("unexpected incident title %q", title)
|
||||
}
|
||||
if teamID := int64(watchedIncidents[0]["team_id"].(float64)); teamID != watched.id {
|
||||
t.Errorf("the incident opened in team %d, want %d", teamID, watched.id)
|
||||
}
|
||||
|
||||
// The unwatched team's alert went stale the ordinary way, so it has the one
|
||||
// incident it always had — not a second, dead man's switch one.
|
||||
if got := len(list(t, unwatched.call(http.MethodGet, "/api/incidents", nil))); got != 1 {
|
||||
t.Errorf("the unwatched team ended with %d incident(s), want 1", got)
|
||||
}
|
||||
}
|
||||
|
||||
// Configuration is an owner's to change and a member's to read, like the rest of
|
||||
// a team's settings.
|
||||
func TestDeadman_ConfigurationIsOwnerOnly(t *testing.T) {
|
||||
s, _ := deadmanTS(t, deadmanCfg())
|
||||
team := newTeam(t, s, "red")
|
||||
|
||||
// A plain member of that team.
|
||||
var user struct {
|
||||
ID int64 `json:"id"`
|
||||
}
|
||||
decode(t, s.req(t, http.MethodPost, "/api/users",
|
||||
map[string]string{"username": "plain", "email": "plain@test.com"}), &user)
|
||||
s.req(t, http.MethodPost, "/api/teams/"+id64(team.id)+"/members",
|
||||
map[string]any{"user_id": user.ID, "role": "member"}).Body.Close()
|
||||
var key struct {
|
||||
Key string `json:"key"`
|
||||
}
|
||||
decode(t, s.req(t, http.MethodPost, "/api/users/"+id64(user.ID)+"/api-keys",
|
||||
map[string]string{"name": "test"}), &key)
|
||||
|
||||
req, _ := http.NewRequest(http.MethodPut,
|
||||
s.URL+"/api/teams/"+id64(team.id)+"/deadman",
|
||||
strings.NewReader(`{"matchers":"alertname=Watchdog","timeout_seconds":60}`))
|
||||
req.Header.Set("Authorization", "Bearer "+key.Key)
|
||||
req.Header.Set("Content-Type", "application/json")
|
||||
resp, err := http.DefaultClient.Do(req)
|
||||
if err != nil {
|
||||
t.Fatalf("put: %v", err)
|
||||
}
|
||||
resp.Body.Close()
|
||||
if resp.StatusCode != http.StatusForbidden {
|
||||
t.Errorf("a member editing the switches: expected 403, got %d", resp.StatusCode)
|
||||
}
|
||||
|
||||
read, _ := http.NewRequest(http.MethodGet, s.URL+"/api/teams/"+id64(team.id)+"/deadman", nil)
|
||||
read.Header.Set("Authorization", "Bearer "+key.Key)
|
||||
got, err := http.DefaultClient.Do(read)
|
||||
if err != nil {
|
||||
t.Fatalf("get: %v", err)
|
||||
}
|
||||
got.Body.Close()
|
||||
if got.StatusCode != http.StatusOK {
|
||||
t.Errorf("a member reading the switches: expected 200, got %d", got.StatusCode)
|
||||
}
|
||||
}
|
||||
|
||||
// A matcher with no alertname watches nothing, silently, which is the failure
|
||||
// this feature exists to prevent — so it is refused at the door.
|
||||
func TestDeadman_UnusableMatchersAreRejected(t *testing.T) {
|
||||
s, _ := deadmanTS(t, deadmanCfg())
|
||||
|
||||
resp := s.req(t, http.MethodPut, "/api/teams/"+defaultTeam+"/deadman", map[string]any{
|
||||
"matchers": "cluster=prod",
|
||||
"timeout_seconds": 900,
|
||||
})
|
||||
resp.Body.Close()
|
||||
if resp.StatusCode != http.StatusBadRequest {
|
||||
t.Errorf("expected 400 for a matcher with no alertname, got %d", resp.StatusCode)
|
||||
}
|
||||
}
|
||||
|
||||
@@ -51,12 +51,13 @@ type querier interface {
|
||||
}
|
||||
|
||||
const incidentSelectFrom = `
|
||||
SELECT i.id, i.group_key, i.title, i.group_labels, i.status, i.severity,
|
||||
SELECT i.id, i.team_id, t.name, i.group_key, i.title, i.group_labels, i.status, i.severity,
|
||||
i.triggered_at,
|
||||
i.acknowledged_by, i.acknowledged_at, ack.username,
|
||||
i.assigned_to, asg.username, i.snoozed_until,
|
||||
i.resolved_at, i.resolution_source, i.archived_at
|
||||
FROM incidents i
|
||||
JOIN teams t ON t.id = i.team_id
|
||||
LEFT JOIN users ack ON ack.id = i.acknowledged_by
|
||||
LEFT JOIN users asg ON asg.id = i.assigned_to`
|
||||
|
||||
@@ -67,7 +68,7 @@ func scanIncident(s scanner) (models.Incident, error) {
|
||||
var ackAt, snoozedUntil, resolvedAt, archivedAt *int64
|
||||
|
||||
if err := s.Scan(
|
||||
&i.ID, &i.GroupKey, &i.Title, &groupLabelsJSON, &i.Status, &i.Severity,
|
||||
&i.ID, &i.TeamID, &i.TeamName, &i.GroupKey, &i.Title, &groupLabelsJSON, &i.Status, &i.Severity,
|
||||
&triggeredAt,
|
||||
&i.AcknowledgedByID, &ackAt, &i.AcknowledgedByUser,
|
||||
&i.AssignedToID, &i.AssignedToUser, &snoozedUntil,
|
||||
@@ -113,12 +114,17 @@ func todayUTC() string {
|
||||
return time.Now().UTC().Format("2006-01-02")
|
||||
}
|
||||
|
||||
// currentOnCall returns today's on-call user, or nil when nobody is scheduled.
|
||||
// A missing schedule entry is not an error — incidents just open unassigned.
|
||||
func currentOnCall(ctx context.Context, q querier) (*int64, error) {
|
||||
// currentOnCall returns a team's on-call user for today, or nil when nobody is
|
||||
// scheduled. A missing schedule entry is not an error — incidents just open
|
||||
// unassigned.
|
||||
//
|
||||
// Per team: each team keeps its own rota, so two teams can have two different
|
||||
// people on call on the same day, which was the point of scoping the schedule.
|
||||
func currentOnCall(ctx context.Context, q querier, teamID int64) (*int64, error) {
|
||||
var userID int64
|
||||
err := q.QueryRowContext(ctx,
|
||||
"SELECT user_id FROM schedule_entries WHERE date = $1", todayUTC()).Scan(&userID)
|
||||
"SELECT user_id FROM schedule_entries WHERE team_id = $1 AND date = $2",
|
||||
teamID, todayUTC()).Scan(&userID)
|
||||
if err == sql.ErrNoRows {
|
||||
return nil, nil
|
||||
}
|
||||
|
||||
+37
-15
@@ -19,6 +19,15 @@ func handleListIncidents(db *sql.DB) http.HandlerFunc {
|
||||
where := []string{}
|
||||
args := &sqlArgs{}
|
||||
|
||||
// The combined queue: every team the caller belongs to, in one list. A
|
||||
// caller in no team sees an empty queue rather than everybody's.
|
||||
where = append(where, "i.team_id = ANY("+args.add(callerTeamIDs(r.Context()))+")")
|
||||
if team := q.Get("team_id"); team != "" {
|
||||
if n, err := strconv.ParseInt(team, 10, 64); err == nil {
|
||||
where = append(where, "i.team_id = "+args.add(n))
|
||||
}
|
||||
}
|
||||
|
||||
// Without an explicit status the queue shows open work, which is what an
|
||||
// on-call person opens the tool to see.
|
||||
if status := q.Get("status"); status != "" {
|
||||
@@ -96,7 +105,7 @@ func handleListIncidents(db *sql.DB) http.HandlerFunc {
|
||||
|
||||
func handleGetIncident(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
id, ok := incidentIDParam(w, r)
|
||||
id, ok := incidentIDParam(w, r, db)
|
||||
if !ok {
|
||||
return
|
||||
}
|
||||
@@ -119,7 +128,7 @@ func handleGetIncident(db *sql.DB) http.HandlerFunc {
|
||||
|
||||
func handleIncidentAlerts(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
id, ok := incidentIDParam(w, r)
|
||||
id, ok := incidentIDParam(w, r, db)
|
||||
if !ok {
|
||||
return
|
||||
}
|
||||
@@ -137,7 +146,7 @@ func handleIncidentAlerts(db *sql.DB) http.HandlerFunc {
|
||||
|
||||
func handleIncidentTimeline(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
id, ok := incidentIDParam(w, r)
|
||||
id, ok := incidentIDParam(w, r, db)
|
||||
if !ok {
|
||||
return
|
||||
}
|
||||
@@ -176,7 +185,7 @@ func handleIncidentTimeline(db *sql.DB) http.HandlerFunc {
|
||||
|
||||
func handleIncidentAcknowledge(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
id, ok := incidentIDParam(w, r)
|
||||
id, ok := incidentIDParam(w, r, db)
|
||||
if !ok {
|
||||
return
|
||||
}
|
||||
@@ -199,7 +208,7 @@ func handleIncidentAcknowledge(db *sql.DB) http.HandlerFunc {
|
||||
|
||||
func handleIncidentUnacknowledge(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
id, ok := incidentIDParam(w, r)
|
||||
id, ok := incidentIDParam(w, r, db)
|
||||
if !ok {
|
||||
return
|
||||
}
|
||||
@@ -223,7 +232,7 @@ func handleIncidentUnacknowledge(db *sql.DB) http.HandlerFunc {
|
||||
// re-send of an alert that never stopped firing. Use snooze for "not now".
|
||||
func handleIncidentResolve(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
id, ok := incidentIDParam(w, r)
|
||||
id, ok := incidentIDParam(w, r, db)
|
||||
if !ok {
|
||||
return
|
||||
}
|
||||
@@ -244,7 +253,7 @@ func handleIncidentResolve(db *sql.DB) http.HandlerFunc {
|
||||
|
||||
func handleIncidentAssign(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
id, ok := incidentIDParam(w, r)
|
||||
id, ok := incidentIDParam(w, r, db)
|
||||
if !ok {
|
||||
return
|
||||
}
|
||||
@@ -285,7 +294,7 @@ func handleIncidentAssign(db *sql.DB) http.HandlerFunc {
|
||||
// {"duration": "2h"}.
|
||||
func handleIncidentSnooze(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
id, ok := incidentIDParam(w, r)
|
||||
id, ok := incidentIDParam(w, r, db)
|
||||
if !ok {
|
||||
return
|
||||
}
|
||||
@@ -340,7 +349,7 @@ func handleIncidentSnooze(db *sql.DB) http.HandlerFunc {
|
||||
|
||||
func handleIncidentUnsnooze(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
id, ok := incidentIDParam(w, r)
|
||||
id, ok := incidentIDParam(w, r, db)
|
||||
if !ok {
|
||||
return
|
||||
}
|
||||
@@ -359,7 +368,7 @@ func handleIncidentUnsnooze(db *sql.DB) http.HandlerFunc {
|
||||
|
||||
func handleIncidentArchive(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
id, ok := incidentIDParam(w, r)
|
||||
id, ok := incidentIDParam(w, r, db)
|
||||
if !ok {
|
||||
return
|
||||
}
|
||||
@@ -379,7 +388,7 @@ func handleIncidentArchive(db *sql.DB) http.HandlerFunc {
|
||||
|
||||
func handleIncidentUnarchive(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
id, ok := incidentIDParam(w, r)
|
||||
id, ok := incidentIDParam(w, r, db)
|
||||
if !ok {
|
||||
return
|
||||
}
|
||||
@@ -401,7 +410,7 @@ func handleIncidentUnarchive(db *sql.DB) http.HandlerFunc {
|
||||
// single query renders the whole story of an incident in order.
|
||||
func handleCreateNote(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
id, ok := incidentIDParam(w, r)
|
||||
id, ok := incidentIDParam(w, r, db)
|
||||
if !ok {
|
||||
return
|
||||
}
|
||||
@@ -448,7 +457,7 @@ func handleCreateNote(db *sql.DB) http.HandlerFunc {
|
||||
// rest of the timeline is what actually happened, and is not editable.
|
||||
func handleDeleteNote(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
id, ok := incidentIDParam(w, r)
|
||||
id, ok := incidentIDParam(w, r, db)
|
||||
if !ok {
|
||||
return
|
||||
}
|
||||
@@ -479,19 +488,32 @@ func handleDeleteNote(db *sql.DB) http.HandlerFunc {
|
||||
// Shared handler plumbing
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
func incidentIDParam(w http.ResponseWriter, r *http.Request) (int64, bool) {
|
||||
// incidentIDParam reads {id} from the path AND confirms the incident belongs to
|
||||
// a team the caller is in. Both in one place, deliberately: every incident route
|
||||
// goes through here, so scoping cannot be forgotten by writing a new handler
|
||||
// that only remembers the first half.
|
||||
//
|
||||
// An incident in somebody else's team is reported as not found rather than
|
||||
// forbidden, because "there is an incident 41 you may not see" is itself
|
||||
// something only that team should know.
|
||||
func incidentIDParam(w http.ResponseWriter, r *http.Request, db *sql.DB) (int64, bool) {
|
||||
id, err := strconv.ParseInt(chi.URLParam(r, "id"), 10, 64)
|
||||
if err != nil {
|
||||
respond(w, http.StatusBadRequest, errResp("invalid incident id"))
|
||||
return 0, false
|
||||
}
|
||||
if !incidentExists(w, r, db, id) {
|
||||
return 0, false
|
||||
}
|
||||
return id, true
|
||||
}
|
||||
|
||||
// incidentExists reports whether the incident is one the caller may see at all.
|
||||
func incidentExists(w http.ResponseWriter, r *http.Request, db *sql.DB, id int64) bool {
|
||||
var exists int
|
||||
if err := db.QueryRowContext(r.Context(),
|
||||
"SELECT 1 FROM incidents WHERE id = $1", id).Scan(&exists); err != nil {
|
||||
"SELECT 1 FROM incidents WHERE id = $1 AND team_id = ANY($2)",
|
||||
id, callerTeamIDs(r.Context())).Scan(&exists); err != nil {
|
||||
respond(w, http.StatusNotFound, errResp("incident not found"))
|
||||
return false
|
||||
}
|
||||
|
||||
@@ -425,7 +425,7 @@ func TestIncident_AutoAssignedToCurrentOnCall(t *testing.T) {
|
||||
s := newTS(t)
|
||||
|
||||
today := time.Now().UTC().Format("2006-01-02")
|
||||
resp := s.req(t, http.MethodPost, "/api/schedule",
|
||||
resp := s.req(t, http.MethodPost, "/api/teams/"+defaultTeam+"/schedule",
|
||||
map[string]any{"user_id": 1, "dates": []string{today}})
|
||||
if resp.StatusCode != http.StatusCreated {
|
||||
t.Fatalf("schedule assignment returned %d", resp.StatusCode)
|
||||
@@ -609,7 +609,7 @@ func TestSweeper_ArchivesResolvedIncidents(t *testing.T) {
|
||||
|
||||
s.exec(t, "UPDATE incidents SET resolved_at = $1 WHERE id = 1",
|
||||
time.Now().Add(-30*24*time.Hour).Unix())
|
||||
api.Sweep(context.Background(), s.db, 7*24*time.Hour, 6*time.Hour, s.deadman, s.notify)
|
||||
api.Sweep(context.Background(), s.db, 7*24*time.Hour, 6*time.Hour, s.notify)
|
||||
|
||||
if inc := getIncident(t, s, 1); inc["archived_at"] == nil {
|
||||
t.Error("expected the sweeper to archive a long-resolved incident")
|
||||
|
||||
+131
-3
@@ -17,6 +17,7 @@ type contextKey string
|
||||
const (
|
||||
ctxUser contextKey = "user"
|
||||
ctxSession contextKey = "session"
|
||||
ctxTeams contextKey = "teams"
|
||||
)
|
||||
|
||||
// AuthMiddleware accepts either of the two credentials the server issues: an
|
||||
@@ -66,6 +67,39 @@ func AuthMiddleware(db *sql.DB) func(http.Handler) http.Handler {
|
||||
}
|
||||
}
|
||||
|
||||
// AdminOnly rejects a caller who is not a system administrator. It runs inside
|
||||
// AuthMiddleware's group, so by the time it sees a request the caller is known.
|
||||
//
|
||||
// 403 and not 404: the route exists and the caller is authenticated, they are
|
||||
// simply not allowed. Hiding the endpoint would buy nothing — every one of them
|
||||
// is in the README.
|
||||
func AdminOnly(next http.Handler) http.Handler {
|
||||
return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
|
||||
caller, ok := userFromContext(r.Context())
|
||||
if !ok || !caller.IsAdmin {
|
||||
respond(w, http.StatusForbidden, errResp("administrator access required"))
|
||||
return
|
||||
}
|
||||
next.ServeHTTP(w, r)
|
||||
})
|
||||
}
|
||||
|
||||
// requireSelfOrAdmin guards the endpoints that are self-service for your own
|
||||
// account and administration for anybody else's: your password, your ntfy
|
||||
// topic, your API keys. Reports whether the request may proceed, and answers it
|
||||
// if not.
|
||||
//
|
||||
// An API key is not an escalation: it carries exactly the rights of the user it
|
||||
// belongs to, so minting your own is no more than signing in again.
|
||||
func requireSelfOrAdmin(w http.ResponseWriter, r *http.Request, targetID int64) bool {
|
||||
caller, ok := userFromContext(r.Context())
|
||||
if !ok || (caller.ID != targetID && !caller.IsAdmin) {
|
||||
respond(w, http.StatusForbidden, errResp("administrator access required"))
|
||||
return false
|
||||
}
|
||||
return true
|
||||
}
|
||||
|
||||
// apiKeyUser resolves an API key to its user and stamps its last use.
|
||||
func apiKeyUser(ctx context.Context, db *sql.DB, token string) (int64, bool) {
|
||||
var keyID, userID int64
|
||||
@@ -111,14 +145,24 @@ func serveAs(w http.ResponseWriter, r *http.Request, next http.Handler, db *sql.
|
||||
var u models.User
|
||||
var createdUnix int64
|
||||
if err := db.QueryRowContext(r.Context(),
|
||||
"SELECT id, username, email, created_at FROM users WHERE id = $1", userID,
|
||||
).Scan(&u.ID, &u.Username, &u.Email, &createdUnix); err != nil {
|
||||
"SELECT id, username, email, created_at, is_admin FROM users WHERE id = $1", userID,
|
||||
).Scan(&u.ID, &u.Username, &u.Email, &createdUnix, &u.IsAdmin); err != nil {
|
||||
respond(w, http.StatusUnauthorized, errResp("unauthorized"))
|
||||
return
|
||||
}
|
||||
u.CreatedAt = time.Unix(createdUnix, 0).UTC()
|
||||
|
||||
ctx := context.WithValue(r.Context(), ctxUser, u)
|
||||
// Every scoped query needs the caller's teams, so they are loaded once here
|
||||
// rather than per handler. One extra round trip per request, against a
|
||||
// table with one row per membership.
|
||||
teams, err := callerMemberships(r.Context(), db, userID)
|
||||
if err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
|
||||
ctx := context.WithValue(r.Context(), ctxTeams, teams)
|
||||
ctx = context.WithValue(ctx, ctxUser, u)
|
||||
if sessionID != 0 {
|
||||
ctx = context.WithValue(ctx, ctxSession, sessionID)
|
||||
}
|
||||
@@ -135,6 +179,90 @@ func userFromContext(ctx context.Context) (models.User, bool) {
|
||||
return u, ok
|
||||
}
|
||||
|
||||
// membership is the caller's role in one team.
|
||||
type membership struct {
|
||||
teamID int64
|
||||
role string
|
||||
}
|
||||
|
||||
func callerMemberships(ctx context.Context, db *sql.DB, userID int64) ([]membership, error) {
|
||||
rows, err := db.QueryContext(ctx,
|
||||
"SELECT team_id, role FROM team_members WHERE user_id = $1 ORDER BY team_id", userID)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
defer rows.Close()
|
||||
|
||||
var out []membership
|
||||
for rows.Next() {
|
||||
var m membership
|
||||
if err := rows.Scan(&m.teamID, &m.role); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
out = append(out, m)
|
||||
}
|
||||
return out, rows.Err()
|
||||
}
|
||||
|
||||
// callerTeamIDs lists the teams the caller belongs to, for the `team_id = ANY`
|
||||
// filter every list query carries. An admin is NOT implicitly in every team:
|
||||
// administration is about accounts, not about reading other people's incidents,
|
||||
// and an admin who needs to see a team's queue can add themselves to it.
|
||||
func callerTeamIDs(ctx context.Context) []int64 {
|
||||
ms, _ := ctx.Value(ctxTeams).([]membership)
|
||||
ids := make([]int64, 0, len(ms))
|
||||
for _, m := range ms {
|
||||
ids = append(ids, m.teamID)
|
||||
}
|
||||
return ids
|
||||
}
|
||||
|
||||
// callerRole reports the caller's role in one team, and whether they are in it
|
||||
// at all.
|
||||
func callerRole(ctx context.Context, teamID int64) (string, bool) {
|
||||
ms, _ := ctx.Value(ctxTeams).([]membership)
|
||||
for _, m := range ms {
|
||||
if m.teamID == teamID {
|
||||
return m.role, true
|
||||
}
|
||||
}
|
||||
return "", false
|
||||
}
|
||||
|
||||
// requireTeamMember answers the request and reports false unless the caller
|
||||
// belongs to teamID.
|
||||
//
|
||||
// 404, not 403: whether a team exists is itself something only its members
|
||||
// should learn, and the same reasoning applies to every incident and alert
|
||||
// under it.
|
||||
func requireTeamMember(w http.ResponseWriter, r *http.Request, teamID int64) bool {
|
||||
if _, ok := callerRole(r.Context(), teamID); !ok {
|
||||
respond(w, http.StatusNotFound, errResp("not found"))
|
||||
return false
|
||||
}
|
||||
return true
|
||||
}
|
||||
|
||||
// requireTeamOwner is requireTeamMember for the things only an owner may change:
|
||||
// the schedule, the integrations and who is in the team. A system administrator
|
||||
// passes without being a member, because somebody has to be able to repair a
|
||||
// team whose owner has left.
|
||||
func requireTeamOwner(w http.ResponseWriter, r *http.Request, teamID int64) bool {
|
||||
role, ok := callerRole(r.Context(), teamID)
|
||||
if ok && role == models.RoleOwner {
|
||||
return true
|
||||
}
|
||||
if caller, _ := userFromContext(r.Context()); caller.IsAdmin {
|
||||
return true
|
||||
}
|
||||
if !ok {
|
||||
respond(w, http.StatusNotFound, errResp("not found"))
|
||||
return false
|
||||
}
|
||||
respond(w, http.StatusForbidden, errResp("team owner access required"))
|
||||
return false
|
||||
}
|
||||
|
||||
// sessionFromContext returns the id of the session a request was authenticated
|
||||
// with, or false for an API-key request.
|
||||
func sessionFromContext(ctx context.Context) (int64, bool) {
|
||||
|
||||
@@ -95,7 +95,7 @@ func notifyTS(t *testing.T, cfg api.NotifyConfig) (*ts, *fakeNtfy) {
|
||||
func putOnCall(t *testing.T, s *ts, userID int) {
|
||||
t.Helper()
|
||||
today := time.Now().UTC().Format("2006-01-02")
|
||||
resp := s.req(t, http.MethodPost, "/api/schedule",
|
||||
resp := s.req(t, http.MethodPost, "/api/teams/"+defaultTeam+"/schedule",
|
||||
map[string]any{"user_id": userID, "dates": []string{today}})
|
||||
defer resp.Body.Close()
|
||||
if resp.StatusCode != http.StatusCreated {
|
||||
@@ -438,7 +438,7 @@ func TestNotify_SweepPurgesExpiredAckTokens(t *testing.T) {
|
||||
s.sweepNotify(t)
|
||||
s.exec(t, "UPDATE incident_ack_tokens SET expires_at = $1", time.Now().Add(-time.Minute).Unix())
|
||||
|
||||
api.Sweep(context.Background(), s.db, 168*time.Hour, 6*time.Hour, s.deadman, s.notify)
|
||||
api.Sweep(context.Background(), s.db, 168*time.Hour, 6*time.Hour, s.notify)
|
||||
|
||||
var n int
|
||||
if err := s.db.QueryRow("SELECT COUNT(*) FROM incident_ack_tokens").Scan(&n); err != nil {
|
||||
|
||||
+58
-12
@@ -9,11 +9,11 @@ import (
|
||||
"github.com/go-chi/chi/v5/middleware"
|
||||
)
|
||||
|
||||
// NewRouter builds the HTTP surface. notify and deadman are passed through to
|
||||
// the webhook, the only handler that has to decide where a new incident's page
|
||||
// goes and which arriving alerts are heartbeats rather than problems. A zero
|
||||
// notify disables notifications; a zero deadman disables dead man's switches.
|
||||
func NewRouter(db *sql.DB, notify NotifyConfig, deadman DeadmanConfig) http.Handler {
|
||||
// NewRouter builds the HTTP surface. notify is passed through to the webhook,
|
||||
// the only handler that has to decide where a new incident's page goes; a zero
|
||||
// notify disables notifications. Dead man's switches are per team and read from
|
||||
// the database, so nothing about them is wired in here.
|
||||
func NewRouter(db *sql.DB, notify NotifyConfig) http.Handler {
|
||||
r := chi.NewRouter()
|
||||
r.Use(middleware.Logger)
|
||||
r.Use(middleware.Recoverer)
|
||||
@@ -27,9 +27,18 @@ func NewRouter(db *sql.DB, notify NotifyConfig, deadman DeadmanConfig) http.Hand
|
||||
// the scoped token in its path rather than an API key, and has to stay
|
||||
// reachable from outside the cluster for the button to work.
|
||||
r.Post("/api/bootstrap", handleBootstrap(db))
|
||||
r.Post("/api/alertmanager/webhook", handleAlertmanagerWebhook(db, notify, deadman))
|
||||
r.Post("/api/notify/ack/{token}", handleNotifyAck(db))
|
||||
|
||||
// Alert ingestion. The key in the path says both that the sender may post
|
||||
// and which team the alerts belong to, which is why it needs no session.
|
||||
r.Post("/api/integrations/{key}/alertmanager", handleIntegrationWebhook(db, notify))
|
||||
|
||||
// DEPRECATED, and unauthenticated: anything that can reach the port can
|
||||
// open an incident here. Kept for one release so an upgrade does not stop
|
||||
// delivering while the Alertmanager config is edited; it routes everything
|
||||
// to the oldest team. Remove it once senders carry a key.
|
||||
r.Post("/api/alertmanager/webhook", handleLegacyWebhook(db, notify))
|
||||
|
||||
// Signing in to the web UI. Login trades a password for a session cookie,
|
||||
// which AuthMiddleware accepts in place of an API key.
|
||||
r.Post("/api/login", handleLogin(db, newLoginLimiter(), notify.PublicURL))
|
||||
@@ -40,14 +49,30 @@ func NewRouter(db *sql.DB, notify NotifyConfig, deadman DeadmanConfig) http.Hand
|
||||
r.Use(AuthMiddleware(db))
|
||||
|
||||
r.Get("/api/me", handleMe(db))
|
||||
|
||||
// Readable by anyone signed in: the queue's assignment control and the
|
||||
// on-call schedule both need to name people.
|
||||
r.Get("/api/users", handleListUsers(db))
|
||||
r.Post("/api/users", handleCreateUser(db))
|
||||
r.Delete("/api/users/{id}", handleDeleteUser(db))
|
||||
|
||||
// Your own account, or anybody's if you are an admin. The handlers call
|
||||
// requireSelfOrAdmin rather than sitting behind AdminOnly, because
|
||||
// which rule applies depends on the {id} in the path.
|
||||
r.Put("/api/users/{id}/notify", handleSetNotifyTarget(db))
|
||||
r.Put("/api/users/{id}/password", handleSetPassword(db))
|
||||
r.Post("/api/users/{id}/api-keys", handleCreateAPIKey(db))
|
||||
r.Delete("/api/users/{id}/api-keys/{keyID}", handleDeleteAPIKey(db))
|
||||
|
||||
// Administration: who exists, and who is an administrator. Until #3
|
||||
// these were open to any authenticated caller, which meant every user
|
||||
// could delete every other one.
|
||||
r.Group(func(r chi.Router) {
|
||||
r.Use(AdminOnly)
|
||||
|
||||
r.Post("/api/users", handleCreateUser(db))
|
||||
r.Delete("/api/users/{id}", handleDeleteUser(db))
|
||||
r.Put("/api/users/{id}/admin", handleSetAdmin(db))
|
||||
})
|
||||
|
||||
// Alerts are read-only: they are Alertmanager's record, not a work
|
||||
// queue. Everything a person does happens on the incident instead.
|
||||
r.Get("/api/alerts", handleListAlerts(db))
|
||||
@@ -68,10 +93,31 @@ func NewRouter(db *sql.DB, notify NotifyConfig, deadman DeadmanConfig) http.Hand
|
||||
r.Post("/api/incidents/{id}/notes", handleCreateNote(db))
|
||||
r.Delete("/api/incidents/{id}/notes/{eventID}", handleDeleteNote(db))
|
||||
|
||||
r.Post("/api/schedule", handleCreateSchedule(db))
|
||||
r.Get("/api/schedule/current", handleCurrentSchedule(db)) // must be before /{id}
|
||||
r.Get("/api/schedule", handleListSchedule(db))
|
||||
r.Delete("/api/schedule/{id}", handleDeleteSchedule(db))
|
||||
// Teams. A user sees the teams they belong to; an owner configures one.
|
||||
r.Get("/api/teams", handleListTeams(db))
|
||||
r.Post("/api/teams", handleCreateTeam(db))
|
||||
r.Delete("/api/teams/{teamID}", handleDeleteTeam(db))
|
||||
r.Get("/api/teams/{teamID}/members", handleListTeamMembers(db))
|
||||
r.Post("/api/teams/{teamID}/members", handleAddTeamMember(db))
|
||||
r.Delete("/api/teams/{teamID}/members/{userID}", handleRemoveTeamMember(db))
|
||||
|
||||
// A team's own dead man's switches: which of its alerts are heartbeats,
|
||||
// and how long a silence has to last before somebody is paged.
|
||||
r.Get("/api/teams/{teamID}/deadman", handleGetTeamDeadman(db))
|
||||
r.Put("/api/teams/{teamID}/deadman", handleSetTeamDeadman(db))
|
||||
|
||||
// Integrations: where a team's alerts come in, and the key that says so.
|
||||
r.Get("/api/teams/{teamID}/integrations", handleListIntegrations(db))
|
||||
r.Post("/api/teams/{teamID}/integrations", handleCreateIntegration(db, notify.PublicURL))
|
||||
r.Delete("/api/teams/{teamID}/integrations/{integrationID}", handleDeleteIntegration(db))
|
||||
|
||||
// The rota is per team. /api/schedule/current is the exception: it
|
||||
// answers across every team the caller is in, which is what somebody on
|
||||
// two rotas wants to see.
|
||||
r.Get("/api/schedule/current", handleCurrentSchedule(db))
|
||||
r.Post("/api/teams/{teamID}/schedule", handleCreateSchedule(db))
|
||||
r.Get("/api/teams/{teamID}/schedule", handleListSchedule(db))
|
||||
r.Delete("/api/teams/{teamID}/schedule/{id}", handleDeleteSchedule(db))
|
||||
|
||||
r.Get("/api/stats/incidents", handleStatsIncidents(db))
|
||||
r.Get("/api/stats/alerts", handleStatsAlerts(db))
|
||||
|
||||
+69
-26
@@ -12,8 +12,18 @@ import (
|
||||
"github.com/go-chi/chi/v5"
|
||||
)
|
||||
|
||||
// The schedule is per team: each team keeps its own rota, so two teams can have
|
||||
// two different people on call on the same day. Editing it is an owner's job,
|
||||
// like the rest of a team's configuration; reading it is any member's.
|
||||
func handleCreateSchedule(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
teamID, ok := teamParam(w, r)
|
||||
if !ok {
|
||||
return
|
||||
}
|
||||
if !requireTeamOwner(w, r, teamID) {
|
||||
return
|
||||
}
|
||||
var req struct {
|
||||
UserID int64 `json:"user_id"`
|
||||
Dates []string `json:"dates"`
|
||||
@@ -43,10 +53,13 @@ func handleCreateSchedule(db *sql.DB) http.HandlerFunc {
|
||||
}
|
||||
}
|
||||
|
||||
// Verify the user exists.
|
||||
// The person taking the shift has to be in the team: paging somebody
|
||||
// who cannot open the incident is worse than paging nobody.
|
||||
var exists int
|
||||
if err := db.QueryRowContext(r.Context(), "SELECT 1 FROM users WHERE id = $1", req.UserID).Scan(&exists); err != nil {
|
||||
respond(w, http.StatusNotFound, errResp("user not found"))
|
||||
if err := db.QueryRowContext(r.Context(),
|
||||
"SELECT 1 FROM team_members WHERE team_id = $1 AND user_id = $2",
|
||||
teamID, req.UserID).Scan(&exists); err != nil {
|
||||
respond(w, http.StatusNotFound, errResp("user is not a member of this team"))
|
||||
return
|
||||
}
|
||||
|
||||
@@ -64,13 +77,15 @@ func handleCreateSchedule(db *sql.DB) http.HandlerFunc {
|
||||
for _, d := range req.Dates {
|
||||
if req.Replace {
|
||||
if _, err := tx.ExecContext(r.Context(),
|
||||
"DELETE FROM schedule_entries WHERE date = $1", d); err != nil {
|
||||
"DELETE FROM schedule_entries WHERE team_id = $1 AND date = $2",
|
||||
teamID, d); err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
}
|
||||
if _, err := tx.ExecContext(r.Context(),
|
||||
"INSERT INTO schedule_entries (user_id, date) VALUES ($1, $2)", req.UserID, d); err != nil {
|
||||
"INSERT INTO schedule_entries (team_id, user_id, date) VALUES ($1, $2, $3)",
|
||||
teamID, req.UserID, d); err != nil {
|
||||
if isUniqueViolation(err) {
|
||||
respond(w, http.StatusConflict,
|
||||
errResp("date already assigned: "+d+" (pass replace to take it)"))
|
||||
@@ -90,7 +105,7 @@ func handleCreateSchedule(db *sql.DB) http.HandlerFunc {
|
||||
for _, d := range req.Dates {
|
||||
dateSet[d] = true
|
||||
}
|
||||
all, err := scheduleRange(r.Context(), db, req.Dates[0], req.Dates[len(req.Dates)-1])
|
||||
all, err := scheduleRange(r.Context(), db, teamID, req.Dates[0], req.Dates[len(req.Dates)-1])
|
||||
if err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
@@ -107,6 +122,13 @@ func handleCreateSchedule(db *sql.DB) http.HandlerFunc {
|
||||
|
||||
func handleListSchedule(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
teamID, ok := teamParam(w, r)
|
||||
if !ok {
|
||||
return
|
||||
}
|
||||
if !requireTeamMember(w, r, teamID) {
|
||||
return
|
||||
}
|
||||
q := r.URL.Query()
|
||||
from, to := q.Get("from"), q.Get("to")
|
||||
|
||||
@@ -123,7 +145,7 @@ func handleListSchedule(db *sql.DB) http.HandlerFunc {
|
||||
}
|
||||
}
|
||||
|
||||
entries, err := scheduleRange(r.Context(), db, from, to)
|
||||
entries, err := scheduleRange(r.Context(), db, teamID, from, to)
|
||||
if err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
@@ -134,12 +156,20 @@ func handleListSchedule(db *sql.DB) http.HandlerFunc {
|
||||
|
||||
func handleDeleteSchedule(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
teamID, ok := teamParam(w, r)
|
||||
if !ok {
|
||||
return
|
||||
}
|
||||
if !requireTeamOwner(w, r, teamID) {
|
||||
return
|
||||
}
|
||||
id, err := strconv.ParseInt(chi.URLParam(r, "id"), 10, 64)
|
||||
if err != nil {
|
||||
respond(w, http.StatusBadRequest, errResp("invalid schedule id"))
|
||||
return
|
||||
}
|
||||
res, err := db.ExecContext(r.Context(), "DELETE FROM schedule_entries WHERE id = $1", id)
|
||||
res, err := db.ExecContext(r.Context(),
|
||||
"DELETE FROM schedule_entries WHERE id = $1 AND team_id = $2", id, teamID)
|
||||
if err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
@@ -152,35 +182,50 @@ func handleDeleteSchedule(db *sql.DB) http.HandlerFunc {
|
||||
}
|
||||
}
|
||||
|
||||
// handleCurrentSchedule answers "who is on call right now" for every team the
|
||||
// caller belongs to — one entry per team, so somebody on two rotas sees both.
|
||||
// A team with nobody scheduled today simply does not appear.
|
||||
func handleCurrentSchedule(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
today := time.Now().UTC().Format("2006-01-02")
|
||||
|
||||
var e models.ScheduleEntry
|
||||
var ts int64
|
||||
err := db.QueryRowContext(r.Context(), `
|
||||
SELECT s.id, s.user_id, u.username, s.date, s.created_at
|
||||
rows, err := db.QueryContext(r.Context(), `
|
||||
SELECT s.id, s.team_id, t.name, s.user_id, u.username, s.date, s.created_at
|
||||
FROM schedule_entries s
|
||||
JOIN users u ON u.id = s.user_id
|
||||
WHERE s.date = $1`, today).Scan(&e.ID, &e.UserID, &e.Username, &e.Date, &ts)
|
||||
if err == sql.ErrNoRows {
|
||||
respond(w, http.StatusNotFound, errResp("no one is on call today"))
|
||||
return
|
||||
}
|
||||
JOIN teams t ON t.id = s.team_id
|
||||
WHERE s.date = $1 AND s.team_id = ANY($2)
|
||||
ORDER BY t.name`, today, callerTeamIDs(r.Context()))
|
||||
if err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
defer rows.Close()
|
||||
|
||||
entries := []models.ScheduleEntry{}
|
||||
for rows.Next() {
|
||||
var e models.ScheduleEntry
|
||||
var ts int64
|
||||
if err := rows.Scan(&e.ID, &e.TeamID, &e.TeamName, &e.UserID, &e.Username, &e.Date, &ts); err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
e.CreatedAt = time.Unix(ts, 0).UTC()
|
||||
respond(w, http.StatusOK, e)
|
||||
entries = append(entries, e)
|
||||
}
|
||||
if err := rows.Err(); err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
respond(w, http.StatusOK, entries)
|
||||
}
|
||||
}
|
||||
|
||||
// scheduleRange returns schedule entries ordered by date.
|
||||
// from and to are YYYY-MM-DD strings; an empty string means unbounded on that side.
|
||||
func scheduleRange(ctx context.Context, db *sql.DB, from, to string) ([]models.ScheduleEntry, error) {
|
||||
where := []string{}
|
||||
func scheduleRange(ctx context.Context, db *sql.DB, teamID int64, from, to string) ([]models.ScheduleEntry, error) {
|
||||
args := &sqlArgs{}
|
||||
where := []string{"s.team_id = " + args.add(teamID)}
|
||||
if from != "" {
|
||||
where = append(where, "s.date >= "+args.add(from))
|
||||
}
|
||||
@@ -188,15 +233,13 @@ func scheduleRange(ctx context.Context, db *sql.DB, from, to string) ([]models.S
|
||||
where = append(where, "s.date <= "+args.add(to))
|
||||
}
|
||||
|
||||
clause := "1=1"
|
||||
if len(where) > 0 {
|
||||
clause = strings.Join(where, " AND ")
|
||||
}
|
||||
clause := strings.Join(where, " AND ")
|
||||
|
||||
rows, err := db.QueryContext(ctx, `
|
||||
SELECT s.id, s.user_id, u.username, s.date, s.created_at
|
||||
SELECT s.id, s.team_id, t.name, s.user_id, u.username, s.date, s.created_at
|
||||
FROM schedule_entries s
|
||||
JOIN users u ON u.id = s.user_id
|
||||
JOIN teams t ON t.id = s.team_id
|
||||
WHERE `+clause+`
|
||||
ORDER BY s.date ASC`, args.all()...)
|
||||
if err != nil {
|
||||
@@ -208,7 +251,7 @@ func scheduleRange(ctx context.Context, db *sql.DB, from, to string) ([]models.S
|
||||
for rows.Next() {
|
||||
var e models.ScheduleEntry
|
||||
var ts int64
|
||||
if err := rows.Scan(&e.ID, &e.UserID, &e.Username, &e.Date, &ts); err != nil {
|
||||
if err := rows.Scan(&e.ID, &e.TeamID, &e.TeamName, &e.UserID, &e.Username, &e.Date, &ts); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
e.CreatedAt = time.Unix(ts, 0).UTC()
|
||||
|
||||
+11
-7
@@ -11,7 +11,7 @@ import (
|
||||
|
||||
func handleStatsAlerts(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
where, args := statsFilter(r.URL.Query(), "received_at")
|
||||
where, args := statsFilter(r.URL.Query(), "received_at", callerTeamIDs(r.Context()))
|
||||
|
||||
// COALESCE because SUM over zero rows is NULL, not 0, and a count of
|
||||
// nothing is 0 — without it an empty window is a 500 rather than a
|
||||
@@ -37,7 +37,7 @@ func handleStatsAlerts(db *sql.DB) http.HandlerFunc {
|
||||
|
||||
func handleStatsTop(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
where, args := statsFilter(r.URL.Query(), "received_at")
|
||||
where, args := statsFilter(r.URL.Query(), "received_at", callerTeamIDs(r.Context()))
|
||||
|
||||
limit := 10
|
||||
if l := r.URL.Query().Get("limit"); l != "" {
|
||||
@@ -79,7 +79,7 @@ func handleStatsTop(db *sql.DB) http.HandlerFunc {
|
||||
|
||||
func handleStatsByHour(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
where, args := statsFilter(r.URL.Query(), "received_at")
|
||||
where, args := statsFilter(r.URL.Query(), "received_at", callerTeamIDs(r.Context()))
|
||||
|
||||
rows, err := db.QueryContext(r.Context(), fmt.Sprintf(`
|
||||
SELECT EXTRACT(HOUR FROM to_timestamp(received_at) AT TIME ZONE 'UTC')::int AS hr,
|
||||
@@ -119,7 +119,7 @@ func handleStatsByHour(db *sql.DB) http.HandlerFunc {
|
||||
|
||||
func handleStatsByDay(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
where, args := statsFilter(r.URL.Query(), "received_at")
|
||||
where, args := statsFilter(r.URL.Query(), "received_at", callerTeamIDs(r.Context()))
|
||||
|
||||
// Postgres EXTRACT(DOW …) → 0=Sunday … 6=Saturday, the same numbering
|
||||
// SQLite's strftime('%w') returned, so the frontend needs no change.
|
||||
@@ -167,7 +167,7 @@ func handleStatsByDay(db *sql.DB) http.HandlerFunc {
|
||||
// mutated in place and carry no acknowledgement or closure time.
|
||||
func handleStatsIncidents(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
where, args := statsFilter(r.URL.Query(), "triggered_at")
|
||||
where, args := statsFilter(r.URL.Query(), "triggered_at", callerTeamIDs(r.Context()))
|
||||
|
||||
// The counts are COALESCEd because SUM over zero rows is NULL, not 0.
|
||||
// The averages are not: mtta and mttr stay null on purpose, since zero
|
||||
@@ -206,9 +206,13 @@ func handleStatsIncidents(db *sql.DB) http.HandlerFunc {
|
||||
// statsFilter builds a WHERE clause and args from optional ?from and ?to query
|
||||
// params, filtering on timeCol. Archived rows are always excluded, matching the
|
||||
// default list views.
|
||||
func statsFilter(q url.Values, timeCol string) (where string, args *sqlArgs) {
|
||||
//
|
||||
// teamIDs scopes every figure to the caller's own teams: a report that counted
|
||||
// other teams' incidents would leak their volume and their names through the
|
||||
// top-alerts list, and would not be a number about the reader's work anyway.
|
||||
func statsFilter(q url.Values, timeCol string, teamIDs []int64) (where string, args *sqlArgs) {
|
||||
args = &sqlArgs{}
|
||||
clauses := []string{"archived_at IS NULL"}
|
||||
clauses := []string{"archived_at IS NULL", "team_id = ANY(" + args.add(teamIDs) + ")"}
|
||||
if from := q.Get("from"); from != "" {
|
||||
if t, err := time.Parse("2006-01-02", from); err == nil {
|
||||
clauses = append(clauses, timeCol+" >= "+args.add(t.UTC().Unix()))
|
||||
|
||||
@@ -0,0 +1,575 @@
|
||||
package api
|
||||
|
||||
import (
|
||||
"context"
|
||||
"database/sql"
|
||||
"errors"
|
||||
"net/http"
|
||||
"strconv"
|
||||
"strings"
|
||||
"time"
|
||||
|
||||
"git.ryuvia.com/niklas/terdut-server/internal/models"
|
||||
"github.com/go-chi/chi/v5"
|
||||
)
|
||||
|
||||
// handleListTeams lists the caller's own teams, each with their role in it. An
|
||||
// administrator listing every team goes through the admin endpoint instead:
|
||||
// this one answers "what am I part of", which is what the UI's team filter and
|
||||
// the combined queue are built from.
|
||||
func handleListTeams(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
caller, _ := userFromContext(r.Context())
|
||||
rows, err := db.QueryContext(r.Context(), `
|
||||
SELECT t.id, t.name, t.created_at, m.role
|
||||
FROM teams t
|
||||
JOIN team_members m ON m.team_id = t.id
|
||||
WHERE m.user_id = $1
|
||||
ORDER BY t.name`, caller.ID)
|
||||
if err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
defer rows.Close()
|
||||
|
||||
teams := []models.Team{}
|
||||
for rows.Next() {
|
||||
var t models.Team
|
||||
var created int64
|
||||
if err := rows.Scan(&t.ID, &t.Name, &created, &t.Role); err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
t.CreatedAt = time.Unix(created, 0).UTC()
|
||||
teams = append(teams, t)
|
||||
}
|
||||
if err := rows.Err(); err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
respond(w, http.StatusOK, teams)
|
||||
}
|
||||
}
|
||||
|
||||
// handleCreateTeam creates a team and makes its creator the first owner. A team
|
||||
// with no owner would need an administrator to repair before anybody could use
|
||||
// it, so the two happen in one transaction.
|
||||
func handleCreateTeam(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
var req struct {
|
||||
Name string `json:"name"`
|
||||
}
|
||||
if err := decodeJSON(r, &req); err != nil {
|
||||
respond(w, http.StatusBadRequest, errResp("invalid request body"))
|
||||
return
|
||||
}
|
||||
req.Name = strings.TrimSpace(req.Name)
|
||||
if req.Name == "" {
|
||||
respond(w, http.StatusBadRequest, errResp("name is required"))
|
||||
return
|
||||
}
|
||||
|
||||
caller, _ := userFromContext(r.Context())
|
||||
|
||||
tx, err := db.BeginTx(r.Context(), nil)
|
||||
if err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
defer tx.Rollback() //nolint:errcheck
|
||||
|
||||
var team models.Team
|
||||
var created int64
|
||||
if err := tx.QueryRowContext(r.Context(),
|
||||
"INSERT INTO teams (name) VALUES ($1) RETURNING id, name, created_at",
|
||||
req.Name).Scan(&team.ID, &team.Name, &created); err != nil {
|
||||
if isUniqueViolation(err) {
|
||||
respond(w, http.StatusConflict, errResp("a team with that name already exists"))
|
||||
return
|
||||
}
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
if _, err := tx.ExecContext(r.Context(),
|
||||
"INSERT INTO team_members (team_id, user_id, role) VALUES ($1, $2, $3)",
|
||||
team.ID, caller.ID, models.RoleOwner); err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
if err := tx.Commit(); err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
|
||||
team.CreatedAt = time.Unix(created, 0).UTC()
|
||||
team.Role = models.RoleOwner
|
||||
respond(w, http.StatusCreated, team)
|
||||
}
|
||||
}
|
||||
|
||||
// handleDeleteTeam removes a team and, by cascade, its incidents, alerts,
|
||||
// schedule and integrations.
|
||||
//
|
||||
// Refused while the team still has open incidents: deleting a team is tidying
|
||||
// up, and tidying up should never be how an unacknowledged page disappears.
|
||||
// Resolve or archive them first, deliberately.
|
||||
func handleDeleteTeam(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
teamID, ok := teamParam(w, r)
|
||||
if !ok {
|
||||
return
|
||||
}
|
||||
if !requireTeamOwner(w, r, teamID) {
|
||||
return
|
||||
}
|
||||
|
||||
var open int
|
||||
if err := db.QueryRowContext(r.Context(),
|
||||
"SELECT COUNT(*) FROM incidents WHERE team_id = $1 AND resolved_at IS NULL", teamID).
|
||||
Scan(&open); err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
if open > 0 {
|
||||
respond(w, http.StatusConflict, errResp("team still has open incidents"))
|
||||
return
|
||||
}
|
||||
|
||||
res, err := db.ExecContext(r.Context(), "DELETE FROM teams WHERE id = $1", teamID)
|
||||
if err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
if n, _ := res.RowsAffected(); n == 0 {
|
||||
respond(w, http.StatusNotFound, errResp("not found"))
|
||||
return
|
||||
}
|
||||
w.WriteHeader(http.StatusNoContent)
|
||||
}
|
||||
}
|
||||
|
||||
// handleListTeamMembers names everybody in a team. Visible to any member: you
|
||||
// can see who else is on the rota you are on.
|
||||
func handleListTeamMembers(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
teamID, ok := teamParam(w, r)
|
||||
if !ok {
|
||||
return
|
||||
}
|
||||
if !requireTeamMember(w, r, teamID) {
|
||||
return
|
||||
}
|
||||
|
||||
rows, err := db.QueryContext(r.Context(), `
|
||||
SELECT m.team_id, m.user_id, u.username, m.role, m.joined_at
|
||||
FROM team_members m
|
||||
JOIN users u ON u.id = m.user_id
|
||||
WHERE m.team_id = $1
|
||||
ORDER BY u.username`, teamID)
|
||||
if err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
defer rows.Close()
|
||||
|
||||
members := []models.TeamMember{}
|
||||
for rows.Next() {
|
||||
var m models.TeamMember
|
||||
var joined int64
|
||||
if err := rows.Scan(&m.TeamID, &m.UserID, &m.Username, &m.Role, &joined); err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
m.JoinedAt = time.Unix(joined, 0).UTC()
|
||||
members = append(members, m)
|
||||
}
|
||||
if err := rows.Err(); err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
respond(w, http.StatusOK, members)
|
||||
}
|
||||
}
|
||||
|
||||
// handleAddTeamMember adds a user to a team, or changes the role of somebody
|
||||
// already in it.
|
||||
func handleAddTeamMember(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
teamID, ok := teamParam(w, r)
|
||||
if !ok {
|
||||
return
|
||||
}
|
||||
if !requireTeamOwner(w, r, teamID) {
|
||||
return
|
||||
}
|
||||
|
||||
var req struct {
|
||||
UserID int64 `json:"user_id"`
|
||||
Role string `json:"role"`
|
||||
}
|
||||
if err := decodeJSON(r, &req); err != nil || req.UserID == 0 {
|
||||
respond(w, http.StatusBadRequest, errResp("user_id is required"))
|
||||
return
|
||||
}
|
||||
if req.Role == "" {
|
||||
req.Role = models.RoleMember
|
||||
}
|
||||
if req.Role != models.RoleOwner && req.Role != models.RoleMember {
|
||||
respond(w, http.StatusBadRequest, errResp("role must be owner or member"))
|
||||
return
|
||||
}
|
||||
|
||||
_, err := db.ExecContext(r.Context(), `
|
||||
INSERT INTO team_members (team_id, user_id, role)
|
||||
VALUES ($1, $2, $3)
|
||||
ON CONFLICT (team_id, user_id) DO UPDATE SET role = excluded.role`,
|
||||
teamID, req.UserID, req.Role)
|
||||
if err != nil {
|
||||
// The only foreign key that can fail here is the user: the team was
|
||||
// resolved from the caller's own membership.
|
||||
respond(w, http.StatusNotFound, errResp("user not found"))
|
||||
return
|
||||
}
|
||||
w.WriteHeader(http.StatusNoContent)
|
||||
}
|
||||
}
|
||||
|
||||
// handleRemoveTeamMember takes a user out of a team.
|
||||
//
|
||||
// A team must keep an owner, for the same reason the install must keep an
|
||||
// administrator: otherwise nobody can configure it, and repairing that needs
|
||||
// somebody with more access than the team has.
|
||||
func handleRemoveTeamMember(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
teamID, ok := teamParam(w, r)
|
||||
if !ok {
|
||||
return
|
||||
}
|
||||
if !requireTeamOwner(w, r, teamID) {
|
||||
return
|
||||
}
|
||||
userID, err := strconv.ParseInt(chi.URLParam(r, "userID"), 10, 64)
|
||||
if err != nil {
|
||||
respond(w, http.StatusBadRequest, errResp("invalid user id"))
|
||||
return
|
||||
}
|
||||
|
||||
last, err := isLastTeamOwner(r.Context(), db, teamID, userID)
|
||||
if err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
if last {
|
||||
respond(w, http.StatusConflict, errResp("cannot remove the last owner of a team"))
|
||||
return
|
||||
}
|
||||
|
||||
res, err := db.ExecContext(r.Context(),
|
||||
"DELETE FROM team_members WHERE team_id = $1 AND user_id = $2", teamID, userID)
|
||||
if err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
if n, _ := res.RowsAffected(); n == 0 {
|
||||
respond(w, http.StatusNotFound, errResp("not a member of this team"))
|
||||
return
|
||||
}
|
||||
w.WriteHeader(http.StatusNoContent)
|
||||
}
|
||||
}
|
||||
|
||||
func isLastTeamOwner(ctx context.Context, db *sql.DB, teamID, userID int64) (bool, error) {
|
||||
var last bool
|
||||
err := db.QueryRowContext(ctx, `
|
||||
SELECT EXISTS (SELECT 1 FROM team_members
|
||||
WHERE team_id = $1 AND user_id = $2 AND role = 'owner')
|
||||
AND NOT EXISTS (SELECT 1 FROM team_members
|
||||
WHERE team_id = $1 AND user_id <> $2 AND role = 'owner')`,
|
||||
teamID, userID).Scan(&last)
|
||||
return last, err
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Integrations
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
// handleListIntegrations lists a team's integrations. Never the keys: those
|
||||
// exist in plaintext only in the response that created them.
|
||||
func handleListIntegrations(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
teamID, ok := teamParam(w, r)
|
||||
if !ok {
|
||||
return
|
||||
}
|
||||
if !requireTeamMember(w, r, teamID) {
|
||||
return
|
||||
}
|
||||
|
||||
rows, err := db.QueryContext(r.Context(), `
|
||||
SELECT id, team_id, kind, name, created_at, last_used_at
|
||||
FROM integrations
|
||||
WHERE team_id = $1
|
||||
ORDER BY id`, teamID)
|
||||
if err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
defer rows.Close()
|
||||
|
||||
integrations := []models.Integration{}
|
||||
for rows.Next() {
|
||||
var i models.Integration
|
||||
var created int64
|
||||
var lastUsed *int64
|
||||
if err := rows.Scan(&i.ID, &i.TeamID, &i.Kind, &i.Name, &created, &lastUsed); err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
i.CreatedAt = time.Unix(created, 0).UTC()
|
||||
i.LastUsedAt = unixPtr(lastUsed)
|
||||
integrations = append(integrations, i)
|
||||
}
|
||||
if err := rows.Err(); err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
respond(w, http.StatusOK, integrations)
|
||||
}
|
||||
}
|
||||
|
||||
// handleCreateIntegration mints an integration key. The key is returned once,
|
||||
// in this response, and only its hash is kept — the same handling as an API key
|
||||
// or an acknowledgement token.
|
||||
func handleCreateIntegration(db *sql.DB, publicURL string) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
teamID, ok := teamParam(w, r)
|
||||
if !ok {
|
||||
return
|
||||
}
|
||||
if !requireTeamOwner(w, r, teamID) {
|
||||
return
|
||||
}
|
||||
|
||||
var req struct {
|
||||
Name string `json:"name"`
|
||||
Kind string `json:"kind"`
|
||||
}
|
||||
if err := decodeJSON(r, &req); err != nil {
|
||||
respond(w, http.StatusBadRequest, errResp("invalid request body"))
|
||||
return
|
||||
}
|
||||
req.Name = strings.TrimSpace(req.Name)
|
||||
if req.Name == "" {
|
||||
respond(w, http.StatusBadRequest, errResp("name is required"))
|
||||
return
|
||||
}
|
||||
if req.Kind == "" {
|
||||
req.Kind = models.IntegrationAlertmanager
|
||||
}
|
||||
if req.Kind != models.IntegrationAlertmanager {
|
||||
respond(w, http.StatusBadRequest, errResp("unsupported integration kind"))
|
||||
return
|
||||
}
|
||||
|
||||
raw, hash, err := randomToken()
|
||||
if err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
|
||||
var i models.Integration
|
||||
var created int64
|
||||
if err := db.QueryRowContext(r.Context(), `
|
||||
INSERT INTO integrations (team_id, kind, name, key_hash)
|
||||
VALUES ($1, $2, $3, $4)
|
||||
RETURNING id, team_id, kind, name, created_at`,
|
||||
teamID, req.Kind, req.Name, hash).
|
||||
Scan(&i.ID, &i.TeamID, &i.Kind, &i.Name, &created); err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
i.CreatedAt = time.Unix(created, 0).UTC()
|
||||
i.Key = raw
|
||||
i.URL = strings.TrimSuffix(publicURL, "/") + integrationPath(raw, i.Kind)
|
||||
respond(w, http.StatusCreated, i)
|
||||
}
|
||||
}
|
||||
|
||||
func handleDeleteIntegration(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
teamID, ok := teamParam(w, r)
|
||||
if !ok {
|
||||
return
|
||||
}
|
||||
if !requireTeamOwner(w, r, teamID) {
|
||||
return
|
||||
}
|
||||
id, err := strconv.ParseInt(chi.URLParam(r, "integrationID"), 10, 64)
|
||||
if err != nil {
|
||||
respond(w, http.StatusBadRequest, errResp("invalid integration id"))
|
||||
return
|
||||
}
|
||||
|
||||
res, err := db.ExecContext(r.Context(),
|
||||
"DELETE FROM integrations WHERE id = $1 AND team_id = $2", id, teamID)
|
||||
if err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
if n, _ := res.RowsAffected(); n == 0 {
|
||||
respond(w, http.StatusNotFound, errResp("not found"))
|
||||
return
|
||||
}
|
||||
w.WriteHeader(http.StatusNoContent)
|
||||
}
|
||||
}
|
||||
|
||||
// integrationPath is where a sender of this kind posts. Built in one place so
|
||||
// the URL handed out at creation and the route the router registers cannot
|
||||
// drift apart.
|
||||
func integrationPath(key, kind string) string {
|
||||
return "/api/integrations/" + key + "/" + kind
|
||||
}
|
||||
|
||||
// teamIDForKey resolves an integration key to its team, and stamps the key's
|
||||
// last use. An unknown key is not an error worth distinguishing: the caller is
|
||||
// told nothing beyond "no".
|
||||
func teamIDForKey(ctx context.Context, db *sql.DB, key string) (int64, error) {
|
||||
var teamID int64
|
||||
err := db.QueryRowContext(ctx,
|
||||
"SELECT team_id FROM integrations WHERE key_hash = $1", hashToken(key)).Scan(&teamID)
|
||||
if errors.Is(err, sql.ErrNoRows) {
|
||||
return 0, errUnknownIntegration
|
||||
}
|
||||
if err != nil {
|
||||
return 0, err
|
||||
}
|
||||
// Best effort, like an API key's: a failed stamp must not reject an alert.
|
||||
db.ExecContext(ctx, //nolint:errcheck
|
||||
"UPDATE integrations SET last_used_at = $1 WHERE key_hash = $2",
|
||||
time.Now().Unix(), hashToken(key))
|
||||
return teamID, nil
|
||||
}
|
||||
|
||||
var errUnknownIntegration = errors.New("unknown integration key")
|
||||
|
||||
// teamParam reads {teamID} from the path.
|
||||
func teamParam(w http.ResponseWriter, r *http.Request) (int64, bool) {
|
||||
id, err := strconv.ParseInt(chi.URLParam(r, "teamID"), 10, 64)
|
||||
if err != nil {
|
||||
respond(w, http.StatusBadRequest, errResp("invalid team id"))
|
||||
return 0, false
|
||||
}
|
||||
return id, true
|
||||
}
|
||||
|
||||
// defaultTeamID is the team the deprecated unauthenticated webhook routes to:
|
||||
// the oldest one, which on an upgraded install is the "Default" team every
|
||||
// pre-teams row was moved into.
|
||||
func defaultTeamID(ctx context.Context, db *sql.DB) (int64, error) {
|
||||
var id int64
|
||||
err := db.QueryRowContext(ctx, "SELECT id FROM teams ORDER BY id LIMIT 1").Scan(&id)
|
||||
return id, err
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// A team's dead man's switches
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
// deadmanResponse is the wire shape of a team's switch configuration. The
|
||||
// timeout is seconds rather than a duration string, because that is what the
|
||||
// column holds and what arithmetic is done on; a client renders it.
|
||||
type deadmanResponse struct {
|
||||
TeamID int64 `json:"team_id"`
|
||||
Matchers string `json:"matchers"`
|
||||
TimeoutSeconds int64 `json:"timeout_seconds"`
|
||||
Severity string `json:"severity"`
|
||||
}
|
||||
|
||||
func handleGetTeamDeadman(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
teamID, ok := teamParam(w, r)
|
||||
if !ok {
|
||||
return
|
||||
}
|
||||
if !requireTeamMember(w, r, teamID) {
|
||||
return
|
||||
}
|
||||
|
||||
out := deadmanResponse{TeamID: teamID, Severity: "critical"}
|
||||
err := db.QueryRowContext(r.Context(),
|
||||
"SELECT matchers, timeout_seconds, severity FROM deadman_configs WHERE team_id = $1",
|
||||
teamID).Scan(&out.Matchers, &out.TimeoutSeconds, &out.Severity)
|
||||
if err != nil && !errors.Is(err, sql.ErrNoRows) {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
// A team with no row watches nothing, which is a configuration and not
|
||||
// an absence: answering 404 would make "off" indistinguishable from
|
||||
// "this server does not do this".
|
||||
respond(w, http.StatusOK, out)
|
||||
}
|
||||
}
|
||||
|
||||
// handleSetTeamDeadman replaces a team's switch configuration.
|
||||
//
|
||||
// Validated by parsing: a matcher string that survives ParseDeadmanConfig with
|
||||
// nothing usable in it is rejected rather than stored, because a switch that
|
||||
// silently watches nothing is the failure this feature exists to prevent.
|
||||
func handleSetTeamDeadman(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
teamID, ok := teamParam(w, r)
|
||||
if !ok {
|
||||
return
|
||||
}
|
||||
if !requireTeamOwner(w, r, teamID) {
|
||||
return
|
||||
}
|
||||
|
||||
var req struct {
|
||||
Matchers string `json:"matchers"`
|
||||
TimeoutSeconds int64 `json:"timeout_seconds"`
|
||||
Severity string `json:"severity"`
|
||||
}
|
||||
if err := decodeJSON(r, &req); err != nil {
|
||||
respond(w, http.StatusBadRequest, errResp("invalid request body"))
|
||||
return
|
||||
}
|
||||
req.Matchers = strings.TrimSpace(req.Matchers)
|
||||
if req.Severity == "" {
|
||||
req.Severity = "critical"
|
||||
}
|
||||
if req.TimeoutSeconds < 0 {
|
||||
respond(w, http.StatusBadRequest, errResp("timeout_seconds must not be negative"))
|
||||
return
|
||||
}
|
||||
if req.Matchers != "" {
|
||||
parsed := parseDeadmanQuietly(req.Matchers, time.Duration(req.TimeoutSeconds)*time.Second, req.Severity)
|
||||
if len(parsed.Matchers) == 0 {
|
||||
respond(w, http.StatusBadRequest, errResp(
|
||||
"no usable matchers: each must name an alertname, as in alertname=Watchdog,cluster=prod"))
|
||||
return
|
||||
}
|
||||
}
|
||||
|
||||
if _, err := db.ExecContext(r.Context(), `
|
||||
INSERT INTO deadman_configs (team_id, matchers, timeout_seconds, severity, updated_at)
|
||||
VALUES ($1, $2, $3, $4, `+nowEpoch+`)
|
||||
ON CONFLICT (team_id) DO UPDATE SET
|
||||
matchers = excluded.matchers,
|
||||
timeout_seconds = excluded.timeout_seconds,
|
||||
severity = excluded.severity,
|
||||
updated_at = excluded.updated_at`,
|
||||
teamID, req.Matchers, req.TimeoutSeconds, req.Severity); err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
|
||||
respond(w, http.StatusOK, deadmanResponse{
|
||||
TeamID: teamID,
|
||||
Matchers: req.Matchers,
|
||||
TimeoutSeconds: req.TimeoutSeconds,
|
||||
Severity: req.Severity,
|
||||
})
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,350 @@
|
||||
package api_test
|
||||
|
||||
import (
|
||||
"bytes"
|
||||
"encoding/json"
|
||||
"io"
|
||||
"net/http"
|
||||
"testing"
|
||||
)
|
||||
|
||||
// The whole point of #4: two teams sharing one server must not see each other's
|
||||
// work. These tests build two of them and check the boundary from both sides.
|
||||
|
||||
type teamFixture struct {
|
||||
id int64
|
||||
key string // integration key: how alerts get in
|
||||
call func(method, path string, body any) *http.Response
|
||||
}
|
||||
|
||||
// newTeam creates a team with its own member, integration key and API key. The
|
||||
// admin does the creating, as an install's first user would.
|
||||
func newTeam(t *testing.T, s *ts, name string) teamFixture {
|
||||
t.Helper()
|
||||
|
||||
var team struct {
|
||||
ID int64 `json:"id"`
|
||||
}
|
||||
decode(t, s.req(t, http.MethodPost, "/api/teams", map[string]string{"name": name}), &team)
|
||||
|
||||
var integration struct {
|
||||
Key string `json:"key"`
|
||||
URL string `json:"url"`
|
||||
}
|
||||
decode(t, s.req(t, http.MethodPost, "/api/teams/"+id64(team.ID)+"/integrations",
|
||||
map[string]string{"name": name + " alertmanager"}), &integration)
|
||||
if integration.Key == "" {
|
||||
t.Fatalf("%s: integration key was not returned", name)
|
||||
}
|
||||
|
||||
// A member of this team and no other.
|
||||
var user struct {
|
||||
ID int64 `json:"id"`
|
||||
}
|
||||
decode(t, s.req(t, http.MethodPost, "/api/users",
|
||||
map[string]string{"username": name + "-user", "email": name + "@test.com"}), &user)
|
||||
|
||||
resp := s.req(t, http.MethodPost, "/api/teams/"+id64(team.ID)+"/members",
|
||||
map[string]any{"user_id": user.ID, "role": "owner"})
|
||||
resp.Body.Close()
|
||||
if resp.StatusCode != http.StatusNoContent {
|
||||
t.Fatalf("%s: add member: %d", name, resp.StatusCode)
|
||||
}
|
||||
|
||||
var key struct {
|
||||
Key string `json:"key"`
|
||||
}
|
||||
decode(t, s.req(t, http.MethodPost, "/api/users/"+id64(user.ID)+"/api-keys",
|
||||
map[string]string{"name": "test"}), &key)
|
||||
|
||||
return teamFixture{
|
||||
id: team.ID,
|
||||
key: integration.Key,
|
||||
call: func(method, path string, body any) *http.Response {
|
||||
t.Helper()
|
||||
var r io.Reader
|
||||
if body != nil {
|
||||
data, _ := json.Marshal(body)
|
||||
r = bytes.NewReader(data)
|
||||
}
|
||||
req, _ := http.NewRequest(method, s.URL+path, r)
|
||||
req.Header.Set("Authorization", "Bearer "+key.Key)
|
||||
if body != nil {
|
||||
req.Header.Set("Content-Type", "application/json")
|
||||
}
|
||||
resp, err := http.DefaultClient.Do(req)
|
||||
if err != nil {
|
||||
t.Fatalf("%s %s: %v", method, path, err)
|
||||
}
|
||||
return resp
|
||||
},
|
||||
}
|
||||
}
|
||||
|
||||
// postToIntegration sends one firing alert on a team's integration key, the way
|
||||
// a real Alertmanager receiver would.
|
||||
func postToIntegration(t *testing.T, s *ts, key, fingerprint, name string) {
|
||||
t.Helper()
|
||||
payload := map[string]any{
|
||||
"version": "4",
|
||||
"status": "firing",
|
||||
"groupKey": "{}:{alertname=\"" + name + "\"}",
|
||||
"groupLabels": map[string]string{"alertname": name},
|
||||
"alerts": []map[string]any{
|
||||
amAlert(fingerprint, name, "firing", "2026-09-20T10:00:00Z", zeroTime, nil),
|
||||
},
|
||||
}
|
||||
data, _ := json.Marshal(payload)
|
||||
resp, err := http.Post(s.URL+"/api/integrations/"+key+"/alertmanager",
|
||||
"application/json", bytes.NewReader(data))
|
||||
if err != nil {
|
||||
t.Fatalf("post alert: %v", err)
|
||||
}
|
||||
defer resp.Body.Close()
|
||||
if resp.StatusCode != http.StatusOK {
|
||||
t.Fatalf("post alert: %d", resp.StatusCode)
|
||||
}
|
||||
}
|
||||
|
||||
func list(t *testing.T, resp *http.Response) []map[string]any {
|
||||
t.Helper()
|
||||
var out []map[string]any
|
||||
decode(t, resp, &out)
|
||||
return out
|
||||
}
|
||||
|
||||
// An alert posted on one team's key opens an incident in that team and nowhere
|
||||
// else, and neither team can read the other's queue.
|
||||
func TestTeams_IncidentsAreScopedToTheReceivingTeam(t *testing.T) {
|
||||
s := newTS(t)
|
||||
red := newTeam(t, s, "red")
|
||||
blue := newTeam(t, s, "blue")
|
||||
|
||||
postToIntegration(t, s, red.key, "fp-red", "RedDiskFull")
|
||||
postToIntegration(t, s, blue.key, "fp-blue", "BlueDiskFull")
|
||||
|
||||
redIncidents := list(t, red.call(http.MethodGet, "/api/incidents", nil))
|
||||
if len(redIncidents) != 1 {
|
||||
t.Fatalf("red should see exactly its own incident, saw %d", len(redIncidents))
|
||||
}
|
||||
if title := redIncidents[0]["title"]; title != "RedDiskFull" {
|
||||
t.Errorf("red saw %v", title)
|
||||
}
|
||||
if teamID := int64(redIncidents[0]["team_id"].(float64)); teamID != red.id {
|
||||
t.Errorf("red's incident belongs to team %d, want %d", teamID, red.id)
|
||||
}
|
||||
|
||||
blueIncidents := list(t, blue.call(http.MethodGet, "/api/incidents", nil))
|
||||
if len(blueIncidents) != 1 || blueIncidents[0]["title"] != "BlueDiskFull" {
|
||||
t.Fatalf("blue should see exactly its own incident, saw %v", blueIncidents)
|
||||
}
|
||||
|
||||
// Reading the other team's incident by id is not found rather than
|
||||
// forbidden: its existence is the other team's business.
|
||||
otherID := int64(blueIncidents[0]["id"].(float64))
|
||||
for _, path := range []string{
|
||||
"/api/incidents/" + id64(otherID),
|
||||
"/api/incidents/" + id64(otherID) + "/alerts",
|
||||
"/api/incidents/" + id64(otherID) + "/timeline",
|
||||
} {
|
||||
resp := red.call(http.MethodGet, path, nil)
|
||||
resp.Body.Close()
|
||||
if resp.StatusCode != http.StatusNotFound {
|
||||
t.Errorf("red reading %s: expected 404, got %d", path, resp.StatusCode)
|
||||
}
|
||||
}
|
||||
|
||||
// And cannot act on it either.
|
||||
for _, path := range []string{"/acknowledge", "/resolve", "/archive"} {
|
||||
resp := red.call(http.MethodPost, "/api/incidents/"+id64(otherID)+path, nil)
|
||||
resp.Body.Close()
|
||||
if resp.StatusCode != http.StatusNotFound {
|
||||
t.Errorf("red posting %s: expected 404, got %d", path, resp.StatusCode)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// Alerts, the raw signal record, are scoped the same way.
|
||||
func TestTeams_AlertsAndStatsAreScoped(t *testing.T) {
|
||||
s := newTS(t)
|
||||
red := newTeam(t, s, "red")
|
||||
blue := newTeam(t, s, "blue")
|
||||
|
||||
postToIntegration(t, s, red.key, "fp-red", "RedDiskFull")
|
||||
postToIntegration(t, s, blue.key, "fp-blue-1", "BlueDiskFull")
|
||||
postToIntegration(t, s, blue.key, "fp-blue-2", "BlueMemory")
|
||||
|
||||
if alerts := list(t, red.call(http.MethodGet, "/api/alerts", nil)); len(alerts) != 1 {
|
||||
t.Errorf("red should see 1 alert, saw %d", len(alerts))
|
||||
}
|
||||
if alerts := list(t, blue.call(http.MethodGet, "/api/alerts", nil)); len(alerts) != 2 {
|
||||
t.Errorf("blue should see 2 alerts, saw %d", len(alerts))
|
||||
}
|
||||
|
||||
// Statistics count your own work only — otherwise a team's volume, and the
|
||||
// names of its alerts, leak through the totals.
|
||||
var stats map[string]any
|
||||
decode(t, red.call(http.MethodGet, "/api/stats/alerts", nil), &stats)
|
||||
if total := stats["total"].(float64); total != 1 {
|
||||
t.Errorf("red's alert stats counted %v alerts, want 1", total)
|
||||
}
|
||||
|
||||
top := list(t, red.call(http.MethodGet, "/api/stats/alerts/top", nil))
|
||||
for _, row := range top {
|
||||
if name := row["name"].(string); name != "RedDiskFull" {
|
||||
t.Errorf("red's top alerts named %q, which is not theirs", name)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// The same fingerprint, the same groupKey and the same date are all legitimate
|
||||
// in two teams at once: two clusters running the same rules, two rotas.
|
||||
func TestTeams_SameFingerprintInTwoTeams(t *testing.T) {
|
||||
s := newTS(t)
|
||||
red := newTeam(t, s, "red")
|
||||
blue := newTeam(t, s, "blue")
|
||||
|
||||
postToIntegration(t, s, red.key, "fp-shared", "DiskFull")
|
||||
postToIntegration(t, s, blue.key, "fp-shared", "DiskFull")
|
||||
|
||||
for _, team := range []struct {
|
||||
name string
|
||||
f teamFixture
|
||||
}{{"red", red}, {"blue", blue}} {
|
||||
incidents := list(t, team.f.call(http.MethodGet, "/api/incidents", nil))
|
||||
if len(incidents) != 1 {
|
||||
t.Errorf("%s: expected its own incident for the shared fingerprint, saw %d",
|
||||
team.name, len(incidents))
|
||||
}
|
||||
}
|
||||
|
||||
// And both rotas can name somebody for the same day.
|
||||
for _, team := range []struct {
|
||||
name string
|
||||
f teamFixture
|
||||
}{{"red", red}, {"blue", blue}} {
|
||||
var members []map[string]any
|
||||
decode(t, team.f.call(http.MethodGet, "/api/teams/"+id64(team.f.id)+"/members", nil), &members)
|
||||
userID := int64(members[0]["user_id"].(float64))
|
||||
|
||||
resp := team.f.call(http.MethodPost, "/api/teams/"+id64(team.f.id)+"/schedule",
|
||||
map[string]any{"user_id": userID, "dates": []string{"2026-10-01"}})
|
||||
resp.Body.Close()
|
||||
if resp.StatusCode != http.StatusCreated {
|
||||
t.Errorf("%s: taking 2026-10-01 returned %d", team.name, resp.StatusCode)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// An unknown key delivers nothing, and says so rather than accepting silently.
|
||||
func TestTeams_UnknownIntegrationKeyIsRejected(t *testing.T) {
|
||||
s := newTS(t)
|
||||
team := newTeam(t, s, "red")
|
||||
|
||||
resp, err := http.Post(s.URL+"/api/integrations/not-a-real-key/alertmanager",
|
||||
"application/json", bytes.NewReader([]byte(`{"version":"4","status":"firing","alerts":[]}`)))
|
||||
if err != nil {
|
||||
t.Fatalf("post: %v", err)
|
||||
}
|
||||
resp.Body.Close()
|
||||
if resp.StatusCode != http.StatusUnauthorized {
|
||||
t.Errorf("expected 401 for an unknown key, got %d", resp.StatusCode)
|
||||
}
|
||||
|
||||
if incidents := list(t, team.call(http.MethodGet, "/api/incidents", nil)); len(incidents) != 0 {
|
||||
t.Errorf("a rejected payload opened %d incident(s)", len(incidents))
|
||||
}
|
||||
}
|
||||
|
||||
// Team configuration is an owner's job; working incidents is a member's.
|
||||
func TestTeams_MemberCannotConfigureTheTeam(t *testing.T) {
|
||||
s := newTS(t)
|
||||
team := newTeam(t, s, "red")
|
||||
|
||||
// A plain member of the same team.
|
||||
var user struct {
|
||||
ID int64 `json:"id"`
|
||||
}
|
||||
decode(t, s.req(t, http.MethodPost, "/api/users",
|
||||
map[string]string{"username": "plain", "email": "plain@test.com"}), &user)
|
||||
resp := s.req(t, http.MethodPost, "/api/teams/"+id64(team.id)+"/members",
|
||||
map[string]any{"user_id": user.ID, "role": "member"})
|
||||
resp.Body.Close()
|
||||
var key struct {
|
||||
Key string `json:"key"`
|
||||
}
|
||||
decode(t, s.req(t, http.MethodPost, "/api/users/"+id64(user.ID)+"/api-keys",
|
||||
map[string]string{"name": "test"}), &key)
|
||||
|
||||
call := func(method, path string, body any) *http.Response {
|
||||
t.Helper()
|
||||
var r io.Reader
|
||||
if body != nil {
|
||||
data, _ := json.Marshal(body)
|
||||
r = bytes.NewReader(data)
|
||||
}
|
||||
req, _ := http.NewRequest(method, s.URL+path, r)
|
||||
req.Header.Set("Authorization", "Bearer "+key.Key)
|
||||
if body != nil {
|
||||
req.Header.Set("Content-Type", "application/json")
|
||||
}
|
||||
resp, err := http.DefaultClient.Do(req)
|
||||
if err != nil {
|
||||
t.Fatalf("%s %s: %v", method, path, err)
|
||||
}
|
||||
return resp
|
||||
}
|
||||
|
||||
base := "/api/teams/" + id64(team.id)
|
||||
for _, c := range []struct {
|
||||
name string
|
||||
method string
|
||||
path string
|
||||
body any
|
||||
}{
|
||||
{"mint an integration key", http.MethodPost, base + "/integrations",
|
||||
map[string]string{"name": "mine"}},
|
||||
{"take a shift", http.MethodPost, base + "/schedule",
|
||||
map[string]any{"user_id": user.ID, "dates": []string{"2026-11-01"}}},
|
||||
{"add a member", http.MethodPost, base + "/members",
|
||||
map[string]any{"user_id": 1}},
|
||||
{"delete the team", http.MethodDelete, base, nil},
|
||||
} {
|
||||
resp := call(c.method, c.path, c.body)
|
||||
resp.Body.Close()
|
||||
if resp.StatusCode != http.StatusForbidden {
|
||||
t.Errorf("%s: expected 403, got %d", c.name, resp.StatusCode)
|
||||
}
|
||||
}
|
||||
|
||||
// But they can read what the team is doing.
|
||||
for _, path := range []string{base + "/members", base + "/integrations", base + "/schedule"} {
|
||||
resp := call(http.MethodGet, path, nil)
|
||||
resp.Body.Close()
|
||||
if resp.StatusCode != http.StatusOK {
|
||||
t.Errorf("reading %s: expected 200, got %d", path, resp.StatusCode)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// A team is not somewhere an outsider can look, whatever they know about it.
|
||||
func TestTeams_OutsiderSeesNothing(t *testing.T) {
|
||||
s := newTS(t)
|
||||
red := newTeam(t, s, "red")
|
||||
blue := newTeam(t, s, "blue")
|
||||
|
||||
base := "/api/teams/" + id64(red.id)
|
||||
for _, path := range []string{base + "/members", base + "/integrations", base + "/schedule"} {
|
||||
resp := blue.call(http.MethodGet, path, nil)
|
||||
resp.Body.Close()
|
||||
if resp.StatusCode != http.StatusNotFound {
|
||||
t.Errorf("blue reading %s: expected 404, got %d", path, resp.StatusCode)
|
||||
}
|
||||
}
|
||||
|
||||
// /api/teams lists your own, never the install's.
|
||||
teams := list(t, blue.call(http.MethodGet, "/api/teams", nil))
|
||||
if len(teams) != 1 || teams[0]["name"] != "blue" {
|
||||
t.Errorf("blue's team list: %v", teams)
|
||||
}
|
||||
}
|
||||
@@ -30,6 +30,10 @@ import (
|
||||
// tests nothing is worse than one that does not run.
|
||||
const testDSNEnv = "TERDUT_TEST_DSN"
|
||||
|
||||
// defaultTeam is the team migration 003 creates and the bootstrap user owns, as
|
||||
// a path segment. Every test that does not say otherwise works inside it.
|
||||
const defaultTeam = "1"
|
||||
|
||||
var schemaSeq int
|
||||
|
||||
// newTestDB returns a migrated database private to this test, and drops it
|
||||
|
||||
+105
-5
@@ -58,7 +58,7 @@ func handleBootstrap(db *sql.DB) http.HandlerFunc {
|
||||
|
||||
var userID int64
|
||||
if err := db.QueryRowContext(r.Context(),
|
||||
"INSERT INTO users (username, email, password_hash) VALUES ($1, $2, $3) RETURNING id",
|
||||
"INSERT INTO users (username, email, password_hash, is_admin) VALUES ($1, $2, $3, true) RETURNING id",
|
||||
req.Username, req.Email, passwordHash).Scan(&userID); err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
@@ -77,6 +77,16 @@ func handleBootstrap(db *sql.DB) http.HandlerFunc {
|
||||
return
|
||||
}
|
||||
|
||||
// The default team exists from migration 003, on a fresh install too.
|
||||
// Without a membership the first user signs in to a working server with
|
||||
// no queue, no schedule and nowhere for an integration to hang off.
|
||||
if teamID, err := defaultTeamID(r.Context(), db); err == nil {
|
||||
db.ExecContext(r.Context(), //nolint:errcheck
|
||||
"INSERT INTO team_members (team_id, user_id, role) VALUES ($1, $2, $3) "+
|
||||
"ON CONFLICT (team_id, user_id) DO NOTHING",
|
||||
teamID, userID, models.RoleOwner)
|
||||
}
|
||||
|
||||
user, _ := fetchUser(r.Context(), db, userID)
|
||||
key := models.APIKey{ID: keyID, UserID: userID, Name: "bootstrap", Key: raw, CreatedAt: user.CreatedAt}
|
||||
respond(w, http.StatusCreated, map[string]any{"user": user, "api_key": key})
|
||||
@@ -86,7 +96,7 @@ func handleBootstrap(db *sql.DB) http.HandlerFunc {
|
||||
func handleListUsers(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
rows, err := db.QueryContext(r.Context(),
|
||||
"SELECT id, username, email, created_at, ntfy_topic FROM users ORDER BY id")
|
||||
"SELECT id, username, email, created_at, ntfy_topic, is_admin FROM users ORDER BY id")
|
||||
if err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
@@ -97,7 +107,7 @@ func handleListUsers(db *sql.DB) http.HandlerFunc {
|
||||
for rows.Next() {
|
||||
var u models.User
|
||||
var ts int64
|
||||
if err := rows.Scan(&u.ID, &u.Username, &u.Email, &ts, &u.NtfyTopic); err != nil {
|
||||
if err := rows.Scan(&u.ID, &u.Username, &u.Email, &ts, &u.NtfyTopic, &u.IsAdmin); err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
@@ -150,6 +160,9 @@ func handleSetNotifyTarget(db *sql.DB) http.HandlerFunc {
|
||||
respond(w, http.StatusBadRequest, errResp("invalid user id"))
|
||||
return
|
||||
}
|
||||
if !requireSelfOrAdmin(w, r, id) {
|
||||
return
|
||||
}
|
||||
var req struct {
|
||||
NtfyTopic string `json:"ntfy_topic"`
|
||||
}
|
||||
@@ -190,6 +203,21 @@ func handleDeleteUser(db *sql.DB) http.HandlerFunc {
|
||||
respond(w, http.StatusBadRequest, errResp("invalid user id"))
|
||||
return
|
||||
}
|
||||
// Deleting yourself is how an install ends up with no administrator at
|
||||
// all, and it is never what somebody meant to do.
|
||||
caller, _ := userFromContext(r.Context())
|
||||
if caller.ID == id {
|
||||
respond(w, http.StatusConflict, errResp("cannot delete your own account"))
|
||||
return
|
||||
}
|
||||
if last, err := isLastAdmin(r.Context(), db, id); err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
} else if last {
|
||||
respond(w, http.StatusConflict, errResp("cannot delete the last administrator"))
|
||||
return
|
||||
}
|
||||
|
||||
res, err := db.ExecContext(r.Context(), "DELETE FROM users WHERE id = $1", id)
|
||||
if err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
@@ -211,6 +239,9 @@ func handleCreateAPIKey(db *sql.DB) http.HandlerFunc {
|
||||
respond(w, http.StatusBadRequest, errResp("invalid user id"))
|
||||
return
|
||||
}
|
||||
if !requireSelfOrAdmin(w, r, userID) {
|
||||
return
|
||||
}
|
||||
|
||||
var req struct {
|
||||
Name string `json:"name"`
|
||||
@@ -254,6 +285,9 @@ func handleDeleteAPIKey(db *sql.DB) http.HandlerFunc {
|
||||
respond(w, http.StatusBadRequest, errResp("invalid user id"))
|
||||
return
|
||||
}
|
||||
if !requireSelfOrAdmin(w, r, userID) {
|
||||
return
|
||||
}
|
||||
keyID, err := strconv.ParseInt(chi.URLParam(r, "keyID"), 10, 64)
|
||||
if err != nil {
|
||||
respond(w, http.StatusBadRequest, errResp("invalid key id"))
|
||||
@@ -292,11 +326,77 @@ func fetchUser(ctx context.Context, db *sql.DB, id int64) (models.User, error) {
|
||||
var u models.User
|
||||
var ts int64
|
||||
err := db.QueryRowContext(ctx,
|
||||
"SELECT id, username, email, created_at, ntfy_topic FROM users WHERE id = $1", id).
|
||||
Scan(&u.ID, &u.Username, &u.Email, &ts, &u.NtfyTopic)
|
||||
"SELECT id, username, email, created_at, ntfy_topic, is_admin FROM users WHERE id = $1", id).
|
||||
Scan(&u.ID, &u.Username, &u.Email, &ts, &u.NtfyTopic, &u.IsAdmin)
|
||||
if err != nil {
|
||||
return u, err
|
||||
}
|
||||
u.CreatedAt = time.Unix(ts, 0).UTC()
|
||||
return u, nil
|
||||
}
|
||||
|
||||
// handleSetAdmin grants or revokes the system administrator flag.
|
||||
//
|
||||
// Revoking is guarded twice: an install must keep at least one administrator,
|
||||
// and you cannot demote yourself. The first stops the flag being lost
|
||||
// altogether; the second stops the likelier accident, where the only admin
|
||||
// clears their own flag while tidying up and locks the door behind them.
|
||||
func handleSetAdmin(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
id, err := strconv.ParseInt(chi.URLParam(r, "id"), 10, 64)
|
||||
if err != nil {
|
||||
respond(w, http.StatusBadRequest, errResp("invalid user id"))
|
||||
return
|
||||
}
|
||||
var req struct {
|
||||
IsAdmin *bool `json:"is_admin"`
|
||||
}
|
||||
if err := decodeJSON(r, &req); err != nil || req.IsAdmin == nil {
|
||||
respond(w, http.StatusBadRequest, errResp("is_admin is required"))
|
||||
return
|
||||
}
|
||||
|
||||
if !*req.IsAdmin {
|
||||
caller, _ := userFromContext(r.Context())
|
||||
if caller.ID == id {
|
||||
respond(w, http.StatusConflict, errResp("cannot revoke your own administrator access"))
|
||||
return
|
||||
}
|
||||
if last, err := isLastAdmin(r.Context(), db, id); err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
} else if last {
|
||||
respond(w, http.StatusConflict, errResp("cannot revoke the last administrator"))
|
||||
return
|
||||
}
|
||||
}
|
||||
|
||||
res, err := db.ExecContext(r.Context(),
|
||||
"UPDATE users SET is_admin = $1 WHERE id = $2", *req.IsAdmin, id)
|
||||
if err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
if n, _ := res.RowsAffected(); n == 0 {
|
||||
respond(w, http.StatusNotFound, errResp("user not found"))
|
||||
return
|
||||
}
|
||||
|
||||
user, err := fetchUser(r.Context(), db, id)
|
||||
if err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
respond(w, http.StatusOK, user)
|
||||
}
|
||||
}
|
||||
|
||||
// isLastAdmin reports whether id is an administrator and no other user is one.
|
||||
// A non-admin id is never the last one, so removing them is always allowed.
|
||||
func isLastAdmin(ctx context.Context, db *sql.DB, id int64) (bool, error) {
|
||||
var last bool
|
||||
err := db.QueryRowContext(ctx, `
|
||||
SELECT EXISTS (SELECT 1 FROM users WHERE id = $1 AND is_admin)
|
||||
AND NOT EXISTS (SELECT 1 FROM users WHERE id <> $1 AND is_admin)`, id).Scan(&last)
|
||||
return last, err
|
||||
}
|
||||
|
||||
@@ -0,0 +1,25 @@
|
||||
-- A system administrator role, and the first thing in this server that one user
|
||||
-- can do and another cannot.
|
||||
--
|
||||
-- Until now every authenticated caller could create and delete users, set
|
||||
-- anybody's password and mint API keys for anybody — auth.go said so in a
|
||||
-- comment. That was defensible with one operator and a hand-made account; it is
|
||||
-- not once people sign themselves up (see #7).
|
||||
--
|
||||
-- EVERY EXISTING USER BECOMES AN ADMIN. They already hold these powers, so
|
||||
-- this migration changes nobody's access: it names what is already true, and
|
||||
-- leaves demotion as a deliberate act somebody performs afterwards. The
|
||||
-- alternative — promoting only user 1 — would silently strip the others, and
|
||||
-- could leave an install whose only admin is an account nobody has a password
|
||||
-- for.
|
||||
--
|
||||
-- New users are not admins: the column defaults to false, and the only ways to
|
||||
-- become one are this backfill, the bootstrap endpoint, or an existing admin
|
||||
-- granting it.
|
||||
ALTER TABLE users ADD COLUMN is_admin BOOLEAN NOT NULL DEFAULT false;
|
||||
|
||||
UPDATE users SET is_admin = true;
|
||||
|
||||
-- The queue's assignment dropdown and the on-call schedule read every user, and
|
||||
-- the admin screens in #5 will filter on this.
|
||||
CREATE INDEX users_is_admin_idx ON users(is_admin) WHERE is_admin;
|
||||
@@ -0,0 +1,103 @@
|
||||
-- Teams: the unit of tenancy. Everything a person works on now belongs to one.
|
||||
--
|
||||
-- Until this migration the install was one shared space — every user saw every
|
||||
-- alert and every incident, and the Alertmanager webhook was unauthenticated, so
|
||||
-- anything that could reach the port could open an incident for everybody.
|
||||
--
|
||||
-- The shape, in one paragraph: a team owns its incidents, alerts, schedule and
|
||||
-- integrations. A user belongs to as many teams as they like, with a role in
|
||||
-- each: an `owner` configures the team, a `member` works its incidents. An
|
||||
-- integration key is what an alert arrives on, and the key is what says which
|
||||
-- team the alert belongs to.
|
||||
--
|
||||
-- EVERYTHING EXISTING MOVES INTO ONE DEFAULT TEAM, and every existing user
|
||||
-- becomes an owner of it. That keeps an upgrade a no-op for the people using it:
|
||||
-- the same queue, the same schedule, the same incidents, with a name on them.
|
||||
|
||||
CREATE TABLE teams (
|
||||
id BIGINT GENERATED BY DEFAULT AS IDENTITY PRIMARY KEY,
|
||||
name TEXT NOT NULL UNIQUE,
|
||||
created_at BIGINT NOT NULL DEFAULT FLOOR(EXTRACT(EPOCH FROM now()))::bigint
|
||||
);
|
||||
|
||||
-- role is free text with a CHECK rather than an enum, so adding a third role
|
||||
-- later is a migration and not a type rewrite.
|
||||
CREATE TABLE team_members (
|
||||
team_id BIGINT NOT NULL REFERENCES teams(id) ON DELETE CASCADE,
|
||||
user_id BIGINT NOT NULL REFERENCES users(id) ON DELETE CASCADE,
|
||||
role TEXT NOT NULL CHECK (role IN ('owner', 'member')),
|
||||
joined_at BIGINT NOT NULL DEFAULT FLOOR(EXTRACT(EPOCH FROM now()))::bigint,
|
||||
PRIMARY KEY (team_id, user_id)
|
||||
);
|
||||
|
||||
CREATE INDEX team_members_user_idx ON team_members(user_id);
|
||||
|
||||
-- How alerts get in, and the only thing that says which team they belong to.
|
||||
-- The key is stored as a SHA-256 hash, like api_keys and the ack tokens: a
|
||||
-- leaked database gives nobody the ability to post alerts.
|
||||
CREATE TABLE integrations (
|
||||
id BIGINT GENERATED BY DEFAULT AS IDENTITY PRIMARY KEY,
|
||||
team_id BIGINT NOT NULL REFERENCES teams(id) ON DELETE CASCADE,
|
||||
kind TEXT NOT NULL CHECK (kind IN ('alertmanager')),
|
||||
name TEXT NOT NULL,
|
||||
key_hash TEXT NOT NULL UNIQUE,
|
||||
created_at BIGINT NOT NULL DEFAULT FLOOR(EXTRACT(EPOCH FROM now()))::bigint,
|
||||
last_used_at BIGINT
|
||||
);
|
||||
|
||||
CREATE INDEX integrations_team_idx ON integrations(team_id);
|
||||
|
||||
-- ---------------------------------------------------------------------------
|
||||
-- The default team, and everything that already exists moving into it.
|
||||
--
|
||||
-- Created unconditionally, even on an empty install, so there is always a team
|
||||
-- for the bootstrap user to land in and for the first integration to hang off.
|
||||
-- ---------------------------------------------------------------------------
|
||||
|
||||
INSERT INTO teams (name) VALUES ('Default');
|
||||
|
||||
INSERT INTO team_members (team_id, user_id, role)
|
||||
SELECT (SELECT id FROM teams WHERE name = 'Default'), id, 'owner' FROM users;
|
||||
|
||||
-- ---------------------------------------------------------------------------
|
||||
-- team_id on everything a team owns.
|
||||
--
|
||||
-- Added nullable, backfilled, then made NOT NULL: adding a NOT NULL column with
|
||||
-- no default to a table with rows is rejected, and a DEFAULT pointing at the
|
||||
-- default team would quietly keep working after the default team is gone.
|
||||
-- ---------------------------------------------------------------------------
|
||||
|
||||
ALTER TABLE alerts ADD COLUMN team_id BIGINT REFERENCES teams(id) ON DELETE CASCADE;
|
||||
ALTER TABLE incidents ADD COLUMN team_id BIGINT REFERENCES teams(id) ON DELETE CASCADE;
|
||||
ALTER TABLE schedule_entries ADD COLUMN team_id BIGINT REFERENCES teams(id) ON DELETE CASCADE;
|
||||
|
||||
UPDATE alerts SET team_id = (SELECT id FROM teams WHERE name = 'Default');
|
||||
UPDATE incidents SET team_id = (SELECT id FROM teams WHERE name = 'Default');
|
||||
UPDATE schedule_entries SET team_id = (SELECT id FROM teams WHERE name = 'Default');
|
||||
|
||||
ALTER TABLE alerts ALTER COLUMN team_id SET NOT NULL;
|
||||
ALTER TABLE incidents ALTER COLUMN team_id SET NOT NULL;
|
||||
ALTER TABLE schedule_entries ALTER COLUMN team_id SET NOT NULL;
|
||||
|
||||
-- ---------------------------------------------------------------------------
|
||||
-- The uniqueness rules were all written for one tenant, and every one of them
|
||||
-- is wrong now: two teams monitoring two clusters legitimately see the same
|
||||
-- fingerprint, the same groupKey, and want somebody on call on the same day.
|
||||
-- ---------------------------------------------------------------------------
|
||||
|
||||
ALTER TABLE alerts DROP CONSTRAINT alerts_fingerprint_key;
|
||||
CREATE UNIQUE INDEX alerts_team_fingerprint_idx ON alerts(team_id, fingerprint);
|
||||
|
||||
DROP INDEX incidents_open_group_key_idx;
|
||||
-- Still load-bearing, now per team: at most one OPEN incident per group_key
|
||||
-- within a team. This is what makes "resolved incident + a new alert occurrence
|
||||
-- = a new incident" work, and what the webhook's find-or-open lookup relies on.
|
||||
CREATE UNIQUE INDEX incidents_open_group_key_idx
|
||||
ON incidents(team_id, group_key) WHERE resolved_at IS NULL;
|
||||
|
||||
ALTER TABLE schedule_entries DROP CONSTRAINT schedule_entries_date_key;
|
||||
CREATE UNIQUE INDEX schedule_entries_team_date_idx ON schedule_entries(team_id, date);
|
||||
|
||||
-- The list views all filter by team first.
|
||||
CREATE INDEX alerts_team_received_idx ON alerts(team_id, received_at DESC);
|
||||
CREATE INDEX incidents_team_triggered_idx ON incidents(team_id, triggered_at DESC);
|
||||
@@ -0,0 +1,39 @@
|
||||
-- Dead man's switches become a team's own configuration.
|
||||
--
|
||||
-- They were three environment variables — TERDUT_DEADMAN_MATCHERS, _TIMEOUT and
|
||||
-- _SEVERITY — which made them one setting for the whole install. That was the
|
||||
-- last piece of the alerting path a team could not control: a team could take
|
||||
-- its own alerts on its own key and still not say which of them were
|
||||
-- heartbeats, or how long a silence had to last before somebody was paged.
|
||||
--
|
||||
-- One row per team rather than one row per switch. The unit of monitoring is
|
||||
-- still the fingerprint, as it always was — two clusters sending the same
|
||||
-- heartbeat alertname are two independent switches — and the matcher string
|
||||
-- keeps the format the environment variable used, so a value can be moved from
|
||||
-- one to the other unchanged.
|
||||
--
|
||||
-- No rows are seeded here: a migration cannot read the environment. The server
|
||||
-- inserts a row per team at startup from its own configuration, and the same
|
||||
-- values therefore carry forward into the first team's row without anybody
|
||||
-- retyping them. See seedDeadmanConfigs.
|
||||
CREATE TABLE deadman_configs (
|
||||
team_id BIGINT PRIMARY KEY REFERENCES teams(id) ON DELETE CASCADE,
|
||||
|
||||
-- ";" separates matchers, "," the label conditions within one, "=" is exact
|
||||
-- equality: `alertname=Watchdog,cluster=prod; alertname=EdgeHeartbeat`.
|
||||
-- Every matcher must name an alertname. Empty watches nothing.
|
||||
matchers TEXT NOT NULL DEFAULT '',
|
||||
|
||||
-- Seconds rather than a Go duration string: the column is compared and
|
||||
-- arithmetic is done on it, and a value that has to be parsed before it can
|
||||
-- be believed is a value that can be stored unparseable. Zero disables the
|
||||
-- team's switches entirely.
|
||||
timeout_seconds BIGINT NOT NULL DEFAULT 0,
|
||||
|
||||
-- The severity these incidents open at. They have no member alerts to
|
||||
-- derive one from, and a heartbeat's own severity label is meaningless —
|
||||
-- Watchdog ships as "none".
|
||||
severity TEXT NOT NULL DEFAULT 'critical',
|
||||
|
||||
updated_at BIGINT NOT NULL DEFAULT FLOOR(EXTRACT(EPOCH FROM now()))::bigint
|
||||
);
|
||||
@@ -8,6 +8,12 @@ import "time"
|
||||
// Incident an alert belongs to.
|
||||
type Alert struct {
|
||||
ID int64 `json:"id"`
|
||||
|
||||
// TeamID is the team whose integration received this alert, and TeamName
|
||||
// rides along so a combined list can label a row without a second request.
|
||||
TeamID int64 `json:"team_id"`
|
||||
TeamName string `json:"team_name,omitempty"`
|
||||
|
||||
Fingerprint string `json:"fingerprint"`
|
||||
Name string `json:"name"`
|
||||
Status string `json:"status"` // "firing" or "resolved"
|
||||
|
||||
@@ -11,6 +11,12 @@ import "time"
|
||||
// the webhook and the sweeper may flip to "resolved" once every member alert has
|
||||
// stopped firing.
|
||||
type Incident struct {
|
||||
// TeamID is the team that owns this incident, fixed when it opens: an
|
||||
// incident never moves between teams. TeamName rides along so the combined
|
||||
// queue can badge each row without a second request.
|
||||
TeamID int64 `json:"team_id"`
|
||||
TeamName string `json:"team_name,omitempty"`
|
||||
|
||||
ID int64 `json:"id"`
|
||||
GroupKey string `json:"group_key"`
|
||||
Title string `json:"title"`
|
||||
|
||||
@@ -4,6 +4,13 @@ import "time"
|
||||
|
||||
type ScheduleEntry struct {
|
||||
ID int64 `json:"id"`
|
||||
|
||||
// TeamID is whose rota this shift belongs to; TeamName rides along so the
|
||||
// combined "who is on call" view can label each entry without a second
|
||||
// request.
|
||||
TeamID int64 `json:"team_id"`
|
||||
TeamName string `json:"team_name,omitempty"`
|
||||
|
||||
UserID int64 `json:"user_id"`
|
||||
Username string `json:"username"`
|
||||
Date string `json:"date"` // YYYY-MM-DD
|
||||
|
||||
@@ -0,0 +1,54 @@
|
||||
package models
|
||||
|
||||
import "time"
|
||||
|
||||
// Team is the unit of tenancy: it owns its incidents, alerts, schedule and
|
||||
// integrations, and a user sees exactly the teams they belong to.
|
||||
type Team struct {
|
||||
ID int64 `json:"id"`
|
||||
Name string `json:"name"`
|
||||
CreatedAt time.Time `json:"created_at"`
|
||||
|
||||
// Role is the caller's own role in this team, populated when a team is
|
||||
// listed for a particular person. Empty when nobody in particular is
|
||||
// asking, as in the admin listing.
|
||||
Role string `json:"role,omitempty"`
|
||||
}
|
||||
|
||||
// Team roles. An owner configures the team — its schedule, its integrations and
|
||||
// who is in it. A member works its incidents.
|
||||
const (
|
||||
RoleOwner = "owner"
|
||||
RoleMember = "member"
|
||||
)
|
||||
|
||||
// TeamMember is one person's membership of one team.
|
||||
type TeamMember struct {
|
||||
TeamID int64 `json:"team_id"`
|
||||
UserID int64 `json:"user_id"`
|
||||
Username string `json:"username"`
|
||||
Role string `json:"role"`
|
||||
JoinedAt time.Time `json:"joined_at"`
|
||||
}
|
||||
|
||||
// Integration is how alerts get in, and the only thing that says which team an
|
||||
// arriving alert belongs to.
|
||||
type Integration struct {
|
||||
ID int64 `json:"id"`
|
||||
TeamID int64 `json:"team_id"`
|
||||
Kind string `json:"kind"`
|
||||
Name string `json:"name"`
|
||||
CreatedAt time.Time `json:"created_at"`
|
||||
LastUsedAt *time.Time `json:"last_used_at,omitempty"`
|
||||
|
||||
// Key is the raw integration key, shown once when the integration is
|
||||
// created and never stored. URL is the address to point the sender at,
|
||||
// likewise only complete at creation time.
|
||||
Key string `json:"key,omitempty"`
|
||||
URL string `json:"url,omitempty"`
|
||||
}
|
||||
|
||||
// Integration kinds.
|
||||
const (
|
||||
IntegrationAlertmanager = "alertmanager"
|
||||
)
|
||||
@@ -12,6 +12,12 @@ type User struct {
|
||||
// none of their own; incidents assigned to them fall back to the configured
|
||||
// fallback topic instead.
|
||||
NtfyTopic *string `json:"ntfy_topic,omitempty"`
|
||||
|
||||
// IsAdmin is the system administrator flag: managing users and API keys.
|
||||
// Not omitempty — a client has to be able to tell "false" from "this server
|
||||
// is too old to have the field", and the web UI decides what to show from
|
||||
// it.
|
||||
IsAdmin bool `json:"is_admin"`
|
||||
}
|
||||
|
||||
type APIKey struct {
|
||||
|
||||
@@ -285,6 +285,16 @@ input:focus, textarea:focus { outline: none; border-color: var(--accent); box-sh
|
||||
min-width: 0;
|
||||
}
|
||||
.row-meta .labels { overflow: hidden; text-overflow: ellipsis; white-space: nowrap; min-width: 0; max-width: 100%; color: var(--faint); }
|
||||
/* Which team's queue a row came from. Only rendered for somebody in more than
|
||||
one team, so it never repeats the same word down the whole list. */
|
||||
/* Separates the status chips from the team chips in the queue's filter row. */
|
||||
.chip-sep { width: 1px; align-self: stretch; background: var(--border); margin: 0 2px; }
|
||||
|
||||
.row-team {
|
||||
padding: 1px 6px; border-radius: 4px;
|
||||
background: var(--surface-2); border: 1px solid var(--border);
|
||||
color: var(--muted); font-size: 12px; white-space: nowrap;
|
||||
}
|
||||
.row.resolved .row-title { color: var(--muted); }
|
||||
|
||||
.sev-critical { --sev: var(--crit); }
|
||||
|
||||
@@ -87,12 +87,11 @@ export const deleteNote = (id, eventID) => call('DELETE', `/incidents/${id}/note
|
||||
export const alerts = (query, opts) => call('GET', '/alerts', { query, ...opts });
|
||||
|
||||
// schedule
|
||||
export const schedule = (from, to) => call('GET', '/schedule', { query: { from, to } });
|
||||
export async function onCallNow() {
|
||||
try {
|
||||
return await call('GET', '/schedule/current');
|
||||
} catch (err) {
|
||||
if (err instanceof ApiError && err.status === 404) return null;
|
||||
throw err;
|
||||
}
|
||||
}
|
||||
export const teams = () => call('GET', '/teams');
|
||||
export const schedule = (teamID, from, to) =>
|
||||
call('GET', `/teams/${teamID}/schedule`, { query: { from, to } });
|
||||
|
||||
// One entry per team the viewer belongs to, for the teams that have somebody
|
||||
// scheduled today. An empty array means nobody anywhere, which is a real answer
|
||||
// rather than an error — unlike the pre-teams endpoint, which 404ed.
|
||||
export const onCallNow = () => call('GET', '/schedule/current');
|
||||
|
||||
@@ -3,7 +3,7 @@
|
||||
import * as api from './api.js';
|
||||
import * as ui from './ui.js';
|
||||
import * as poll from './poll.js';
|
||||
import { state, reset } from './state.js';
|
||||
import { state, reset, loadTeams } from './state.js';
|
||||
import * as queue from './queue.js';
|
||||
import * as incident from './incident.js';
|
||||
import * as oncall from './oncall.js';
|
||||
@@ -141,6 +141,7 @@ async function boot() {
|
||||
|
||||
try {
|
||||
state.me = await api.me();
|
||||
await loadTeams();
|
||||
showApp();
|
||||
} catch (err) {
|
||||
if (err.status === 401) showLogin();
|
||||
|
||||
@@ -1,10 +1,14 @@
|
||||
// On-call: who is on duty now, the week around it, and your own next shifts.
|
||||
// Read-only for now; the TUI edits the schedule.
|
||||
//
|
||||
// One team's rota at a time — the viewer's first team, since a viewer in one
|
||||
// team has nothing to choose between. "On call now" is the exception and shows
|
||||
// every team the viewer is in, because somebody on two rotas wants both.
|
||||
|
||||
import * as api from './api.js';
|
||||
import { h, clear, icon, spinner } from './ui.js';
|
||||
import { isoDate, mondayOf, addDays, isoWeek, initial } from './format.js';
|
||||
import { myID } from './state.js';
|
||||
import { myID, currentTeam } from './state.js';
|
||||
|
||||
const view = () => document.getElementById('view-oncall');
|
||||
|
||||
@@ -24,10 +28,17 @@ export async function refresh() {
|
||||
const start = weekStart;
|
||||
const today = new Date();
|
||||
try {
|
||||
const team = currentTeam();
|
||||
if (!team) {
|
||||
data = { now: [], week: [], upcoming: [] };
|
||||
error = null;
|
||||
render();
|
||||
return;
|
||||
}
|
||||
const [now, week, upcoming] = await Promise.all([
|
||||
api.onCallNow(),
|
||||
api.schedule(isoDate(start), isoDate(addDays(start, 6))),
|
||||
api.schedule(isoDate(today), isoDate(addDays(today, 60))),
|
||||
api.schedule(team.id, isoDate(start), isoDate(addDays(start, 6))),
|
||||
api.schedule(team.id, isoDate(today), isoDate(addDays(today, 60))),
|
||||
]);
|
||||
if (start !== weekStart) return;
|
||||
data = { now, week, upcoming };
|
||||
@@ -60,16 +71,33 @@ function you(userID) {
|
||||
return userID === myID() ? h('span', { class: 'you', text: 'you' }) : null;
|
||||
}
|
||||
|
||||
// One card per team with somebody on call, and a single empty card when there
|
||||
// is nobody anywhere. The team's name is shown only when the viewer is in more
|
||||
// than one, so the common case reads exactly as it did before teams existed.
|
||||
function nowCard() {
|
||||
const n = data.now;
|
||||
const entries = data.now || [];
|
||||
const showTeam = entries.length > 1;
|
||||
if (entries.length === 0) {
|
||||
return h('div', { class: 'card now-card' },
|
||||
h('div', { class: `avatar ${n ? '' : 'none'}`, text: n ? initial(n.username) : '–' }),
|
||||
h('div', { class: 'avatar none', text: '–' }),
|
||||
h('div', {},
|
||||
h('div', { class: 'now-label', text: 'On call now' }),
|
||||
h('div', { class: 'now-name' }, n ? n.username : 'Nobody', n && you(n.user_id)),
|
||||
h('div', { class: 'now-name', text: 'Nobody' }),
|
||||
),
|
||||
);
|
||||
}
|
||||
return h('div', {}, ...entries.map((n) =>
|
||||
h('div', { class: 'card now-card' },
|
||||
h('div', { class: 'avatar', text: initial(n.username) }),
|
||||
h('div', {},
|
||||
h('div', {
|
||||
class: 'now-label',
|
||||
text: showTeam ? `On call now · ${n.team_name}` : 'On call now',
|
||||
}),
|
||||
h('div', { class: 'now-name' }, n.username, you(n.user_id)),
|
||||
),
|
||||
)));
|
||||
}
|
||||
|
||||
function weekCard() {
|
||||
const byDate = new Map(data.week.map((e) => [e.date, e]));
|
||||
|
||||
@@ -26,12 +26,32 @@ const EMPTY = {
|
||||
};
|
||||
|
||||
let filter = loadFilter();
|
||||
let teamFilter = loadTeamFilter(); // '' for every team the viewer is in
|
||||
let items = null; // null while loading
|
||||
let error = null;
|
||||
let selected = null;
|
||||
let cursor = -1; // keyboard position in the list
|
||||
let built = false;
|
||||
|
||||
function loadTeamFilter() {
|
||||
try {
|
||||
return sessionStorage.getItem('terdut.queue.team') || '';
|
||||
} catch {
|
||||
return '';
|
||||
}
|
||||
}
|
||||
|
||||
function setTeamFilter(id) {
|
||||
teamFilter = id;
|
||||
try {
|
||||
sessionStorage.setItem('terdut.queue.team', id);
|
||||
} catch {
|
||||
/* storage unavailable */
|
||||
}
|
||||
renderChips();
|
||||
refresh({ fresh: true });
|
||||
}
|
||||
|
||||
function loadFilter() {
|
||||
try {
|
||||
const f = sessionStorage.getItem('terdut.queue.filter');
|
||||
@@ -64,7 +84,11 @@ export async function refresh({ fresh = false } = {}) {
|
||||
const requested = filter;
|
||||
try {
|
||||
// The open list is already fetched for the badges; no need to ask twice.
|
||||
const result = filter === 'open' && !fresh ? state.open : await api.incidents(f.query);
|
||||
// The cached open queue covers every team, so it can only be reused when
|
||||
// no team filter is applied.
|
||||
const query = teamFilter ? { ...f.query, team_id: teamFilter } : f.query;
|
||||
const cached = filter === 'open' && !fresh && !teamFilter;
|
||||
const result = cached ? state.open : await api.incidents(query);
|
||||
if (requested !== filter) return;
|
||||
items = result;
|
||||
error = null;
|
||||
@@ -88,7 +112,7 @@ function setFilter(id) {
|
||||
|
||||
function renderChips() {
|
||||
const el = document.getElementById('queue-filters');
|
||||
clear(el, FILTERS.map((f) =>
|
||||
const chips = FILTERS.map((f) =>
|
||||
h('button', {
|
||||
class: 'chip',
|
||||
type: 'button',
|
||||
@@ -97,7 +121,34 @@ function renderChips() {
|
||||
onclick: () => setFilter(f.id),
|
||||
text: f.label,
|
||||
}),
|
||||
));
|
||||
);
|
||||
|
||||
// Somebody in one team has nothing to choose between, so the row of team
|
||||
// chips appears only when there is more than one. The default is all of
|
||||
// them: the combined queue is the point.
|
||||
if (state.teams.length > 1) {
|
||||
chips.push(h('span', { class: 'chip-sep' }));
|
||||
chips.push(h('button', {
|
||||
class: 'chip',
|
||||
type: 'button',
|
||||
role: 'tab',
|
||||
'aria-selected': String(teamFilter === ''),
|
||||
onclick: () => setTeamFilter(''),
|
||||
text: 'All teams',
|
||||
}));
|
||||
for (const team of state.teams) {
|
||||
chips.push(h('button', {
|
||||
class: 'chip',
|
||||
type: 'button',
|
||||
role: 'tab',
|
||||
'aria-selected': String(teamFilter === String(team.id)),
|
||||
onclick: () => setTeamFilter(String(team.id)),
|
||||
text: team.name,
|
||||
}));
|
||||
}
|
||||
}
|
||||
|
||||
clear(el, chips);
|
||||
}
|
||||
|
||||
function renderList() {
|
||||
@@ -142,6 +193,13 @@ function row(inc, index) {
|
||||
// The server already puts the group labels in the title; show only the rest.
|
||||
const labels = labelSummary(Object.fromEntries(
|
||||
Object.entries(inc.group_labels || {}).filter(([k, v]) => !inc.title.includes(`${k}=${v}`))));
|
||||
// The team is shown only to somebody who is in more than one. For everybody
|
||||
// else it is the same word on every row, which is noise rather than
|
||||
// information.
|
||||
const team = state.teams.length > 1 && inc.team_name
|
||||
? h('span', { class: 'row-team', text: inc.team_name })
|
||||
: null;
|
||||
|
||||
return h('a', {
|
||||
class: `row ${severityClass(inc.severity)} ${resolved ? 'resolved' : ''} ${index === cursor ? 'kbd-focus' : ''}`,
|
||||
href: `/incidents/${inc.id}`,
|
||||
@@ -153,6 +211,7 @@ function row(inc, index) {
|
||||
h('div', { class: 'row-meta' },
|
||||
status,
|
||||
assignee,
|
||||
team,
|
||||
labels && h('span', { class: 'labels', text: labels }),
|
||||
),
|
||||
);
|
||||
|
||||
@@ -6,8 +6,15 @@ import * as api from './api.js';
|
||||
export const state = {
|
||||
me: null, // { user, has_password }
|
||||
open: [], // the default queue: open, not snoozed
|
||||
teams: [], // the teams the viewer belongs to, each with their role
|
||||
};
|
||||
|
||||
// The team whose schedule and settings the views act on. A viewer in one team —
|
||||
// which is everybody until somebody makes a second — never has to choose.
|
||||
export function currentTeam() {
|
||||
return state.teams[0] || null;
|
||||
}
|
||||
|
||||
export function myID() {
|
||||
return state.me ? state.me.user.id : null;
|
||||
}
|
||||
@@ -24,8 +31,14 @@ export async function users() {
|
||||
return usersCache;
|
||||
}
|
||||
|
||||
export async function loadTeams() {
|
||||
state.teams = await api.teams();
|
||||
return state.teams;
|
||||
}
|
||||
|
||||
export function reset() {
|
||||
state.me = null;
|
||||
state.open = [];
|
||||
state.teams = [];
|
||||
usersCache = null;
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user