Compare commits
1
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
0b1d766765 |
@@ -161,8 +161,7 @@ and when it stops. Its routes:
|
|||||||
- `GET /.well-known/healthcheck` — answers 200 with `"status":"ok"`, the
|
- `GET /.well-known/healthcheck` — answers 200 with `"status":"ok"`, the
|
||||||
server's version and its uptime
|
server's version and its uptime
|
||||||
- `GET /metrics` — Prometheus metrics behind basic auth, only when
|
- `GET /metrics` — Prometheus metrics behind basic auth, only when
|
||||||
`METRICS_USERNAME` and `METRICS_PASSWORD` are set; each client address may
|
`METRICS_USERNAME` and `METRICS_PASSWORD` are set
|
||||||
make a limited number of requests to it a minute
|
|
||||||
|
|
||||||
In the image, the `builder` stage of `Dockerfile` tests it and builds it with
|
In the image, the `builder` stage of `Dockerfile` tests it and builds it with
|
||||||
`backend/script/build`, and `bin/entrypoint.sh` runs it as user `netwatch` on
|
`backend/script/build`, and `bin/entrypoint.sh` runs it as user `netwatch` on
|
||||||
|
|||||||
@@ -28,19 +28,11 @@ Write the latency statistics once and move the thresholds written inline in
|
|||||||
the image builds and runs it, and points to `backend/README.md` for its
|
the image builds and runs it, and points to `backend/README.md` for its
|
||||||
settings; the checks are GET requests; the 26 WAN hosts, the four health
|
settings; the checks are GET requests; the 26 WAN hosts, the four health
|
||||||
states, the summary's figures and the features the list lacked are described
|
states, the summary's figures and the features the list lacked are described
|
||||||
as the page has them; and its TODO section points here, as does the one in
|
as the page has them; and its TODO section points here. This file's Workflow
|
||||||
`backend/README.md`, whose open items moved to Future Steps. This file's
|
branches from `next` and opens the PR against `next`, Status says where the
|
||||||
Workflow branches from `next` and opens the PR against `next`, Status says
|
repo stands, and Next Step and Future Steps hold only open work, linked to its
|
||||||
where the repo stands, and Next Step and Future Steps hold only open work,
|
issue where one exists. The viewport harness README names Node's test runner,
|
||||||
linked to its issue where one exists. The viewport harness README names Node's
|
not `vitest`
|
||||||
test runner, not `vitest`
|
|
||||||
- 2026-10-04: password guesses at `/metrics` are rate limited (issue #104): each
|
|
||||||
client address, resolved through `TRUSTED_PROXIES` as for reports, may make 60
|
|
||||||
requests to `/metrics` a minute, counted by `go-chi/httprate` apart from its
|
|
||||||
reports; past that it gets 429 and its basic auth credentials are not checked.
|
|
||||||
The limit is a constant in `backend/internal/server/routes.go`. A test uses up
|
|
||||||
one client's allowance on wrong passwords, gets 429 with the right one, and
|
|
||||||
checks that another client behind the same nginx still gets in
|
|
||||||
- 2026-10-04: the backend reports errors to Sentry (issue #95). With
|
- 2026-10-04: the backend reports errors to Sentry (issue #95). With
|
||||||
`SENTRY_DSN` set, it sets up `sentry-go` with the release `netwatch-server-`
|
`SENTRY_DSN` set, it sets up `sentry-go` with the release `netwatch-server-`
|
||||||
and its version, reports each panic in a handler through `sentryhttp`, the
|
and its version, reports each panic in a handler through `sentryhttp`, the
|
||||||
@@ -375,19 +367,13 @@ Write the latency statistics once and move the thresholds written inline in
|
|||||||
|
|
||||||
# Future Steps
|
# Future Steps
|
||||||
|
|
||||||
- Take "IPv4 only" out of the page's footer, as nothing in the page limits a
|
- Rate limit password attempts on `/metrics`
|
||||||
check to IPv4 ([#111](https://git.eeqj.de/sneak/netwatch/issues/111))
|
([#104](https://git.eeqj.de/sneak/netwatch/issues/104))
|
||||||
- Decide whether the repo moves to the layout `REPO_POLICIES.md` gives, with
|
- Decide whether the repo moves to the layout `REPO_POLICIES.md` gives, with
|
||||||
`backend/` no longer repeating files from the root
|
`backend/` no longer repeating files from the root
|
||||||
([#30](https://git.eeqj.de/sneak/netwatch/issues/30))
|
([#30](https://git.eeqj.de/sneak/netwatch/issues/30))
|
||||||
- Run `make frontend-viewport-test` in CI as its own step; it is not part of
|
- Run `make frontend-viewport-test` in CI as its own step; it is not part of
|
||||||
`make check`, as it needs Docker and takes minutes
|
`make check`, as it needs Docker and takes minutes
|
||||||
- A backend test that posts a report to `POST /api/v1/reports` and checks the
|
|
||||||
compressed file it is written to
|
|
||||||
- A backend route that decompresses the stored reports and answers queries on
|
|
||||||
them
|
|
||||||
- Prometheus metrics for the backend's in-memory buffer: its size, the number of
|
|
||||||
flushes and the number of reports
|
|
||||||
- A configurable host list (an environment variable or a config file)
|
- A configurable host list (an environment variable or a config file)
|
||||||
- Export of the latency history (CSV or JSON)
|
- Export of the latency history (CSV or JSON)
|
||||||
- A notification when the health status changes to DEGRADED
|
- A notification when the health status changes to DEGRADED
|
||||||
|
|||||||
+3
-10
@@ -199,14 +199,6 @@ is recorded and `/metrics` answers 404. One without the other stops the server
|
|||||||
from starting, with an error naming both; so does a `METRICS_USERNAME`
|
from starting, with an error naming both; so does a `METRICS_USERNAME`
|
||||||
containing `:`, which basic auth cannot carry, with an error naming it.
|
containing `:`, which basic auth cannot carry, with an error naming it.
|
||||||
|
|
||||||
`/metrics` is rate limited, so that its password cannot be guessed quickly: each
|
|
||||||
client address, resolved through `TRUSTED_PROXIES`, may make 60 requests to it a
|
|
||||||
minute, whatever their credentials. Past that it gets 429 with
|
|
||||||
`Retry-After: 60`, and its credentials are not checked. The minute slides as it
|
|
||||||
does for reports (see [Report limits](#report-limits)), so a scraper polling
|
|
||||||
every 2 seconds or less often is never refused. This allowance is apart from the
|
|
||||||
one for reports.
|
|
||||||
|
|
||||||
### Sentry
|
### Sentry
|
||||||
|
|
||||||
With `SENTRY_DSN` set, the server sends its errors to that Sentry project: each
|
With `SENTRY_DSN` set, the server sends its errors to that Sentry project: each
|
||||||
@@ -219,8 +211,9 @@ sent to it.
|
|||||||
|
|
||||||
## TODO
|
## TODO
|
||||||
|
|
||||||
The to-do list, this backend's open work included, is [TODO.md](../TODO.md) at
|
- Add integration test that POSTs a report and verifies the compressed output
|
||||||
the repo root.
|
- Add report decompression/query endpoint
|
||||||
|
- Add metrics (Prometheus) for buffer size, flush count, report count
|
||||||
|
|
||||||
## License
|
## License
|
||||||
|
|
||||||
|
|||||||
@@ -12,10 +12,6 @@ func (s *Server) Router() *chi.Mux {
|
|||||||
// external tests.
|
// external tests.
|
||||||
const MaxRequestBodyBytes = maxRequestBodyBytes
|
const MaxRequestBodyBytes = maxRequestBodyBytes
|
||||||
|
|
||||||
// MetricsRequestsPerMinute exposes the /metrics rate limit to the
|
|
||||||
// external tests.
|
|
||||||
const MetricsRequestsPerMinute = metricsRequestsPerMinute
|
|
||||||
|
|
||||||
// ListenAddr exposes the address the server listens on to the
|
// ListenAddr exposes the address the server listens on to the
|
||||||
// external tests.
|
// external tests.
|
||||||
func (s *Server) ListenAddr() string {
|
func (s *Server) ListenAddr() string {
|
||||||
|
|||||||
@@ -18,12 +18,6 @@ const (
|
|||||||
// can mount s.mw.MaxBodyBytes with a smaller value to lower
|
// can mount s.mw.MaxBodyBytes with a smaller value to lower
|
||||||
// its bound, but cannot raise it: this cap runs first.
|
// its bound, but cannot raise it: this cap runs first.
|
||||||
maxRequestBodyBytes int64 = 1 << 20 // 1 MiB
|
maxRequestBodyBytes int64 = 1 << 20 // 1 MiB
|
||||||
|
|
||||||
// metricsRequestsPerMinute is how many requests to /metrics each
|
|
||||||
// client address may make a minute, whatever their credentials. A
|
|
||||||
// scraper polling every 2 seconds sends half of it, which httprate
|
|
||||||
// never refuses.
|
|
||||||
metricsRequestsPerMinute = 60
|
|
||||||
)
|
)
|
||||||
|
|
||||||
// SetupRoutes configures the chi router with middleware and
|
// SetupRoutes configures the chi router with middleware and
|
||||||
@@ -72,14 +66,10 @@ func (s *Server) SetupRoutes() {
|
|||||||
Post("/api/v1/reports", s.h.HandleReport())
|
Post("/api/v1/reports", s.h.HandleReport())
|
||||||
})
|
})
|
||||||
|
|
||||||
// The rate limit comes before the basic auth, so a client past it
|
|
||||||
// gets 429 and its password is not checked.
|
|
||||||
if s.params.Config.MetricsUsername != "" {
|
if s.params.Config.MetricsUsername != "" {
|
||||||
s.router.With(
|
s.router.With(s.mw.MetricsAuth()).
|
||||||
s.mw.RateLimit(metricsRequestsPerMinute),
|
Get("/metrics", promhttp.HandlerFor(
|
||||||
s.mw.MetricsAuth(),
|
registry, promhttp.HandlerOpts{},
|
||||||
).Get("/metrics", promhttp.HandlerFor(
|
).ServeHTTP)
|
||||||
registry, promhttp.HandlerOpts{},
|
|
||||||
).ServeHTTP)
|
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -188,50 +188,6 @@ func TestMetricsBehindBasicAuth(t *testing.T) {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
// TestMetricsAreRateLimited: a client that has used up its /metrics
|
|
||||||
// allowance on wrong passwords gets 429 even with the right one, which
|
|
||||||
// is then not checked, while another client behind the same nginx
|
|
||||||
// still gets in.
|
|
||||||
func TestMetricsAreRateLimited(t *testing.T) {
|
|
||||||
t.Setenv("METRICS_USERNAME", "prometheus")
|
|
||||||
t.Setenv("METRICS_PASSWORD", "right")
|
|
||||||
// As in the container: nginx connects from loopback and names the
|
|
||||||
// client in X-Forwarded-For.
|
|
||||||
t.Setenv("TRUSTED_PROXIES", "127.0.0.1/32")
|
|
||||||
|
|
||||||
srv := newServer(t)
|
|
||||||
srv.SetupRoutes()
|
|
||||||
|
|
||||||
get := func(client, password string) int {
|
|
||||||
rec := httptest.NewRecorder()
|
|
||||||
req := httptest.NewRequestWithContext(t.Context(),
|
|
||||||
http.MethodGet, "/metrics", http.NoBody)
|
|
||||||
req.RemoteAddr = "127.0.0.1:40000"
|
|
||||||
req.Header.Set("X-Forwarded-For", client)
|
|
||||||
req.SetBasicAuth("prometheus", password)
|
|
||||||
srv.ServeHTTP(rec, req)
|
|
||||||
|
|
||||||
return rec.Code
|
|
||||||
}
|
|
||||||
|
|
||||||
for i := range server.MetricsRequestsPerMinute {
|
|
||||||
if code := get("203.0.113.7", "wrong"); code != http.StatusUnauthorized {
|
|
||||||
t.Fatalf("guess %d: status = %d, want %d",
|
|
||||||
i+1, code, http.StatusUnauthorized)
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
if code := get("203.0.113.7", "right"); code != http.StatusTooManyRequests {
|
|
||||||
t.Fatalf("right password past the limit: status = %d, want %d",
|
|
||||||
code, http.StatusTooManyRequests)
|
|
||||||
}
|
|
||||||
|
|
||||||
if code := get("203.0.113.8", "right"); code != http.StatusOK {
|
|
||||||
t.Fatalf("another client: status = %d, want %d",
|
|
||||||
code, http.StatusOK)
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
// TestMetricsInTwoServers: two servers in one process can both have
|
// TestMetricsInTwoServers: two servers in one process can both have
|
||||||
// metrics on.
|
// metrics on.
|
||||||
func TestMetricsInTwoServers(t *testing.T) {
|
func TestMetricsInTwoServers(t *testing.T) {
|
||||||
|
|||||||
Reference in New Issue
Block a user