2 Commits
Author SHA1 Message Date
sneak 7a86284e2b Latency statistics written once, thresholds read from CONFIG (closes #102)
check / check (push) Successful in 3m24s
A target's min, max, median and average latency now come from one list
of its answers, through latencyStats(), which the summary's figures use
too, so the median is written once. latencyHex() and latencyClass() read
one table of color limits in CONFIG. The health thresholds, the debug
log's length, the gateway check's timeout, the recovery probe's number
of hosts and interval, how often the rows are sorted and the delay
before the first sparkline resize are CONFIG entries, read where the
numbers were. HostState's minLatency(), maxLatency(), averageLatency()
and medianLatency() are gone; the statistics test reads historyStats().
Nothing the page does or shows changes. A unit test now covers the
summary's figures.

Model: opus-5-5
2026-10-04 04:32:10 +00:00
clawbot d40e67d4ab Rate limit password attempts on /metrics (closes #104)
check / check (push) Successful in 3m32s
Each client address may make 60 requests to /metrics a minute,
through the same httprate middleware and TRUSTED_PROXIES
resolution the report route uses, with an allowance of its own.
The limit runs before the basic auth, so past it the answer is 429
and the password is not checked. backend/README.md says so; a test
uses up one client's allowance on wrong passwords, gets 429 with
the right one, and checks that another client behind the same
nginx still gets in.

Model: opus-5-5
2026-10-04 06:19:06 +02:00
7 changed files with 206 additions and 96 deletions
+17
View File
@@ -23,6 +23,23 @@ latest run passes.
# Completed Steps # Completed Steps
- 2026-10-04: in `src/main.js` (issue #102), a target's min, max, median and
average latency come from one list of its answers, through the same function
the summary's figures use, so the median is written once. The latency color
limits are one table in `CONFIG`, read by both the figure's and the
sparkline's color. The health thresholds, the debug log's length, the gateway
check's timeout, the recovery probe's number of hosts and interval, how often
the rows are sorted and the delay before the first sparkline resize are
`CONFIG` entries too. A unit test checks the summary's figures. Nothing the
page does or shows changed; the footer's color legend still writes the limits
out as text
- 2026-10-04: password guesses at `/metrics` are rate limited (issue #104): each
client address, resolved through `TRUSTED_PROXIES` as for reports, may make 60
requests to `/metrics` a minute, counted by `go-chi/httprate` apart from its
reports; past that it gets 429 and its basic auth credentials are not checked.
The limit is a constant in `backend/internal/server/routes.go`. A test uses up
one client's allowance on wrong passwords, gets 429 with the right one, and
checks that another client behind the same nginx still gets in
- 2026-10-04: the backend reports errors to Sentry (issue #95). With - 2026-10-04: the backend reports errors to Sentry (issue #95). With
`SENTRY_DSN` set, it sets up `sentry-go` with the release `netwatch-server-` `SENTRY_DSN` set, it sets up `sentry-go` with the release `netwatch-server-`
and its version, reports each panic in a handler through `sentryhttp`, the and its version, reports each panic in a handler through `sentryhttp`, the
+8
View File
@@ -199,6 +199,14 @@ is recorded and `/metrics` answers 404. One without the other stops the server
from starting, with an error naming both; so does a `METRICS_USERNAME` from starting, with an error naming both; so does a `METRICS_USERNAME`
containing `:`, which basic auth cannot carry, with an error naming it. containing `:`, which basic auth cannot carry, with an error naming it.
`/metrics` is rate limited, so that its password cannot be guessed quickly: each
client address, resolved through `TRUSTED_PROXIES`, may make 60 requests to it a
minute, whatever their credentials. Past that it gets 429 with
`Retry-After: 60`, and its credentials are not checked. The minute slides as it
does for reports (see [Report limits](#report-limits)), so a scraper polling
every 2 seconds or less often is never refused. This allowance is apart from the
one for reports.
### Sentry ### Sentry
With `SENTRY_DSN` set, the server sends its errors to that Sentry project: each With `SENTRY_DSN` set, the server sends its errors to that Sentry project: each
+4
View File
@@ -12,6 +12,10 @@ func (s *Server) Router() *chi.Mux {
// external tests. // external tests.
const MaxRequestBodyBytes = maxRequestBodyBytes const MaxRequestBodyBytes = maxRequestBodyBytes
// MetricsRequestsPerMinute exposes the /metrics rate limit to the
// external tests.
const MetricsRequestsPerMinute = metricsRequestsPerMinute
// ListenAddr exposes the address the server listens on to the // ListenAddr exposes the address the server listens on to the
// external tests. // external tests.
func (s *Server) ListenAddr() string { func (s *Server) ListenAddr() string {
+14 -4
View File
@@ -18,6 +18,12 @@ const (
// can mount s.mw.MaxBodyBytes with a smaller value to lower // can mount s.mw.MaxBodyBytes with a smaller value to lower
// its bound, but cannot raise it: this cap runs first. // its bound, but cannot raise it: this cap runs first.
maxRequestBodyBytes int64 = 1 << 20 // 1 MiB maxRequestBodyBytes int64 = 1 << 20 // 1 MiB
// metricsRequestsPerMinute is how many requests to /metrics each
// client address may make a minute, whatever their credentials. A
// scraper polling every 2 seconds sends half of it, which httprate
// never refuses.
metricsRequestsPerMinute = 60
) )
// SetupRoutes configures the chi router with middleware and // SetupRoutes configures the chi router with middleware and
@@ -66,10 +72,14 @@ func (s *Server) SetupRoutes() {
Post("/api/v1/reports", s.h.HandleReport()) Post("/api/v1/reports", s.h.HandleReport())
}) })
// The rate limit comes before the basic auth, so a client past it
// gets 429 and its password is not checked.
if s.params.Config.MetricsUsername != "" { if s.params.Config.MetricsUsername != "" {
s.router.With(s.mw.MetricsAuth()). s.router.With(
Get("/metrics", promhttp.HandlerFor( s.mw.RateLimit(metricsRequestsPerMinute),
registry, promhttp.HandlerOpts{}, s.mw.MetricsAuth(),
).ServeHTTP) ).Get("/metrics", promhttp.HandlerFor(
registry, promhttp.HandlerOpts{},
).ServeHTTP)
} }
} }
+44
View File
@@ -188,6 +188,50 @@ func TestMetricsBehindBasicAuth(t *testing.T) {
} }
} }
// TestMetricsAreRateLimited: a client that has used up its /metrics
// allowance on wrong passwords gets 429 even with the right one, which
// is then not checked, while another client behind the same nginx
// still gets in.
func TestMetricsAreRateLimited(t *testing.T) {
t.Setenv("METRICS_USERNAME", "prometheus")
t.Setenv("METRICS_PASSWORD", "right")
// As in the container: nginx connects from loopback and names the
// client in X-Forwarded-For.
t.Setenv("TRUSTED_PROXIES", "127.0.0.1/32")
srv := newServer(t)
srv.SetupRoutes()
get := func(client, password string) int {
rec := httptest.NewRecorder()
req := httptest.NewRequestWithContext(t.Context(),
http.MethodGet, "/metrics", http.NoBody)
req.RemoteAddr = "127.0.0.1:40000"
req.Header.Set("X-Forwarded-For", client)
req.SetBasicAuth("prometheus", password)
srv.ServeHTTP(rec, req)
return rec.Code
}
for i := range server.MetricsRequestsPerMinute {
if code := get("203.0.113.7", "wrong"); code != http.StatusUnauthorized {
t.Fatalf("guess %d: status = %d, want %d",
i+1, code, http.StatusUnauthorized)
}
}
if code := get("203.0.113.7", "right"); code != http.StatusTooManyRequests {
t.Fatalf("right password past the limit: status = %d, want %d",
code, http.StatusTooManyRequests)
}
if code := get("203.0.113.8", "right"); code != http.StatusOK {
t.Fatalf("another client: status = %d, want %d",
code, http.StatusOK)
}
}
// TestMetricsInTwoServers: two servers in one process can both have // TestMetricsInTwoServers: two servers in one process can both have
// metrics on. // metrics on.
func TestMetricsInTwoServers(t *testing.T) { func TestMetricsInTwoServers(t *testing.T) {
+93 -82
View File
@@ -30,6 +30,39 @@ export const CONFIG = {
return [0, 1, 2, 3, 4, 5].map((i) => Math.round((d * i) / 5)); return [0, 1, 2, 3, 4, 5].map((i) => Math.round((d * i) / 5));
}, },
canvasHeight: 96, canvasHeight: 96,
// A latency figure and its sparkline take the color of the first entry
// whose limit, in ms, the latency is below.
latencyColors: [
{ below: 50, hex: "#22c55e", className: "text-green-500" },
{ below: 100, hex: "#84cc16", className: "text-lime-500" },
{ below: 200, hex: "#eab308", className: "text-yellow-500" },
{ below: 500, hex: "#f97316", className: "text-orange-500" },
{ below: Infinity, hex: "#ef4444", className: "text-red-500" },
],
// The health is offline when more than offlineTimeouts WAN hosts timed
// out or were unreachable and at most offlineReachable answered;
// otherwise degraded when more than degradedTimeouts timed out or were
// unreachable; otherwise slow when more than slowHosts answered after
// more than slowLatency ms.
offlineTimeouts: 10,
offlineReachable: 4,
degradedTimeouts: 4,
slowHosts: 3,
slowLatency: 1000,
// The debug log keeps its last maxLogEntries lines.
maxLogEntries: 1000,
// A gateway candidate that has not answered after gatewayTimeout ms is
// passed over.
gatewayTimeout: 1500,
// When no WAN host answers, the recovery probe checks recoveryProbeHosts
// random ones every recoveryProbeInterval ms.
recoveryProbeHosts: 4,
recoveryProbeInterval: 500,
// The rows are sorted after the first round that is not discarded, then
// every roundsPerSort rounds.
roundsPerSort: 10,
// The sparklines are sized and drawn again resizeDelay ms after start.
resizeDelay: 100,
}; };
// WAN endpoints to monitor. These are used for the aggregate health/stats // WAN endpoints to monitor. These are used for the aggregate health/stats
@@ -114,7 +147,8 @@ const debugLog = [];
const log = (() => { const log = (() => {
function append(level, message) { function append(level, message) {
debugLog.push({ timestamp: new Date(), level, message }); debugLog.push({ timestamp: new Date(), level, message });
if (debugLog.length > 1000) debugLog.splice(0, debugLog.length - 1000); if (debugLog.length > CONFIG.maxLogEntries)
debugLog.splice(0, debugLog.length - CONFIG.maxLogEntries);
const panel = document.getElementById("debug-panel"); const panel = document.getElementById("debug-panel");
if (panel && !panel.classList.contains("hidden")) renderDebugLog(); if (panel && !panel.classList.contains("hidden")) renderDebugLog();
} }
@@ -174,7 +208,10 @@ async function detectGateway() {
const result = await Promise.any( const result = await Promise.any(
GATEWAY_CANDIDATES.map(async (url) => { GATEWAY_CANDIDATES.map(async (url) => {
const controller = new AbortController(); const controller = new AbortController();
const timeoutId = setTimeout(() => controller.abort(), 1500); const timeoutId = setTimeout(
() => controller.abort(),
CONFIG.gatewayTimeout,
);
try { try {
await fetch(url, { await fetch(url, {
method: "GET", method: "GET",
@@ -199,6 +236,27 @@ async function detectGateway() {
// --- App State --------------------------------------------------------------- // --- App State ---------------------------------------------------------------
// The min, max, median and average of latencies, a list of numbers, or all
// null when it is empty. The median of an even count is the mean of the
// middle two; it and the average are rounded.
function latencyStats(latencies) {
if (latencies.length === 0)
return { min: null, max: null, med: null, avg: null };
const sorted = [...latencies].sort((a, b) => a - b);
const mid = Math.floor(sorted.length / 2);
return {
min: sorted[0],
max: sorted[sorted.length - 1],
med:
sorted.length % 2
? sorted[mid]
: Math.round((sorted[mid - 1] + sorted[mid]) / 2),
avg: Math.round(
latencies.reduce((a, b) => a + b, 0) / latencies.length,
),
};
}
export class HostState { export class HostState {
constructor(host, pinned = false) { constructor(host, pinned = false) {
this.name = host.name; this.name = host.name;
@@ -230,38 +288,16 @@ export class HostState {
this._trim(); this._trim();
} }
averageLatency() { // The min, max, median and average latency of the checks in the history
const valid = this.history.filter((p) => p.latency !== null); // that got an answer.
if (valid.length === 0) return null; historyStats() {
return Math.round( return latencyStats(
valid.reduce((s, p) => s + p.latency, 0) / valid.length, this.history
.filter((p) => p.latency !== null)
.map((p) => p.latency),
); );
} }
minLatency() {
const valid = this.history.filter((p) => p.latency !== null);
if (valid.length === 0) return null;
return Math.min(...valid.map((p) => p.latency));
}
maxLatency() {
const valid = this.history.filter((p) => p.latency !== null);
if (valid.length === 0) return null;
return Math.max(...valid.map((p) => p.latency));
}
medianLatency() {
const sorted = this.history
.filter((p) => p.latency !== null)
.map((p) => p.latency)
.sort((a, b) => a - b);
if (sorted.length === 0) return null;
const mid = Math.floor(sorted.length / 2);
return sorted.length % 2
? sorted[mid]
: Math.round((sorted[mid - 1] + sorted[mid]) / 2);
}
_trim() { _trim() {
while (this.history.length > CONFIG.maxHistoryPoints) while (this.history.length > CONFIG.maxHistoryPoints)
this.history.shift(); this.history.shift();
@@ -288,33 +324,13 @@ export class AppState {
/** WAN-only stats from latest sample (excludes local) */ /** WAN-only stats from latest sample (excludes local) */
wanStats() { wanStats() {
const reachable = this.wan.filter((h) => h.lastLatency !== null); const latencies = this.wan
const latencies = reachable.map((h) => h.lastLatency); .filter((h) => h.lastLatency !== null)
const total = this.wan.length; .map((h) => h.lastLatency);
if (latencies.length === 0)
return {
reachable: 0,
total,
min: null,
max: null,
med: null,
avg: null,
};
const sorted = [...latencies].sort((a, b) => a - b);
const mid = Math.floor(sorted.length / 2);
const med =
sorted.length % 2
? sorted[mid]
: Math.round((sorted[mid - 1] + sorted[mid]) / 2);
return { return {
reachable: latencies.length, reachable: latencies.length,
total, total: this.wan.length,
min: Math.min(...latencies), ...latencyStats(latencies),
max: Math.max(...latencies),
med,
avg: Math.round(
latencies.reduce((a, b) => a + b, 0) / latencies.length,
),
}; };
} }
@@ -340,12 +356,16 @@ export class AppState {
const timeouts = this.wan.filter( const timeouts = this.wan.filter(
(h) => h.status === "error" || h.status === "offline", (h) => h.status === "error" || h.status === "offline",
).length; ).length;
if (timeouts > 10 && reachable <= 4) return "offline"; if (
if (timeouts > 4) return "degraded"; timeouts > CONFIG.offlineTimeouts &&
reachable <= CONFIG.offlineReachable
)
return "offline";
if (timeouts > CONFIG.degradedTimeouts) return "degraded";
const slow = this.wan.filter( const slow = this.wan.filter(
(h) => h.lastLatency !== null && h.lastLatency > 1000, (h) => h.lastLatency !== null && h.lastLatency > CONFIG.slowLatency,
).length; ).length;
if (slow > 3) return "slow"; if (slow > CONFIG.slowHosts) return "slow";
return "healthy"; return "healthy";
} }
@@ -557,21 +577,13 @@ export async function measureLatency(url, signal) {
export function latencyHex(latency) { export function latencyHex(latency) {
if (latency === null) return "#6b7280"; if (latency === null) return "#6b7280";
if (latency < 50) return "#22c55e"; return CONFIG.latencyColors.find((c) => latency < c.below).hex;
if (latency < 100) return "#84cc16";
if (latency < 200) return "#eab308";
if (latency < 500) return "#f97316";
return "#ef4444";
} }
export function latencyClass(latency, status) { export function latencyClass(latency, status) {
if (status === "offline" || status === "error" || latency === null) if (status === "offline" || status === "error" || latency === null)
return "text-gray-500"; return "text-gray-500";
if (latency < 50) return "text-green-500"; return CONFIG.latencyColors.find((c) => latency < c.below).className;
if (latency < 100) return "text-lime-500";
if (latency < 200) return "text-yellow-500";
if (latency < 500) return "text-orange-500";
return "text-red-500";
} }
// --- Sparkline Renderer ------------------------------------------------------ // --- Sparkline Renderer ------------------------------------------------------
@@ -912,10 +924,7 @@ function updateHostRow(host, index) {
latencyEl.innerHTML = `<span class="text-gray-500">---</span>`; latencyEl.innerHTML = `<span class="text-gray-500">---</span>`;
} }
const avg = host.averageLatency(); const { min, med, avg, max } = host.historyStats();
const med = host.medianLatency();
const min = host.minLatency();
const max = host.maxLatency();
if (host.status === "online" && avg !== null) { if (host.status === "online" && avg !== null) {
statusEl.innerHTML = statusStatsHTML([ statusEl.innerHTML = statusStatsHTML([
["min", min], ["min", min],
@@ -1182,8 +1191,9 @@ export async function tick(state, signal, onOffline) {
// rows whose check ended before the resume still read "paused" // rows whose check ended before the resume still read "paused"
state.allHosts.forEach((host, i) => updateHostRow(host, i)); state.allHosts.forEach((host, i) => updateHostRow(host, i));
// Sort after the first real check, then every 10 ticks thereafter // Sort after the first real check, then every CONFIG.roundsPerSort
if (state.tickCount === 2 || state.tickCount % 10 === 1) { // ticks thereafter
if (state.tickCount === 2 || state.tickCount % CONFIG.roundsPerSort === 1) {
sortAndRebuildWAN(state); sortAndRebuildWAN(state);
} }
@@ -1206,9 +1216,10 @@ export async function tick(state, signal, onOffline) {
// --- Recovery Probe ---------------------------------------------------------- // --- Recovery Probe ----------------------------------------------------------
// When offline, check 4 random WAN hosts every 500ms, giving up the checks // When offline, check CONFIG.recoveryProbeHosts random WAN hosts every
// started 500ms before, so at most 4 are ever waiting. As soon as one // CONFIG.recoveryProbeInterval ms, giving up the checks started one interval
// answers, stop probing and start a new round at once. // before, so at most that many are ever waiting. As soon as one answers,
// stop probing and start a new round at once.
function startRecoveryProbe(state, startRounds) { function startRecoveryProbe(state, startRounds) {
if (state._recoveryProbeId) return; // already running if (state._recoveryProbeId) return; // already running
const candidates = [...state.wan]; const candidates = [...state.wan];
@@ -1216,7 +1227,7 @@ function startRecoveryProbe(state, startRounds) {
const j = Math.floor(Math.random() * (i + 1)); const j = Math.floor(Math.random() * (i + 1));
[candidates[i], candidates[j]] = [candidates[j], candidates[i]]; [candidates[i], candidates[j]] = [candidates[j], candidates[i]];
} }
const canaries = candidates.slice(0, 4); const canaries = candidates.slice(0, CONFIG.recoveryProbeHosts);
log.notice( log.notice(
`Recovery probe started (${canaries.map((h) => h.name).join(", ")})`, `Recovery probe started (${canaries.map((h) => h.name).join(", ")})`,
); );
@@ -1233,7 +1244,7 @@ function startRecoveryProbe(state, startRounds) {
startRounds(); startRounds();
}); });
} }
}, 500); }, CONFIG.recoveryProbeInterval);
} }
function stopRecoveryProbe(state) { function stopRecoveryProbe(state) {
@@ -1497,7 +1508,7 @@ async function init() {
}); });
window.addEventListener("resize", () => handleResize(state)); window.addEventListener("resize", () => handleResize(state));
setTimeout(() => handleResize(state), 100); setTimeout(() => handleResize(state), CONFIG.resizeDelay);
} }
// Bootstrap only when loaded as the page: a real DOM containing the #app // Bootstrap only when loaded as the page: a real DOM containing the #app
+26 -10
View File
@@ -347,19 +347,35 @@ for (const { history, latencies, statistics } of [
}, },
]) { ]) {
test(`a target's min, max, average and median latency over ${history}`, () => { test(`a target's min, max, average and median latency over ${history}`, () => {
const host = hostAfter(latencies); const { min, max, avg, med } = hostAfter(latencies).historyStats();
assert.deepEqual( assert.deepEqual({ min, max, average: avg, median: med }, statistics);
{
min: host.minLatency(),
max: host.maxLatency(),
average: host.averageLatency(),
median: host.medianLatency(),
},
statistics,
);
}); });
} }
// The summary's figures come from each WAN target's last check, by the same
// rules as a target's own: here four answered, one was found unreachable
// and the rest have not been checked yet. The median, 22.5, and the
// average, 21.25, are rounded.
test("the summary's min, max, median and average latency over the WAN targets' last checks", () => {
const state = new AppState([]);
[30, 10, null, 25, 20].forEach((latency, i) =>
state.wan[i].pushSample(
Date.now(),
latency === null
? { latency: null, error: "unreachable" }
: { latency, error: null },
),
);
assert.deepEqual(state.wanStats(), {
reachable: 4,
total: state.wan.length,
min: 10,
max: 30,
med: 23,
avg: 21,
});
});
// An app state in which, of the WAN targets, the first timedOut timed out, // An app state in which, of the WAN targets, the first timedOut timed out,
// the next unreachable were found unreachable, the next answered answered // the next unreachable were found unreachable, the next answered answered
// after latency ms, and the rest have not been checked yet. // after latency ms, and the rest have not been checked yet.