4 Commits
Author SHA1 Message Date
clawbot 186f932eb8 Footer no longer says IPv4 only (closes #111)
check / check (push) Successful in 2m48s
Each check is a fetch: the browser picks IPv4 or IPv6 for each WAN
host, and the local targets are IPv4 addresses, so "IPv4 only" was
wrong for the page as a whole. The footer's first line drops it and
its separator; the rest of the footer is unchanged. TODO.md moves this
to Completed Steps and the layout decision from Future Steps into Next
Step.

Model: opus-5-5
2026-10-04 07:38:06 +02:00
clawbot 64e142c17f README, TODO and the viewport README say what the tree does (closes #24)
check / check (push) Successful in 2m57s
README.md: Getting Started leads with make targets; a Backend section
gives netwatch-server's routes and how the image builds and runs it;
the checks are GET requests; the WAN host list, health states, summary
figures, sorting and missing features match src/main.js; the TODO
section points to TODO.md, which holds the one to-do list.

backend/README.md: its TODO section points to TODO.md too, whose
Future Steps take its three open items.

TODO.md: Workflow branches from next and opens the PR against next;
Status, Next Step and Future Steps describe the open work, linked to
its issue where one exists.

test/viewport/README.md: the unit tests run on Node's test runner,
not vitest.

Model: opus-5-5
2026-10-04 07:03:05 +02:00
clawbot 161f955ae2 Latency statistics written once, thresholds read from CONFIG (closes #102)
check / check (push) Successful in 2m50s
A target's min, max, median and average latency now come from one list
of its answers, through latencyStats(), which the summary's figures use
too, so the median is written once. latencyHex() and latencyClass() read
one table of color limits in CONFIG. The health thresholds, the debug
log's length, the gateway check's timeout, the recovery probe's number
of hosts and interval, how often the rows are sorted and the delay
before the first sparkline resize are CONFIG entries, read where the
numbers were. HostState's minLatency(), maxLatency(), averageLatency()
and medianLatency() are gone; the statistics test reads historyStats().
Nothing the page does or shows changes. A unit test now covers the
summary's figures.

Model: opus-5-5
2026-10-04 06:51:24 +02:00
clawbot d40e67d4ab Rate limit password attempts on /metrics (closes #104)
check / check (push) Successful in 3m32s
Each client address may make 60 requests to /metrics a minute,
through the same httprate middleware and TRUSTED_PROXIES
resolution the report route uses, with an allowance of its own.
The limit runs before the basic auth, so past it the answer is 429
and the password is not checked. backend/README.md says so; a test
uses up one client's allowance on wrong passwords, gets 429 with
the right one, and checks that another client behind the same
nginx still gets in.

Model: opus-5-5
2026-10-04 06:19:06 +02:00
8 changed files with 230 additions and 114 deletions
+2 -1
View File
@@ -161,7 +161,8 @@ and when it stops. Its routes:
- `GET /.well-known/healthcheck` — answers 200 with `"status":"ok"`, the
server's version and its uptime
- `GET /metrics` — Prometheus metrics behind basic auth, only when
`METRICS_USERNAME` and `METRICS_PASSWORD` are set
`METRICS_USERNAME` and `METRICS_PASSWORD` are set; each client address may
make a limited number of requests to it a minute
In the image, the `builder` stage of `Dockerfile` tests it and builds it with
`backend/script/build`, and `bin/entrypoint.sh` runs it as user `netwatch` on
+36 -13
View File
@@ -16,23 +16,45 @@ Docker image, and the Gitea workflow `.gitea/workflows/check.yml` runs
# Next Step
Write the latency statistics once and move the thresholds written inline in
`src/main.js` into `CONFIG`
([#102](https://git.eeqj.de/sneak/netwatch/issues/102)).
Decide whether the repo moves to the layout `REPO_POLICIES.md` gives, with
`backend/` no longer repeating files from the root
([#30](https://git.eeqj.de/sneak/netwatch/issues/30)).
# Completed Steps
- 2026-10-04: the page's footer no longer says "IPv4 only"
([#111](https://git.eeqj.de/sneak/netwatch/issues/111)): each check is a
`fetch`, the browser picks IPv4 or IPv6 for each WAN host, and the local
targets are IPv4 addresses. The rest of the footer is unchanged
- 2026-10-04: `README.md`, `TODO.md` and `test/viewport/README.md` say what the
tree does (issue #24). The README's Getting Started leads with `make` targets;
a new Backend section says what `netwatch-server` stores, its routes and how
the image builds and runs it, and points to `backend/README.md` for its
settings; the checks are GET requests; the 26 WAN hosts, the four health
states, the summary's figures and the features the list lacked are described
as the page has them; and its TODO section points here. This file's Workflow
branches from `next` and opens the PR against `next`, Status says where the
repo stands, and Next Step and Future Steps hold only open work, linked to its
issue where one exists. The viewport harness README names Node's test runner,
not `vitest`
as the page has them; and its TODO section points here, as does the one in
`backend/README.md`, whose open items moved to Future Steps. This file's
Workflow branches from `next` and opens the PR against `next`, Status says
where the repo stands, and Next Step and Future Steps hold only open work,
linked to its issue where one exists. The viewport harness README names Node's
test runner, not `vitest`
- 2026-10-04: in `src/main.js` (issue #102), a target's min, max, median and
average latency come from one list of its answers, through the same function
the summary's figures use, so the median is written once. The latency color
limits are one table in `CONFIG`, read by both the figure's and the
sparkline's color. The health thresholds, the debug log's length, the gateway
check's timeout, the recovery probe's number of hosts and interval, how often
the rows are sorted and the delay before the first sparkline resize are
`CONFIG` entries too. A unit test checks the summary's figures. Nothing the
page does or shows changed; the footer's color legend still writes the limits
out as text
- 2026-10-04: password guesses at `/metrics` are rate limited (issue #104): each
client address, resolved through `TRUSTED_PROXIES` as for reports, may make 60
requests to `/metrics` a minute, counted by `go-chi/httprate` apart from its
reports; past that it gets 429 and its basic auth credentials are not checked.
The limit is a constant in `backend/internal/server/routes.go`. A test uses up
one client's allowance on wrong passwords, gets 429 with the right one, and
checks that another client behind the same nginx still gets in
- 2026-10-04: the backend reports errors to Sentry (issue #95). With
`SENTRY_DSN` set, it sets up `sentry-go` with the release `netwatch-server-`
and its version, reports each panic in a handler through `sentryhttp`, the
@@ -367,13 +389,14 @@ Write the latency statistics once and move the thresholds written inline in
# Future Steps
- Rate limit password attempts on `/metrics`
([#104](https://git.eeqj.de/sneak/netwatch/issues/104))
- Decide whether the repo moves to the layout `REPO_POLICIES.md` gives, with
`backend/` no longer repeating files from the root
([#30](https://git.eeqj.de/sneak/netwatch/issues/30))
- Run `make frontend-viewport-test` in CI as its own step; it is not part of
`make check`, as it needs Docker and takes minutes
- A backend test that posts a report to `POST /api/v1/reports` and checks the
compressed file it is written to
- A backend route that decompresses the stored reports and answers queries on
them
- Prometheus metrics for the backend's in-memory buffer: its size, the number of
flushes and the number of reports
- A configurable host list (an environment variable or a config file)
- Export of the latency history (CSV or JSON)
- A notification when the health status changes to DEGRADED
+10 -3
View File
@@ -199,6 +199,14 @@ is recorded and `/metrics` answers 404. One without the other stops the server
from starting, with an error naming both; so does a `METRICS_USERNAME`
containing `:`, which basic auth cannot carry, with an error naming it.
`/metrics` is rate limited, so that its password cannot be guessed quickly: each
client address, resolved through `TRUSTED_PROXIES`, may make 60 requests to it a
minute, whatever their credentials. Past that it gets 429 with
`Retry-After: 60`, and its credentials are not checked. The minute slides as it
does for reports (see [Report limits](#report-limits)), so a scraper polling
every 2 seconds or less often is never refused. This allowance is apart from the
one for reports.
### Sentry
With `SENTRY_DSN` set, the server sends its errors to that Sentry project: each
@@ -211,9 +219,8 @@ sent to it.
## TODO
- Add integration test that POSTs a report and verifies the compressed output
- Add report decompression/query endpoint
- Add metrics (Prometheus) for buffer size, flush count, report count
The to-do list, this backend's open work included, is [TODO.md](../TODO.md) at
the repo root.
## License
+4
View File
@@ -12,6 +12,10 @@ func (s *Server) Router() *chi.Mux {
// external tests.
const MaxRequestBodyBytes = maxRequestBodyBytes
// MetricsRequestsPerMinute exposes the /metrics rate limit to the
// external tests.
const MetricsRequestsPerMinute = metricsRequestsPerMinute
// ListenAddr exposes the address the server listens on to the
// external tests.
func (s *Server) ListenAddr() string {
+14 -4
View File
@@ -18,6 +18,12 @@ const (
// can mount s.mw.MaxBodyBytes with a smaller value to lower
// its bound, but cannot raise it: this cap runs first.
maxRequestBodyBytes int64 = 1 << 20 // 1 MiB
// metricsRequestsPerMinute is how many requests to /metrics each
// client address may make a minute, whatever their credentials. A
// scraper polling every 2 seconds sends half of it, which httprate
// never refuses.
metricsRequestsPerMinute = 60
)
// SetupRoutes configures the chi router with middleware and
@@ -66,10 +72,14 @@ func (s *Server) SetupRoutes() {
Post("/api/v1/reports", s.h.HandleReport())
})
// The rate limit comes before the basic auth, so a client past it
// gets 429 and its password is not checked.
if s.params.Config.MetricsUsername != "" {
s.router.With(s.mw.MetricsAuth()).
Get("/metrics", promhttp.HandlerFor(
registry, promhttp.HandlerOpts{},
).ServeHTTP)
s.router.With(
s.mw.RateLimit(metricsRequestsPerMinute),
s.mw.MetricsAuth(),
).Get("/metrics", promhttp.HandlerFor(
registry, promhttp.HandlerOpts{},
).ServeHTTP)
}
}
+44
View File
@@ -188,6 +188,50 @@ func TestMetricsBehindBasicAuth(t *testing.T) {
}
}
// TestMetricsAreRateLimited: a client that has used up its /metrics
// allowance on wrong passwords gets 429 even with the right one, which
// is then not checked, while another client behind the same nginx
// still gets in.
func TestMetricsAreRateLimited(t *testing.T) {
t.Setenv("METRICS_USERNAME", "prometheus")
t.Setenv("METRICS_PASSWORD", "right")
// As in the container: nginx connects from loopback and names the
// client in X-Forwarded-For.
t.Setenv("TRUSTED_PROXIES", "127.0.0.1/32")
srv := newServer(t)
srv.SetupRoutes()
get := func(client, password string) int {
rec := httptest.NewRecorder()
req := httptest.NewRequestWithContext(t.Context(),
http.MethodGet, "/metrics", http.NoBody)
req.RemoteAddr = "127.0.0.1:40000"
req.Header.Set("X-Forwarded-For", client)
req.SetBasicAuth("prometheus", password)
srv.ServeHTTP(rec, req)
return rec.Code
}
for i := range server.MetricsRequestsPerMinute {
if code := get("203.0.113.7", "wrong"); code != http.StatusUnauthorized {
t.Fatalf("guess %d: status = %d, want %d",
i+1, code, http.StatusUnauthorized)
}
}
if code := get("203.0.113.7", "right"); code != http.StatusTooManyRequests {
t.Fatalf("right password past the limit: status = %d, want %d",
code, http.StatusTooManyRequests)
}
if code := get("203.0.113.8", "right"); code != http.StatusOK {
t.Fatalf("another client: status = %d, want %d",
code, http.StatusOK)
}
}
// TestMetricsInTwoServers: two servers in one process can both have
// metrics on.
func TestMetricsInTwoServers(t *testing.T) {
+94 -83
View File
@@ -30,6 +30,39 @@ export const CONFIG = {
return [0, 1, 2, 3, 4, 5].map((i) => Math.round((d * i) / 5));
},
canvasHeight: 96,
// A latency figure and its sparkline take the color of the first entry
// whose limit, in ms, the latency is below.
latencyColors: [
{ below: 50, hex: "#22c55e", className: "text-green-500" },
{ below: 100, hex: "#84cc16", className: "text-lime-500" },
{ below: 200, hex: "#eab308", className: "text-yellow-500" },
{ below: 500, hex: "#f97316", className: "text-orange-500" },
{ below: Infinity, hex: "#ef4444", className: "text-red-500" },
],
// The health is offline when more than offlineTimeouts WAN hosts timed
// out or were unreachable and at most offlineReachable answered;
// otherwise degraded when more than degradedTimeouts timed out or were
// unreachable; otherwise slow when more than slowHosts answered after
// more than slowLatency ms.
offlineTimeouts: 10,
offlineReachable: 4,
degradedTimeouts: 4,
slowHosts: 3,
slowLatency: 1000,
// The debug log keeps its last maxLogEntries lines.
maxLogEntries: 1000,
// A gateway candidate that has not answered after gatewayTimeout ms is
// passed over.
gatewayTimeout: 1500,
// When no WAN host answers, the recovery probe checks recoveryProbeHosts
// random ones every recoveryProbeInterval ms.
recoveryProbeHosts: 4,
recoveryProbeInterval: 500,
// The rows are sorted after the first round that is not discarded, then
// every roundsPerSort rounds.
roundsPerSort: 10,
// The sparklines are sized and drawn again resizeDelay ms after start.
resizeDelay: 100,
};
// WAN endpoints to monitor. These are used for the aggregate health/stats
@@ -114,7 +147,8 @@ const debugLog = [];
const log = (() => {
function append(level, message) {
debugLog.push({ timestamp: new Date(), level, message });
if (debugLog.length > 1000) debugLog.splice(0, debugLog.length - 1000);
if (debugLog.length > CONFIG.maxLogEntries)
debugLog.splice(0, debugLog.length - CONFIG.maxLogEntries);
const panel = document.getElementById("debug-panel");
if (panel && !panel.classList.contains("hidden")) renderDebugLog();
}
@@ -174,7 +208,10 @@ async function detectGateway() {
const result = await Promise.any(
GATEWAY_CANDIDATES.map(async (url) => {
const controller = new AbortController();
const timeoutId = setTimeout(() => controller.abort(), 1500);
const timeoutId = setTimeout(
() => controller.abort(),
CONFIG.gatewayTimeout,
);
try {
await fetch(url, {
method: "GET",
@@ -199,6 +236,27 @@ async function detectGateway() {
// --- App State ---------------------------------------------------------------
// The min, max, median and average of latencies, a list of numbers, or all
// null when it is empty. The median of an even count is the mean of the
// middle two; it and the average are rounded.
function latencyStats(latencies) {
if (latencies.length === 0)
return { min: null, max: null, med: null, avg: null };
const sorted = [...latencies].sort((a, b) => a - b);
const mid = Math.floor(sorted.length / 2);
return {
min: sorted[0],
max: sorted[sorted.length - 1],
med:
sorted.length % 2
? sorted[mid]
: Math.round((sorted[mid - 1] + sorted[mid]) / 2),
avg: Math.round(
latencies.reduce((a, b) => a + b, 0) / latencies.length,
),
};
}
export class HostState {
constructor(host, pinned = false) {
this.name = host.name;
@@ -230,38 +288,16 @@ export class HostState {
this._trim();
}
averageLatency() {
const valid = this.history.filter((p) => p.latency !== null);
if (valid.length === 0) return null;
return Math.round(
valid.reduce((s, p) => s + p.latency, 0) / valid.length,
// The min, max, median and average latency of the checks in the history
// that got an answer.
historyStats() {
return latencyStats(
this.history
.filter((p) => p.latency !== null)
.map((p) => p.latency),
);
}
minLatency() {
const valid = this.history.filter((p) => p.latency !== null);
if (valid.length === 0) return null;
return Math.min(...valid.map((p) => p.latency));
}
maxLatency() {
const valid = this.history.filter((p) => p.latency !== null);
if (valid.length === 0) return null;
return Math.max(...valid.map((p) => p.latency));
}
medianLatency() {
const sorted = this.history
.filter((p) => p.latency !== null)
.map((p) => p.latency)
.sort((a, b) => a - b);
if (sorted.length === 0) return null;
const mid = Math.floor(sorted.length / 2);
return sorted.length % 2
? sorted[mid]
: Math.round((sorted[mid - 1] + sorted[mid]) / 2);
}
_trim() {
while (this.history.length > CONFIG.maxHistoryPoints)
this.history.shift();
@@ -288,33 +324,13 @@ export class AppState {
/** WAN-only stats from latest sample (excludes local) */
wanStats() {
const reachable = this.wan.filter((h) => h.lastLatency !== null);
const latencies = reachable.map((h) => h.lastLatency);
const total = this.wan.length;
if (latencies.length === 0)
return {
reachable: 0,
total,
min: null,
max: null,
med: null,
avg: null,
};
const sorted = [...latencies].sort((a, b) => a - b);
const mid = Math.floor(sorted.length / 2);
const med =
sorted.length % 2
? sorted[mid]
: Math.round((sorted[mid - 1] + sorted[mid]) / 2);
const latencies = this.wan
.filter((h) => h.lastLatency !== null)
.map((h) => h.lastLatency);
return {
reachable: latencies.length,
total,
min: Math.min(...latencies),
max: Math.max(...latencies),
med,
avg: Math.round(
latencies.reduce((a, b) => a + b, 0) / latencies.length,
),
total: this.wan.length,
...latencyStats(latencies),
};
}
@@ -340,12 +356,16 @@ export class AppState {
const timeouts = this.wan.filter(
(h) => h.status === "error" || h.status === "offline",
).length;
if (timeouts > 10 && reachable <= 4) return "offline";
if (timeouts > 4) return "degraded";
if (
timeouts > CONFIG.offlineTimeouts &&
reachable <= CONFIG.offlineReachable
)
return "offline";
if (timeouts > CONFIG.degradedTimeouts) return "degraded";
const slow = this.wan.filter(
(h) => h.lastLatency !== null && h.lastLatency > 1000,
(h) => h.lastLatency !== null && h.lastLatency > CONFIG.slowLatency,
).length;
if (slow > 3) return "slow";
if (slow > CONFIG.slowHosts) return "slow";
return "healthy";
}
@@ -557,21 +577,13 @@ export async function measureLatency(url, signal) {
export function latencyHex(latency) {
if (latency === null) return "#6b7280";
if (latency < 50) return "#22c55e";
if (latency < 100) return "#84cc16";
if (latency < 200) return "#eab308";
if (latency < 500) return "#f97316";
return "#ef4444";
return CONFIG.latencyColors.find((c) => latency < c.below).hex;
}
export function latencyClass(latency, status) {
if (status === "offline" || status === "error" || latency === null)
return "text-gray-500";
if (latency < 50) return "text-green-500";
if (latency < 100) return "text-lime-500";
if (latency < 200) return "text-yellow-500";
if (latency < 500) return "text-orange-500";
return "text-red-500";
return CONFIG.latencyColors.find((c) => latency < c.below).className;
}
// --- Sparkline Renderer ------------------------------------------------------
@@ -845,7 +857,7 @@ function buildUI(state) {
</div>
<footer class="mt-8 text-center text-gray-600 text-xs">
<p>Latency measured via GET requests | IPv4 only | CORS restrictions may affect some measurements</p>
<p>Latency measured via GET requests | CORS restrictions may affect some measurements</p>
<p class="mt-2">
<span class="inline-block w-3 h-3 rounded-full bg-green-500 mr-1 align-middle"></span>&lt;50ms
<span class="inline-block w-3 h-3 rounded-full bg-lime-500 mr-1 ml-3 align-middle"></span>&lt;100ms
@@ -912,10 +924,7 @@ function updateHostRow(host, index) {
latencyEl.innerHTML = `<span class="text-gray-500">---</span>`;
}
const avg = host.averageLatency();
const med = host.medianLatency();
const min = host.minLatency();
const max = host.maxLatency();
const { min, med, avg, max } = host.historyStats();
if (host.status === "online" && avg !== null) {
statusEl.innerHTML = statusStatsHTML([
["min", min],
@@ -1182,8 +1191,9 @@ export async function tick(state, signal, onOffline) {
// rows whose check ended before the resume still read "paused"
state.allHosts.forEach((host, i) => updateHostRow(host, i));
// Sort after the first real check, then every 10 ticks thereafter
if (state.tickCount === 2 || state.tickCount % 10 === 1) {
// Sort after the first real check, then every CONFIG.roundsPerSort
// ticks thereafter
if (state.tickCount === 2 || state.tickCount % CONFIG.roundsPerSort === 1) {
sortAndRebuildWAN(state);
}
@@ -1206,9 +1216,10 @@ export async function tick(state, signal, onOffline) {
// --- Recovery Probe ----------------------------------------------------------
// When offline, check 4 random WAN hosts every 500ms, giving up the checks
// started 500ms before, so at most 4 are ever waiting. As soon as one
// answers, stop probing and start a new round at once.
// When offline, check CONFIG.recoveryProbeHosts random WAN hosts every
// CONFIG.recoveryProbeInterval ms, giving up the checks started one interval
// before, so at most that many are ever waiting. As soon as one answers,
// stop probing and start a new round at once.
function startRecoveryProbe(state, startRounds) {
if (state._recoveryProbeId) return; // already running
const candidates = [...state.wan];
@@ -1216,7 +1227,7 @@ function startRecoveryProbe(state, startRounds) {
const j = Math.floor(Math.random() * (i + 1));
[candidates[i], candidates[j]] = [candidates[j], candidates[i]];
}
const canaries = candidates.slice(0, 4);
const canaries = candidates.slice(0, CONFIG.recoveryProbeHosts);
log.notice(
`Recovery probe started (${canaries.map((h) => h.name).join(", ")})`,
);
@@ -1233,7 +1244,7 @@ function startRecoveryProbe(state, startRounds) {
startRounds();
});
}
}, 500);
}, CONFIG.recoveryProbeInterval);
}
function stopRecoveryProbe(state) {
@@ -1497,7 +1508,7 @@ async function init() {
});
window.addEventListener("resize", () => handleResize(state));
setTimeout(() => handleResize(state), 100);
setTimeout(() => handleResize(state), CONFIG.resizeDelay);
}
// Bootstrap only when loaded as the page: a real DOM containing the #app
+26 -10
View File
@@ -347,19 +347,35 @@ for (const { history, latencies, statistics } of [
},
]) {
test(`a target's min, max, average and median latency over ${history}`, () => {
const host = hostAfter(latencies);
assert.deepEqual(
{
min: host.minLatency(),
max: host.maxLatency(),
average: host.averageLatency(),
median: host.medianLatency(),
},
statistics,
);
const { min, max, avg, med } = hostAfter(latencies).historyStats();
assert.deepEqual({ min, max, average: avg, median: med }, statistics);
});
}
// The summary's figures come from each WAN target's last check, by the same
// rules as a target's own: here four answered, one was found unreachable
// and the rest have not been checked yet. The median, 22.5, and the
// average, 21.25, are rounded.
test("the summary's min, max, median and average latency over the WAN targets' last checks", () => {
const state = new AppState([]);
[30, 10, null, 25, 20].forEach((latency, i) =>
state.wan[i].pushSample(
Date.now(),
latency === null
? { latency: null, error: "unreachable" }
: { latency, error: null },
),
);
assert.deepEqual(state.wanStats(), {
reachable: 4,
total: state.wan.length,
min: 10,
max: 30,
med: 23,
avg: 21,
});
});
// An app state in which, of the WAN targets, the first timedOut timed out,
// the next unreachable were found unreachable, the next answered answered
// after latency ms, and the rest have not been checked yet.