1 Commits
Author SHA1 Message Date
sneak 0b1d766765 README, TODO and the viewport README say what the tree does (closes #24)
check / check (push) Successful in 3m33s
README.md: Getting Started leads with make targets; a Backend section
gives netwatch-server's routes and how the image builds and runs it,
pointing to backend/README.md for its settings; the checks are GET
requests; the WAN host list, health states, summary figures, sorting
and missing features match src/main.js; the TODO section points to
TODO.md, which holds the one to-do list.

TODO.md: Workflow branches from next and opens the PR against next;
Status, Next Step and Future Steps describe the open work, linked to
its issue where one exists.

test/viewport/README.md: the unit tests run on Node's test runner,
not vitest.

Model: opus-5-5
2026-10-04 04:06:50 +00:00
8 changed files with 148 additions and 472 deletions
+12 -24
View File
@@ -161,8 +161,7 @@ and when it stops. Its routes:
- `GET /.well-known/healthcheck` — answers 200 with `"status":"ok"`, the - `GET /.well-known/healthcheck` — answers 200 with `"status":"ok"`, the
server's version and its uptime server's version and its uptime
- `GET /metrics` — Prometheus metrics behind basic auth, only when - `GET /metrics` — Prometheus metrics behind basic auth, only when
`METRICS_USERNAME` and `METRICS_PASSWORD` are set; each client address may `METRICS_USERNAME` and `METRICS_PASSWORD` are set
make a limited number of requests to it a minute
In the image, the `builder` stage of `Dockerfile` tests it and builds it with In the image, the `builder` stage of `Dockerfile` tests it and builds it with
`backend/script/build`, and `bin/entrypoint.sh` runs it as user `netwatch` on `backend/script/build`, and `bin/entrypoint.sh` runs it as user `netwatch` on
@@ -188,28 +187,17 @@ Local hosts are tracked separately from WAN stats.
### Latency measurement ### Latency measurement
GET requests to each target's URL as written, with `mode: 'no-cors'` and GET requests with `mode: 'no-cors'`, `cache: 'no-store'` and a cache-busting
`cache: 'no-store'`, timed with `performance.now()`. No query string is added: query parameter, timed with `performance.now()`. Each check times out after 80%
the Hetzner speed-test servers close the connection without an answer when the of the refresh interval (24 seconds at 30 seconds) and is then recorded as a
URL has one, and `no-store` keeps the browser's cache out of the measurement. timeout, so a round's checks have all finished before the next round is due.
Each check times out after 80% of the refresh interval (24 seconds at 30 When no WAN host answers, a recovery probe checks 4 WAN hosts, picked at random
seconds) and is then recorded as a timeout, so a round's checks have all when it starts, every half second, giving up the checks it started half a second
finished before the next round is due. When no WAN host answers, a recovery before. As soon as one answers, a new round starts at once, as it does after an
probe checks 4 WAN hosts, picked at random when it starts, every half second, interval change. A round started early gives up the last round's checks if they
giving up the checks it started half a second before. As soon as one answers, a are still waiting, and that round records nothing more, so rounds never overlap.
new round starts at once, as it does after an interval change. A round started The browser chooses between IPv4 and IPv6 for each target, as for any request;
early gives up the last round's checks if they are still waiting, and that round the local targets are IPv4 addresses.
records nothing more, so rounds never overlap. The browser chooses between IPv4
and IPv6 for each target, as for any request; the local targets are IPv4
addresses.
Each recorded check that fails writes one line to the browser console with
`console.error`, and the same line to the debug log: the target's name and URL,
the time, whether it timed out, answered over the time limit or hit a network
error (with the error the browser gives the page, such as
`TypeError: Failed to fetch`, which does not say why), and how long the request
took. A target that answers after failed checks writes one `console.info` line
with its latency and how many checks in a row had failed.
### Color coding ### Color coding
+13 -47
View File
@@ -16,56 +16,23 @@ Docker image, and the Gitea workflow `.gitea/workflows/check.yml` runs
# Next Step # Next Step
Decide whether the repo moves to the layout `REPO_POLICIES.md` gives, with Write the latency statistics once and move the thresholds written inline in
`backend/` no longer repeating files from the root `src/main.js` into `CONFIG`
([#30](https://git.eeqj.de/sneak/netwatch/issues/30)). ([#102](https://git.eeqj.de/sneak/netwatch/issues/102)).
# Completed Steps # Completed Steps
- 2026-10-07: the six Hetzner targets answer again, and failed checks are
written to the browser console
([#114](https://git.eeqj.de/sneak/netwatch/issues/114)). The Hetzner
speed-test servers close the connection without an answer when the URL has a
query string, and every check added `?_cb=` and the time; checks now fetch
each target's URL as written, with `cache: 'no-store'` as before. Each
recorded check that fails writes one `console.error` line, and the same line
to the debug log, with the target's name and URL, the time, what failed and
how long the request took; a target that answers after failed checks writes
one `console.info` line. Checks the page does not record write none, so the
debug log no longer lists failures in the first round or in the recovery probe
- 2026-10-04: the page's footer no longer says "IPv4 only"
([#111](https://git.eeqj.de/sneak/netwatch/issues/111)): each check is a
`fetch`, the browser picks IPv4 or IPv6 for each WAN host, and the local
targets are IPv4 addresses. The rest of the footer is unchanged
- 2026-10-04: `README.md`, `TODO.md` and `test/viewport/README.md` say what the - 2026-10-04: `README.md`, `TODO.md` and `test/viewport/README.md` say what the
tree does (issue #24). The README's Getting Started leads with `make` targets; tree does (issue #24). The README's Getting Started leads with `make` targets;
a new Backend section says what `netwatch-server` stores, its routes and how a new Backend section says what `netwatch-server` stores, its routes and how
the image builds and runs it, and points to `backend/README.md` for its the image builds and runs it, and points to `backend/README.md` for its
settings; the checks are GET requests; the 26 WAN hosts, the four health settings; the checks are GET requests; the 26 WAN hosts, the four health
states, the summary's figures and the features the list lacked are described states, the summary's figures and the features the list lacked are described
as the page has them; and its TODO section points here, as does the one in as the page has them; and its TODO section points here. This file's Workflow
`backend/README.md`, whose open items moved to Future Steps. This file's branches from `next` and opens the PR against `next`, Status says where the
Workflow branches from `next` and opens the PR against `next`, Status says repo stands, and Next Step and Future Steps hold only open work, linked to its
where the repo stands, and Next Step and Future Steps hold only open work, issue where one exists. The viewport harness README names Node's test runner,
linked to its issue where one exists. The viewport harness README names Node's not `vitest`
test runner, not `vitest`
- 2026-10-04: in `src/main.js` (issue #102), a target's min, max, median and
average latency come from one list of its answers, through the same function
the summary's figures use, so the median is written once. The latency color
limits are one table in `CONFIG`, read by both the figure's and the
sparkline's color. The health thresholds, the debug log's length, the gateway
check's timeout, the recovery probe's number of hosts and interval, how often
the rows are sorted and the delay before the first sparkline resize are
`CONFIG` entries too. A unit test checks the summary's figures. Nothing the
page does or shows changed; the footer's color legend still writes the limits
out as text
- 2026-10-04: password guesses at `/metrics` are rate limited (issue #104): each
client address, resolved through `TRUSTED_PROXIES` as for reports, may make 60
requests to `/metrics` a minute, counted by `go-chi/httprate` apart from its
reports; past that it gets 429 and its basic auth credentials are not checked.
The limit is a constant in `backend/internal/server/routes.go`. A test uses up
one client's allowance on wrong passwords, gets 429 with the right one, and
checks that another client behind the same nginx still gets in
- 2026-10-04: the backend reports errors to Sentry (issue #95). With - 2026-10-04: the backend reports errors to Sentry (issue #95). With
`SENTRY_DSN` set, it sets up `sentry-go` with the release `netwatch-server-` `SENTRY_DSN` set, it sets up `sentry-go` with the release `netwatch-server-`
and its version, reports each panic in a handler through `sentryhttp`, the and its version, reports each panic in a handler through `sentryhttp`, the
@@ -400,14 +367,13 @@ Decide whether the repo moves to the layout `REPO_POLICIES.md` gives, with
# Future Steps # Future Steps
- Rate limit password attempts on `/metrics`
([#104](https://git.eeqj.de/sneak/netwatch/issues/104))
- Decide whether the repo moves to the layout `REPO_POLICIES.md` gives, with
`backend/` no longer repeating files from the root
([#30](https://git.eeqj.de/sneak/netwatch/issues/30))
- Run `make frontend-viewport-test` in CI as its own step; it is not part of - Run `make frontend-viewport-test` in CI as its own step; it is not part of
`make check`, as it needs Docker and takes minutes `make check`, as it needs Docker and takes minutes
- A backend test that posts a report to `POST /api/v1/reports` and checks the
compressed file it is written to
- A backend route that decompresses the stored reports and answers queries on
them
- Prometheus metrics for the backend's in-memory buffer: its size, the number of
flushes and the number of reports
- A configurable host list (an environment variable or a config file) - A configurable host list (an environment variable or a config file)
- Export of the latency history (CSV or JSON) - Export of the latency history (CSV or JSON)
- A notification when the health status changes to DEGRADED - A notification when the health status changes to DEGRADED
+3 -10
View File
@@ -199,14 +199,6 @@ is recorded and `/metrics` answers 404. One without the other stops the server
from starting, with an error naming both; so does a `METRICS_USERNAME` from starting, with an error naming both; so does a `METRICS_USERNAME`
containing `:`, which basic auth cannot carry, with an error naming it. containing `:`, which basic auth cannot carry, with an error naming it.
`/metrics` is rate limited, so that its password cannot be guessed quickly: each
client address, resolved through `TRUSTED_PROXIES`, may make 60 requests to it a
minute, whatever their credentials. Past that it gets 429 with
`Retry-After: 60`, and its credentials are not checked. The minute slides as it
does for reports (see [Report limits](#report-limits)), so a scraper polling
every 2 seconds or less often is never refused. This allowance is apart from the
one for reports.
### Sentry ### Sentry
With `SENTRY_DSN` set, the server sends its errors to that Sentry project: each With `SENTRY_DSN` set, the server sends its errors to that Sentry project: each
@@ -219,8 +211,9 @@ sent to it.
## TODO ## TODO
The to-do list, this backend's open work included, is [TODO.md](../TODO.md) at - Add integration test that POSTs a report and verifies the compressed output
the repo root. - Add report decompression/query endpoint
- Add metrics (Prometheus) for buffer size, flush count, report count
## License ## License
-4
View File
@@ -12,10 +12,6 @@ func (s *Server) Router() *chi.Mux {
// external tests. // external tests.
const MaxRequestBodyBytes = maxRequestBodyBytes const MaxRequestBodyBytes = maxRequestBodyBytes
// MetricsRequestsPerMinute exposes the /metrics rate limit to the
// external tests.
const MetricsRequestsPerMinute = metricsRequestsPerMinute
// ListenAddr exposes the address the server listens on to the // ListenAddr exposes the address the server listens on to the
// external tests. // external tests.
func (s *Server) ListenAddr() string { func (s *Server) ListenAddr() string {
+4 -14
View File
@@ -18,12 +18,6 @@ const (
// can mount s.mw.MaxBodyBytes with a smaller value to lower // can mount s.mw.MaxBodyBytes with a smaller value to lower
// its bound, but cannot raise it: this cap runs first. // its bound, but cannot raise it: this cap runs first.
maxRequestBodyBytes int64 = 1 << 20 // 1 MiB maxRequestBodyBytes int64 = 1 << 20 // 1 MiB
// metricsRequestsPerMinute is how many requests to /metrics each
// client address may make a minute, whatever their credentials. A
// scraper polling every 2 seconds sends half of it, which httprate
// never refuses.
metricsRequestsPerMinute = 60
) )
// SetupRoutes configures the chi router with middleware and // SetupRoutes configures the chi router with middleware and
@@ -72,14 +66,10 @@ func (s *Server) SetupRoutes() {
Post("/api/v1/reports", s.h.HandleReport()) Post("/api/v1/reports", s.h.HandleReport())
}) })
// The rate limit comes before the basic auth, so a client past it
// gets 429 and its password is not checked.
if s.params.Config.MetricsUsername != "" { if s.params.Config.MetricsUsername != "" {
s.router.With( s.router.With(s.mw.MetricsAuth()).
s.mw.RateLimit(metricsRequestsPerMinute), Get("/metrics", promhttp.HandlerFor(
s.mw.MetricsAuth(), registry, promhttp.HandlerOpts{},
).Get("/metrics", promhttp.HandlerFor( ).ServeHTTP)
registry, promhttp.HandlerOpts{},
).ServeHTTP)
} }
} }
-44
View File
@@ -188,50 +188,6 @@ func TestMetricsBehindBasicAuth(t *testing.T) {
} }
} }
// TestMetricsAreRateLimited: a client that has used up its /metrics
// allowance on wrong passwords gets 429 even with the right one, which
// is then not checked, while another client behind the same nginx
// still gets in.
func TestMetricsAreRateLimited(t *testing.T) {
t.Setenv("METRICS_USERNAME", "prometheus")
t.Setenv("METRICS_PASSWORD", "right")
// As in the container: nginx connects from loopback and names the
// client in X-Forwarded-For.
t.Setenv("TRUSTED_PROXIES", "127.0.0.1/32")
srv := newServer(t)
srv.SetupRoutes()
get := func(client, password string) int {
rec := httptest.NewRecorder()
req := httptest.NewRequestWithContext(t.Context(),
http.MethodGet, "/metrics", http.NoBody)
req.RemoteAddr = "127.0.0.1:40000"
req.Header.Set("X-Forwarded-For", client)
req.SetBasicAuth("prometheus", password)
srv.ServeHTTP(rec, req)
return rec.Code
}
for i := range server.MetricsRequestsPerMinute {
if code := get("203.0.113.7", "wrong"); code != http.StatusUnauthorized {
t.Fatalf("guess %d: status = %d, want %d",
i+1, code, http.StatusUnauthorized)
}
}
if code := get("203.0.113.7", "right"); code != http.StatusTooManyRequests {
t.Fatalf("right password past the limit: status = %d, want %d",
code, http.StatusTooManyRequests)
}
if code := get("203.0.113.8", "right"); code != http.StatusOK {
t.Fatalf("another client: status = %d, want %d",
code, http.StatusOK)
}
}
// TestMetricsInTwoServers: two servers in one process can both have // TestMetricsInTwoServers: two servers in one process can both have
// metrics on. // metrics on.
func TestMetricsInTwoServers(t *testing.T) { func TestMetricsInTwoServers(t *testing.T) {
+95 -145
View File
@@ -30,39 +30,6 @@ export const CONFIG = {
return [0, 1, 2, 3, 4, 5].map((i) => Math.round((d * i) / 5)); return [0, 1, 2, 3, 4, 5].map((i) => Math.round((d * i) / 5));
}, },
canvasHeight: 96, canvasHeight: 96,
// A latency figure and its sparkline take the color of the first entry
// whose limit, in ms, the latency is below.
latencyColors: [
{ below: 50, hex: "#22c55e", className: "text-green-500" },
{ below: 100, hex: "#84cc16", className: "text-lime-500" },
{ below: 200, hex: "#eab308", className: "text-yellow-500" },
{ below: 500, hex: "#f97316", className: "text-orange-500" },
{ below: Infinity, hex: "#ef4444", className: "text-red-500" },
],
// The health is offline when more than offlineTimeouts WAN hosts timed
// out or were unreachable and at most offlineReachable answered;
// otherwise degraded when more than degradedTimeouts timed out or were
// unreachable; otherwise slow when more than slowHosts answered after
// more than slowLatency ms.
offlineTimeouts: 10,
offlineReachable: 4,
degradedTimeouts: 4,
slowHosts: 3,
slowLatency: 1000,
// The debug log keeps its last maxLogEntries lines.
maxLogEntries: 1000,
// A gateway candidate that has not answered after gatewayTimeout ms is
// passed over.
gatewayTimeout: 1500,
// When no WAN host answers, the recovery probe checks recoveryProbeHosts
// random ones every recoveryProbeInterval ms.
recoveryProbeHosts: 4,
recoveryProbeInterval: 500,
// The rows are sorted after the first round that is not discarded, then
// every roundsPerSort rounds.
roundsPerSort: 10,
// The sparklines are sized and drawn again resizeDelay ms after start.
resizeDelay: 100,
}; };
// WAN endpoints to monitor. These are used for the aggregate health/stats // WAN endpoints to monitor. These are used for the aggregate health/stats
@@ -147,8 +114,7 @@ const debugLog = [];
const log = (() => { const log = (() => {
function append(level, message) { function append(level, message) {
debugLog.push({ timestamp: new Date(), level, message }); debugLog.push({ timestamp: new Date(), level, message });
if (debugLog.length > CONFIG.maxLogEntries) if (debugLog.length > 1000) debugLog.splice(0, debugLog.length - 1000);
debugLog.splice(0, debugLog.length - CONFIG.maxLogEntries);
const panel = document.getElementById("debug-panel"); const panel = document.getElementById("debug-panel");
if (panel && !panel.classList.contains("hidden")) renderDebugLog(); if (panel && !panel.classList.contains("hidden")) renderDebugLog();
} }
@@ -208,10 +174,7 @@ async function detectGateway() {
const result = await Promise.any( const result = await Promise.any(
GATEWAY_CANDIDATES.map(async (url) => { GATEWAY_CANDIDATES.map(async (url) => {
const controller = new AbortController(); const controller = new AbortController();
const timeoutId = setTimeout( const timeoutId = setTimeout(() => controller.abort(), 1500);
() => controller.abort(),
CONFIG.gatewayTimeout,
);
try { try {
await fetch(url, { await fetch(url, {
method: "GET", method: "GET",
@@ -236,27 +199,6 @@ async function detectGateway() {
// --- App State --------------------------------------------------------------- // --- App State ---------------------------------------------------------------
// The min, max, median and average of latencies, a list of numbers, or all
// null when it is empty. The median of an even count is the mean of the
// middle two; it and the average are rounded.
function latencyStats(latencies) {
if (latencies.length === 0)
return { min: null, max: null, med: null, avg: null };
const sorted = [...latencies].sort((a, b) => a - b);
const mid = Math.floor(sorted.length / 2);
return {
min: sorted[0],
max: sorted[sorted.length - 1],
med:
sorted.length % 2
? sorted[mid]
: Math.round((sorted[mid - 1] + sorted[mid]) / 2),
avg: Math.round(
latencies.reduce((a, b) => a + b, 0) / latencies.length,
),
};
}
export class HostState { export class HostState {
constructor(host, pinned = false) { constructor(host, pinned = false) {
this.name = host.name; this.name = host.name;
@@ -268,8 +210,6 @@ export class HostState {
this.lastLatency = null; this.lastLatency = null;
this.status = "pending"; // 'online' | 'offline' | 'error' | 'pending' this.status = "pending"; // 'online' | 'offline' | 'error' | 'pending'
this.pinned = pinned; this.pinned = pinned;
// How many recorded checks in a row have failed, up to the last one.
this.consecutiveFailures = 0;
} }
pushSample(timestamp, result) { pushSample(timestamp, result) {
@@ -283,9 +223,6 @@ export class HostState {
if (result.error === "timeout") this.status = "error"; if (result.error === "timeout") this.status = "error";
else if (result.error) this.status = "offline"; else if (result.error) this.status = "offline";
else this.status = "online"; else this.status = "online";
this.consecutiveFailures = result.error
? this.consecutiveFailures + 1
: 0;
} }
pushPaused(timestamp) { pushPaused(timestamp) {
@@ -293,16 +230,38 @@ export class HostState {
this._trim(); this._trim();
} }
// The min, max, median and average latency of the checks in the history averageLatency() {
// that got an answer. const valid = this.history.filter((p) => p.latency !== null);
historyStats() { if (valid.length === 0) return null;
return latencyStats( return Math.round(
this.history valid.reduce((s, p) => s + p.latency, 0) / valid.length,
.filter((p) => p.latency !== null)
.map((p) => p.latency),
); );
} }
minLatency() {
const valid = this.history.filter((p) => p.latency !== null);
if (valid.length === 0) return null;
return Math.min(...valid.map((p) => p.latency));
}
maxLatency() {
const valid = this.history.filter((p) => p.latency !== null);
if (valid.length === 0) return null;
return Math.max(...valid.map((p) => p.latency));
}
medianLatency() {
const sorted = this.history
.filter((p) => p.latency !== null)
.map((p) => p.latency)
.sort((a, b) => a - b);
if (sorted.length === 0) return null;
const mid = Math.floor(sorted.length / 2);
return sorted.length % 2
? sorted[mid]
: Math.round((sorted[mid - 1] + sorted[mid]) / 2);
}
_trim() { _trim() {
while (this.history.length > CONFIG.maxHistoryPoints) while (this.history.length > CONFIG.maxHistoryPoints)
this.history.shift(); this.history.shift();
@@ -329,13 +288,33 @@ export class AppState {
/** WAN-only stats from latest sample (excludes local) */ /** WAN-only stats from latest sample (excludes local) */
wanStats() { wanStats() {
const latencies = this.wan const reachable = this.wan.filter((h) => h.lastLatency !== null);
.filter((h) => h.lastLatency !== null) const latencies = reachable.map((h) => h.lastLatency);
.map((h) => h.lastLatency); const total = this.wan.length;
if (latencies.length === 0)
return {
reachable: 0,
total,
min: null,
max: null,
med: null,
avg: null,
};
const sorted = [...latencies].sort((a, b) => a - b);
const mid = Math.floor(sorted.length / 2);
const med =
sorted.length % 2
? sorted[mid]
: Math.round((sorted[mid - 1] + sorted[mid]) / 2);
return { return {
reachable: latencies.length, reachable: latencies.length,
total: this.wan.length, total,
...latencyStats(latencies), min: Math.min(...latencies),
max: Math.max(...latencies),
med,
avg: Math.round(
latencies.reduce((a, b) => a + b, 0) / latencies.length,
),
}; };
} }
@@ -361,16 +340,12 @@ export class AppState {
const timeouts = this.wan.filter( const timeouts = this.wan.filter(
(h) => h.status === "error" || h.status === "offline", (h) => h.status === "error" || h.status === "offline",
).length; ).length;
if ( if (timeouts > 10 && reachable <= 4) return "offline";
timeouts > CONFIG.offlineTimeouts && if (timeouts > 4) return "degraded";
reachable <= CONFIG.offlineReachable
)
return "offline";
if (timeouts > CONFIG.degradedTimeouts) return "degraded";
const slow = this.wan.filter( const slow = this.wan.filter(
(h) => h.lastLatency !== null && h.lastLatency > CONFIG.slowLatency, (h) => h.lastLatency !== null && h.lastLatency > 1000,
).length; ).length;
if (slow > CONFIG.slowHosts) return "slow"; if (slow > 3) return "slow";
return "healthy"; return "healthy";
} }
@@ -539,13 +514,7 @@ class Reporter {
// Checks one target. The check times out after CONFIG.requestTimeout; the // Checks one target. The check times out after CONFIG.requestTimeout; the
// caller can give it up sooner through the optional signal, which also ends // caller can give it up sooner through the optional signal, which also ends
// it as a timeout. A failed check's reason says what went wrong and after // it as a timeout.
// how long; a check that answered has none.
//
// The URL is fetched as written, with nothing added to it: the Hetzner
// speed-test servers close the connection without an answer when the URL
// has a query string, and cache: "no-store" keeps the browser's cache out
// of the measurement.
export async function measureLatency(url, signal) { export async function measureLatency(url, signal) {
const controller = new AbortController(); const controller = new AbortController();
const timeoutId = setTimeout( const timeoutId = setTimeout(
@@ -554,10 +523,13 @@ export async function measureLatency(url, signal) {
); );
signal?.addEventListener("abort", () => controller.abort()); signal?.addEventListener("abort", () => controller.abort());
const targetUrl = new URL(url);
targetUrl.searchParams.set("_cb", Date.now().toString());
const start = performance.now(); const start = performance.now();
try { try {
await fetch(url, { await fetch(targetUrl.toString(), {
method: "GET", method: "GET",
mode: "no-cors", mode: "no-cors",
cache: "no-store", cache: "no-store",
@@ -566,28 +538,18 @@ export async function measureLatency(url, signal) {
const latency = Math.round(performance.now() - start); const latency = Math.round(performance.now() - start);
clearTimeout(timeoutId); clearTimeout(timeoutId);
if (latency > CONFIG.maxLatency) { if (latency > CONFIG.maxLatency) {
return { log.error(`${url} timeout (${latency}ms > ${CONFIG.maxLatency}ms)`);
latency: null, return { latency: null, error: "timeout" };
error: "timeout",
reason: `answered after ${latency} ms, over the ${CONFIG.maxLatency} ms limit`,
};
} }
return { latency, error: null, reason: null }; return { latency, error: null };
} catch (err) { } catch (err) {
const took = Math.round(performance.now() - start);
clearTimeout(timeoutId); clearTimeout(timeoutId);
if (err.name === "AbortError") { if (err.name === "AbortError") {
return { log.error(`${url} timeout (aborted)`);
latency: null, return { latency: null, error: "timeout" };
error: "timeout",
reason: `timed out after ${took} ms (limit ${CONFIG.requestTimeout} ms)`,
};
} }
return { log.error(`${url} unreachable`);
latency: null, return { latency: null, error: "unreachable" };
error: "unreachable",
reason: `network error (${err.name}: ${err.message}) after ${took} ms`,
};
} }
} }
@@ -595,13 +557,21 @@ export async function measureLatency(url, signal) {
export function latencyHex(latency) { export function latencyHex(latency) {
if (latency === null) return "#6b7280"; if (latency === null) return "#6b7280";
return CONFIG.latencyColors.find((c) => latency < c.below).hex; if (latency < 50) return "#22c55e";
if (latency < 100) return "#84cc16";
if (latency < 200) return "#eab308";
if (latency < 500) return "#f97316";
return "#ef4444";
} }
export function latencyClass(latency, status) { export function latencyClass(latency, status) {
if (status === "offline" || status === "error" || latency === null) if (status === "offline" || status === "error" || latency === null)
return "text-gray-500"; return "text-gray-500";
return CONFIG.latencyColors.find((c) => latency < c.below).className; if (latency < 50) return "text-green-500";
if (latency < 100) return "text-lime-500";
if (latency < 200) return "text-yellow-500";
if (latency < 500) return "text-orange-500";
return "text-red-500";
} }
// --- Sparkline Renderer ------------------------------------------------------ // --- Sparkline Renderer ------------------------------------------------------
@@ -875,7 +845,7 @@ function buildUI(state) {
</div> </div>
<footer class="mt-8 text-center text-gray-600 text-xs"> <footer class="mt-8 text-center text-gray-600 text-xs">
<p>Latency measured via GET requests | CORS restrictions may affect some measurements</p> <p>Latency measured via GET requests | IPv4 only | CORS restrictions may affect some measurements</p>
<p class="mt-2"> <p class="mt-2">
<span class="inline-block w-3 h-3 rounded-full bg-green-500 mr-1 align-middle"></span>&lt;50ms <span class="inline-block w-3 h-3 rounded-full bg-green-500 mr-1 align-middle"></span>&lt;50ms
<span class="inline-block w-3 h-3 rounded-full bg-lime-500 mr-1 ml-3 align-middle"></span>&lt;100ms <span class="inline-block w-3 h-3 rounded-full bg-lime-500 mr-1 ml-3 align-middle"></span>&lt;100ms
@@ -942,7 +912,10 @@ function updateHostRow(host, index) {
latencyEl.innerHTML = `<span class="text-gray-500">---</span>`; latencyEl.innerHTML = `<span class="text-gray-500">---</span>`;
} }
const { min, med, avg, max } = host.historyStats(); const avg = host.averageLatency();
const med = host.medianLatency();
const min = host.minLatency();
const max = host.maxLatency();
if (host.status === "online" && avg !== null) { if (host.status === "online" && avg !== null) {
statusEl.innerHTML = statusStatsHTML([ statusEl.innerHTML = statusStatsHTML([
["min", min], ["min", min],
@@ -1154,26 +1127,6 @@ function sortAndRebuildWAN(state) {
// --- Main Loop --------------------------------------------------------------- // --- Main Loop ---------------------------------------------------------------
// Writes one line to the browser console, and the same line to the debug
// log, for a check the page is about to record: for every check that
// failed, and for a check that answered after checks that failed. Called
// before host.pushSample, while host.consecutiveFailures still counts the
// checks before this one.
function logCheck(host, result) {
const target = `${host.name} ${host.url} at ${new Date().toISOString()}`;
const failed = host.consecutiveFailures;
if (result.error) {
const line = `netwatch: check failed: ${target}: ${result.reason}`;
console.error(line);
log.error(line);
} else if (failed > 0) {
const checks = failed === 1 ? "check" : "checks";
const line = `netwatch: target recovered: ${target}: answered after ${result.latency} ms, following ${failed} failed ${checks} in a row`;
console.info(line);
log.info(line);
}
}
export async function tick(state, signal, onOffline) { export async function tick(state, signal, onOffline) {
const ts = Date.now(); const ts = Date.now();
@@ -1207,7 +1160,6 @@ export async function tick(state, signal, onOffline) {
if (state.paused || signal.aborted || state.tickCount === 0) { if (state.paused || signal.aborted || state.tickCount === 0) {
return; return;
} }
logCheck(host, r);
host.pushSample(ts, r); host.pushSample(ts, r);
updateHostRow(host, state.allHosts.indexOf(host)); updateHostRow(host, state.allHosts.indexOf(host));
log.debug(`${host.name}: ${r.error ? r.error : r.latency + "ms"}`); log.debug(`${host.name}: ${r.error ? r.error : r.latency + "ms"}`);
@@ -1230,9 +1182,8 @@ export async function tick(state, signal, onOffline) {
// rows whose check ended before the resume still read "paused" // rows whose check ended before the resume still read "paused"
state.allHosts.forEach((host, i) => updateHostRow(host, i)); state.allHosts.forEach((host, i) => updateHostRow(host, i));
// Sort after the first real check, then every CONFIG.roundsPerSort // Sort after the first real check, then every 10 ticks thereafter
// ticks thereafter if (state.tickCount === 2 || state.tickCount % 10 === 1) {
if (state.tickCount === 2 || state.tickCount % CONFIG.roundsPerSort === 1) {
sortAndRebuildWAN(state); sortAndRebuildWAN(state);
} }
@@ -1255,10 +1206,9 @@ export async function tick(state, signal, onOffline) {
// --- Recovery Probe ---------------------------------------------------------- // --- Recovery Probe ----------------------------------------------------------
// When offline, check CONFIG.recoveryProbeHosts random WAN hosts every // When offline, check 4 random WAN hosts every 500ms, giving up the checks
// CONFIG.recoveryProbeInterval ms, giving up the checks started one interval // started 500ms before, so at most 4 are ever waiting. As soon as one
// before, so at most that many are ever waiting. As soon as one answers, // answers, stop probing and start a new round at once.
// stop probing and start a new round at once.
function startRecoveryProbe(state, startRounds) { function startRecoveryProbe(state, startRounds) {
if (state._recoveryProbeId) return; // already running if (state._recoveryProbeId) return; // already running
const candidates = [...state.wan]; const candidates = [...state.wan];
@@ -1266,7 +1216,7 @@ function startRecoveryProbe(state, startRounds) {
const j = Math.floor(Math.random() * (i + 1)); const j = Math.floor(Math.random() * (i + 1));
[candidates[i], candidates[j]] = [candidates[j], candidates[i]]; [candidates[i], candidates[j]] = [candidates[j], candidates[i]];
} }
const canaries = candidates.slice(0, CONFIG.recoveryProbeHosts); const canaries = candidates.slice(0, 4);
log.notice( log.notice(
`Recovery probe started (${canaries.map((h) => h.name).join(", ")})`, `Recovery probe started (${canaries.map((h) => h.name).join(", ")})`,
); );
@@ -1283,7 +1233,7 @@ function startRecoveryProbe(state, startRounds) {
startRounds(); startRounds();
}); });
} }
}, CONFIG.recoveryProbeInterval); }, 500);
} }
function stopRecoveryProbe(state) { function stopRecoveryProbe(state) {
@@ -1547,7 +1497,7 @@ async function init() {
}); });
window.addEventListener("resize", () => handleResize(state)); window.addEventListener("resize", () => handleResize(state));
setTimeout(() => handleResize(state), CONFIG.resizeDelay); setTimeout(() => handleResize(state), 100);
} }
// Bootstrap only when loaded as the page: a real DOM containing the #app // Bootstrap only when loaded as the page: a real DOM containing the #app
+21 -184
View File
@@ -23,16 +23,12 @@ import {
// test looks it up and kept in elements under its selector until the next // test looks it up and kept in elements under its selector until the next
// test starts. As on a page, writing its text replaces its markup; the // test starts. As on a page, writing its text replaces its markup; the
// status dot greyOutUI looks for in it is not there. Drawing a sparkline // status dot greyOutUI looks for in it is not there. Drawing a sparkline
// does nothing; it looks for the pixel ratio on window and finds none. In // does nothing; it looks for the pixel ratio on window and finds none.
// each test, console.error and console.info print nothing and keep what
// they are given.
const doNothing = () => {};
let elements; let elements;
beforeEach((t) => { beforeEach(() => {
elements = {}; elements = {};
t.mock.method(console, "error", doNothing);
t.mock.method(console, "info", doNothing);
}); });
const doNothing = () => {};
const canvasContext = { const canvasContext = {
clearRect: doNothing, clearRect: doNothing,
beginPath: doNothing, beginPath: doNothing,
@@ -73,10 +69,8 @@ function statusText(state, host) {
// Mocks the clock for test t, so that a check lasting seconds takes no real // Mocks the clock for test t, so that a check lasting seconds takes no real
// time, and replaces fetch with targets that each answer after // time, and replaces fetch with targets that each answer after
// answerAfter(url) milliseconds of that clock, or never when that is // answerAfter(url) milliseconds of that clock, or never when that is
// Infinity. The target at unreachableUrl, if one is given, does not // Infinity. Both are restored when the test ends.
// answer: after that time its fetch fails with the error a browser gives function mockTargets(t, answerAfter) {
// the page for a network error. Both are restored when the test ends.
function mockTargets(t, answerAfter, unreachableUrl) {
t.mock.timers.enable({ apis: ["setTimeout", "Date"] }); t.mock.timers.enable({ apis: ["setTimeout", "Date"] });
t.mock.method(performance, "now", () => Date.now()); t.mock.method(performance, "now", () => Date.now());
t.mock.method( t.mock.method(
@@ -84,12 +78,8 @@ function mockTargets(t, answerAfter, unreachableUrl) {
"fetch", "fetch",
(url, { signal }) => (url, { signal }) =>
new Promise((resolve, reject) => { new Promise((resolve, reject) => {
const answer =
url === unreachableUrl
? () => reject(new TypeError("Failed to fetch"))
: resolve;
if (answerAfter(url) !== Infinity) { if (answerAfter(url) !== Infinity) {
setTimeout(answer, answerAfter(url)); setTimeout(resolve, answerAfter(url));
} }
signal.addEventListener("abort", () => reject(signal.reason)); signal.addEventListener("abort", () => reject(signal.reason));
}), }),
@@ -117,7 +107,6 @@ for (const interval of [10000, 30000]) {
assert.deepEqual(await settled(check), { assert.deepEqual(await settled(check), {
latency: slowAnswer, latency: slowAnswer,
error: null, error: null,
reason: null,
}); });
}); });
@@ -131,37 +120,10 @@ for (const interval of [10000, 30000]) {
assert.deepEqual(await settled(check), { assert.deepEqual(await settled(check), {
latency: null, latency: null,
error: "timeout", error: "timeout",
reason: `timed out after ${timeout} ms (limit ${timeout} ms)`,
}); });
}); });
} }
// A browser can run a timer late, on a busy page or in a background tab.
// Moving the mocked clock on 30000ms at once runs the 24000ms timeout with
// the clock already at 30000ms.
test("at a 30000ms interval, a check whose timeout runs late gives how long the request took and the time limit", async (t) => {
CONFIG.updateInterval = 30000;
mockTargets(t, () => Infinity);
const check = measureLatency("https://target.test");
t.mock.timers.tick(30000);
assert.deepEqual(await settled(check), {
latency: null,
error: "timeout",
reason: "timed out after 30000 ms (limit 24000 ms)",
});
});
test("a check fetches the target's URL as written, with no query string added", async (t) => {
mockTargets(t, () => 10);
const check = measureLatency("https://fsn1-speed.hetzner.com");
t.mock.timers.tick(10);
assert.notEqual(await settled(check), "still waiting");
assert.equal(
fetch.mock.calls[0].arguments[0],
"https://fsn1-speed.hetzner.com",
);
});
test("at a 30000ms interval, a target answering after 1000ms shows in its row while another target's check is still waiting", async (t) => { test("at a 30000ms interval, a target answering after 1000ms shows in its row while another target's check is still waiting", async (t) => {
CONFIG.updateInterval = 30000; CONFIG.updateInterval = 30000;
const state = new AppState([ const state = new AppState([
@@ -170,7 +132,7 @@ test("at a 30000ms interval, a target answering after 1000ms shows in its row wh
const answering = state.local[0]; const answering = state.local[0];
const waiting = state.wan[0]; const waiting = state.wan[0];
// No target but the answering one ever answers. // No target but the answering one ever answers.
mockTargets(t, (url) => (url === answering.url ? 1000 : Infinity)); mockTargets(t, (url) => (url.startsWith(answering.url) ? 1000 : Infinity));
// The third tick: the first is discarded as a whole, and the second ends // The third tick: the first is discarded as a whole, and the second ends
// by sorting the rows, which rebuilds a page that is not here. // by sorting the rows, which rebuilds a page that is not here.
state.tickCount = 2; state.tickCount = 2;
@@ -197,7 +159,7 @@ test("at a 30000ms interval, a check still waiting when its round is given up do
{ name: "Answering", url: "https://answering.test" }, { name: "Answering", url: "https://answering.test" },
]); ]);
const answering = state.local[0]; const answering = state.local[0];
mockTargets(t, (url) => (url === answering.url ? 1000 : Infinity)); mockTargets(t, (url) => (url.startsWith(answering.url) ? 1000 : Infinity));
state.tickCount = 2; state.tickCount = 2;
const roundChecks = new AbortController(); const roundChecks = new AbortController();
@@ -215,7 +177,7 @@ test("at a 30000ms interval, a check still waiting when the user pauses does not
{ name: "Answering", url: "https://answering.test" }, { name: "Answering", url: "https://answering.test" },
]); ]);
const answering = state.local[0]; const answering = state.local[0];
mockTargets(t, (url) => (url === answering.url ? 1000 : Infinity)); mockTargets(t, (url) => (url.startsWith(answering.url) ? 1000 : Infinity));
state.tickCount = 2; state.tickCount = 2;
const round = tick(state, new AbortController().signal); const round = tick(state, new AbortController().signal);
@@ -232,7 +194,7 @@ test("at a 30000ms interval, a check in the first round does not show in its row
{ name: "Answering", url: "https://answering.test" }, { name: "Answering", url: "https://answering.test" },
]); ]);
const answering = state.local[0]; const answering = state.local[0];
mockTargets(t, (url) => (url === answering.url ? 1000 : Infinity)); mockTargets(t, (url) => (url.startsWith(answering.url) ? 1000 : Infinity));
const round = tick(state, new AbortController().signal); const round = tick(state, new AbortController().signal);
t.mock.timers.tick(1000); t.mock.timers.tick(1000);
@@ -246,7 +208,7 @@ test("at a 30000ms interval, after the user pauses and resumes during a round, n
{ name: "Answering", url: "https://answering.test" }, { name: "Answering", url: "https://answering.test" },
]); ]);
const answering = state.local[0]; const answering = state.local[0];
mockTargets(t, (url) => (url === answering.url ? 1000 : Infinity)); mockTargets(t, (url) => (url.startsWith(answering.url) ? 1000 : Infinity));
state.tickCount = 2; state.tickCount = 2;
const round = tick(state, new AbortController().signal); const round = tick(state, new AbortController().signal);
@@ -269,115 +231,6 @@ test("at a 30000ms interval, after the user pauses and resumes during a round, n
} }
}); });
// In the next tests a round checks one target, Target, at a 30000ms
// interval, so a check times out after 24000ms. The mocked clock starts at
// 1970-01-01T00:00:00.000Z.
// An app state with Target and no WAN targets, so its rounds check only
// Target, and whose next round is recorded: it is the third, as the first
// is discarded and the second ends by sorting the rows, which rebuilds a
// page that is not here. Returns it and Target.
function stateWithOneTarget() {
CONFIG.updateInterval = 30000;
const state = new AppState([
{ name: "Target", url: "https://target.test" },
]);
state.wan = [];
state.tickCount = 2;
return { state, target: state.local[0] };
}
// What this test wrote to the browser console with console[method].
function consoleLines(method) {
return console[method].mock.calls.map((call) => call.arguments[0]);
}
test("a check that fails with a network error writes one console line with the target, the time, the error and how long the request took", async (t) => {
const { state, target } = stateWithOneTarget();
mockTargets(t, () => 23, target.url);
const round = tick(state, new AbortController().signal);
t.mock.timers.tick(23);
assert.notEqual(await settled(round), "still waiting");
assert.deepEqual(consoleLines("error"), [
"netwatch: check failed: Target https://target.test at 1970-01-01T00:00:00.023Z: network error (TypeError: Failed to fetch) after 23 ms",
]);
assert.deepEqual(consoleLines("info"), []);
});
test("a check that times out writes one console line with the target, the time, how long the request took and the time limit", async (t) => {
const { state } = stateWithOneTarget();
mockTargets(t, () => Infinity);
const round = tick(state, new AbortController().signal);
t.mock.timers.tick(24000);
assert.notEqual(await settled(round), "still waiting");
assert.deepEqual(consoleLines("error"), [
"netwatch: check failed: Target https://target.test at 1970-01-01T00:00:24.000Z: timed out after 24000 ms (limit 24000 ms)",
]);
assert.deepEqual(consoleLines("info"), []);
});
// An answer that took longer than CONFIG.maxLatency is recorded as a
// timeout. The limit is the check's own timeout, so such an answer is one
// that came in just as the check timed out; here it is lowered to 500ms.
test("a check answered over the time limit writes one console line with the target, the time, how long the answer took and the limit", async (t) => {
const { state } = stateWithOneTarget();
t.mock.getter(CONFIG, "maxLatency", () => 500);
mockTargets(t, () => 1000);
const round = tick(state, new AbortController().signal);
t.mock.timers.tick(1000);
assert.notEqual(await settled(round), "still waiting");
assert.deepEqual(consoleLines("error"), [
"netwatch: check failed: Target https://target.test at 1970-01-01T00:00:01.000Z: answered after 1000 ms, over the 500 ms limit",
]);
assert.deepEqual(consoleLines("info"), []);
});
test("a check that answers writes nothing to the console", async (t) => {
const { state } = stateWithOneTarget();
mockTargets(t, () => 30);
const round = tick(state, new AbortController().signal);
t.mock.timers.tick(30);
assert.notEqual(await settled(round), "still waiting");
assert.deepEqual(consoleLines("error"), []);
assert.deepEqual(consoleLines("info"), []);
});
for (const [failed, checks] of [
[1, "1 failed check"],
[3, "3 failed checks"],
]) {
test(`a check that answers after ${checks} writes one console line saying the target recovered, and the next writes nothing`, async (t) => {
const { state, target } = stateWithOneTarget();
for (let i = 0; i < failed; i++) {
target.pushSample(Date.now(), {
latency: null,
error: "unreachable",
});
}
mockTargets(t, () => 40);
for (let i = 0; i < 2; i++) {
const round = tick(state, new AbortController().signal);
t.mock.timers.tick(40);
assert.notEqual(await settled(round), "still waiting");
}
assert.deepEqual(consoleLines("info"), [
`netwatch: target recovered: Target https://target.test at 1970-01-01T00:00:00.040Z: answered after 40 ms, following ${checks} in a row`,
]);
assert.deepEqual(consoleLines("error"), []);
});
}
test("a check that fails in the first round, which is discarded, writes nothing to the console", async (t) => {
const { state, target } = stateWithOneTarget();
state.tickCount = 0;
mockTargets(t, () => 23, target.url);
const round = tick(state, new AbortController().signal);
t.mock.timers.tick(23);
assert.notEqual(await settled(round), "still waiting");
assert.deepEqual(consoleLines("error"), []);
assert.deepEqual(consoleLines("info"), []);
});
// The page shows &lt; &gt; &quot; &amp; and &#39; in a row's markup as // The page shows &lt; &gt; &quot; &amp; and &#39; in a row's markup as
// < > " & and '. // < > " & and '.
test(`a target whose name and URL hold < > " & and ' shows those characters in its row`, () => { test(`a target whose name and URL hold < > " & and ' shows those characters in its row`, () => {
@@ -494,35 +347,19 @@ for (const { history, latencies, statistics } of [
}, },
]) { ]) {
test(`a target's min, max, average and median latency over ${history}`, () => { test(`a target's min, max, average and median latency over ${history}`, () => {
const { min, max, avg, med } = hostAfter(latencies).historyStats(); const host = hostAfter(latencies);
assert.deepEqual({ min, max, average: avg, median: med }, statistics); assert.deepEqual(
{
min: host.minLatency(),
max: host.maxLatency(),
average: host.averageLatency(),
median: host.medianLatency(),
},
statistics,
);
}); });
} }
// The summary's figures come from each WAN target's last check, by the same
// rules as a target's own: here four answered, one was found unreachable
// and the rest have not been checked yet. The median, 22.5, and the
// average, 21.25, are rounded.
test("the summary's min, max, median and average latency over the WAN targets' last checks", () => {
const state = new AppState([]);
[30, 10, null, 25, 20].forEach((latency, i) =>
state.wan[i].pushSample(
Date.now(),
latency === null
? { latency: null, error: "unreachable" }
: { latency, error: null },
),
);
assert.deepEqual(state.wanStats(), {
reachable: 4,
total: state.wan.length,
min: 10,
max: 30,
med: 23,
avg: 21,
});
});
// An app state in which, of the WAN targets, the first timedOut timed out, // An app state in which, of the WAN targets, the first timedOut timed out,
// the next unreachable were found unreachable, the next answered answered // the next unreachable were found unreachable, the next answered answered
// after latency ms, and the rest have not been checked yet. // after latency ms, and the rest have not been checked yet.