Author SHA1 Message Date
clawbot 6f8f42f2bf Target check timeout is 80% of the refresh interval (closes #78)
check / check (push) Successful in 1m54s
Each target check now times out after 80% of the refresh interval, 24
seconds at 30 seconds, where it was capped at 3 seconds, so slow, far
targets are recorded with their real time. A round asked for while the
last one's checks are still waiting, after an interval change or by the
recovery probe, is skipped, so rounds never overlap; the recovery probe
starts no new checks while its last ones wait.

The frontend has its first unit tests, run by script/frontend-test with
Node's built-in test runner. index.html now links src/styles.css, which
src/main.js imported, since Node cannot import CSS.

Model: opus-5-5
2026-09-29 10:32:03 +00:00
sneak 7dfb3b8c32 Merge branch 'main' into next
check / check (push) Successful in 2m6s
2026-09-29 12:04:10 +02:00
clawbot a5ca73c585 Container sets up its own data directory (closes #75) (#76)
check / check (push) Successful in 2m6s
Closes #75.

`bin/entrypoint.sh`, which already runs as root, now makes the data directory usable before the backend starts: it creates `DATA_DIR` if missing, gives it and `/data` to the `netwatch` user (`chown -R`), and sets mode 750 on both, the mode the backend gives a directory it creates. The backend still runs as `netwatch`. The README "Running under upaas" section loses its first-run step that created and chowned the host directory and names only the path to mount. The Dockerfile's build-time `mkdir` and `chown` of `/data` are gone, since the entrypoint now does this on every start.

What the diff does not show:

- The host directory mounted at `/data` ends up owned by uid 1000 with mode 750, and everything under `DATA_DIR` is chowned to uid 1000 on every start.
- If the directory cannot be created or chowned, the container stops with that tool's error before either process starts.

Recorded runs with `--mount type=bind`: an empty directory owned by root (mode 755, and again mode 700), and one holding a `reports` directory and report file owned by uid 1001 with mode 700. Each time the container turned healthy, `netwatch-server` ran as `netwatch`, and a posted report was written to `DATA_DIR`; a second start on the root-owned and the uid 1001 directories did the same.

Judgement call: `/data` itself is given to `netwatch` as well as `DATA_DIR`, so the backend can reach `DATA_DIR` inside a host directory with mode 700.

Model: opus-5-5
Reviewed-on: #76
Co-authored-by: clawbot <35+clawbot@noreply.example.org>
2026-09-29 12:03:59 +02:00
sneak bd08e901ee next into main: netwatch as one container, ready for upaas (#49)
check / check (push) Successful in 12s
Reviewed-on: #49
2026-09-29 10:43:12 +02:00
9 changed files with 138 additions and 32 deletions
+2 -1
View File
@@ -77,8 +77,9 @@ COPY --from=frontend /app/dist /usr/share/nginx/html
COPY --from=builder /src/netwatch-server /usr/local/bin/netwatch-server
COPY bin/entrypoint.sh /usr/local/bin/entrypoint.sh
# bin/entrypoint.sh creates DATA_DIR at start and gives it and /data to
# the netwatch user, whatever is mounted there.
ENV DATA_DIR=/data/reports
RUN mkdir -p /data/reports && chown -R netwatch:netwatch /data
VOLUME /data
# The default public port; PORT changes it.
+10 -16
View File
@@ -52,8 +52,8 @@ halves, so the root `make check` fails if either one is broken. We provide:
- `script/fmt` — format all files (writes): prettier, then gofmt over `backend/`
- `script/fmt-check` — check formatting (read-only): prettier, then gofmt
- `script/check` — run test, lint, and fmt-check
- `script/frontend-test` — run the production build as the frontend's test (no
unit tests yet)
- `script/frontend-test` — run the unit tests in `test/unit/` with Node's
built-in test runner, then the production build
- `script/frontend-lint` — run prettier in check mode
- `script/frontend-fmt` — format everything prettier understands (writes)
- `script/frontend-fmt-check` — check prettier formatting (read-only)
@@ -136,8 +136,11 @@ Local hosts are tracked separately from WAN stats.
### Latency measurement
HEAD requests with `mode: 'no-cors'` and `cache: 'no-store'`, timed with
`performance.now()`. 1-second timeout; anything over 1000ms is clamped to
unreachable. IPv4 only.
`performance.now()`. Each check times out after 80% of the refresh interval (24
seconds at 30 seconds) and is then recorded as a timeout, so a round's checks
have all finished before the next round is due. A round due while the last one
is still waiting, which happens only after an interval change or when the
recovery probe starts one, is skipped. IPv4 only.
### Color coding
@@ -192,8 +195,9 @@ only inside the container, on `127.0.0.1:8081`. The image:
- Sends the security headers `REPO_POLICIES.md` requires on every response, as
`security-headers.conf` sets them, in place of the backend's own
- Stores reports in `DATA_DIR`, `/data/reports` by default, on the `/data`
volume. The backend runs as user `netwatch` (uid 1000), so a directory
bind-mounted at `/data` must be writable by uid 1000
volume. Before the backend starts, the image creates `DATA_DIR` and gives it
and `/data` to user `netwatch` (uid 1000), which the backend runs as, so a
host directory bind-mounted at `/data` ends up owned by uid 1000
- Writes buffered reports to disk on `docker stop`, and exits non-zero if nginx
or the backend exits on its own, so the platform restarts it
@@ -203,16 +207,6 @@ What the [upaas](https://git.eeqj.de/sneak/upaas) app for netwatch needs:
- **Port:** container port `8080`.
- **Volume:** container path `/data`; the reports are kept in `/data/reports`.
- **First run:** upaas bind-mounts the host directory it is given and does not
create it, and the backend, which runs as uid 1000, does not start unless it
can write there. Create the directory, owned by uid 1000, before the first
deploy:
```bash
mkdir -p /path/to/data
chown 1000:1000 /path/to/data
```
- **Environment variables:** none is required. An empty one counts as unset, and
one set to a value netwatch cannot use stops the container at start, with the
reason in its log.
+16
View File
@@ -23,6 +23,22 @@ latest run passes.
# Completed Steps
- 2026-09-29: each target check times out after 80% of the refresh interval
(issue #78), 24 seconds at 30 seconds, where it was capped at 3 seconds. A
round due while the last one's checks are still waiting is skipped, so rounds
never overlap, and the recovery probe starts no new checks while its last ones
are waiting. The frontend has its first unit tests, run by
`script/frontend-test` with Node's built-in test runner; for them,
`index.html` now links `src/styles.css`, which `src/main.js` used to import
- 2026-09-29: the container sets up its own data directory (issue #75):
`bin/entrypoint.sh`, still as root, creates `DATA_DIR` if missing and gives it
and `/data` to the `netwatch` user with mode 750 before starting the backend
as that user, so an empty host directory owned by root, or one holding files
from another uid, works with no step on the host. It stops the start instead
when a symbolic link is on the path to `DATA_DIR`, since root would change
whatever the link points to. The `README.md` first-run step that created and
chowned the host directory is gone, and the image no longer sets that
ownership at build time
- 2026-09-29: CI can no longer pass on checks that did not run (issue #37):
`script/cibuild` is now the org model, byte for byte. It runs
`script/bootstrap` and `script/check`, then builds the image with `--no-cache`
+2 -1
View File
@@ -104,7 +104,8 @@ this server. The image's entrypoint, `bin/entrypoint.sh`, starts the server as
user `netwatch` (uid 1000) with `BIND_ADDRESS=127.0.0.1` and `PORT=8081`, so
only nginx reaches it, and with `TRUSTED_PROXIES=127.0.0.1/32`, so it takes the
client address nginx passes on and no other. `DATA_DIR` is `/data/reports`, on
the `/data` volume, which `netwatch` owns. nginx replaces the security headers
the `/data` volume; the entrypoint creates it and gives it and `/data` to
`netwatch` before starting the server. nginx replaces the security headers
this server sets with those in the root `security-headers.conf`, so those are
what clients of the image see.
+23
View File
@@ -61,6 +61,29 @@ for proxy in $(printf '%s' "$TRUSTED_PROXIES" | tr ',' ' '); do
echo "set_real_ip_from $cidr;"
done > /etc/nginx/trusted-proxies.conf
# netwatch-server keeps its report files in DATA_DIR, on the /data
# volume, which may be a host directory owned by root or by another
# uid. Both are given to the netwatch user here, with the mode the
# server gives a directory it creates, so the host directory needs no
# preparing.
#
# chown and chmod, run as root, change whatever a symbolic link on the
# path points to, anywhere in the container, and the netwatch user can
# put one in /data. So the start stops unless readlink -f, which
# follows every link on a path, gives /data and DATA_DIR back as they
# are. It also writes a path in full, so a DATA_DIR with '.', '..' or
# an extra '/' in it is refused too.
export DATA_DIR="${DATA_DIR:-/data/reports}"
mkdir -p "$DATA_DIR" || exit 1
if [ "$(readlink -f /data)" != /data ] ||
[ "$(readlink -f "$DATA_DIR")" != "$DATA_DIR" ]; then
echo "entrypoint: DATA_DIR must be a full path with no '.', '..'," \
"extra '/' or symbolic link on it or on /data, not '$DATA_DIR'" >&2
exit 1
fi
chown -R netwatch:netwatch /data "$DATA_DIR" || exit 1
chmod 750 /data "$DATA_DIR" || exit 1
# A stop signal is only noted here; the loop below acts on it.
stop_requested=""
trap 'stop_requested=yes' TERM INT
+3
View File
@@ -9,6 +9,9 @@
type="image/svg+xml"
href="data:image/svg+xml,<svg xmlns='http://www.w3.org/2000/svg' viewBox='0 0 100 100'><text y='.9em' font-size='90'>📡</text></svg>"
/>
<!-- Linked here, not imported by src/main.js, so the unit tests can
import that module in Node, which cannot import CSS. -->
<link rel="stylesheet" href="/src/styles.css" />
</head>
<body class="bg-gray-900 text-white min-h-screen">
<div id="app"></div>
+4 -3
View File
@@ -1,13 +1,14 @@
#!/bin/sh
# script/frontend-test: run the frontend test suite. The frontend has no
# unit tests; the production build serves as the test (fails on broken
# code).
# script/frontend-test: run the frontend test suite: the unit tests in
# test/unit/ with Node's built-in test runner, then the production
# build, which fails on broken code.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
main() {
cd "$ROOT"
timeout 30 node --test test/unit/*.test.js
timeout 30 yarn build
}
+26 -11
View File
@@ -1,14 +1,14 @@
import "./styles.css";
// --- Configuration -----------------------------------------------------------
// Timing, axis labels, and display constants. Latency above maxLatency is
// clamped to "unreachable". The sparkline Y-axis is capped at
// Timing, axis labels, and display constants. A target check times out
// after requestTimeout, 80% of updateInterval, so a round's checks have
// all finished before the next round is due; latency above maxLatency is
// recorded as a timeout. The sparkline Y-axis is capped at
// graphMaxLatency — values above it pin to the top of the chart but still
// display their real value in the latency figure. The history buffer holds
// maxHistoryPoints samples (historyDuration / updateInterval).
// reportInterval is how often collected samples are POSTed to the backend.
const CONFIG = {
export const CONFIG = {
updateInterval: 3000,
maxHistoryPoints: 100,
reportInterval: 60000,
@@ -16,7 +16,7 @@ const CONFIG = {
return (this.maxHistoryPoints * this.updateInterval) / 1000;
},
get requestTimeout() {
return Math.min(this.updateInterval - 100, 3000);
return this.updateInterval * 0.8;
},
get maxLatency() {
return this.requestTimeout;
@@ -503,7 +503,7 @@ class Reporter {
// --- Latency Measurement -----------------------------------------------------
async function measureLatency(url) {
export async function measureLatency(url) {
const controller = new AbortController();
const timeoutId = setTimeout(
() => controller.abort(),
@@ -1181,11 +1181,16 @@ function startRecoveryProbe(state, triggerTick) {
log.notice(
`Recovery probe started (${canaries.map((h) => h.name).join(", ")})`,
);
// A check can wait up to CONFIG.requestTimeout, far longer than
// 500ms, so no new checks start while the last ones are waiting.
let checking = false;
state._recoveryProbeId = setInterval(async () => {
if (state.paused) return;
if (state.paused || checking) return;
checking = true;
const results = await Promise.all(
canaries.map((h) => measureLatency(h.url)),
);
checking = false;
if (results.some((r) => r.error === null)) {
log.notice("Recovery probe: connectivity detected");
stopRecoveryProbe(state);
@@ -1374,8 +1379,18 @@ async function init() {
updateClocks();
setInterval(updateClocks, 1000);
function doTick() {
tick(state, () => startRecoveryProbe(state, doTick));
// A round waits up to CONFIG.requestTimeout for its checks. A round
// asked for while one is still waiting, after an interval change or
// by the recovery probe, is skipped, so rounds never overlap.
let roundRunning = false;
async function doTick() {
if (roundRunning) return;
roundRunning = true;
try {
await tick(state, () => startRecoveryProbe(state, doTick));
} finally {
roundRunning = false;
}
}
doTick();
@@ -1444,7 +1459,7 @@ async function init() {
// Bootstrap only when loaded as the page: a real DOM containing the #app
// mount point this module renders into. Importing the module in a unit test
// (which has no #app) runs nothing, so buildReport can be tested in isolation.
// (which has no #app) runs nothing, so its exports can be tested in isolation.
if (typeof document !== "undefined" && document.getElementById("app")) {
if (document.readyState === "loading") {
document.addEventListener("DOMContentLoaded", init);
+52
View File
@@ -0,0 +1,52 @@
// Unit tests for src/main.js, run by script/frontend-test with Node's
// built-in test runner. Importing the module does not start the page.
import { after, before, test } from "node:test";
import assert from "node:assert/strict";
import { createServer } from "node:http";
import { CONFIG, measureLatency } from "../../src/main.js";
// measureLatency writes timeouts to the debug log, which looks for its
// panel in the page. There is no page here.
globalThis.document = { getElementById: () => null };
// A target that answers after the number of milliseconds in the path,
// e.g. /600.
let server;
let target;
before(async () => {
server = createServer((req, res) => {
const delay = Number(new URL(req.url, "http://x").pathname.slice(1));
setTimeout(() => res.end(), delay);
});
await new Promise((resolve) => server.listen(0, "127.0.0.1", resolve));
target = `http://127.0.0.1:${server.address().port}`;
});
after(() => {
server.closeAllConnections();
server.close();
});
test("the timeout is 80% of the refresh interval", () => {
CONFIG.updateInterval = 30000;
assert.equal(CONFIG.requestTimeout, 24000);
CONFIG.updateInterval = 3000;
assert.equal(CONFIG.requestTimeout, 2400);
});
test("an answer within the timeout is recorded with its real time, a later one as a timeout", async () => {
// 600ms is past the 400ms timeout of a 500ms interval...
CONFIG.updateInterval = 500;
assert.deepEqual(await measureLatency(`${target}/600`), {
latency: null,
error: "timeout",
});
// ...and within the 1200ms timeout of a 1500ms interval.
CONFIG.updateInterval = 1500;
const { latency, error } = await measureLatency(`${target}/600`);
assert.equal(error, null);
assert.ok(latency >= 600 && latency < 1200, `latency ${latency}ms`);
});