Author SHA1 Message Date
clawbot a9c647a13b cibuild: the org model, which runs every check uncached (closes #37)
check / check (push) Successful in 2m2s
script/cibuild was a plain docker build ., so on a tree Docker had
seen before every check step came from the build cache and the build
still passed. It is now the org model from sneak/prompts, byte for
byte: script/bootstrap, script/check, then docker build --no-cache
with the git describe version as the VERSION build argument.

The workflow puts ~/.local/bin, where bootstrap links what it
installs, on the step's PATH. Bootstrap now installs its pinned node
when the installed one is older than 22.12.0, the oldest the
frontend's dependencies accept (puppeteer-core's engines field), as
it already does for Go against backend/go.mod.

Model: opus-5-5
2026-09-29 09:14:18 +00:00
9 changed files with 32 additions and 138 deletions
+1 -2
View File
@@ -77,9 +77,8 @@ COPY --from=frontend /app/dist /usr/share/nginx/html
COPY --from=builder /src/netwatch-server /usr/local/bin/netwatch-server COPY --from=builder /src/netwatch-server /usr/local/bin/netwatch-server
COPY bin/entrypoint.sh /usr/local/bin/entrypoint.sh COPY bin/entrypoint.sh /usr/local/bin/entrypoint.sh
# bin/entrypoint.sh creates DATA_DIR at start and gives it and /data to
# the netwatch user, whatever is mounted there.
ENV DATA_DIR=/data/reports ENV DATA_DIR=/data/reports
RUN mkdir -p /data/reports && chown -R netwatch:netwatch /data
VOLUME /data VOLUME /data
# The default public port; PORT changes it. # The default public port; PORT changes it.
+16 -10
View File
@@ -52,8 +52,8 @@ halves, so the root `make check` fails if either one is broken. We provide:
- `script/fmt` — format all files (writes): prettier, then gofmt over `backend/` - `script/fmt` — format all files (writes): prettier, then gofmt over `backend/`
- `script/fmt-check` — check formatting (read-only): prettier, then gofmt - `script/fmt-check` — check formatting (read-only): prettier, then gofmt
- `script/check` — run test, lint, and fmt-check - `script/check` — run test, lint, and fmt-check
- `script/frontend-test` — run the unit tests in `test/unit/` with Node's - `script/frontend-test` — run the production build as the frontend's test (no
built-in test runner, then the production build unit tests yet)
- `script/frontend-lint` — run prettier in check mode - `script/frontend-lint` — run prettier in check mode
- `script/frontend-fmt` — format everything prettier understands (writes) - `script/frontend-fmt` — format everything prettier understands (writes)
- `script/frontend-fmt-check` — check prettier formatting (read-only) - `script/frontend-fmt-check` — check prettier formatting (read-only)
@@ -136,11 +136,8 @@ Local hosts are tracked separately from WAN stats.
### Latency measurement ### Latency measurement
HEAD requests with `mode: 'no-cors'` and `cache: 'no-store'`, timed with HEAD requests with `mode: 'no-cors'` and `cache: 'no-store'`, timed with
`performance.now()`. Each check times out after 80% of the refresh interval (24 `performance.now()`. 1-second timeout; anything over 1000ms is clamped to
seconds at 30 seconds) and is then recorded as a timeout, so a round's checks unreachable. IPv4 only.
have all finished before the next round is due. A round due while the last one
is still waiting, which happens only after an interval change or when the
recovery probe starts one, is skipped. IPv4 only.
### Color coding ### Color coding
@@ -195,9 +192,8 @@ only inside the container, on `127.0.0.1:8081`. The image:
- Sends the security headers `REPO_POLICIES.md` requires on every response, as - Sends the security headers `REPO_POLICIES.md` requires on every response, as
`security-headers.conf` sets them, in place of the backend's own `security-headers.conf` sets them, in place of the backend's own
- Stores reports in `DATA_DIR`, `/data/reports` by default, on the `/data` - Stores reports in `DATA_DIR`, `/data/reports` by default, on the `/data`
volume. Before the backend starts, the image creates `DATA_DIR` and gives it volume. The backend runs as user `netwatch` (uid 1000), so a directory
and `/data` to user `netwatch` (uid 1000), which the backend runs as, so a bind-mounted at `/data` must be writable by uid 1000
host directory bind-mounted at `/data` ends up owned by uid 1000
- Writes buffered reports to disk on `docker stop`, and exits non-zero if nginx - Writes buffered reports to disk on `docker stop`, and exits non-zero if nginx
or the backend exits on its own, so the platform restarts it or the backend exits on its own, so the platform restarts it
@@ -207,6 +203,16 @@ What the [upaas](https://git.eeqj.de/sneak/upaas) app for netwatch needs:
- **Port:** container port `8080`. - **Port:** container port `8080`.
- **Volume:** container path `/data`; the reports are kept in `/data/reports`. - **Volume:** container path `/data`; the reports are kept in `/data/reports`.
- **First run:** upaas bind-mounts the host directory it is given and does not
create it, and the backend, which runs as uid 1000, does not start unless it
can write there. Create the directory, owned by uid 1000, before the first
deploy:
```bash
mkdir -p /path/to/data
chown 1000:1000 /path/to/data
```
- **Environment variables:** none is required. An empty one counts as unset, and - **Environment variables:** none is required. An empty one counts as unset, and
one set to a value netwatch cannot use stops the container at start, with the one set to a value netwatch cannot use stops the container at start, with the
reason in its log. reason in its log.
-16
View File
@@ -23,22 +23,6 @@ latest run passes.
# Completed Steps # Completed Steps
- 2026-09-29: each target check times out after 80% of the refresh interval
(issue #78), 24 seconds at 30 seconds, where it was capped at 3 seconds. A
round due while the last one's checks are still waiting is skipped, so rounds
never overlap, and the recovery probe starts no new checks while its last ones
are waiting. The frontend has its first unit tests, run by
`script/frontend-test` with Node's built-in test runner; for them,
`index.html` now links `src/styles.css`, which `src/main.js` used to import
- 2026-09-29: the container sets up its own data directory (issue #75):
`bin/entrypoint.sh`, still as root, creates `DATA_DIR` if missing and gives it
and `/data` to the `netwatch` user with mode 750 before starting the backend
as that user, so an empty host directory owned by root, or one holding files
from another uid, works with no step on the host. It stops the start instead
when a symbolic link is on the path to `DATA_DIR`, since root would change
whatever the link points to. The `README.md` first-run step that created and
chowned the host directory is gone, and the image no longer sets that
ownership at build time
- 2026-09-29: CI can no longer pass on checks that did not run (issue #37): - 2026-09-29: CI can no longer pass on checks that did not run (issue #37):
`script/cibuild` is now the org model, byte for byte. It runs `script/cibuild` is now the org model, byte for byte. It runs
`script/bootstrap` and `script/check`, then builds the image with `--no-cache` `script/bootstrap` and `script/check`, then builds the image with `--no-cache`
+1 -2
View File
@@ -104,8 +104,7 @@ this server. The image's entrypoint, `bin/entrypoint.sh`, starts the server as
user `netwatch` (uid 1000) with `BIND_ADDRESS=127.0.0.1` and `PORT=8081`, so user `netwatch` (uid 1000) with `BIND_ADDRESS=127.0.0.1` and `PORT=8081`, so
only nginx reaches it, and with `TRUSTED_PROXIES=127.0.0.1/32`, so it takes the only nginx reaches it, and with `TRUSTED_PROXIES=127.0.0.1/32`, so it takes the
client address nginx passes on and no other. `DATA_DIR` is `/data/reports`, on client address nginx passes on and no other. `DATA_DIR` is `/data/reports`, on
the `/data` volume; the entrypoint creates it and gives it and `/data` to the `/data` volume, which `netwatch` owns. nginx replaces the security headers
`netwatch` before starting the server. nginx replaces the security headers
this server sets with those in the root `security-headers.conf`, so those are this server sets with those in the root `security-headers.conf`, so those are
what clients of the image see. what clients of the image see.
-23
View File
@@ -61,29 +61,6 @@ for proxy in $(printf '%s' "$TRUSTED_PROXIES" | tr ',' ' '); do
echo "set_real_ip_from $cidr;" echo "set_real_ip_from $cidr;"
done > /etc/nginx/trusted-proxies.conf done > /etc/nginx/trusted-proxies.conf
# netwatch-server keeps its report files in DATA_DIR, on the /data
# volume, which may be a host directory owned by root or by another
# uid. Both are given to the netwatch user here, with the mode the
# server gives a directory it creates, so the host directory needs no
# preparing.
#
# chown and chmod, run as root, change whatever a symbolic link on the
# path points to, anywhere in the container, and the netwatch user can
# put one in /data. So the start stops unless readlink -f, which
# follows every link on a path, gives /data and DATA_DIR back as they
# are. It also writes a path in full, so a DATA_DIR with '.', '..' or
# an extra '/' in it is refused too.
export DATA_DIR="${DATA_DIR:-/data/reports}"
mkdir -p "$DATA_DIR" || exit 1
if [ "$(readlink -f /data)" != /data ] ||
[ "$(readlink -f "$DATA_DIR")" != "$DATA_DIR" ]; then
echo "entrypoint: DATA_DIR must be a full path with no '.', '..'," \
"extra '/' or symbolic link on it or on /data, not '$DATA_DIR'" >&2
exit 1
fi
chown -R netwatch:netwatch /data "$DATA_DIR" || exit 1
chmod 750 /data "$DATA_DIR" || exit 1
# A stop signal is only noted here; the loop below acts on it. # A stop signal is only noted here; the loop below acts on it.
stop_requested="" stop_requested=""
trap 'stop_requested=yes' TERM INT trap 'stop_requested=yes' TERM INT
-3
View File
@@ -9,9 +9,6 @@
type="image/svg+xml" type="image/svg+xml"
href="data:image/svg+xml,<svg xmlns='http://www.w3.org/2000/svg' viewBox='0 0 100 100'><text y='.9em' font-size='90'>📡</text></svg>" href="data:image/svg+xml,<svg xmlns='http://www.w3.org/2000/svg' viewBox='0 0 100 100'><text y='.9em' font-size='90'>📡</text></svg>"
/> />
<!-- Linked here, not imported by src/main.js, so the unit tests can
import that module in Node, which cannot import CSS. -->
<link rel="stylesheet" href="/src/styles.css" />
</head> </head>
<body class="bg-gray-900 text-white min-h-screen"> <body class="bg-gray-900 text-white min-h-screen">
<div id="app"></div> <div id="app"></div>
+3 -4
View File
@@ -1,14 +1,13 @@
#!/bin/sh #!/bin/sh
# script/frontend-test: run the frontend test suite: the unit tests in # script/frontend-test: run the frontend test suite. The frontend has no
# test/unit/ with Node's built-in test runner, then the production # unit tests; the production build serves as the test (fails on broken
# build, which fails on broken code. # code).
set -eu set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)" ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
main() { main() {
cd "$ROOT" cd "$ROOT"
timeout 30 node --test test/unit/*.test.js
timeout 30 yarn build timeout 30 yarn build
} }
+11 -26
View File
@@ -1,14 +1,14 @@
import "./styles.css";
// --- Configuration ----------------------------------------------------------- // --- Configuration -----------------------------------------------------------
// Timing, axis labels, and display constants. A target check times out // Timing, axis labels, and display constants. Latency above maxLatency is
// after requestTimeout, 80% of updateInterval, so a round's checks have // clamped to "unreachable". The sparkline Y-axis is capped at
// all finished before the next round is due; latency above maxLatency is
// recorded as a timeout. The sparkline Y-axis is capped at
// graphMaxLatency — values above it pin to the top of the chart but still // graphMaxLatency — values above it pin to the top of the chart but still
// display their real value in the latency figure. The history buffer holds // display their real value in the latency figure. The history buffer holds
// maxHistoryPoints samples (historyDuration / updateInterval). // maxHistoryPoints samples (historyDuration / updateInterval).
// reportInterval is how often collected samples are POSTed to the backend. // reportInterval is how often collected samples are POSTed to the backend.
export const CONFIG = { const CONFIG = {
updateInterval: 3000, updateInterval: 3000,
maxHistoryPoints: 100, maxHistoryPoints: 100,
reportInterval: 60000, reportInterval: 60000,
@@ -16,7 +16,7 @@ export const CONFIG = {
return (this.maxHistoryPoints * this.updateInterval) / 1000; return (this.maxHistoryPoints * this.updateInterval) / 1000;
}, },
get requestTimeout() { get requestTimeout() {
return this.updateInterval * 0.8; return Math.min(this.updateInterval - 100, 3000);
}, },
get maxLatency() { get maxLatency() {
return this.requestTimeout; return this.requestTimeout;
@@ -503,7 +503,7 @@ class Reporter {
// --- Latency Measurement ----------------------------------------------------- // --- Latency Measurement -----------------------------------------------------
export async function measureLatency(url) { async function measureLatency(url) {
const controller = new AbortController(); const controller = new AbortController();
const timeoutId = setTimeout( const timeoutId = setTimeout(
() => controller.abort(), () => controller.abort(),
@@ -1181,16 +1181,11 @@ function startRecoveryProbe(state, triggerTick) {
log.notice( log.notice(
`Recovery probe started (${canaries.map((h) => h.name).join(", ")})`, `Recovery probe started (${canaries.map((h) => h.name).join(", ")})`,
); );
// A check can wait up to CONFIG.requestTimeout, far longer than
// 500ms, so no new checks start while the last ones are waiting.
let checking = false;
state._recoveryProbeId = setInterval(async () => { state._recoveryProbeId = setInterval(async () => {
if (state.paused || checking) return; if (state.paused) return;
checking = true;
const results = await Promise.all( const results = await Promise.all(
canaries.map((h) => measureLatency(h.url)), canaries.map((h) => measureLatency(h.url)),
); );
checking = false;
if (results.some((r) => r.error === null)) { if (results.some((r) => r.error === null)) {
log.notice("Recovery probe: connectivity detected"); log.notice("Recovery probe: connectivity detected");
stopRecoveryProbe(state); stopRecoveryProbe(state);
@@ -1379,18 +1374,8 @@ async function init() {
updateClocks(); updateClocks();
setInterval(updateClocks, 1000); setInterval(updateClocks, 1000);
// A round waits up to CONFIG.requestTimeout for its checks. A round function doTick() {
// asked for while one is still waiting, after an interval change or tick(state, () => startRecoveryProbe(state, doTick));
// by the recovery probe, is skipped, so rounds never overlap.
let roundRunning = false;
async function doTick() {
if (roundRunning) return;
roundRunning = true;
try {
await tick(state, () => startRecoveryProbe(state, doTick));
} finally {
roundRunning = false;
}
} }
doTick(); doTick();
@@ -1459,7 +1444,7 @@ async function init() {
// Bootstrap only when loaded as the page: a real DOM containing the #app // Bootstrap only when loaded as the page: a real DOM containing the #app
// mount point this module renders into. Importing the module in a unit test // mount point this module renders into. Importing the module in a unit test
// (which has no #app) runs nothing, so its exports can be tested in isolation. // (which has no #app) runs nothing, so buildReport can be tested in isolation.
if (typeof document !== "undefined" && document.getElementById("app")) { if (typeof document !== "undefined" && document.getElementById("app")) {
if (document.readyState === "loading") { if (document.readyState === "loading") {
document.addEventListener("DOMContentLoaded", init); document.addEventListener("DOMContentLoaded", init);
-52
View File
@@ -1,52 +0,0 @@
// Unit tests for src/main.js, run by script/frontend-test with Node's
// built-in test runner. Importing the module does not start the page.
import { after, before, test } from "node:test";
import assert from "node:assert/strict";
import { createServer } from "node:http";
import { CONFIG, measureLatency } from "../../src/main.js";
// measureLatency writes timeouts to the debug log, which looks for its
// panel in the page. There is no page here.
globalThis.document = { getElementById: () => null };
// A target that answers after the number of milliseconds in the path,
// e.g. /600.
let server;
let target;
before(async () => {
server = createServer((req, res) => {
const delay = Number(new URL(req.url, "http://x").pathname.slice(1));
setTimeout(() => res.end(), delay);
});
await new Promise((resolve) => server.listen(0, "127.0.0.1", resolve));
target = `http://127.0.0.1:${server.address().port}`;
});
after(() => {
server.closeAllConnections();
server.close();
});
test("the timeout is 80% of the refresh interval", () => {
CONFIG.updateInterval = 30000;
assert.equal(CONFIG.requestTimeout, 24000);
CONFIG.updateInterval = 3000;
assert.equal(CONFIG.requestTimeout, 2400);
});
test("an answer within the timeout is recorded with its real time, a later one as a timeout", async () => {
// 600ms is past the 400ms timeout of a 500ms interval...
CONFIG.updateInterval = 500;
assert.deepEqual(await measureLatency(`${target}/600`), {
latency: null,
error: "timeout",
});
// ...and within the 1200ms timeout of a 1500ms interval.
CONFIG.updateInterval = 1500;
const { latency, error } = await measureLatency(`${target}/600`);
assert.equal(error, null);
assert.ok(latency >= 600 && latency < 1200, `latency ${latency}ms`);
});