From 17470dc4a4b80a030866be33b67087236dda2993 Mon Sep 17 00:00:00 2001 From: sneak Date: Sun, 4 Oct 2026 04:06:50 +0000 Subject: [PATCH] README, TODO and the viewport README say what the tree does (closes #24) README.md: Getting Started leads with make targets; a Backend section gives netwatch-server's routes and how the image builds and runs it; the checks are GET requests; the WAN host list, health states, summary figures, sorting and missing features match src/main.js; the TODO section points to TODO.md, which holds the one to-do list. backend/README.md: its TODO section points to TODO.md too, whose Future Steps take its three open items. TODO.md: Workflow branches from next and opens the PR against next; Status, Next Step and Future Steps describe the open work, linked to its issue where one exists. test/viewport/README.md: the unit tests run on Node's test runner, not vitest. Model: opus-5-5 --- README.md | 158 ++++++++++++++++++++++++++-------------- TODO.md | 53 +++++++++----- backend/README.md | 5 +- test/viewport/README.md | 15 ++-- 4 files changed, 147 insertions(+), 84 deletions(-) diff --git a/README.md b/README.md index 05184c1..c3a11e0 100644 --- a/README.md +++ b/README.md @@ -1,30 +1,33 @@ NetWatch is an MIT-licensed JavaScript single-page application by [@sneak](https://sneak.berlin) that provides real-time network latency monitoring to common internet hosts, displayed with color-coded figures and -sparkline graphs, served from a static bucket or Docker container. +sparkline graphs, served from a static bucket or from its Docker image, where a +small Go backend stores the measurements the page reports. ## Getting Started ```bash -# Install dependencies -yarn install +# Install the dependencies and the git pre-commit hook +make setup -# Development server -yarn dev +# Run the page on the Vite dev server +make dev -# Production build -yarn build +# Run the tests, both linters and the format check +make check -# Preview production build -yarn preview +# Build the page into dist/ +make build -# Docker -docker build -t netwatch . +# Build the image and run it +make docker docker run -p 8080:8080 netwatch ``` -`yarn dev` proxies `/api` to `http://127.0.0.1:8080`, so a locally running -`netwatch-server` (see `backend/`) receives the reports the page posts. +`make check` and `make docker` need Docker. `make dev` passes `/api` to +`http://127.0.0.1:8080`, where `make run` in `backend/` starts `netwatch-server` +with its defaults, so the reports the page posts are stored in +`backend/data/reports`. ## Entrypoints @@ -94,22 +97,25 @@ halves, so the root `make check` fails if either one is broken. We provide: The narrow-viewport layout lives in the `max-width: 768px` media block in `src/styles.css`. It is verified automatically by `make frontend-viewport-test`, which drives a digest-pinned headless Chrome against the built `dist/` and -asserts on computed layout at widths derived from that CSS — one pixel either -side of every breakpoint it declares, plus a 320px floor, a desktop baseline and -two landscape sizes. See [test/viewport/README.md](test/viewport/README.md) for -what it covers and what it genuinely cannot. +asserts on computed layout at widths derived from that CSS — on every breakpoint +it declares and one pixel either side of it, plus a 320px floor, a desktop +baseline and two landscape sizes. See +[test/viewport/README.md](test/viewport/README.md) for what it covers and what +it genuinely cannot. ## Rationale When debugging network issues, it's useful to have a persistent at-a-glance view of latency and reachability to multiple well-known internet endpoints. NetWatch -provides this as a zero-dependency SPA that can be deployed anywhere static -files are served, with no backend required. +provides this as a single page that does all its measuring in the browser, so it +can be served from anywhere static files are served. The backend in its Docker +image only stores the measurements the page reports; without it, the page works +the same and nothing is stored. ## Design -The application is a single-page app built with Vite and Tailwind CSS v4. All -code lives in `src/main.js` with a class-based architecture: +The page is built with Vite and Tailwind CSS v4. Its code is all in +`src/main.js`, with a class-based architecture: - **`CONFIG`**: Configuration object (update interval, timeouts, axis ticks, etc.). The interval menu sets `updateInterval`, the one value the page writes @@ -125,10 +131,10 @@ code lives in `src/main.js` with a class-based architecture: `updateSummary()` / `updateHealthBox()` handle incremental updates - **`tick()`**: Main loop — measures all hosts in parallel, pushing each host's sample and redrawing its row as soon as its check ends, then redraws every - row, the summary and the health box once the last check ends. The rows are - sorted then too, after the first round that is not discarded and every tenth - round after that. When paused, pushes blank markers (no probes, no false - outage) + row, the summary and the health box once the last check ends. The first round, + after loading or an interval change, is discarded. The rows are sorted when + the last check ends in round 2, the first one kept, and in rounds 11, 21, 31 + and so on. When paused, pushes blank markers (no probes, no false outage) - **`Reporter`**: Posts collected samples to the backend ### Reporting @@ -142,12 +148,35 @@ delivered report, and while paused nothing is sent. Delivery failure is quiet one debug-log line per outage, retried at the next interval, never blocking probing. The report-building step is a pure function of host state. +### Backend + +`netwatch-server`, in `backend/`, is a small Go HTTP server that stores the +reports the page posts. It keeps them in memory and writes them to `DATA_DIR` as +zstd-compressed files of JSON lines: every minute, whenever 10 MiB are waiting, +and when it stops. Its routes: + +- `POST /api/v1/reports` — takes a report, without credentials; each client + address may send a limited number a minute, and the report files are capped in + size, the oldest deleted first +- `GET /.well-known/healthcheck` — answers 200 with `"status":"ok"`, the + server's version and its uptime +- `GET /metrics` — Prometheus metrics behind basic auth, only when + `METRICS_USERNAME` and `METRICS_PASSWORD` are set; each client address may + make a limited number of requests to it a minute + +In the image, the `builder` stage of `Dockerfile` tests it and builds it with +`backend/script/build`, and `bin/entrypoint.sh` runs it as user `netwatch` on +`127.0.0.1:8081`, behind nginx. Outside the image, `make run` in `backend/` +builds it and runs it on port 8080. Its settings, report storage and limits are +in [backend/README.md](backend/README.md). + ### Monitoring targets -- **22 WAN hosts**: datavi.be, Anthropic API, OpenAI API, AWS Console, GCP - Console, Azure, Cloudflare, Fastly, Akamai, GitHub, B2, 7 S3 regional - endpoints (Cape Town, London, Bahrain, Tokyo, Sydney, Oregon, São Paulo), 4 - GCS locational endpoints (Iowa, Belgium, Singapore, Sydney) +- **26 WAN hosts**: datavi.be (pinned at start), Anthropic API, OpenAI API, AWS + Console, Google Cloud Console, Microsoft Azure, Cloudflare, Fastly CDN, + Akamai, Google, GitHub, B2, 8 S3 regional endpoints (Cape Town, London, + Bahrain, Tokyo, Singapore, Sydney, Oregon, São Paulo) and 6 Hetzner speed test + servers (Nuremberg, Falkenstein, Helsinki, Ashburn, Hillsboro, Singapore) - **Local CPE**: Cable modem at 192.168.100.1 (always monitored) - **Local Gateway**: Auto-detected on startup by probing common default gateway addresses (192.168.1.1, 192.168.0.1, 192.168.8.1, 10.0.0.1); first responder @@ -159,15 +188,17 @@ Local hosts are tracked separately from WAN stats. ### Latency measurement -HEAD requests with `mode: 'no-cors'` and `cache: 'no-store'`, timed with -`performance.now()`. Each check times out after 80% of the refresh interval (24 -seconds at 30 seconds) and is then recorded as a timeout, so a round's checks -have all finished before the next round is due. When no WAN host answers, a -recovery probe checks 4 random WAN hosts every half second, giving up the checks -it started half a second before. As soon as one answers, a new round starts at -once, as it does after an interval change. A round started early gives up the -last round's checks if they are still waiting, and that round records nothing -more, so rounds never overlap. IPv4 only. +GET requests with `mode: 'no-cors'`, `cache: 'no-store'` and a cache-busting +query parameter, timed with `performance.now()`. Each check times out after 80% +of the refresh interval (24 seconds at 30 seconds) and is then recorded as a +timeout, so a round's checks have all finished before the next round is due. +When no WAN host answers, a recovery probe checks 4 WAN hosts, picked at random +when it starts, every half second, giving up the checks it started half a second +before. As soon as one answers, a new round starts at once, as it does after an +interval change. A round started early gives up the last round's checks if they +are still waiting, and that round records nothing more, so rounds never overlap. +The browser chooses between IPv4 and IPv6 for each target, as for any request; +the local targets are IPv4 addresses. ### Color coding @@ -192,25 +223,42 @@ dist/ ## Features -- Real-time monitoring with 2s update interval and 300s history sparklines -- Health indicator: green (HEALTHY) or red (DEGRADED) based on WAN reachability -- Summary stats: reachable count, min/max/avg latency across WAN hosts only -- Fixed chart axes: Y-axis 0–1000ms, X-axis 0–300s +- A round of checks every 3 seconds by default; the interval menu sets 1, 2, 3, + 5, 10, 15, 30 or 60 seconds and clears the history +- Sparklines of each target's last 100 rounds: 300 seconds at 3 seconds +- The first round after loading or an interval change is discarded, as DNS and + TLS setup inflate its latencies +- Health indicator from the WAN hosts' latest results: OFFLINE (red) when more + than 10 fail and at most 4 answer, otherwise DEGRADED (orange) when more than + 4 fail, otherwise SLOW (yellow) when more than 3 take over 1000ms, otherwise + HEALTHY (green) +- Summary stats across WAN hosts only: how many answered, the min, median, + average and max of their latest latencies, the min and max over the whole + history, and the number of rounds run (`Checks`) +- Fixed chart axes: Y-axis 0–1000ms, higher latencies drawn at the top; X-axis + the time the history spans - Color-coded latency figures and sparkline line segments +- WAN host rows sorted by latest latency, unreachable last; pinned rows stay on + top, in name order - Play/pause: pause stops probes but history keeps scrolling (blank gaps, no false outage) +- Debug log panel, behind a checkbox in the footer, with five levels (error, + warning, notice, info, debug) and the last 1000 lines +- Local and UTC clocks - Clickable service URLs +- A footer link to the commit the page was built from - Canvas-based sparkline rendering with devicePixelRatio scaling - Zero runtime dependencies: all resources bundled into build artifacts ## Deployment -After running `yarn build`, deploy the contents of the `dist/` directory to any -static file host (S3, GCS, Cloudflare Pages, Vercel, Netlify, GitHub Pages) or -use the Docker image behind a reverse proxy. +`make build` writes the page to `dist/`, which any static file host (S3, GCS, +Cloudflare Pages, Vercel, Netlify, GitHub Pages) can serve; with no backend +there, its reports fail quietly and nothing is stored. Or run the Docker image +behind a reverse proxy. -The Docker image, built from `Dockerfile`, is the whole service in one -container: nginx serves the built frontend and passes `/api/`, +The Docker image, built from `Dockerfile` by `make docker`, is the whole service +in one container: nginx serves the built frontend and passes `/api/`, `/.well-known/healthcheck` and `/metrics` to the Go backend, `netwatch-server`, which listens only inside the container, on `127.0.0.1:8081`. The image: @@ -283,19 +331,19 @@ properties. ## Limitations -- **CORS**: Some hosts may block cross-origin HEAD requests. The app uses - `no-cors` mode which allows the request but provides opaque responses. Latency - is still measurable based on request timing. -- **Local gateway**: The 192.168.100.1 endpoint requires the host to be - accessible from your network. +- **CORS**: The checks are cross-origin requests in `no-cors` mode, so the page + cannot read the answer, only time it: any answer counts as reachable, an error + page included. +- **Local targets**: The cable modem at 192.168.100.1 and the detected gateway + answer only on a network that has them, and only when NetWatch is served from + localhost or a private address (see Monitoring targets). - **Network conditions**: Measurements reflect browser-to-endpoint latency, which includes your local network, ISP, and internet routing. ## TODO -- Add configurable host list (environment variable or config file) -- Add latency history export (CSV/JSON) -- Add notification/alert when status changes to DEGRADED +The to-do list is [TODO.md](TODO.md): where the work stands, the next step, the +open work, and what has been done. ## License diff --git a/TODO.md b/TODO.md index 4a457e3..eea61eb 100644 --- a/TODO.md +++ b/TODO.md @@ -1,28 +1,38 @@ # Workflow -- branch (from `main`) +- branch from `next` - do the work in Next Step - move Next Step to the top of Completed Steps - move the top item of Future Steps into Next Step - commit (`TODO.md` changes in the same commit as the work) -- merge to `main` if the branch is not protected, otherwise open a PR -- push +- push the branch and open a PR against `next` # Status -pre-1.0. No git tags. `feat/reportbuf-storage` is merged; the backend, the CI -workflow, and the backend repo standard files are all on `main`. Frontend and -backend are both functional. Working toward the 1.0.0 milestone by closing the -remaining repo-compliance issues on the tracker. +pre-1.0. No git tags. `main` is the stable branch and `next` the development +branch, which every PR targets. The frontend and the Go backend ship as one +Docker image, and the Gitea workflow `.gitea/workflows/check.yml` runs +`script/cibuild` on every push. Working toward 1.0.0. # Next Step -Confirm the `.gitea/workflows/check.yml` run is green (main always green -policy). The workflow file is already on `main`; what is unverified is that its -latest run passes. +Take "IPv4 only" out of the page's footer, as nothing in the page limits a check +to IPv4 ([#111](https://git.eeqj.de/sneak/netwatch/issues/111)). # Completed Steps +- 2026-10-04: `README.md`, `TODO.md` and `test/viewport/README.md` say what the + tree does (issue #24). The README's Getting Started leads with `make` targets; + a new Backend section says what `netwatch-server` stores, its routes and how + the image builds and runs it, and points to `backend/README.md` for its + settings; the checks are GET requests; the 26 WAN hosts, the four health + states, the summary's figures and the features the list lacked are described + as the page has them; and its TODO section points here, as does the one in + `backend/README.md`, whose open items moved to Future Steps. This file's + Workflow branches from `next` and opens the PR against `next`, Status says + where the repo stands, and Next Step and Future Steps hold only open work, + linked to its issue where one exists. The viewport harness README names Node's + test runner, not `vitest` - 2026-10-04: in `src/main.js` (issue #102), a target's min, max, median and average latency come from one list of its answers, through the same function the summary's figures use, so the median is written once. The latency color @@ -374,12 +384,17 @@ latest run passes. # Future Steps -- Wire `script/frontend-viewport-test` into CI as its own step (deliberately not - part of `make check` today; the decision has real CI-runtime cost and is - tracked separately) -- Compliance top-up as one small commit: add .editorconfig and add the hooks - target to the Makefile -- After merge, confirm .gitea/workflows/check.yml is on main and CI is green - (main always green policy) -- Decide what to do with untracked resume.sh: commit it, gitignore it, or delete - it +- Decide whether the repo moves to the layout `REPO_POLICIES.md` gives, with + `backend/` no longer repeating files from the root + ([#30](https://git.eeqj.de/sneak/netwatch/issues/30)) +- Run `make frontend-viewport-test` in CI as its own step; it is not part of + `make check`, as it needs Docker and takes minutes +- A backend test that posts a report to `POST /api/v1/reports` and checks the + compressed file it is written to +- A backend route that decompresses the stored reports and answers queries on + them +- Prometheus metrics for the backend's in-memory buffer: its size, the number of + flushes and the number of reports +- A configurable host list (an environment variable or a config file) +- Export of the latency history (CSV or JSON) +- A notification when the health status changes to DEGRADED diff --git a/backend/README.md b/backend/README.md index 0cbefe9..47f74ee 100644 --- a/backend/README.md +++ b/backend/README.md @@ -219,9 +219,8 @@ sent to it. ## TODO -- Add integration test that POSTs a report and verifies the compressed output -- Add report decompression/query endpoint -- Add metrics (Prometheus) for buffer size, flush count, report count +The to-do list, this backend's open work included, is [TODO.md](../TODO.md) at +the repo root. ## License diff --git a/test/viewport/README.md b/test/viewport/README.md index 9ee28a5..8e89805 100644 --- a/test/viewport/README.md +++ b/test/viewport/README.md @@ -77,7 +77,7 @@ the internet, so the app's latency probes cannot reach anything real. The harness answers them itself from a fixed delay table, with a deterministic fraction failed outright, so the rows render a realistic spread of one-, two- and three-digit latencies plus some unreachable rows. That spread is what the -layout has to survive; 24 identical `---` placeholders would not exercise it. +layout has to survive; a `---` placeholder in every row would not exercise it. ## What this cannot verify @@ -104,10 +104,11 @@ Everything else this issue was actually about — does the layout reflow, does anything overflow, is content clipped, are the controls big enough — is a function of viewport width and CSS, and is covered above. -## Relation to the unit test framework (#21) +## Relation to the unit tests -Complementary layers, not two stacks. `vitest` (#21) will exercise module-level -logic in-process with no browser. This harness exercises rendered layout in a -real engine and is the only thing here that can see a media query. Neither -replaces the other; assertions about computed styles and element geometry belong -here, assertions about functions belong in `vitest`. +Complementary layers, not two stacks. The unit tests in `test/unit/`, which +`make test` runs with Node's built-in test runner, exercise the functions +`src/main.js` exports in-process with no browser. This harness exercises +rendered layout in a real engine and is the only thing here that can see a media +query. Neither replaces the other; assertions about computed styles and element +geometry belong here, assertions about functions belong in `test/unit/`. -- 2.54.0