check / check (push) Successful in 2m42s
README.md: Getting Started leads with make targets; a Backend section gives netwatch-server's routes and how the image builds and runs it; the checks are GET requests; the WAN host list, health states, summary figures, sorting and missing features match src/main.js; the TODO section points to TODO.md, which holds the one to-do list. backend/README.md: its TODO section points to TODO.md too, whose Future Steps take its three open items. TODO.md: Workflow branches from next and opens the PR against next; Status, Next Step and Future Steps describe the open work, linked to its issue where one exists. test/viewport/README.md: the unit tests run on Node's test runner, not vitest. Model: opus-5-5
355 lines
18 KiB
Markdown
355 lines
18 KiB
Markdown
NetWatch is an MIT-licensed JavaScript single-page application by
|
||
[@sneak](https://sneak.berlin) that provides real-time network latency
|
||
monitoring to common internet hosts, displayed with color-coded figures and
|
||
sparkline graphs, served from a static bucket or from its Docker image, where a
|
||
small Go backend stores the measurements the page reports.
|
||
|
||
## Getting Started
|
||
|
||
```bash
|
||
# Install the dependencies and the git pre-commit hook
|
||
make setup
|
||
|
||
# Run the page on the Vite dev server
|
||
make dev
|
||
|
||
# Run the tests, both linters and the format check
|
||
make check
|
||
|
||
# Build the page into dist/
|
||
make build
|
||
|
||
# Build the image and run it
|
||
make docker
|
||
docker run -p 8080:8080 netwatch
|
||
```
|
||
|
||
`make check` and `make docker` need Docker. `make dev` passes `/api` to
|
||
`http://127.0.0.1:8080`, where `make run` in `backend/` starts `netwatch-server`
|
||
with its defaults, so the reports the page posts are stored in
|
||
`backend/data/reports`.
|
||
|
||
## Entrypoints
|
||
|
||
This repository adheres to the
|
||
[Scripts to Rule Them All](https://github.com/github/scripts-to-rule-them-all)
|
||
standard: normalized scripts in `script/` are the entrypoints for the
|
||
development workflow, and the Makefile targets are thin shims that call them.
|
||
The Go backend in `backend/` has its own `script/` directory and shim Makefile
|
||
(see [backend/README.md](backend/README.md)). The root scripts cover both
|
||
halves, so the root `make check` fails if either one is broken. We provide:
|
||
|
||
- `script/bootstrap` — install all dependencies (the pinned node via nvm unless
|
||
one new enough for the frontend's dependencies is installed, yarn via
|
||
corepack, `yarn install --frozen-lockfile`, the pinned Go unless one at least
|
||
as new as `backend/go.mod` asks for is installed, the Go modules, and gcc with
|
||
the C library headers unless gcc is installed, for the race detector in
|
||
`make test`), linking what it installs itself into `~/.local/bin`, which has
|
||
to be on `PATH`. It installs no Go linter and not Docker: `make lint` runs
|
||
both linters in Docker
|
||
- `script/setup` — make a fresh clone ready for development: bootstrap plus the
|
||
git pre-commit hook
|
||
- `script/dev` — run the Vite dev server, which proxies `/api` to a locally
|
||
running `netwatch-server`
|
||
- `script/build` — build the frontend for production into `dist/`;
|
||
`backend/script/build` builds the Go server
|
||
- `script/projectname` — print the project name (used for the Docker image tag)
|
||
- `script/test` — run `script/frontend-test`, then `backend/script/test`, the
|
||
backend's Go tests with the race detector and coverage
|
||
- `script/lint` — run eslint, then golangci-lint, both in Docker, by building
|
||
the `frontend-lint` and `lint` stages of `Dockerfile` without the cache
|
||
- `script/fmt` — format all files (writes): prettier, then gofmt over `backend/`
|
||
- `script/fmt-check` — check formatting (read-only): prettier, then gofmt
|
||
- `script/check` — run test, lint, and fmt-check
|
||
- `script/add-dependency` — add a frontend package, or move one to another
|
||
version: `make add-dependency PACKAGE=<name>@<version>` runs `yarn add --dev`,
|
||
which changes `package.json` and `yarn.lock` together, then
|
||
`yarn install --frozen-lockfile`
|
||
- `script/tidy` — run `go mod tidy` in `backend/`: to add a Go module, import it
|
||
and run `make tidy`; to move one to another version, edit its `require` line
|
||
in `backend/go.mod`, then run `make tidy`
|
||
- `script/frontend-test` — run the unit tests in `test/unit/` with Node's
|
||
built-in test runner, through the `test` script in `package.json`, and if any
|
||
fails, run them again listing every test, and fail; then the production build.
|
||
Each run has a 30-second timeout
|
||
- `script/frontend-lint` — run eslint with the rules in `eslint.config.js`; it
|
||
runs inside the `frontend-lint` stage of `Dockerfile`, which `make lint`
|
||
builds
|
||
- `script/frontend-fmt` — format everything prettier understands (writes), the
|
||
markdown in `backend/` included
|
||
- `script/frontend-fmt-check` — check prettier formatting (read-only)
|
||
- `script/frontend-check` — run `script/frontend-test` and
|
||
`script/frontend-fmt-check`, for the frontend stage of `Dockerfile`, which has
|
||
neither Go nor Docker
|
||
- `script/frontend-viewport-test` — responsive-layout verification of the built
|
||
frontend in a containerised headless Chrome (see
|
||
[test/viewport/README.md](test/viewport/README.md)). Not part of
|
||
`script/check`: it needs Docker and takes minutes.
|
||
- `script/docker` — build the image from `Dockerfile` without the build cache,
|
||
tagged `netwatch` via `script/projectname`
|
||
- `script/cibuild` — CI entrypoint: runs `script/bootstrap` and `script/check`,
|
||
then builds the image as `script/docker` does, without the build cache
|
||
- `script/precommit` — run by the git pre-commit hook; runs `script/check`
|
||
- `script/install-precommit` — install the git pre-commit hook
|
||
|
||
## Responsive layout
|
||
|
||
The narrow-viewport layout lives in the `max-width: 768px` media block in
|
||
`src/styles.css`. It is verified automatically by `make frontend-viewport-test`,
|
||
which drives a digest-pinned headless Chrome against the built `dist/` and
|
||
asserts on computed layout at widths derived from that CSS — on every breakpoint
|
||
it declares and one pixel either side of it, plus a 320px floor, a desktop
|
||
baseline and two landscape sizes. See
|
||
[test/viewport/README.md](test/viewport/README.md) for what it covers and what
|
||
it genuinely cannot.
|
||
|
||
## Rationale
|
||
|
||
When debugging network issues, it's useful to have a persistent at-a-glance view
|
||
of latency and reachability to multiple well-known internet endpoints. NetWatch
|
||
provides this as a single page that does all its measuring in the browser, so it
|
||
can be served from anywhere static files are served. The backend in its Docker
|
||
image only stores the measurements the page reports; without it, the page works
|
||
the same and nothing is stored.
|
||
|
||
## Design
|
||
|
||
The page is built with Vite and Tailwind CSS v4. Its code is all in
|
||
`src/main.js`, with a class-based architecture:
|
||
|
||
- **`CONFIG`**: Configuration object (update interval, timeouts, axis ticks,
|
||
etc.). The interval menu sets `updateInterval`, the one value the page writes
|
||
into `CONFIG`; the timeouts, the time the history spans and the x-axis ticks
|
||
are computed from it
|
||
- **`HostState`**: Per-host state management — history buffer, latency tracking,
|
||
status transitions
|
||
- **`AppState`**: Top-level state container — WAN hosts, local hosts, pause
|
||
state, aggregate stats
|
||
- **`SparklineRenderer`**: Canvas 2D sparkline drawing with fixed axes,
|
||
color-coded line segments, error regions, and DPR-aware scaling
|
||
- **UI functions**: `buildUI()` constructs the DOM, `updateHostRow()` /
|
||
`updateSummary()` / `updateHealthBox()` handle incremental updates
|
||
- **`tick()`**: Main loop — measures all hosts in parallel, pushing each host's
|
||
sample and redrawing its row as soon as its check ends, then redraws every
|
||
row, the summary and the health box once the last check ends. The first round,
|
||
after loading or an interval change, is discarded. The rows are sorted when
|
||
the last check ends in round 2, the first one kept, and in rounds 11, 21, 31
|
||
and so on. When paused, pushes blank markers (no probes, no false outage)
|
||
- **`Reporter`**: Posts collected samples to the backend
|
||
|
||
### Reporting
|
||
|
||
Every `reportInterval` (default 60s) the page POSTs a JSON report to the
|
||
same-origin path `/api/v1/reports`: a random per-browser `clientId` kept in
|
||
`localStorage`, `geo` sent as null, and each host's unreported, non-paused
|
||
samples (timestamp, latency, error). A per-host high-water mark makes every
|
||
report a delta, so only new samples are sent; the mark advances only on a
|
||
delivered report, and while paused nothing is sent. Delivery failure is quiet —
|
||
one debug-log line per outage, retried at the next interval, never blocking
|
||
probing. The report-building step is a pure function of host state.
|
||
|
||
### Backend
|
||
|
||
`netwatch-server`, in `backend/`, is a small Go HTTP server that stores the
|
||
reports the page posts. It keeps them in memory and writes them to `DATA_DIR` as
|
||
zstd-compressed files of JSON lines: every minute, whenever 10 MiB are waiting,
|
||
and when it stops. Its routes:
|
||
|
||
- `POST /api/v1/reports` — takes a report, without credentials; each client
|
||
address may send a limited number a minute, and the report files are capped in
|
||
size, the oldest deleted first
|
||
- `GET /.well-known/healthcheck` — answers 200 with `"status":"ok"`, the
|
||
server's version and its uptime
|
||
- `GET /metrics` — Prometheus metrics behind basic auth, only when
|
||
`METRICS_USERNAME` and `METRICS_PASSWORD` are set; each client address may
|
||
make a limited number of requests to it a minute
|
||
|
||
In the image, the `builder` stage of `Dockerfile` tests it and builds it with
|
||
`backend/script/build`, and `bin/entrypoint.sh` runs it as user `netwatch` on
|
||
`127.0.0.1:8081`, behind nginx. Outside the image, `make run` in `backend/`
|
||
builds it and runs it on port 8080. Its settings, report storage and limits are
|
||
in [backend/README.md](backend/README.md).
|
||
|
||
### Monitoring targets
|
||
|
||
- **26 WAN hosts**: datavi.be (pinned at start), Anthropic API, OpenAI API, AWS
|
||
Console, Google Cloud Console, Microsoft Azure, Cloudflare, Fastly CDN,
|
||
Akamai, Google, GitHub, B2, 8 S3 regional endpoints (Cape Town, London,
|
||
Bahrain, Tokyo, Singapore, Sydney, Oregon, São Paulo) and 6 Hetzner speed test
|
||
servers (Nuremberg, Falkenstein, Helsinki, Ashburn, Hillsboro, Singapore)
|
||
- **Local CPE**: Cable modem at 192.168.100.1 (always monitored)
|
||
- **Local Gateway**: Auto-detected on startup by probing common default gateway
|
||
addresses (192.168.1.1, 192.168.0.1, 192.168.8.1, 10.0.0.1); first responder
|
||
wins. Note: modern browsers enforce Private Network Access restrictions that
|
||
block public-origin pages from reaching RFC1918 addresses, so local targets
|
||
only work when NetWatch is served from localhost or a private address.
|
||
|
||
Local hosts are tracked separately from WAN stats.
|
||
|
||
### Latency measurement
|
||
|
||
GET requests with `mode: 'no-cors'`, `cache: 'no-store'` and a cache-busting
|
||
query parameter, timed with `performance.now()`. Each check times out after 80%
|
||
of the refresh interval (24 seconds at 30 seconds) and is then recorded as a
|
||
timeout, so a round's checks have all finished before the next round is due.
|
||
When no WAN host answers, a recovery probe checks 4 WAN hosts, picked at random
|
||
when it starts, every half second, giving up the checks it started half a second
|
||
before. As soon as one answers, a new round starts at once, as it does after an
|
||
interval change. A round started early gives up the last round's checks if they
|
||
are still waiting, and that round records nothing more, so rounds never overlap.
|
||
The browser chooses between IPv4 and IPv6 for each target, as for any request;
|
||
the local targets are IPv4 addresses.
|
||
|
||
### Color coding
|
||
|
||
| Latency | Color |
|
||
| ----------- | ------ |
|
||
| < 50ms | Green |
|
||
| < 100ms | Lime |
|
||
| < 200ms | Yellow |
|
||
| < 500ms | Orange |
|
||
| >= 500ms | Red |
|
||
| Unreachable | Gray |
|
||
|
||
### Output structure
|
||
|
||
```
|
||
dist/
|
||
├── index.html
|
||
└── assets/
|
||
├── index-*.css
|
||
└── index-*.js
|
||
```
|
||
|
||
## Features
|
||
|
||
- A round of checks every 3 seconds by default; the interval menu sets 1, 2, 3,
|
||
5, 10, 15, 30 or 60 seconds and clears the history
|
||
- Sparklines of each target's last 100 rounds: 300 seconds at 3 seconds
|
||
- The first round after loading or an interval change is discarded, as DNS and
|
||
TLS setup inflate its latencies
|
||
- Health indicator from the WAN hosts' latest results: OFFLINE (red) when more
|
||
than 10 fail and at most 4 answer, otherwise DEGRADED (orange) when more than
|
||
4 fail, otherwise SLOW (yellow) when more than 3 take over 1000ms, otherwise
|
||
HEALTHY (green)
|
||
- Summary stats across WAN hosts only: how many answered, the min, median,
|
||
average and max of their latest latencies, the min and max over the whole
|
||
history, and the number of rounds run (`Checks`)
|
||
- Fixed chart axes: Y-axis 0–1000ms, higher latencies drawn at the top; X-axis
|
||
the time the history spans
|
||
- Color-coded latency figures and sparkline line segments
|
||
- WAN host rows sorted by latest latency, unreachable last; pinned rows stay on
|
||
top, in name order
|
||
- Play/pause: pause stops probes but history keeps scrolling (blank gaps, no
|
||
false outage)
|
||
- Debug log panel, behind a checkbox in the footer, with five levels (error,
|
||
warning, notice, info, debug) and the last 1000 lines
|
||
- Local and UTC clocks
|
||
- Clickable service URLs
|
||
- A footer link to the commit the page was built from
|
||
- Canvas-based sparkline rendering with devicePixelRatio scaling
|
||
- Zero runtime dependencies: all resources bundled into build artifacts
|
||
|
||
## Deployment
|
||
|
||
`make build` writes the page to `dist/`, which any static file host (S3, GCS,
|
||
Cloudflare Pages, Vercel, Netlify, GitHub Pages) can serve; with no backend
|
||
there, its reports fail quietly and nothing is stored. Or run the Docker image
|
||
behind a reverse proxy.
|
||
|
||
The Docker image, built from `Dockerfile` by `make docker`, is the whole service
|
||
in one container: nginx serves the built frontend and passes `/api/`,
|
||
`/.well-known/healthcheck` and `/metrics` to the Go backend, `netwatch-server`,
|
||
which listens only inside the container, on `127.0.0.1:8081`. The image:
|
||
|
||
- Listens on port 8080 by default (override with `PORT` env var)
|
||
- Takes the client address from `X-Forwarded-For` only on requests from the
|
||
reverse proxies named in `TRUSTED_PROXIES`, and by default from none
|
||
- Sends access logs to stdout
|
||
- Caches static assets with immutable headers
|
||
- Sends the security headers `REPO_POLICIES.md` requires on every response, as
|
||
`security-headers.conf` sets them, in place of the backend's own
|
||
- Stores reports in `DATA_DIR`, `/data/reports` by default, on the `/data`
|
||
volume. Before the backend starts, the image creates `DATA_DIR` and gives it
|
||
and `/data` to user `netwatch` (uid 1000), which the backend runs as, so a
|
||
host directory bind-mounted at `/data` ends up owned by uid 1000
|
||
- Writes buffered reports to disk on `docker stop`, and exits non-zero if nginx
|
||
or the backend exits on its own, so the platform restarts it
|
||
|
||
## Running under upaas
|
||
|
||
What the [upaas](https://git.eeqj.de/sneak/upaas) app for netwatch needs:
|
||
|
||
- **Port:** container port `8080`.
|
||
- **Volume:** container path `/data`; the reports are kept in `/data/reports`.
|
||
- **Environment variables:** none is required. An empty one counts as unset, and
|
||
one set to a value netwatch cannot use stops the container at start, with the
|
||
reason in its log.
|
||
- `PORT`, default `8080`: the container port, from 1 to 65535. `8081` cannot
|
||
be used: the backend listens on it inside the container
|
||
- `REPORTS_PER_MINUTE`, default `60`: reports each client address may send a
|
||
minute
|
||
- `DATA_DIR_MAX_BYTES`, default `1073741824` (1 GiB): the most room the
|
||
report files may take; the oldest are deleted to stay under it
|
||
- `CORS_ALLOWED_ORIGINS`, default empty: other origins whose pages may call
|
||
the API
|
||
- `DEBUG`, default `false`: debug logging
|
||
- `DATA_DIR`, default `/data/reports`: the directory the reports are kept
|
||
in: `/data` or a path below it, with no `.` or `..` part and no extra `/`.
|
||
The container also stops if the path goes through a symbolic link that
|
||
leads out of `/data` or is written as a full path
|
||
- `TRUSTED_PROXIES`, default empty: set it to the address the reverse proxy
|
||
in front of the container connects from, as an IP address or CIDR; several
|
||
are separated by commas. nginx takes the client address from
|
||
`X-Forwarded-For` only on a request from one of them, and the rate limit
|
||
counts that address. Unset, `X-Forwarded-For` is ignored and every client
|
||
behind the proxy shares the proxy's one allowance of `REPORTS_PER_MINUTE`.
|
||
Name only addresses nothing but the proxy connects from: any client that
|
||
connects from one can write its own `X-Forwarded-For`, and through a port
|
||
Docker publishes, every client may connect from the Docker network's
|
||
gateway, such as `172.17.0.1`.
|
||
- `METRICS_USERNAME` and `METRICS_PASSWORD`, default empty: with both set,
|
||
the backend records Prometheus metrics of its requests and serves them at
|
||
`/metrics` on the container port, to requests with this user name and
|
||
password as their basic auth credentials. With neither set, there are no
|
||
metrics and `/metrics` is not found. One set without the other, or a user
|
||
name containing `:`, stops the container
|
||
- `SENTRY_DSN`, default empty: set to a Sentry project's DSN, the backend
|
||
sends its errors to that Sentry project: each request whose handling
|
||
crashes, which still gets a 500 response. A value Sentry does not accept
|
||
stops the container. Empty, the backend sends nothing to Sentry
|
||
- **Health check:** the image's `HEALTHCHECK` requests
|
||
`/.well-known/healthcheck` through nginx every 30 seconds, so it fails unless
|
||
both nginx and the backend answer. upaas reads the container's health 60
|
||
seconds after a deploy and fails the deploy unless it is `healthy`. The
|
||
container also stops when either process exits.
|
||
|
||
## Browser Compatibility
|
||
|
||
Requires a modern browser with ES modules, Fetch API, Canvas API, and CSS custom
|
||
properties.
|
||
|
||
## Limitations
|
||
|
||
- **CORS**: The checks are cross-origin requests in `no-cors` mode, so the page
|
||
cannot read the answer, only time it: any answer counts as reachable, an error
|
||
page included.
|
||
- **Local targets**: The cable modem at 192.168.100.1 and the detected gateway
|
||
answer only on a network that has them, and only when NetWatch is served from
|
||
localhost or a private address (see Monitoring targets).
|
||
- **Network conditions**: Measurements reflect browser-to-endpoint latency,
|
||
which includes your local network, ISP, and internet routing.
|
||
|
||
## TODO
|
||
|
||
The to-do list is [TODO.md](TODO.md): where the work stands, the next step, the
|
||
open work, and what has been done.
|
||
|
||
## License
|
||
|
||
MIT. See [LICENSE](LICENSE).
|
||
|
||
## Author
|
||
|
||
[@sneak](https://sneak.berlin)
|