check / check (push) Waiting to run
The shared files are the sneak/prompts copies at dd4027b, plus this repository's own entries. make lint and make test each build one Dockerfile phase without the cache, both covering the frontend; the builder stage waits on both and takes its version from git describe unless VERSION is given. The test phase keeps Go's module and build caches in memory, out of the image make test tags. golangci-lint moves to v2.14.0 with the new .golangci.yml; one test spells X-Request-ID as canonicalheader asks. prettier formats only JavaScript, CSS, HTML and Markdown, so .golangci.yml stays as fetched. script/fmt and script/fmt-check put ~/.local/bin on PATH. script/bootstrap keeps a Go only if it is exactly GO_VERSION, and re-checks the go on PATH after installing. Model: opus-5-5
364 lines
18 KiB
Markdown
364 lines
18 KiB
Markdown
NetWatch is an MIT-licensed JavaScript single-page application by
|
||
[@sneak](https://sneak.berlin) that provides real-time network latency
|
||
monitoring to common internet hosts, displayed with color-coded figures and
|
||
sparkline graphs, served from a static bucket or from its Docker image, where a
|
||
small Go backend stores the measurements the page reports.
|
||
|
||
## Getting Started
|
||
|
||
```bash
|
||
# Install the dependencies and the git pre-commit hook
|
||
make setup
|
||
|
||
# Run the page on the Vite dev server
|
||
make dev
|
||
|
||
# Run the tests, both linters and the format check
|
||
make check
|
||
|
||
# Build the page into dist/
|
||
make build
|
||
|
||
# Build the image and run it
|
||
make docker
|
||
docker run -p 8080:8080 netwatch
|
||
```
|
||
|
||
`make check` and `make docker` need Docker. `make dev` passes `/api` to
|
||
`http://127.0.0.1:8080`, where `make run` in `backend/` starts `netwatch-server`
|
||
with its defaults, so the reports the page posts are stored in
|
||
`backend/data/reports`.
|
||
|
||
## Entrypoints
|
||
|
||
This repository adheres to the
|
||
[Scripts to Rule Them All](https://github.com/github/scripts-to-rule-them-all)
|
||
standard: normalized scripts in `script/` are the entrypoints for the
|
||
development workflow, and the Makefile targets are thin shims that call them.
|
||
The Go backend in `backend/` has its own `script/` directory and shim Makefile
|
||
(see [backend/README.md](backend/README.md)). The root scripts cover both
|
||
halves, so the root `make check` fails if either one is broken. We provide:
|
||
|
||
- `script/bootstrap` — install all dependencies (the pinned node via nvm unless
|
||
one new enough for the frontend's dependencies is installed, yarn via
|
||
corepack, `yarn install --frozen-lockfile`, Go `GO_VERSION` unless the `go` on
|
||
`PATH` is exactly that version, the Go modules, and gcc with the C library
|
||
headers unless gcc is installed, for the race detector in `make test` in
|
||
`backend/`), linking what it installs itself into `~/.local/bin`, which has to
|
||
be on `PATH`. It installs no Go linter and not Docker: `make test` and
|
||
`make lint` run in Docker
|
||
- `script/setup` — make a fresh clone ready for development: bootstrap plus the
|
||
git pre-commit hook
|
||
- `script/dev` — run the Vite dev server, which proxies `/api` to a locally
|
||
running `netwatch-server`
|
||
- `script/build` — build the frontend for production into `dist/`;
|
||
`backend/script/build` builds the Go server
|
||
- `script/projectname` — print the project name (used for the Docker image tag)
|
||
- `script/test` — build the `test` phase of `Dockerfile` without the cache: the
|
||
frontend's unit tests and production build in its `frontend` stage, and the
|
||
backend's Go tests with the race detector and coverage
|
||
- `script/lint` — build the `lint` phase of `Dockerfile` without the cache:
|
||
eslint in its `frontend-lint` stage, and golangci-lint over `backend/`
|
||
- `script/fmt` — format all files (writes): prettier over the JavaScript, CSS,
|
||
HTML and Markdown, then gofmt over `backend/`. It runs on the host, as
|
||
`script/fmt-check` does, with `~/.local/bin` put on `PATH`
|
||
- `script/fmt-check` — check formatting (read-only): prettier, then gofmt
|
||
- `script/check` — run test, lint, and fmt-check
|
||
- `script/add-dependency` — add a frontend package, or move one to another
|
||
version: `make add-dependency PACKAGE=<name>@<version>` runs `yarn add --dev`,
|
||
which changes `package.json` and `yarn.lock` together, then
|
||
`yarn install --frozen-lockfile`
|
||
- `script/tidy` — run `go mod tidy` in `backend/`: to add a Go module, import it
|
||
and run `make tidy`; to move one to another version, edit its `require` line
|
||
in `backend/go.mod`, then run `make tidy`
|
||
- `script/frontend-test` — run the unit tests in `test/unit/` on the host with
|
||
Node's built-in test runner, through the `test` script in `package.json`, and
|
||
if any fails, run them again listing every test, and fail; then the production
|
||
build. Each test run has a 90-second timeout. `make test` runs the same in
|
||
Docker
|
||
- `script/frontend-fmt` — format the JavaScript, CSS, HTML and Markdown with
|
||
prettier (writes), the markdown in `backend/` included
|
||
- `script/frontend-fmt-check` — check prettier formatting (read-only)
|
||
- `script/frontend-viewport-test` — responsive-layout verification of the built
|
||
frontend in a containerised headless Chrome (see
|
||
[test/viewport/README.md](test/viewport/README.md)). Not part of
|
||
`script/check`: it takes minutes.
|
||
- `script/docker` — build the image from `Dockerfile` without the build cache,
|
||
tagged `netwatch` via `script/projectname`
|
||
- `script/cibuild` — CI entrypoint: runs `script/bootstrap` and `script/check`,
|
||
then builds the image as `script/docker` does, without the build cache
|
||
- `script/precommit` — run by the git pre-commit hook; runs `script/check`
|
||
- `script/install-precommit` — install the git pre-commit hook
|
||
|
||
## Responsive layout
|
||
|
||
The narrow-viewport layout lives in the `max-width: 768px` media block in
|
||
`src/styles.css`. It is verified automatically by `make frontend-viewport-test`,
|
||
which drives a digest-pinned headless Chrome against the built `dist/` and
|
||
asserts on computed layout at widths derived from that CSS — on every breakpoint
|
||
it declares and one pixel either side of it, plus a 320px floor, a desktop
|
||
baseline and two landscape sizes. See
|
||
[test/viewport/README.md](test/viewport/README.md) for what it covers and what
|
||
it genuinely cannot.
|
||
|
||
## Rationale
|
||
|
||
When debugging network issues, it's useful to have a persistent at-a-glance view
|
||
of latency and reachability to multiple well-known internet endpoints. NetWatch
|
||
provides this as a single page that does all its measuring in the browser, so it
|
||
can be served from anywhere static files are served. The backend in its Docker
|
||
image only stores the measurements the page reports; without it, the page works
|
||
the same and nothing is stored.
|
||
|
||
## Design
|
||
|
||
The page is built with Vite and Tailwind CSS v4. Its code is all in
|
||
`src/main.js`, with a class-based architecture:
|
||
|
||
- **`CONFIG`**: Configuration object (update interval, timeouts, axis ticks,
|
||
etc.). The interval menu sets `updateInterval`, the one value the page writes
|
||
into `CONFIG`; the timeouts, the time the history spans and the x-axis ticks
|
||
are computed from it
|
||
- **`HostState`**: Per-host state management — history buffer, latency tracking,
|
||
status transitions
|
||
- **`AppState`**: Top-level state container — WAN hosts, local hosts, pause
|
||
state, aggregate stats
|
||
- **`SparklineRenderer`**: Canvas 2D sparkline drawing with fixed axes,
|
||
color-coded line segments, error regions, and DPR-aware scaling
|
||
- **UI functions**: `buildUI()` constructs the DOM, `updateHostRow()` /
|
||
`updateSummary()` / `updateHealthBox()` handle incremental updates
|
||
- **`tick()`**: Main loop — measures all hosts in parallel, pushing each host's
|
||
sample and redrawing its row as soon as its check ends, then redraws every
|
||
row, the summary and the health box once the last check ends. The first round,
|
||
after loading or an interval change, is discarded. The rows are sorted when
|
||
the last check ends in round 2, the first one kept, and in rounds 11, 21, 31
|
||
and so on. When paused, pushes blank markers (no probes, no false outage)
|
||
- **`Reporter`**: Posts collected samples to the backend
|
||
|
||
### Reporting
|
||
|
||
Every `reportInterval` (default 60s) the page POSTs a JSON report to the
|
||
same-origin path `/api/v1/reports`: a random per-browser `clientId` kept in
|
||
`localStorage`, `geo` sent as null, and each host's unreported, non-paused
|
||
samples (timestamp, latency, error). A per-host high-water mark makes every
|
||
report a delta, so only new samples are sent; the mark advances only on a
|
||
delivered report, and while paused nothing is sent. Delivery failure is quiet —
|
||
one debug-log line per outage, retried at the next interval, never blocking
|
||
probing. The report-building step is a pure function of host state.
|
||
|
||
### Backend
|
||
|
||
`netwatch-server`, in `backend/`, is a small Go HTTP server that stores the
|
||
reports the page posts. It keeps them in memory and writes them to `DATA_DIR` as
|
||
zstd-compressed files of JSON lines: every minute, whenever 10 MiB are waiting,
|
||
and when it stops. Its routes:
|
||
|
||
- `POST /api/v1/reports` — takes a report, without credentials; each client
|
||
address may send a limited number a minute, and the report files are capped in
|
||
size, the oldest deleted first
|
||
- `GET /.well-known/healthcheck` — answers 200 with `"status":"ok"`, the
|
||
server's version and its uptime
|
||
- `GET /metrics` — Prometheus metrics behind basic auth, only when
|
||
`METRICS_USERNAME` and `METRICS_PASSWORD` are set; each client address may
|
||
make a limited number of requests to it a minute
|
||
|
||
In the image, the `test` phase of `Dockerfile` tests it, the `builder` stage
|
||
builds it with `backend/script/build`, and `bin/entrypoint.sh` runs it as user
|
||
`netwatch` on `127.0.0.1:8081`, behind nginx. Outside the image, `make run` in
|
||
`backend/` builds it and runs it on port 8080. Its settings, report storage and
|
||
limits are in [backend/README.md](backend/README.md).
|
||
|
||
### Monitoring targets
|
||
|
||
- **26 WAN hosts**: datavi.be (pinned at start), Anthropic API, OpenAI API, AWS
|
||
Console, Google Cloud Console, Microsoft Azure, Cloudflare, Fastly CDN,
|
||
Akamai, Google, GitHub, B2, 8 S3 regional endpoints (Cape Town, London,
|
||
Bahrain, Tokyo, Singapore, Sydney, Oregon, São Paulo) and 6 Hetzner speed test
|
||
servers (Nuremberg, Falkenstein, Helsinki, Ashburn, Hillsboro, Singapore)
|
||
- **Local CPE**: Cable modem at 192.168.100.1 (always monitored)
|
||
- **Local Gateway**: Auto-detected on startup by probing common default gateway
|
||
addresses (192.168.1.1, 192.168.0.1, 192.168.8.1, 10.0.0.1); first responder
|
||
wins. Note: modern browsers enforce Private Network Access restrictions that
|
||
block public-origin pages from reaching RFC1918 addresses, so local targets
|
||
only work when NetWatch is served from localhost or a private address.
|
||
|
||
Local hosts are tracked separately from WAN stats.
|
||
|
||
### Latency measurement
|
||
|
||
GET requests to each target's URL as written, with `mode: 'no-cors'` and
|
||
`cache: 'no-store'`, timed with `performance.now()`. No query string is added:
|
||
the Hetzner speed-test servers close the connection without an answer when the
|
||
URL has one, and `no-store` keeps the browser's cache out of the measurement.
|
||
Each check times out after 80% of the refresh interval (24 seconds at 30
|
||
seconds) and is then recorded as a timeout, so a round's checks have all
|
||
finished before the next round is due. When no WAN host answers, a recovery
|
||
probe checks 4 WAN hosts, picked at random when it starts, every half second,
|
||
giving up the checks it started half a second before. As soon as one answers, a
|
||
new round starts at once, as it does after an interval change. A round started
|
||
early gives up the last round's checks if they are still waiting, and that round
|
||
records nothing more, so rounds never overlap. The browser chooses between IPv4
|
||
and IPv6 for each target, as for any request; the local targets are IPv4
|
||
addresses.
|
||
|
||
Each recorded check that fails writes one line to the browser console with
|
||
`console.error`, and the same line to the debug log: the target's name and URL,
|
||
the time, whether it timed out, answered over the time limit or hit a network
|
||
error (with the error the browser gives the page, such as
|
||
`TypeError: Failed to fetch`, which does not say why), and how long the request
|
||
took. A target that answers after failed checks writes one `console.info` line
|
||
with its latency and how many checks in a row had failed.
|
||
|
||
### Color coding
|
||
|
||
| Latency | Color |
|
||
| ----------- | ------ |
|
||
| < 50ms | Green |
|
||
| < 100ms | Lime |
|
||
| < 200ms | Yellow |
|
||
| < 500ms | Orange |
|
||
| >= 500ms | Red |
|
||
| Unreachable | Gray |
|
||
|
||
### Output structure
|
||
|
||
```
|
||
dist/
|
||
├── index.html
|
||
└── assets/
|
||
├── index-*.css
|
||
└── index-*.js
|
||
```
|
||
|
||
## Features
|
||
|
||
- A round of checks every 3 seconds by default; the interval menu sets 1, 2, 3,
|
||
5, 10, 15, 30 or 60 seconds and clears the history
|
||
- Sparklines of each target's last 100 rounds: 300 seconds at 3 seconds
|
||
- The first round after loading or an interval change is discarded, as DNS and
|
||
TLS setup inflate its latencies
|
||
- Health indicator from the WAN hosts' latest results: OFFLINE (red) when more
|
||
than 10 fail and at most 4 answer, otherwise DEGRADED (orange) when more than
|
||
4 fail, otherwise SLOW (yellow) when more than 3 take over 1000ms, otherwise
|
||
HEALTHY (green)
|
||
- Summary stats across WAN hosts only: how many answered, the min, median,
|
||
average and max of their latest latencies, the min and max over the whole
|
||
history, and the number of rounds run (`Checks`)
|
||
- Fixed chart axes: Y-axis 0–1000ms, higher latencies drawn at the top; X-axis
|
||
the time the history spans
|
||
- Color-coded latency figures and sparkline line segments
|
||
- WAN host rows sorted by latest latency, unreachable last; pinned rows stay on
|
||
top, in name order
|
||
- Play/pause: pause stops probes but history keeps scrolling (blank gaps, no
|
||
false outage)
|
||
- Debug log panel, behind a checkbox in the footer, with five levels (error,
|
||
warning, notice, info, debug) and the last 1000 lines
|
||
- Local and UTC clocks
|
||
- Clickable service URLs
|
||
- A footer link to the commit the page was built from
|
||
- Canvas-based sparkline rendering with devicePixelRatio scaling
|
||
- Zero runtime dependencies: all resources bundled into build artifacts
|
||
|
||
## Deployment
|
||
|
||
`make build` writes the page to `dist/`, which any static file host (S3, GCS,
|
||
Cloudflare Pages, Vercel, Netlify, GitHub Pages) can serve; with no backend
|
||
there, its reports fail quietly and nothing is stored. Or run the Docker image
|
||
behind a reverse proxy.
|
||
|
||
The Docker image, built from `Dockerfile` by `make docker`, is the whole service
|
||
in one container: nginx serves the built frontend and passes `/api/`,
|
||
`/.well-known/healthcheck` and `/metrics` to the Go backend, `netwatch-server`,
|
||
which listens only inside the container, on `127.0.0.1:8081`. The image:
|
||
|
||
- Listens on port 8080 by default (override with `PORT` env var)
|
||
- Takes the client address from `X-Forwarded-For` only on requests from the
|
||
reverse proxies named in `TRUSTED_PROXIES`, and by default from none
|
||
- Sends access logs to stdout
|
||
- Caches static assets with immutable headers
|
||
- Sends the security headers `REPO_POLICIES.md` requires on every response, as
|
||
`security-headers.conf` sets them, in place of the backend's own
|
||
- Stores reports in `DATA_DIR`, `/data/reports` by default, on the `/data`
|
||
volume. Before the backend starts, the image creates `DATA_DIR` and gives it
|
||
and `/data` to user `netwatch` (uid 1000), which the backend runs as, so a
|
||
host directory bind-mounted at `/data` ends up owned by uid 1000
|
||
- Writes buffered reports to disk on `docker stop`, and exits non-zero if nginx
|
||
or the backend exits on its own, so the platform restarts it
|
||
|
||
## Running under upaas
|
||
|
||
What the [upaas](https://git.eeqj.de/sneak/upaas) app for netwatch needs:
|
||
|
||
- **Port:** container port `8080`.
|
||
- **Volume:** container path `/data`; the reports are kept in `/data/reports`.
|
||
- **Environment variables:** none is required. An empty one counts as unset, and
|
||
one set to a value netwatch cannot use stops the container at start, with the
|
||
reason in its log.
|
||
- `PORT`, default `8080`: the container port, from 1 to 65535. `8081` cannot
|
||
be used: the backend listens on it inside the container
|
||
- `REPORTS_PER_MINUTE`, default `60`: reports each client address may send a
|
||
minute
|
||
- `DATA_DIR_MAX_BYTES`, default `1073741824` (1 GiB): the most room the
|
||
report files may take; the oldest are deleted to stay under it
|
||
- `CORS_ALLOWED_ORIGINS`, default empty: other origins whose pages may call
|
||
the API
|
||
- `DEBUG`, default `false`: debug logging
|
||
- `DATA_DIR`, default `/data/reports`: the directory the reports are kept
|
||
in: `/data` or a path below it, with no `.` or `..` part and no extra `/`.
|
||
The container also stops if the path goes through a symbolic link that
|
||
leads out of `/data` or is written as a full path
|
||
- `TRUSTED_PROXIES`, default empty: set it to the address the reverse proxy
|
||
in front of the container connects from, as an IP address or CIDR; several
|
||
are separated by commas. nginx takes the client address from
|
||
`X-Forwarded-For` only on a request from one of them, and the rate limit
|
||
counts that address. Unset, `X-Forwarded-For` is ignored and every client
|
||
behind the proxy shares the proxy's one allowance of `REPORTS_PER_MINUTE`.
|
||
Name only addresses nothing but the proxy connects from: any client that
|
||
connects from one can write its own `X-Forwarded-For`, and through a port
|
||
Docker publishes, every client may connect from the Docker network's
|
||
gateway, such as `172.17.0.1`.
|
||
- `METRICS_USERNAME` and `METRICS_PASSWORD`, default empty: with both set,
|
||
the backend records Prometheus metrics of its requests and serves them at
|
||
`/metrics` on the container port, to requests with this user name and
|
||
password as their basic auth credentials. With neither set, there are no
|
||
metrics and `/metrics` is not found. One set without the other, or a user
|
||
name containing `:`, stops the container
|
||
- `SENTRY_DSN`, default empty: set to a Sentry project's DSN, the backend
|
||
sends its errors to that Sentry project: each request whose handling
|
||
crashes, which still gets a 500 response. A value Sentry does not accept
|
||
stops the container. Empty, the backend sends nothing to Sentry
|
||
- **Health check:** the image's `HEALTHCHECK` requests
|
||
`/.well-known/healthcheck` through nginx every 30 seconds, so it fails unless
|
||
both nginx and the backend answer. upaas reads the container's health 60
|
||
seconds after a deploy and fails the deploy unless it is `healthy`. The
|
||
container also stops when either process exits.
|
||
|
||
## Browser Compatibility
|
||
|
||
Requires a modern browser with ES modules, Fetch API, Canvas API, and CSS custom
|
||
properties.
|
||
|
||
## Limitations
|
||
|
||
- **CORS**: The checks are cross-origin requests in `no-cors` mode, so the page
|
||
cannot read the answer, only time it: any answer counts as reachable, an error
|
||
page included.
|
||
- **Local targets**: The cable modem at 192.168.100.1 and the detected gateway
|
||
answer only on a network that has them, and only when NetWatch is served from
|
||
localhost or a private address (see Monitoring targets).
|
||
- **Network conditions**: Measurements reflect browser-to-endpoint latency,
|
||
which includes your local network, ISP, and internet routing.
|
||
|
||
## TODO
|
||
|
||
The to-do list is [TODO.md](TODO.md): where the work stands, the next step, the
|
||
open work, and what has been done.
|
||
|
||
## License
|
||
|
||
MIT. See [LICENSE](LICENSE).
|
||
|
||
## Author
|
||
|
||
[@sneak](https://sneak.berlin)
|