@@ -1,30 +1,33 @@
NetWatch is an MIT-licensed JavaScript single-page application by
NetWatch is an MIT-licensed JavaScript single-page application by
[@sneak ](https://sneak.berlin ) that provides real-time network latency
[@sneak ](https://sneak.berlin ) that provides real-time network latency
monitoring to common internet hosts, displayed with color-coded figures and
monitoring to common internet hosts, displayed with color-coded figures and
sparkline graphs, served from a static bucket or Docker container.
sparkline graphs, served from a static bucket or from its Docker image, where a
small Go backend stores the measurements the page reports.
## Getting Started
## Getting Started
``` bash
``` bash
# Install dependencies
# Install the dependencies and the git pre-commit hook
yarn install
make setup
# Development server
# Run the page on the Vite dev server
yarn dev
make dev
# Production build
# Run the tests, both linters and the format check
yarn build
make check
# Preview production build
# Build the page into dist/
yarn preview
make build
# Docker
# Build the image and run it
docker build -t netwatch .
make docker
docker run -p 8080:8080 netwatch
docker run -p 8080:8080 netwatch
```
```
`yarn dev` proxi es `/api` to `http://127.0.0.1:8080` , so a locally running
`make check` and `make docker` need Docker. `make dev` pass es `/api` to
`netwatch-server` (see `backend/` ) receives the reports the page posts.
`http://127.0.0.1:8080` , where `make run` in `backend/` starts `netwatch-server`
with its defaults, so the reports the page posts are stored in
`backend/data/reports` .
## Entrypoints
## Entrypoints
@@ -94,22 +97,25 @@ halves, so the root `make check` fails if either one is broken. We provide:
The narrow-viewport layout lives in the `max-width: 768px` media block in
The narrow-viewport layout lives in the `max-width: 768px` media block in
`src/styles.css` . It is verified automatically by `make frontend-viewport-test` ,
`src/styles.css` . It is verified automatically by `make frontend-viewport-test` ,
which drives a digest-pinned headless Chrome against the built `dist/` and
which drives a digest-pinned headless Chrome against the built `dist/` and
asserts on computed layout at widths derived from that CSS — one pixel either
asserts on computed layout at widths derived from that CSS — on every breakpoint
side of every breakpoint it declares , plus a 320px floor, a desktop baseline and
it declares and one pixel either side of it , plus a 320px floor, a desktop
two landscape sizes. See [ test/viewport/README.md ]( test/viewport/README.md ) for
baseline and two landscape sizes. See
what it covers and what it genuinely cannot.
[ test/viewport/README.md ]( test/viewport/README.md ) for what it covers and what
it genuinely cannot.
## Rationale
## Rationale
When debugging network issues, it's useful to have a persistent at-a-glance view
When debugging network issues, it's useful to have a persistent at-a-glance view
of latency and reachability to multiple well-known internet endpoints. NetWatch
of latency and reachability to multiple well-known internet endpoints. NetWatch
provides this as a zero-dependency SPA that can be deployed anywhere static
provides this as a single page that does all its measuring in the browser, so it
files are served, with no backend required.
can be served from anywhere static files are served. The backend in its Docker
image only stores the measurements the page reports; without it, the page works
the same and nothing is stored.
## Design
## Design
The application is a single-page app built with Vite and Tailwind CSS v4. All
The page is built with Vite and Tailwind CSS v4. Its code is all in
code lives in `src/main.js` with a class-based architecture:
`src/main.js` , with a class-based architecture:
- **`CONFIG` **: Configuration object (update interval, timeouts, axis ticks,
- **`CONFIG` **: Configuration object (update interval, timeouts, axis ticks,
etc.). The interval menu sets `updateInterval` , the one value the page writes
etc.). The interval menu sets `updateInterval` , the one value the page writes
@@ -125,10 +131,10 @@ code lives in `src/main.js` with a class-based architecture:
`updateSummary()` / `updateHealthBox()` handle incremental updates
`updateSummary()` / `updateHealthBox()` handle incremental updates
- **`tick()` **: Main loop — measures all hosts in parallel, pushing each host's
- **`tick()` **: Main loop — measures all hosts in parallel, pushing each host's
sample and redrawing its row as soon as its check ends, then redraws every
sample and redrawing its row as soon as its check ends, then redraws every
row, the summary and the health box once the last check ends. The rows are
row, the summary and the health box once the last check ends. The first round,
sorted then too, after the first round that is not discarded and every tenth
after loading or an interval change, is discarded. The rows are sorted when
round after that. When paused, pushes blank markers (no probes, no false
the last check ends in round 2, the first one kept, and in rounds 11, 21, 31
outage)
and so on. When paused, pushes blank markers (no probes, no false outage)
- **`Reporter` **: Posts collected samples to the backend
- **`Reporter` **: Posts collected samples to the backend
### Reporting
### Reporting
@@ -142,12 +148,35 @@ delivered report, and while paused nothing is sent. Delivery failure is quiet
one debug-log line per outage, retried at the next interval, never blocking
one debug-log line per outage, retried at the next interval, never blocking
probing. The report-building step is a pure function of host state.
probing. The report-building step is a pure function of host state.
### Backend
`netwatch-server` , in `backend/` , is a small Go HTTP server that stores the
reports the page posts. It keeps them in memory and writes them to `DATA_DIR` as
zstd-compressed files of JSON lines: every minute, whenever 10 MiB are waiting,
and when it stops. Its routes:
- `POST /api/v1/reports` — takes a report, without credentials; each client
address may send a limited number a minute, and the report files are capped in
size, the oldest deleted first
- `GET /.well-known/healthcheck` — answers 200 with `"status":"ok"` , the
server's version and its uptime
- `GET /metrics` — Prometheus metrics behind basic auth, only when
`METRICS_USERNAME` and `METRICS_PASSWORD` are set; each client address may
make a limited number of requests to it a minute
In the image, the `builder` stage of `Dockerfile` tests it and builds it with
`backend/script/build` , and `bin/entrypoint.sh` runs it as user `netwatch` on
`127.0.0.1:8081` , behind nginx. Outside the image, `make run` in `backend/`
builds it and runs it on port 8080. Its settings, report storage and limits are
in [backend/README.md ](backend/README.md ).
### Monitoring targets
### Monitoring targets
- **22 WAN hosts**: datavi.be, Anthropic API, OpenAI API, AWS Console, GCP
- **26 WAN hosts**: datavi.be (pinned at start) , Anthropic API, OpenAI API, AWS
Console, Azure, Cloudflare, Fastly, Akamai, GitHub, B2, 7 S3 regional
Console, Google Cloud Console, Microsoft Azure, Cloudflare, Fastly CDN,
endpoints (Cape Town, London, Bahrain, Tokyo, Sydney, Oregon, São Paulo), 4
Akamai, Google, GitHub, B2, 8 S3 regional endpoints (Cape Town, London,
GCS locational endpoints (Iowa, Belgium, Singapore, Sydney)
Bahrain, Tokyo, Singapore, Sydney, Oregon, São Paulo) and 6 Hetzner speed test
servers (Nuremberg, Falkenstein, Helsinki, Ashburn, Hillsboro, Singapore)
- **Local CPE**: Cable modem at 192.168.100.1 (always monitored)
- **Local CPE**: Cable modem at 192.168.100.1 (always monitored)
- **Local Gateway**: Auto-detected on startup by probing common default gateway
- **Local Gateway**: Auto-detected on startup by probing common default gateway
addresses (192.168.1.1, 192.168.0.1, 192.168.8.1, 10.0.0.1); first responder
addresses (192.168.1.1, 192.168.0.1, 192.168.8.1, 10.0.0.1); first responder
@@ -159,15 +188,17 @@ Local hosts are tracked separately from WAN stats.
### Latency measurement
### Latency measurement
HEAD requests with `mode: 'no-cors'` and `cache: 'no-store'` , timed with
GET requests with `mode: 'no-cors'` , `cache: 'no-store'` and a cache-busting
`performance.now()` . Each check times out after 80% of the refresh interval (24
query parameter, timed with `performance.now()` . Each check times out after 80%
seconds at 30 seconds) and is then recorded as a timeout, so a round's checks
of the refresh interval (24 seconds at 30 seconds) and is then recorded as a
have all finished before the next round is due. When no WAN host answers, a
timeout, so a round's checks have all finished before the next round is due.
recovery probe checks 4 random WAN hosts every half second, giving up the checks
When no WAN host answers, a recovery probe checks 4 WAN hosts, picked at random
it started half a second before. As soon as one answers, a new round starts at
when it starts, every half second, giving up the checks it started half a second
once, as it does after an interval change. A round started early gives up the
before. As soon as one answers, a new round starts at once, as it does after an
last round's checks if they are still waiting, and that round records nothing
interval change. A round started early gives up the last round's checks if they
more, so rounds never overlap. IPv4 only.
are still waiting, and that round records nothing more, so rounds never overlap.
The browser chooses between IPv4 and IPv6 for each target, as for any request;
the local targets are IPv4 addresses.
### Color coding
### Color coding
@@ -192,25 +223,42 @@ dist/
## Features
## Features
- Real-time monitoring with 2s update interval and 300s history sparklines
- A round of checks every 3 seconds by default; the interval menu sets 1, 2, 3,
- Health indicator: green (HEALTHY) or red (DEGRADED) based on WAN reachabilit y
5, 10, 15, 30 or 60 seconds and clears the histor y
- Summary stats: reachable count, min/max/avg latency across WAN hosts only
- Sparklines of each target's last 100 rounds: 300 seconds at 3 seconds
- Fixed chart axes: Y-axis 0– 1000ms, X-axis 0– 300s
- The first round after loading or an interval change is discarded, as DNS and
TLS setup inflate its latencies
- Health indicator from the WAN hosts' latest results: OFFLINE (red) when more
than 10 fail and at most 4 answer, otherwise DEGRADED (orange) when more than
4 fail, otherwise SLOW (yellow) when more than 3 take over 1000ms, otherwise
HEALTHY (green)
- Summary stats across WAN hosts only: how many answered, the min, median,
average and max of their latest latencies, the min and max over the whole
history, and the number of rounds run (`Checks` )
- Fixed chart axes: Y-axis 0– 1000ms, higher latencies drawn at the top; X-axis
the time the history spans
- Color-coded latency figures and sparkline line segments
- Color-coded latency figures and sparkline line segments
- WAN host rows sorted by latest latency, unreachable last; pinned rows stay on
top, in name order
- Play/pause: pause stops probes but history keeps scrolling (blank gaps, no
- Play/pause: pause stops probes but history keeps scrolling (blank gaps, no
false outage)
false outage)
- Debug log panel, behind a checkbox in the footer, with five levels (error,
warning, notice, info, debug) and the last 1000 lines
- Local and UTC clocks
- Clickable service URLs
- Clickable service URLs
- A footer link to the commit the page was built from
- Canvas-based sparkline rendering with devicePixelRatio scaling
- Canvas-based sparkline rendering with devicePixelRatio scaling
- Zero runtime dependencies: all resources bundled into build artifacts
- Zero runtime dependencies: all resources bundled into build artifacts
## Deployment
## Deployment
After running `yarn build` , deploy the contents of the `dist/` directory to any
`make build` writes the page to `dist/` , which any static file host (S3, GCS,
static file host (S3, GCS, Cloudflare Pages, Vercel, Netlify, GitHub Pages) or
Cloudflare Pages, Vercel, Netlify, GitHub Pages) can serve; with no backend
use the Docker image behind a reverse proxy.
there, its reports fail quietly and nothing is stored. Or run the Docker image
behind a reverse proxy.
The Docker image, built from `Dockerfile` , is the whole service in one
The Docker image, built from `Dockerfile` by `make docker` , is the whole service
container: nginx serves the built frontend and passes `/api/` ,
in one container: nginx serves the built frontend and passes `/api/` ,
`/.well-known/healthcheck` and `/metrics` to the Go backend, `netwatch-server` ,
`/.well-known/healthcheck` and `/metrics` to the Go backend, `netwatch-server` ,
which listens only inside the container, on `127.0.0.1:8081` . The image:
which listens only inside the container, on `127.0.0.1:8081` . The image:
@@ -283,19 +331,19 @@ properties.
## Limitations
## Limitations
- **CORS**: Some hosts may block cross-origin HEAD requests. The app uses
- **CORS**: The checks are cross-origin requests in `no-cors` mode, so the page
`no-cors` mode which allows the request but provides opaque responses. Latency
cannot read the answer, only time it: any answer counts as reachable, an error
is still measurable based on request timing .
page included .
- **Local gateway **: The 192.168.100.1 e ndpoint requires the host to be
- **Local targets **: The cable modem at 192.168.100.1 a nd the detected gateway
accessible from your network.
answer only on a network that has them, and only when NetWatch is served from
localhost or a private address (see Monitoring targets).
- **Network conditions**: Measurements reflect browser-to-endpoint latency,
- **Network conditions**: Measurements reflect browser-to-endpoint latency,
which includes your local network, ISP, and internet routing.
which includes your local network, ISP, and internet routing.
## TODO
## TODO
- Add configurable host list (environment variable or config file)
The to-do list is [TODO.md ](TODO.md ): where the work stands, the next step, the
- Add latency history export (CSV/JSON)
open work, and what has been done.
- Add notification/alert when status changes to DEGRADED
## License
## License