1 Commits
Author SHA1 Message Date
sneak 1b4fbb032b Add make add-dependency and make tidy (closes #45)
check / check (push) Successful in 3m5s
No entrypoint could change yarn.lock or go.mod: script/bootstrap
installs with --frozen-lockfile, so adding a package meant running
yarn by hand. make add-dependency PACKAGE=<name>@<version> shims to
the new script/add-dependency: yarn add --dev, then yarn install
--frozen-lockfile. make tidy shims to the new script/tidy, go mod
tidy in backend/; a Go module is added by importing it, or moved by
editing its require line, then make tidy. script/bootstrap is
unchanged.

Model: opus-5-5
2026-10-04 02:59:06 +00:00
13 changed files with 206 additions and 616 deletions
+57 -111
View File
@@ -1,33 +1,30 @@
NetWatch is an MIT-licensed JavaScript single-page application by
[@sneak](https://sneak.berlin) that provides real-time network latency
monitoring to common internet hosts, displayed with color-coded figures and
sparkline graphs, served from a static bucket or from its Docker image, where a
small Go backend stores the measurements the page reports.
sparkline graphs, served from a static bucket or Docker container.
## Getting Started
```bash
# Install the dependencies and the git pre-commit hook
make setup
# Install dependencies
yarn install
# Run the page on the Vite dev server
make dev
# Development server
yarn dev
# Run the tests, both linters and the format check
make check
# Production build
yarn build
# Build the page into dist/
make build
# Preview production build
yarn preview
# Build the image and run it
make docker
# Docker
docker build -t netwatch .
docker run -p 8080:8080 netwatch
```
`make check` and `make docker` need Docker. `make dev` passes `/api` to
`http://127.0.0.1:8080`, where `make run` in `backend/` starts `netwatch-server`
with its defaults, so the reports the page posts are stored in
`backend/data/reports`.
`yarn dev` proxies `/api` to `http://127.0.0.1:8080`, so a locally running
`netwatch-server` (see `backend/`) receives the reports the page posts.
## Entrypoints
@@ -97,30 +94,25 @@ halves, so the root `make check` fails if either one is broken. We provide:
The narrow-viewport layout lives in the `max-width: 768px` media block in
`src/styles.css`. It is verified automatically by `make frontend-viewport-test`,
which drives a digest-pinned headless Chrome against the built `dist/` and
asserts on computed layout at widths derived from that CSS — on every breakpoint
it declares and one pixel either side of it, plus a 320px floor, a desktop
baseline and two landscape sizes. See
[test/viewport/README.md](test/viewport/README.md) for what it covers and what
it genuinely cannot.
asserts on computed layout at widths derived from that CSS — one pixel either
side of every breakpoint it declares, plus a 320px floor, a desktop baseline and
two landscape sizes. See [test/viewport/README.md](test/viewport/README.md) for
what it covers and what it genuinely cannot.
## Rationale
When debugging network issues, it's useful to have a persistent at-a-glance view
of latency and reachability to multiple well-known internet endpoints. NetWatch
provides this as a single page that does all its measuring in the browser, so it
can be served from anywhere static files are served. The backend in its Docker
image only stores the measurements the page reports; without it, the page works
the same and nothing is stored.
provides this as a zero-dependency SPA that can be deployed anywhere static
files are served, with no backend required.
## Design
The page is built with Vite and Tailwind CSS v4. Its code is all in
`src/main.js`, with a class-based architecture:
The application is a single-page app built with Vite and Tailwind CSS v4. All
code lives in `src/main.js` with a class-based architecture:
- **`CONFIG`**: Configuration object (update interval, timeouts, axis ticks,
etc.). The interval menu sets `updateInterval`, the one value the page writes
into `CONFIG`; the timeouts, the time the history spans and the x-axis ticks
are computed from it
- **`CONFIG`**: Frozen configuration object (update interval, timeouts, axis
ticks, etc.)
- **`HostState`**: Per-host state management — history buffer, latency tracking,
status transitions
- **`AppState`**: Top-level state container — WAN hosts, local hosts, pause
@@ -131,10 +123,10 @@ The page is built with Vite and Tailwind CSS v4. Its code is all in
`updateSummary()` / `updateHealthBox()` handle incremental updates
- **`tick()`**: Main loop — measures all hosts in parallel, pushing each host's
sample and redrawing its row as soon as its check ends, then redraws every
row, the summary and the health box once the last check ends. The first round,
after loading or an interval change, is discarded. The rows are sorted when
the last check ends in round 2, the first one kept, and in rounds 11, 21, 31
and so on. When paused, pushes blank markers (no probes, no false outage)
row, the summary and the health box once the last check ends. The rows are
sorted then too, after the first round that is not discarded and every tenth
round after that. When paused, pushes blank markers (no probes, no false
outage)
- **`Reporter`**: Posts collected samples to the backend
### Reporting
@@ -148,35 +140,12 @@ delivered report, and while paused nothing is sent. Delivery failure is quiet
one debug-log line per outage, retried at the next interval, never blocking
probing. The report-building step is a pure function of host state.
### Backend
`netwatch-server`, in `backend/`, is a small Go HTTP server that stores the
reports the page posts. It keeps them in memory and writes them to `DATA_DIR` as
zstd-compressed files of JSON lines: every minute, whenever 10 MiB are waiting,
and when it stops. Its routes:
- `POST /api/v1/reports` — takes a report, without credentials; each client
address may send a limited number a minute, and the report files are capped in
size, the oldest deleted first
- `GET /.well-known/healthcheck` — answers 200 with `"status":"ok"`, the
server's version and its uptime
- `GET /metrics` — Prometheus metrics behind basic auth, only when
`METRICS_USERNAME` and `METRICS_PASSWORD` are set; each client address may
make a limited number of requests to it a minute
In the image, the `builder` stage of `Dockerfile` tests it and builds it with
`backend/script/build`, and `bin/entrypoint.sh` runs it as user `netwatch` on
`127.0.0.1:8081`, behind nginx. Outside the image, `make run` in `backend/`
builds it and runs it on port 8080. Its settings, report storage and limits are
in [backend/README.md](backend/README.md).
### Monitoring targets
- **26 WAN hosts**: datavi.be (pinned at start), Anthropic API, OpenAI API, AWS
Console, Google Cloud Console, Microsoft Azure, Cloudflare, Fastly CDN,
Akamai, Google, GitHub, B2, 8 S3 regional endpoints (Cape Town, London,
Bahrain, Tokyo, Singapore, Sydney, Oregon, São Paulo) and 6 Hetzner speed test
servers (Nuremberg, Falkenstein, Helsinki, Ashburn, Hillsboro, Singapore)
- **22 WAN hosts**: datavi.be, Anthropic API, OpenAI API, AWS Console, GCP
Console, Azure, Cloudflare, Fastly, Akamai, GitHub, B2, 7 S3 regional
endpoints (Cape Town, London, Bahrain, Tokyo, Sydney, Oregon, São Paulo), 4
GCS locational endpoints (Iowa, Belgium, Singapore, Sydney)
- **Local CPE**: Cable modem at 192.168.100.1 (always monitored)
- **Local Gateway**: Auto-detected on startup by probing common default gateway
addresses (192.168.1.1, 192.168.0.1, 192.168.8.1, 10.0.0.1); first responder
@@ -188,17 +157,15 @@ Local hosts are tracked separately from WAN stats.
### Latency measurement
GET requests with `mode: 'no-cors'`, `cache: 'no-store'` and a cache-busting
query parameter, timed with `performance.now()`. Each check times out after 80%
of the refresh interval (24 seconds at 30 seconds) and is then recorded as a
timeout, so a round's checks have all finished before the next round is due.
When no WAN host answers, a recovery probe checks 4 WAN hosts, picked at random
when it starts, every half second, giving up the checks it started half a second
before. As soon as one answers, a new round starts at once, as it does after an
interval change. A round started early gives up the last round's checks if they
are still waiting, and that round records nothing more, so rounds never overlap.
The browser chooses between IPv4 and IPv6 for each target, as for any request;
the local targets are IPv4 addresses.
HEAD requests with `mode: 'no-cors'` and `cache: 'no-store'`, timed with
`performance.now()`. Each check times out after 80% of the refresh interval (24
seconds at 30 seconds) and is then recorded as a timeout, so a round's checks
have all finished before the next round is due. When no WAN host answers, a
recovery probe checks 4 random WAN hosts every half second, giving up the checks
it started half a second before. As soon as one answers, a new round starts at
once, as it does after an interval change. A round started early gives up the
last round's checks if they are still waiting, and that round records nothing
more, so rounds never overlap. IPv4 only.
### Color coding
@@ -223,42 +190,25 @@ dist/
## Features
- A round of checks every 3 seconds by default; the interval menu sets 1, 2, 3,
5, 10, 15, 30 or 60 seconds and clears the history
- Sparklines of each target's last 100 rounds: 300 seconds at 3 seconds
- The first round after loading or an interval change is discarded, as DNS and
TLS setup inflate its latencies
- Health indicator from the WAN hosts' latest results: OFFLINE (red) when more
than 10 fail and at most 4 answer, otherwise DEGRADED (orange) when more than
4 fail, otherwise SLOW (yellow) when more than 3 take over 1000ms, otherwise
HEALTHY (green)
- Summary stats across WAN hosts only: how many answered, the min, median,
average and max of their latest latencies, the min and max over the whole
history, and the number of rounds run (`Checks`)
- Fixed chart axes: Y-axis 0–1000ms, higher latencies drawn at the top; X-axis
the time the history spans
- Real-time monitoring with 2s update interval and 300s history sparklines
- Health indicator: green (HEALTHY) or red (DEGRADED) based on WAN reachability
- Summary stats: reachable count, min/max/avg latency across WAN hosts only
- Fixed chart axes: Y-axis 0–1000ms, X-axis 0–300s
- Color-coded latency figures and sparkline line segments
- WAN host rows sorted by latest latency, unreachable last; pinned rows stay on
top, in name order
- Play/pause: pause stops probes but history keeps scrolling (blank gaps, no
false outage)
- Debug log panel, behind a checkbox in the footer, with five levels (error,
warning, notice, info, debug) and the last 1000 lines
- Local and UTC clocks
- Clickable service URLs
- A footer link to the commit the page was built from
- Canvas-based sparkline rendering with devicePixelRatio scaling
- Zero runtime dependencies: all resources bundled into build artifacts
## Deployment
`make build` writes the page to `dist/`, which any static file host (S3, GCS,
Cloudflare Pages, Vercel, Netlify, GitHub Pages) can serve; with no backend
there, its reports fail quietly and nothing is stored. Or run the Docker image
behind a reverse proxy.
After running `yarn build`, deploy the contents of the `dist/` directory to any
static file host (S3, GCS, Cloudflare Pages, Vercel, Netlify, GitHub Pages) or
use the Docker image behind a reverse proxy.
The Docker image, built from `Dockerfile` by `make docker`, is the whole service
in one container: nginx serves the built frontend and passes `/api/`,
The Docker image, built from `Dockerfile`, is the whole service in one
container: nginx serves the built frontend and passes `/api/`,
`/.well-known/healthcheck` and `/metrics` to the Go backend, `netwatch-server`,
which listens only inside the container, on `127.0.0.1:8081`. The image:
@@ -314,10 +264,6 @@ What the [upaas](https://git.eeqj.de/sneak/upaas) app for netwatch needs:
password as their basic auth credentials. With neither set, there are no
metrics and `/metrics` is not found. One set without the other, or a user
name containing `:`, stops the container
- `SENTRY_DSN`, default empty: set to a Sentry project's DSN, the backend
sends its errors to that Sentry project: each request whose handling
crashes, which still gets a 500 response. A value Sentry does not accept
stops the container. Empty, the backend sends nothing to Sentry
- **Health check:** the image's `HEALTHCHECK` requests
`/.well-known/healthcheck` through nginx every 30 seconds, so it fails unless
both nginx and the backend answer. upaas reads the container's health 60
@@ -331,19 +277,19 @@ properties.
## Limitations
- **CORS**: The checks are cross-origin requests in `no-cors` mode, so the page
cannot read the answer, only time it: any answer counts as reachable, an error
page included.
- **Local targets**: The cable modem at 192.168.100.1 and the detected gateway
answer only on a network that has them, and only when NetWatch is served from
localhost or a private address (see Monitoring targets).
- **CORS**: Some hosts may block cross-origin HEAD requests. The app uses
`no-cors` mode which allows the request but provides opaque responses. Latency
is still measurable based on request timing.
- **Local gateway**: The 192.168.100.1 endpoint requires the host to be
accessible from your network.
- **Network conditions**: Measurements reflect browser-to-endpoint latency,
which includes your local network, ISP, and internet routing.
## TODO
The to-do list is [TODO.md](TODO.md): where the work stands, the next step, the
open work, and what has been done.
- Add configurable host list (environment variable or config file)
- Add latency history export (CSV/JSON)
- Add notification/alert when status changes to DEGRADED
## License
+19 -70
View File
@@ -1,77 +1,28 @@
# Workflow
- branch from `next`
- branch (from `main`)
- do the work in Next Step
- move Next Step to the top of Completed Steps
- move the top item of Future Steps into Next Step
- commit (`TODO.md` changes in the same commit as the work)
- push the branch and open a PR against `next`
- merge to `main` if the branch is not protected, otherwise open a PR
- push
# Status
pre-1.0. No git tags. `main` is the stable branch and `next` the development
branch, which every PR targets. The frontend and the Go backend ship as one
Docker image, and the Gitea workflow `.gitea/workflows/check.yml` runs
`script/cibuild` on every push. Working toward 1.0.0.
pre-1.0. No git tags. `feat/reportbuf-storage` is merged; the backend, the CI
workflow, and the backend repo standard files are all on `main`. Frontend and
backend are both functional. Working toward the 1.0.0 milestone by closing the
remaining repo-compliance issues on the tracker.
# Next Step
Decide whether the repo moves to the layout `REPO_POLICIES.md` gives, with
`backend/` no longer repeating files from the root
([#30](https://git.eeqj.de/sneak/netwatch/issues/30)).
Confirm the `.gitea/workflows/check.yml` run is green (main always green
policy). The workflow file is already on `main`; what is unverified is that its
latest run passes.
# Completed Steps
- 2026-10-04: the page's footer no longer says "IPv4 only"
([#111](https://git.eeqj.de/sneak/netwatch/issues/111)): each check is a
`fetch`, the browser picks IPv4 or IPv6 for each WAN host, and the local
targets are IPv4 addresses. The rest of the footer is unchanged
- 2026-10-04: `README.md`, `TODO.md` and `test/viewport/README.md` say what the
tree does (issue #24). The README's Getting Started leads with `make` targets;
a new Backend section says what `netwatch-server` stores, its routes and how
the image builds and runs it, and points to `backend/README.md` for its
settings; the checks are GET requests; the 26 WAN hosts, the four health
states, the summary's figures and the features the list lacked are described
as the page has them; and its TODO section points here, as does the one in
`backend/README.md`, whose open items moved to Future Steps. This file's
Workflow branches from `next` and opens the PR against `next`, Status says
where the repo stands, and Next Step and Future Steps hold only open work,
linked to its issue where one exists. The viewport harness README names Node's
test runner, not `vitest`
- 2026-10-04: in `src/main.js` (issue #102), a target's min, max, median and
average latency come from one list of its answers, through the same function
the summary's figures use, so the median is written once. The latency color
limits are one table in `CONFIG`, read by both the figure's and the
sparkline's color. The health thresholds, the debug log's length, the gateway
check's timeout, the recovery probe's number of hosts and interval, how often
the rows are sorted and the delay before the first sparkline resize are
`CONFIG` entries too. A unit test checks the summary's figures. Nothing the
page does or shows changed; the footer's color legend still writes the limits
out as text
- 2026-10-04: password guesses at `/metrics` are rate limited (issue #104): each
client address, resolved through `TRUSTED_PROXIES` as for reports, may make 60
requests to `/metrics` a minute, counted by `go-chi/httprate` apart from its
reports; past that it gets 429 and its basic auth credentials are not checked.
The limit is a constant in `backend/internal/server/routes.go`. A test uses up
one client's allowance on wrong passwords, gets 429 with the right one, and
checks that another client behind the same nginx still gets in
- 2026-10-04: the backend reports errors to Sentry (issue #95). With
`SENTRY_DSN` set, it sets up `sentry-go` with the release `netwatch-server-`
and its version, reports each panic in a handler through `sentryhttp`, the
last of the middleware every request goes through, which panics again so the
request still gets the 500 from the panic recovery, and waits up to 2 seconds
on shutdown for Sentry to finish sending. A DSN Sentry refuses stops the start
with an error naming `SENTRY_DSN`. With it empty, Sentry is not set up and
nothing is sent to it
- 2026-10-04: a target's name and URL and a debug log message show as the
characters they are and are never read as HTML (issue #29): a host row escapes
the name and URL it writes into its markup, and the debug log sets each line
as text. A unit test checks a target whose name and URL hold `<`, `>`, `"`,
`&` and `'`. `README.md` no longer calls `CONFIG` frozen: the interval menu
sets its `updateInterval`, and the values computed from it follow. `AppState`
declares the recovery probe's two properties, the sparkline axis functions
lose the parameters they did not use, and the comment on a target's history
names both kinds of entry it holds. Nothing the page does changed
- 2026-10-04: a dependency can be added without running yarn or go by hand
(issue #45): `make add-dependency PACKAGE=<name>@<version>` shims to the new
`script/add-dependency`, which runs `yarn add --dev`, so `package.json` and
@@ -389,14 +340,12 @@ Decide whether the repo moves to the layout `REPO_POLICIES.md` gives, with
# Future Steps
- Run `make frontend-viewport-test` in CI as its own step; it is not part of
`make check`, as it needs Docker and takes minutes
- A backend test that posts a report to `POST /api/v1/reports` and checks the
compressed file it is written to
- A backend route that decompresses the stored reports and answers queries on
them
- Prometheus metrics for the backend's in-memory buffer: its size, the number of
flushes and the number of reports
- A configurable host list (an environment variable or a config file)
- Export of the latency history (CSV or JSON)
- A notification when the health status changes to DEGRADED
- Wire `script/frontend-viewport-test` into CI as its own step (deliberately not
part of `make check` today; the decision has real CI-runtime cost and is
tracked separately)
- Compliance top-up as one small commit: add .editorconfig and add the hooks
target to the Makefile
- After merge, confirm .gitea/workflows/check.yml is on main and CI is green
(main always green policy)
- Decide what to do with untracked resume.sh: commit it, gitignore it, or delete
it
+3 -21
View File
@@ -90,7 +90,6 @@ project layout:
| `CORS_ALLOWED_ORIGINS` | empty | Comma-separated origins whose pages may call the API; see [CORS](#cors) |
| `METRICS_USERNAME` | empty | Basic auth user name for `/metrics`; see [Metrics](#metrics) |
| `METRICS_PASSWORD` | empty | Basic auth password for `/metrics`; see [Metrics](#metrics) |
| `SENTRY_DSN` | empty | DSN of the Sentry project to send errors to; see [Sentry](#sentry) |
`TRUSTED_PROXIES` defaults to
`127.0.0.1/32,::1/128,10.0.0.0/8,172.16.0.0/12,192.168.0.0/16`. The loopback
@@ -199,28 +198,11 @@ is recorded and `/metrics` answers 404. One without the other stops the server
from starting, with an error naming both; so does a `METRICS_USERNAME`
containing `:`, which basic auth cannot carry, with an error naming it.
`/metrics` is rate limited, so that its password cannot be guessed quickly: each
client address, resolved through `TRUSTED_PROXIES`, may make 60 requests to it a
minute, whatever their credentials. Past that it gets 429 with
`Retry-After: 60`, and its credentials are not checked. The minute slides as it
does for reports (see [Report limits](#report-limits)), so a scraper polling
every 2 seconds or less often is never refused. This allowance is apart from the
one for reports.
### Sentry
With `SENTRY_DSN` set, the server sends its errors to that Sentry project: each
panic in a handler is reported there, under the release `netwatch-server-`
followed by the server's version, and the request still gets 500 from the
server's panic recovery. On shutdown the server waits up to 2 seconds for Sentry
to finish sending. A DSN Sentry refuses stops the server from starting, with an
error naming `SENTRY_DSN`. With it empty, Sentry is not set up, and nothing is
sent to it.
## TODO
The to-do list, this backend's open work included, is [TODO.md](../TODO.md) at
the repo root.
- Add integration test that POSTs a report and verifies the compressed output
- Add report decompression/query endpoint
- Add metrics (Prometheus) for buffer size, flush count, report count
## License
-22
View File
@@ -115,28 +115,6 @@ func TestMalformedConfigFileStopsTheStart(t *testing.T) {
}
}
// TestRefusedSentryDSNStopsTheStart: a SENTRY_DSN that Sentry refuses
// stops the start, and the error, naming SENTRY_DSN, is logged as JSON.
func TestRefusedSentryDSNStopsTheStart(t *testing.T) {
t.Setenv("SENTRY_DSN", "not-a-dsn")
ctx, cancel := context.WithTimeout(t.Context(), childTimeout)
defer cancel()
child, stdout, stderr := startServer(ctx, t, t.TempDir(), freePort(ctx, t))
err := child.Wait()
if child.ProcessState.ExitCode() != 1 {
t.Fatalf("server exit = %v, want exit status 1", err)
}
requireJSONLines(t, stdout, stderr)
if !strings.Contains(stdout.String(), "SENTRY_DSN") {
t.Fatalf("no error naming SENTRY_DSN in stdout:\n%s", stdout)
}
}
// startServer runs main() in a child process listening on
// 127.0.0.1:port, with home as its HOME and working directory and its
// data directory in home, so it touches nothing outside home. Its
-1
View File
@@ -4,7 +4,6 @@ go 1.25.5
require (
github.com/99designs/basicauth-go v0.0.0-20230316000542-bf6f9cbbf0f8
github.com/getsentry/sentry-go v0.49.0
github.com/go-chi/chi/v5 v5.2.5
github.com/go-chi/cors v1.2.2
github.com/go-chi/httprate v0.16.0
+8 -16
View File
@@ -4,22 +4,18 @@ github.com/beorn7/perks v1.0.1 h1:VlbKKnNfV8bJzeqoa4cOKqO6bYr3WgKZxO8Z16+hsOM=
github.com/beorn7/perks v1.0.1/go.mod h1:G2ZrVWU2WbWT9wwq4/hrbKbnv/1ERSJQ0ibhJ6rlkpw=
github.com/cespare/xxhash/v2 v2.3.0 h1:UL815xU9SqsFlibzuggzjXhog7bL6oX9BbNZnL2UFvs=
github.com/cespare/xxhash/v2 v2.3.0/go.mod h1:VGX0DQ3Q6kWi7AoAeZDth3/j3BFtOZR5XLFGgcrjCOs=
github.com/davecgh/go-spew v1.1.2-0.20180830191138-d8f796af33cc h1:U9qPSI2PIWSS1VwoXQT9A3Wy9MM3WgvqSxFWenqJduM=
github.com/davecgh/go-spew v1.1.2-0.20180830191138-d8f796af33cc/go.mod h1:J7Y8YcW2NihsgmVo/mv3lAwl/skON4iLHjSsI+c5H38=
github.com/davecgh/go-spew v1.1.1 h1:vj9j/u1bqnvCEfJOwUhtlOARqs3+rkHYY13jYWTU97c=
github.com/davecgh/go-spew v1.1.1/go.mod h1:J7Y8YcW2NihsgmVo/mv3lAwl/skON4iLHjSsI+c5H38=
github.com/frankban/quicktest v1.14.6 h1:7Xjx+VpznH+oBnejlPUj8oUpdxnVs4f8XU8WnHkI4W8=
github.com/frankban/quicktest v1.14.6/go.mod h1:4ptaffx2x8+WTWXmUCuVU6aPUX1/Mz7zb5vbUoiM6w0=
github.com/fsnotify/fsnotify v1.9.0 h1:2Ml+OJNzbYCTzsxtv8vKSFD9PbJjmhYF14k/jKC7S9k=
github.com/fsnotify/fsnotify v1.9.0/go.mod h1:8jBTzvmWwFyi3Pb8djgCCO5IBqzKJ/Jwo8TRcHyHii0=
github.com/getsentry/sentry-go v0.49.0 h1:Ehejknu1l023Ub7QoRBVLAI7g3Jnhqku4oWx4B4Sh5s=
github.com/getsentry/sentry-go v0.49.0/go.mod h1:nuMJAoCfe1u0Bts2ocyNI+TW8HT84vRMqwA5Qq/SKUI=
github.com/go-chi/chi/v5 v5.2.5 h1:Eg4myHZBjyvJmAFjFvWgrqDTXFyOzjj7YIm3L3mu6Ug=
github.com/go-chi/chi/v5 v5.2.5/go.mod h1:X7Gx4mteadT3eDOMTsXzmI4/rwUpOwBHLpAfupzFJP0=
github.com/go-chi/cors v1.2.2 h1:Jmey33TE+b+rB7fT8MUy1u0I4L+NARQlK6LhzKPSyQE=
github.com/go-chi/cors v1.2.2/go.mod h1:sSbTewc+6wYHBBCW7ytsFSn836hqM7JxpglAy2Vzc58=
github.com/go-chi/httprate v0.16.0 h1:8V5DH9j6pSK6UQoBsTpvMyFxycqaKEIToyPKzHJjUa8=
github.com/go-chi/httprate v0.16.0/go.mod h1:A8lo+qRhk+s9LiuP5saS7XCGDXRXMcrueq0NfIuCa/I=
github.com/go-errors/errors v1.4.2 h1:J6MZopCL4uSllY1OfXM374weqZFFItUbrImctkmUxIA=
github.com/go-errors/errors v1.4.2/go.mod h1:sIVyrIiJhuEF+Pj9Ebtd6P/rEYROXFi3BopGUQ5a5Og=
github.com/go-viper/mapstructure/v2 v2.4.0 h1:EBsztssimR/CONLSZZ04E8qAkxNYq4Qp9LvH92wZUgs=
github.com/go-viper/mapstructure/v2 v2.4.0/go.mod h1:oJDH3BJKyqBA2TXFhDsKDGDTlndYOZ6rGS0BRZIxGhM=
github.com/google/go-cmp v0.7.0 h1:wk8382ETsv4JYUZwIsn6YpYiWiBsYLSJiTsyBybVuN8=
@@ -40,12 +36,8 @@ github.com/munnerz/goautoneg v0.0.0-20191010083416-a7dc8b61c822 h1:C3w9PqII01/Oq
github.com/munnerz/goautoneg v0.0.0-20191010083416-a7dc8b61c822/go.mod h1:+n7T8mK8HuQTcFwEeznm/DIxMOiR9yIdICNftLE1DvQ=
github.com/pelletier/go-toml/v2 v2.2.4 h1:mye9XuhQ6gvn5h28+VilKrrPoQVanw5PMw/TB0t5Ec4=
github.com/pelletier/go-toml/v2 v2.2.4/go.mod h1:2gIqNv+qfxSVS7cM2xJQKtLSTLUE9V8t9Stt+h56mCY=
github.com/pingcap/errors v0.11.4 h1:lFuQV/oaUMGcD2tqt+01ROSmJs75VG1ToEOkZIZ4nE4=
github.com/pingcap/errors v0.11.4/go.mod h1:Oi8TUi2kEtXXLMJk9l1cGmz20kV3TaQ0usTwv5KuLY8=
github.com/pkg/errors v0.9.1 h1:FEBLx1zS214owpjy7qsBeixbURkuhQAwrK5UwLGTwt4=
github.com/pkg/errors v0.9.1/go.mod h1:bwawxfHBFNV+L2hUp1rHADufV3IMtnDRdf1r5NINEl0=
github.com/pmezard/go-difflib v1.0.1-0.20181226105442-5d4384ee4fb2 h1:Jamvg5psRIccs7FGNTlIRMkT8wgtp5eCXdBlqhYGL6U=
github.com/pmezard/go-difflib v1.0.1-0.20181226105442-5d4384ee4fb2/go.mod h1:iKH77koFhYxTK1pcRnkKkqfTogsbg7gZNVY4sRDYZ/4=
github.com/pmezard/go-difflib v1.0.0 h1:4DBwDE0NGyQoBHbLQYPwSUPoCMWR5BEzIk/f1lZbAQM=
github.com/pmezard/go-difflib v1.0.0/go.mod h1:iKH77koFhYxTK1pcRnkKkqfTogsbg7gZNVY4sRDYZ/4=
github.com/prometheus/client_golang v1.24.1 h1:JnJkREXzWxUdCuPFpIWZiPispT9xVV59uiuyR2bPlnU=
github.com/prometheus/client_golang v1.24.1/go.mod h1:F+oSRECHg4sse5ucfYpYDeIv/hu68Zo0uoHKetWnzcE=
github.com/prometheus/client_model v0.6.2 h1:oBsgwpGs7iVziMvrGhE53c/GrLUsZdHnqNwqPLxwZyk=
@@ -54,8 +46,8 @@ github.com/prometheus/common v0.70.1 h1:1HvjP4D5oL3t8RsPlwxA9onvvStjtIHYE5XuuwOi
github.com/prometheus/common v0.70.1/go.mod h1:VdFUQDMZK3VLkurFUVhia6uys/0suUp86TJz5qbJRhc=
github.com/prometheus/procfs v0.21.1 h1:GljZCt+zSTS+NZq88cyQ1LjZ+RCHp3uVuabBWA5+OJI=
github.com/prometheus/procfs v0.21.1/go.mod h1:aB55Cww9pdSJVHk0hUf0inxWyyjPogFIjmHKYgMKmtY=
github.com/rogpeppe/go-internal v1.14.1 h1:UQB4HGPB6osV0SQTLymcB4TgvyWu6ZyliaW0tI/otEQ=
github.com/rogpeppe/go-internal v1.14.1/go.mod h1:MaRKkUm5W0goXpeCfT7UZI6fk/L7L7so1lCWt35ZSgc=
github.com/rogpeppe/go-internal v1.9.0 h1:73kH8U+JUqXU8lRuOHeVHaa/SZPifC7BkcraZVejAe8=
github.com/rogpeppe/go-internal v1.9.0/go.mod h1:WtVeX8xhTBvf0smdhujwtBcq4Qrzq/fJaraNFVN+nFs=
github.com/sagikazarmark/locafero v0.11.0 h1:1iurJgmM9G3PA/I+wWYIOw/5SyBtxapeHDcg+AAIFXc=
github.com/sagikazarmark/locafero v0.11.0/go.mod h1:nVIGvgyzw595SUSUE6tvCp3YYTeHs15MvlmU87WwIik=
github.com/slok/go-http-metrics v0.13.0 h1:lQDyJJx9wKhmbliyUsZ2l6peGnXRHjsjoqPt5VYzcP8=
@@ -101,7 +93,7 @@ golang.org/x/text v0.40.0/go.mod h1:hpnzDAfGV753zIKo+wk3u1bVKCGPbrnF7+7LBF/UHVY=
google.golang.org/protobuf v1.36.11 h1:fV6ZwhNocDyBLK0dj+fg8ektcVegBBuEolpbTQyBNVE=
google.golang.org/protobuf v1.36.11/go.mod h1:HTf+CrKn2C3g5S8VImy6tdcUvCska2kB7j23XfzDpco=
gopkg.in/check.v1 v0.0.0-20161208181325-20d25e280405/go.mod h1:Co6ibVJAznAaIkqp8huTwlJQCZ016jof/cbN4VW5Yz0=
gopkg.in/check.v1 v1.0.0-20201130134442-10cb98267c6c h1:Hei/4ADfdWqJk1ZMxUNpqntNwaWcugrBjAiHlqqRiVk=
gopkg.in/check.v1 v1.0.0-20201130134442-10cb98267c6c/go.mod h1:JHkPIbrfpd72SG/EVd6muEfDQjcINNoR0C8j2r3qZ4Q=
gopkg.in/check.v1 v1.0.0-20190902080502-41f04d3bba15 h1:YR8cESwS4TdDjEe65xsg0ogRM/Nc3DYOhEAlW+xobZo=
gopkg.in/check.v1 v1.0.0-20190902080502-41f04d3bba15/go.mod h1:Co6ibVJAznAaIkqp8huTwlJQCZ016jof/cbN4VW5Yz0=
gopkg.in/yaml.v3 v3.0.1 h1:fxVm/GzAzEWqLHuvctI91KS9hhNmmWOoWu0XTYJS7CA=
gopkg.in/yaml.v3 v3.0.1/go.mod h1:K4uyk7z7BCEPqu6E+C64Yfv1cQ7kz7rIZviUmN+EgEM=
-12
View File
@@ -1,21 +1,9 @@
package server
import "github.com/go-chi/chi/v5"
// Router exposes the router to the external tests, which add routes
// of their own to it after SetupRoutes.
func (s *Server) Router() *chi.Mux {
return s.router
}
// MaxRequestBodyBytes exposes the router-wide body limit to the
// external tests.
const MaxRequestBodyBytes = maxRequestBodyBytes
// MetricsRequestsPerMinute exposes the /metrics rate limit to the
// external tests.
const MetricsRequestsPerMinute = metricsRequestsPerMinute
// ListenAddr exposes the address the server listens on to the
// external tests.
func (s *Server) ListenAddr() string {
+4 -21
View File
@@ -3,7 +3,6 @@ package server
import (
"time"
sentryhttp "github.com/getsentry/sentry-go/http"
"github.com/go-chi/chi/v5"
"github.com/go-chi/chi/v5/middleware"
"github.com/prometheus/client_golang/prometheus"
@@ -18,12 +17,6 @@ const (
// can mount s.mw.MaxBodyBytes with a smaller value to lower
// its bound, but cannot raise it: this cap runs first.
maxRequestBodyBytes int64 = 1 << 20 // 1 MiB
// metricsRequestsPerMinute is how many requests to /metrics each
// client address may make a minute, whatever their credentials. A
// scraper polling every 2 seconds sends half of it, which httprate
// never refuses.
metricsRequestsPerMinute = 60
)
// SetupRoutes configures the chi router with middleware and
@@ -39,12 +32,6 @@ func (s *Server) SetupRoutes() {
s.router.Use(s.mw.MaxBodyBytes(maxRequestBodyBytes))
s.router.Use(middleware.Timeout(requestTimeout))
// Sentry reports a panic, then panics again, so that s.mw.Recoverer
// still answers 500.
if s.params.Config.SentryDSN != "" {
s.router.Use(sentryhttp.New(sentryhttp.Options{Repanic: true}).Handle)
}
// The metrics go in a registry of this server's own, not in
// Prometheus' default one, which takes them only once per process.
registry := prometheus.NewRegistry()
@@ -72,14 +59,10 @@ func (s *Server) SetupRoutes() {
Post("/api/v1/reports", s.h.HandleReport())
})
// The rate limit comes before the basic auth, so a client past it
// gets 429 and its password is not checked.
if s.params.Config.MetricsUsername != "" {
s.router.With(
s.mw.RateLimit(metricsRequestsPerMinute),
s.mw.MetricsAuth(),
).Get("/metrics", promhttp.HandlerFor(
registry, promhttp.HandlerOpts{},
).ServeHTTP)
s.router.With(s.mw.MetricsAuth()).
Get("/metrics", promhttp.HandlerFor(
registry, promhttp.HandlerOpts{},
).ServeHTTP)
}
}
-111
View File
@@ -1,12 +1,10 @@
package server_test
import (
"io"
"net/http"
"net/http/httptest"
"strings"
"testing"
"time"
"sneak.berlin/go/netwatch/internal/config"
"sneak.berlin/go/netwatch/internal/globals"
@@ -17,7 +15,6 @@ import (
"sneak.berlin/go/netwatch/internal/reportbuf"
"sneak.berlin/go/netwatch/internal/server"
"github.com/getsentry/sentry-go"
"go.uber.org/fx"
"go.uber.org/fx/fxtest"
)
@@ -188,50 +185,6 @@ func TestMetricsBehindBasicAuth(t *testing.T) {
}
}
// TestMetricsAreRateLimited: a client that has used up its /metrics
// allowance on wrong passwords gets 429 even with the right one, which
// is then not checked, while another client behind the same nginx
// still gets in.
func TestMetricsAreRateLimited(t *testing.T) {
t.Setenv("METRICS_USERNAME", "prometheus")
t.Setenv("METRICS_PASSWORD", "right")
// As in the container: nginx connects from loopback and names the
// client in X-Forwarded-For.
t.Setenv("TRUSTED_PROXIES", "127.0.0.1/32")
srv := newServer(t)
srv.SetupRoutes()
get := func(client, password string) int {
rec := httptest.NewRecorder()
req := httptest.NewRequestWithContext(t.Context(),
http.MethodGet, "/metrics", http.NoBody)
req.RemoteAddr = "127.0.0.1:40000"
req.Header.Set("X-Forwarded-For", client)
req.SetBasicAuth("prometheus", password)
srv.ServeHTTP(rec, req)
return rec.Code
}
for i := range server.MetricsRequestsPerMinute {
if code := get("203.0.113.7", "wrong"); code != http.StatusUnauthorized {
t.Fatalf("guess %d: status = %d, want %d",
i+1, code, http.StatusUnauthorized)
}
}
if code := get("203.0.113.7", "right"); code != http.StatusTooManyRequests {
t.Fatalf("right password past the limit: status = %d, want %d",
code, http.StatusTooManyRequests)
}
if code := get("203.0.113.8", "right"); code != http.StatusOK {
t.Fatalf("another client: status = %d, want %d",
code, http.StatusOK)
}
}
// TestMetricsInTwoServers: two servers in one process can both have
// metrics on.
func TestMetricsInTwoServers(t *testing.T) {
@@ -243,70 +196,6 @@ func TestMetricsInTwoServers(t *testing.T) {
}
}
// TestSentry: with SENTRY_DSN empty there is no Sentry client. With it
// pointing at a local server standing in for Sentry, a panic in a
// handler reaches that server, and the request still gets the 500 from
// the panic recovery.
func TestSentry(t *testing.T) {
const panicMessage = "handler panic for TestSentry"
// sentry.Init sets the client for the whole process; take it away
// again so that no other test reports to Sentry.
t.Cleanup(func() { sentry.CurrentHub().BindClient(nil) })
t.Setenv("SENTRY_DSN", "")
newServer(t)
if sentry.CurrentHub().Client() != nil {
t.Fatal("a Sentry client exists with SENTRY_DSN empty")
}
// The body of the first request the stand-in for Sentry receives.
received := make(chan string, 1)
sentryServer := httptest.NewServer(http.HandlerFunc(
func(_ http.ResponseWriter, r *http.Request) {
body, _ := io.ReadAll(r.Body)
select {
case received <- string(body):
default:
}
},
))
defer sentryServer.Close()
t.Setenv("SENTRY_DSN",
"http://key@"+sentryServer.Listener.Addr().String()+"/1")
srv := newServer(t)
srv.SetupRoutes()
srv.Router().Get("/panic", func(http.ResponseWriter, *http.Request) {
panic(panicMessage)
})
rec := httptest.NewRecorder()
req := httptest.NewRequestWithContext(t.Context(),
http.MethodGet, "/panic", http.NoBody)
srv.ServeHTTP(rec, req)
if rec.Code != http.StatusInternalServerError {
t.Fatalf("status = %d, want %d",
rec.Code, http.StatusInternalServerError)
}
// Sentry sends from a goroutine of its own.
select {
case body := <-received:
if !strings.Contains(body, panicMessage) {
t.Fatalf("the Sentry server received no report of the panic:\n%s",
body)
}
case <-time.After(5 * time.Second):
t.Fatal("nothing reached the Sentry server")
}
}
// TestHealthCheckRejectsOversizeBody sends the health check, which
// never reads its body, a body one byte over the limit. Only the
// router-wide body limit can reject it.
+1 -40
View File
@@ -7,10 +7,8 @@ package server
import (
"context"
"fmt"
"log/slog"
"net/http"
"time"
"sneak.berlin/go/netwatch/internal/config"
"sneak.berlin/go/netwatch/internal/globals"
@@ -18,15 +16,10 @@ import (
"sneak.berlin/go/netwatch/internal/logger"
"sneak.berlin/go/netwatch/internal/middleware"
"github.com/getsentry/sentry-go"
"github.com/go-chi/chi/v5"
"go.uber.org/fx"
)
// sentryFlushTimeout is how long shutdown waits for Sentry to send
// what it still holds.
const sentryFlushTimeout = 2 * time.Second
// Params defines the dependencies for Server.
type Params struct {
fx.In
@@ -63,11 +56,6 @@ func New(
s.log = params.Logger.Get()
s.shutdowner = params.Shutdowner
err := s.enableSentry()
if err != nil {
return nil, err
}
lc.Append(fx.Hook{
OnStart: func(_ context.Context) error {
// Build the router and http.Server synchronously
@@ -99,37 +87,10 @@ func (s *Server) ServeHTTP(
s.router.ServeHTTP(w, r)
}
// enableSentry sets Sentry up when SENTRY_DSN is set, so that
// SetupRoutes can report panics to it. With SENTRY_DSN empty it does
// nothing. A DSN Sentry refuses stops the start.
func (s *Server) enableSentry() error {
if s.params.Config.SentryDSN == "" {
return nil
}
err := sentry.Init(sentry.ClientOptions{
Dsn: s.params.Config.SentryDSN,
Release: s.params.Globals.Appname + "-" + s.params.Globals.Version,
})
if err != nil {
return fmt.Errorf("SENTRY_DSN: %w", err)
}
s.log.Info("sentry error reporting activated")
return nil
}
// shutdown gracefully stops the HTTP server within the
// deadline of the context fx provides for OnStop, then gives
// Sentry, if set up, time to send what it still holds.
// deadline of the context fx provides for OnStop.
func (s *Server) shutdown(ctx context.Context) error {
err := s.httpServer.Shutdown(ctx)
if s.params.Config.SentryDSN != "" {
sentry.Flush(sentryFlushTimeout)
}
if err != nil {
s.log.Error("server clean shutdown failed", "error", err)
+97 -130
View File
@@ -8,8 +8,6 @@
// display their real value in the latency figure. The history buffer holds
// maxHistoryPoints samples (historyDuration / updateInterval).
// reportInterval is how often collected samples are POSTed to the backend.
// The interval menu changes updateInterval while the page runs; the
// getters compute their values from it each time they are read.
export const CONFIG = {
updateInterval: 3000,
maxHistoryPoints: 100,
@@ -30,39 +28,6 @@ export const CONFIG = {
return [0, 1, 2, 3, 4, 5].map((i) => Math.round((d * i) / 5));
},
canvasHeight: 96,
// A latency figure and its sparkline take the color of the first entry
// whose limit, in ms, the latency is below.
latencyColors: [
{ below: 50, hex: "#22c55e", className: "text-green-500" },
{ below: 100, hex: "#84cc16", className: "text-lime-500" },
{ below: 200, hex: "#eab308", className: "text-yellow-500" },
{ below: 500, hex: "#f97316", className: "text-orange-500" },
{ below: Infinity, hex: "#ef4444", className: "text-red-500" },
],
// The health is offline when more than offlineTimeouts WAN hosts timed
// out or were unreachable and at most offlineReachable answered;
// otherwise degraded when more than degradedTimeouts timed out or were
// unreachable; otherwise slow when more than slowHosts answered after
// more than slowLatency ms.
offlineTimeouts: 10,
offlineReachable: 4,
degradedTimeouts: 4,
slowHosts: 3,
slowLatency: 1000,
// The debug log keeps its last maxLogEntries lines.
maxLogEntries: 1000,
// A gateway candidate that has not answered after gatewayTimeout ms is
// passed over.
gatewayTimeout: 1500,
// When no WAN host answers, the recovery probe checks recoveryProbeHosts
// random ones every recoveryProbeInterval ms.
recoveryProbeHosts: 4,
recoveryProbeInterval: 500,
// The rows are sorted after the first round that is not discarded, then
// every roundsPerSort rounds.
roundsPerSort: 10,
// The sparklines are sized and drawn again resizeDelay ms after start.
resizeDelay: 100,
};
// WAN endpoints to monitor. These are used for the aggregate health/stats
@@ -147,8 +112,7 @@ const debugLog = [];
const log = (() => {
function append(level, message) {
debugLog.push({ timestamp: new Date(), level, message });
if (debugLog.length > CONFIG.maxLogEntries)
debugLog.splice(0, debugLog.length - CONFIG.maxLogEntries);
if (debugLog.length > 1000) debugLog.splice(0, debugLog.length - 1000);
const panel = document.getElementById("debug-panel");
if (panel && !panel.classList.contains("hidden")) renderDebugLog();
}
@@ -208,10 +172,7 @@ async function detectGateway() {
const result = await Promise.any(
GATEWAY_CANDIDATES.map(async (url) => {
const controller = new AbortController();
const timeoutId = setTimeout(
() => controller.abort(),
CONFIG.gatewayTimeout,
);
const timeoutId = setTimeout(() => controller.abort(), 1500);
try {
await fetch(url, {
method: "GET",
@@ -236,35 +197,11 @@ async function detectGateway() {
// --- App State ---------------------------------------------------------------
// The min, max, median and average of latencies, a list of numbers, or all
// null when it is empty. The median of an even count is the mean of the
// middle two; it and the average are rounded.
function latencyStats(latencies) {
if (latencies.length === 0)
return { min: null, max: null, med: null, avg: null };
const sorted = [...latencies].sort((a, b) => a - b);
const mid = Math.floor(sorted.length / 2);
return {
min: sorted[0],
max: sorted[sorted.length - 1],
med:
sorted.length % 2
? sorted[mid]
: Math.round((sorted[mid - 1] + sorted[mid]) / 2),
avg: Math.round(
latencies.reduce((a, b) => a + b, 0) / latencies.length,
),
};
}
export class HostState {
constructor(host, pinned = false) {
this.name = host.name;
this.url = host.url;
// Each entry is either a check's result, { timestamp, latency,
// error }, or a round skipped while paused, { timestamp,
// latency: null, paused: true }.
this.history = [];
this.history = []; // { timestamp, latency, paused }
this.lastLatency = null;
this.status = "pending"; // 'online' | 'offline' | 'error' | 'pending'
this.pinned = pinned;
@@ -288,16 +225,38 @@ export class HostState {
this._trim();
}
// The min, max, median and average latency of the checks in the history
// that got an answer.
historyStats() {
return latencyStats(
this.history
.filter((p) => p.latency !== null)
.map((p) => p.latency),
averageLatency() {
const valid = this.history.filter((p) => p.latency !== null);
if (valid.length === 0) return null;
return Math.round(
valid.reduce((s, p) => s + p.latency, 0) / valid.length,
);
}
minLatency() {
const valid = this.history.filter((p) => p.latency !== null);
if (valid.length === 0) return null;
return Math.min(...valid.map((p) => p.latency));
}
maxLatency() {
const valid = this.history.filter((p) => p.latency !== null);
if (valid.length === 0) return null;
return Math.max(...valid.map((p) => p.latency));
}
medianLatency() {
const sorted = this.history
.filter((p) => p.latency !== null)
.map((p) => p.latency)
.sort((a, b) => a - b);
if (sorted.length === 0) return null;
const mid = Math.floor(sorted.length / 2);
return sorted.length % 2
? sorted[mid]
: Math.round((sorted[mid - 1] + sorted[mid]) / 2);
}
_trim() {
while (this.history.length > CONFIG.maxHistoryPoints)
this.history.shift();
@@ -312,10 +271,6 @@ export class AppState {
this.local = localHosts.map((h) => new HostState(h));
this.paused = false;
this.tickCount = 0;
// The recovery probe's timer, null while it is not running, and the
// checks it started last.
this._recoveryProbeId = null;
this._recoveryProbeChecks = null;
}
get allHosts() {
@@ -324,13 +279,33 @@ export class AppState {
/** WAN-only stats from latest sample (excludes local) */
wanStats() {
const latencies = this.wan
.filter((h) => h.lastLatency !== null)
.map((h) => h.lastLatency);
const reachable = this.wan.filter((h) => h.lastLatency !== null);
const latencies = reachable.map((h) => h.lastLatency);
const total = this.wan.length;
if (latencies.length === 0)
return {
reachable: 0,
total,
min: null,
max: null,
med: null,
avg: null,
};
const sorted = [...latencies].sort((a, b) => a - b);
const mid = Math.floor(sorted.length / 2);
const med =
sorted.length % 2
? sorted[mid]
: Math.round((sorted[mid - 1] + sorted[mid]) / 2);
return {
reachable: latencies.length,
total: this.wan.length,
...latencyStats(latencies),
total,
min: Math.min(...latencies),
max: Math.max(...latencies),
med,
avg: Math.round(
latencies.reduce((a, b) => a + b, 0) / latencies.length,
),
};
}
@@ -356,16 +331,12 @@ export class AppState {
const timeouts = this.wan.filter(
(h) => h.status === "error" || h.status === "offline",
).length;
if (
timeouts > CONFIG.offlineTimeouts &&
reachable <= CONFIG.offlineReachable
)
return "offline";
if (timeouts > CONFIG.degradedTimeouts) return "degraded";
if (timeouts > 10 && reachable <= 4) return "offline";
if (timeouts > 4) return "degraded";
const slow = this.wan.filter(
(h) => h.lastLatency !== null && h.lastLatency > CONFIG.slowLatency,
(h) => h.lastLatency !== null && h.lastLatency > 1000,
).length;
if (slow > CONFIG.slowHosts) return "slow";
if (slow > 3) return "slow";
return "healthy";
}
@@ -577,13 +548,21 @@ export async function measureLatency(url, signal) {
export function latencyHex(latency) {
if (latency === null) return "#6b7280";
return CONFIG.latencyColors.find((c) => latency < c.below).hex;
if (latency < 50) return "#22c55e";
if (latency < 100) return "#84cc16";
if (latency < 200) return "#eab308";
if (latency < 500) return "#f97316";
return "#ef4444";
}
export function latencyClass(latency, status) {
if (status === "offline" || status === "error" || latency === null)
return "text-gray-500";
return CONFIG.latencyColors.find((c) => latency < c.below).className;
if (latency < 50) return "text-green-500";
if (latency < 100) return "text-lime-500";
if (latency < 200) return "text-yellow-500";
if (latency < 500) return "text-orange-500";
return "text-red-500";
}
// --- Sparkline Renderer ------------------------------------------------------
@@ -601,8 +580,8 @@ class SparklineRenderer {
const ch = h - m.top - m.bottom;
ctx.clearRect(0, 0, w, h);
SparklineRenderer._drawYAxis(ctx, w, m, ch);
SparklineRenderer._drawXAxis(ctx, h, m, cw);
SparklineRenderer._drawYAxis(ctx, w, h, m, ch);
SparklineRenderer._drawXAxis(ctx, w, h, m, cw);
const len = history.length;
const pw = cw / (CONFIG.maxHistoryPoints - 1);
@@ -618,7 +597,7 @@ class SparklineRenderer {
SparklineRenderer._drawTip(ctx, history, getX, getY);
}
static _drawYAxis(ctx, w, m, ch) {
static _drawYAxis(ctx, w, h, m, ch) {
ctx.font = "300 12px monospace";
ctx.textAlign = "right";
ctx.textBaseline = "middle";
@@ -635,7 +614,7 @@ class SparklineRenderer {
}
}
static _drawXAxis(ctx, h, m, cw) {
static _drawXAxis(ctx, w, h, m, cw) {
ctx.textAlign = "center";
ctx.textBaseline = "top";
for (const tick of CONFIG.xAxisTicks) {
@@ -723,18 +702,7 @@ class SparklineRenderer {
// horizontally.
const STATUS_TEXT_CLASS = "status-text text-xs text-right col-span-2 mt-5";
// Escapes text for HTML, so it shows as written inside an element or a
// quoted attribute and is never read as markup.
function escapeHTML(text) {
return text
.replaceAll("&", "&amp;")
.replaceAll("<", "&lt;")
.replaceAll(">", "&gt;")
.replaceAll('"', "&quot;")
.replaceAll("'", "&#39;");
}
export function hostRowHTML(host, index, showPin = true) {
function hostRowHTML(host, index, showPin = true) {
const pinColor = host.pinned
? "text-blue-500"
: "text-gray-600 hover:text-gray-400";
@@ -753,12 +721,12 @@ export function hostRowHTML(host, index, showPin = true) {
<div class="w-[420px] flex-shrink-0 grid grid-cols-[minmax(0,1fr)_auto] items-center">
<div class="flex items-center gap-2 min-w-[200px]">
<div class="w-3 h-3 rounded-full flex-shrink-0 bg-[#6b7280]"></div>
<span class="font-medium text-white truncate">${escapeHTML(host.name)}</span>
<span class="font-medium text-white truncate">${host.name}</span>
</div>
<div class="latency-value text-4xl font-bold tabular-nums text-right mt-3" data-host="${index}">
<span class="text-gray-500">---</span>
</div>
<a href="${escapeHTML(host.url)}" target="_blank" rel="noopener" class="text-xs text-gray-500 truncate block col-span-2 -mt-2">${escapeHTML(host.url)}</a>
<a href="${host.url}" target="_blank" rel="noopener" class="text-xs text-gray-500 truncate block col-span-2 -mt-2">${host.url}</a>
<div class="${STATUS_TEXT_CLASS} text-gray-500" data-host="${index}">waiting...</div>
</div>
<div class="flex-grow sparkline-container rounded overflow-hidden border border-gray-700/30">
@@ -857,7 +825,7 @@ function buildUI(state) {
</div>
<footer class="mt-8 text-center text-gray-600 text-xs">
<p>Latency measured via GET requests | CORS restrictions may affect some measurements</p>
<p>Latency measured via GET requests | IPv4 only | CORS restrictions may affect some measurements</p>
<p class="mt-2">
<span class="inline-block w-3 h-3 rounded-full bg-green-500 mr-1 align-middle"></span>&lt;50ms
<span class="inline-block w-3 h-3 rounded-full bg-lime-500 mr-1 ml-3 align-middle"></span>&lt;100ms
@@ -924,7 +892,10 @@ function updateHostRow(host, index) {
latencyEl.innerHTML = `<span class="text-gray-500">---</span>`;
}
const { min, med, avg, max } = host.historyStats();
const avg = host.averageLatency();
const med = host.medianLatency();
const min = host.minLatency();
const max = host.maxLatency();
if (host.status === "online" && avg !== null) {
statusEl.innerHTML = statusStatsHTML([
["min", min],
@@ -1071,16 +1042,14 @@ function renderDebugLog() {
info: "text-gray-300",
debug: "text-gray-500",
};
el.replaceChildren(
...debugLog.map((entry) => {
el.innerHTML = debugLog
.map((entry) => {
const ts = formatUTCTimestamp(entry.timestamp);
const cls = levelColors[entry.level] || "text-gray-400";
const lvl = entry.level.toUpperCase().padEnd(7);
const line = document.createElement("div");
line.className = levelColors[entry.level] || "text-gray-400";
line.textContent = `${ts} ${lvl} ${entry.message}`;
return line;
}),
);
return `<div class="${cls}">${ts} ${lvl} ${entry.message}</div>`;
})
.join("");
el.scrollTop = el.scrollHeight;
}
@@ -1191,9 +1160,8 @@ export async function tick(state, signal, onOffline) {
// rows whose check ended before the resume still read "paused"
state.allHosts.forEach((host, i) => updateHostRow(host, i));
// Sort after the first real check, then every CONFIG.roundsPerSort
// ticks thereafter
if (state.tickCount === 2 || state.tickCount % CONFIG.roundsPerSort === 1) {
// Sort after the first real check, then every 10 ticks thereafter
if (state.tickCount === 2 || state.tickCount % 10 === 1) {
sortAndRebuildWAN(state);
}
@@ -1216,10 +1184,9 @@ export async function tick(state, signal, onOffline) {
// --- Recovery Probe ----------------------------------------------------------
// When offline, check CONFIG.recoveryProbeHosts random WAN hosts every
// CONFIG.recoveryProbeInterval ms, giving up the checks started one interval
// before, so at most that many are ever waiting. As soon as one answers,
// stop probing and start a new round at once.
// When offline, check 4 random WAN hosts every 500ms, giving up the checks
// started 500ms before, so at most 4 are ever waiting. As soon as one
// answers, stop probing and start a new round at once.
function startRecoveryProbe(state, startRounds) {
if (state._recoveryProbeId) return; // already running
const candidates = [...state.wan];
@@ -1227,7 +1194,7 @@ function startRecoveryProbe(state, startRounds) {
const j = Math.floor(Math.random() * (i + 1));
[candidates[i], candidates[j]] = [candidates[j], candidates[i]];
}
const canaries = candidates.slice(0, CONFIG.recoveryProbeHosts);
const canaries = candidates.slice(0, 4);
log.notice(
`Recovery probe started (${canaries.map((h) => h.name).join(", ")})`,
);
@@ -1244,7 +1211,7 @@ function startRecoveryProbe(state, startRounds) {
startRounds();
});
}
}, CONFIG.recoveryProbeInterval);
}, 500);
}
function stopRecoveryProbe(state) {
@@ -1508,7 +1475,7 @@ async function init() {
});
window.addEventListener("resize", () => handleResize(state));
setTimeout(() => handleResize(state), CONFIG.resizeDelay);
setTimeout(() => handleResize(state), 100);
}
// Bootstrap only when loaded as the page: a real DOM containing the #app
+10 -53
View File
@@ -8,7 +8,6 @@ import {
AppState,
CONFIG,
greyOutUI,
hostRowHTML,
HostState,
humanDuration,
latencyClass,
@@ -231,32 +230,6 @@ test("at a 30000ms interval, after the user pauses and resumes during a round, n
}
});
// The page shows &lt; &gt; &quot; &amp; and &#39; in a row's markup as
// < > " & and '.
test(`a target whose name and URL hold < > " & and ' shows those characters in its row`, () => {
const host = new HostState({
name: `<b>"x" & 'y'</b>`,
url: `https://x.test/<b>?a="x"&b='y'`,
});
const row = hostRowHTML(host, 0);
assert.doesNotMatch(row, /<b>/);
assert.ok(
row.includes(
">&lt;b&gt;&quot;x&quot; &amp; &#39;y&#39;&lt;/b&gt;</span>",
),
);
assert.ok(
row.includes(
'href="https://x.test/&lt;b&gt;?a=&quot;x&quot;&amp;b=&#39;y&#39;"',
),
);
assert.ok(
row.includes(
">https://x.test/&lt;b&gt;?a=&quot;x&quot;&amp;b=&#39;y&#39;</a>",
),
);
});
for (const [seconds, text] of [
[0, "0s"],
[1, "1s"],
@@ -347,35 +320,19 @@ for (const { history, latencies, statistics } of [
},
]) {
test(`a target's min, max, average and median latency over ${history}`, () => {
const { min, max, avg, med } = hostAfter(latencies).historyStats();
assert.deepEqual({ min, max, average: avg, median: med }, statistics);
const host = hostAfter(latencies);
assert.deepEqual(
{
min: host.minLatency(),
max: host.maxLatency(),
average: host.averageLatency(),
median: host.medianLatency(),
},
statistics,
);
});
}
// The summary's figures come from each WAN target's last check, by the same
// rules as a target's own: here four answered, one was found unreachable
// and the rest have not been checked yet. The median, 22.5, and the
// average, 21.25, are rounded.
test("the summary's min, max, median and average latency over the WAN targets' last checks", () => {
const state = new AppState([]);
[30, 10, null, 25, 20].forEach((latency, i) =>
state.wan[i].pushSample(
Date.now(),
latency === null
? { latency: null, error: "unreachable" }
: { latency, error: null },
),
);
assert.deepEqual(state.wanStats(), {
reachable: 4,
total: state.wan.length,
min: 10,
max: 30,
med: 23,
avg: 21,
});
});
// An app state in which, of the WAN targets, the first timedOut timed out,
// the next unreachable were found unreachable, the next answered answered
// after latency ms, and the rest have not been checked yet.
+7 -8
View File
@@ -77,7 +77,7 @@ the internet, so the app's latency probes cannot reach anything real. The
harness answers them itself from a fixed delay table, with a deterministic
fraction failed outright, so the rows render a realistic spread of one-, two-
and three-digit latencies plus some unreachable rows. That spread is what the
layout has to survive; a `---` placeholder in every row would not exercise it.
layout has to survive; 24 identical `---` placeholders would not exercise it.
## What this cannot verify
@@ -104,11 +104,10 @@ Everything else this issue was actually about — does the layout reflow, does
anything overflow, is content clipped, are the controls big enough — is a
function of viewport width and CSS, and is covered above.
## Relation to the unit tests
## Relation to the unit test framework (#21)
Complementary layers, not two stacks. The unit tests in `test/unit/`, which
`make test` runs with Node's built-in test runner, exercise the functions
`src/main.js` exports in-process with no browser. This harness exercises
rendered layout in a real engine and is the only thing here that can see a media
query. Neither replaces the other; assertions about computed styles and element
geometry belong here, assertions about functions belong in `test/unit/`.
Complementary layers, not two stacks. `vitest` (#21) will exercise module-level
logic in-process with no browser. This harness exercises rendered layout in a
real engine and is the only thing here that can see a media query. Neither
replaces the other; assertions about computed styles and element geometry belong
here, assertions about functions belong in `vitest`.