19 Commits
Author SHA1 Message Date
clawbot 186f932eb8 Footer no longer says IPv4 only (closes #111)
check / check (push) Successful in 2m48s
Each check is a fetch: the browser picks IPv4 or IPv6 for each WAN
host, and the local targets are IPv4 addresses, so "IPv4 only" was
wrong for the page as a whole. The footer's first line drops it and
its separator; the rest of the footer is unchanged. TODO.md moves this
to Completed Steps and the layout decision from Future Steps into Next
Step.

Model: opus-5-5
2026-10-04 07:38:06 +02:00
clawbot 64e142c17f README, TODO and the viewport README say what the tree does (closes #24)
check / check (push) Successful in 2m57s
README.md: Getting Started leads with make targets; a Backend section
gives netwatch-server's routes and how the image builds and runs it;
the checks are GET requests; the WAN host list, health states, summary
figures, sorting and missing features match src/main.js; the TODO
section points to TODO.md, which holds the one to-do list.

backend/README.md: its TODO section points to TODO.md too, whose
Future Steps take its three open items.

TODO.md: Workflow branches from next and opens the PR against next;
Status, Next Step and Future Steps describe the open work, linked to
its issue where one exists.

test/viewport/README.md: the unit tests run on Node's test runner,
not vitest.

Model: opus-5-5
2026-10-04 07:03:05 +02:00
clawbot 161f955ae2 Latency statistics written once, thresholds read from CONFIG (closes #102)
check / check (push) Successful in 2m50s
A target's min, max, median and average latency now come from one list
of its answers, through latencyStats(), which the summary's figures use
too, so the median is written once. latencyHex() and latencyClass() read
one table of color limits in CONFIG. The health thresholds, the debug
log's length, the gateway check's timeout, the recovery probe's number
of hosts and interval, how often the rows are sorted and the delay
before the first sparkline resize are CONFIG entries, read where the
numbers were. HostState's minLatency(), maxLatency(), averageLatency()
and medianLatency() are gone; the statistics test reads historyStats().
Nothing the page does or shows changes. A unit test now covers the
summary's figures.

Model: opus-5-5
2026-10-04 06:51:24 +02:00
clawbot d40e67d4ab Rate limit password attempts on /metrics (closes #104)
check / check (push) Successful in 3m32s
Each client address may make 60 requests to /metrics a minute,
through the same httprate middleware and TRUSTED_PROXIES
resolution the report route uses, with an allowance of its own.
The limit runs before the basic auth, so past it the answer is 429
and the password is not checked. backend/README.md says so; a test
uses up one client's allowance on wrong passwords, gets 429 with
the right one, and checks that another client behind the same
nginx still gets in.

Model: opus-5-5
2026-10-04 06:19:06 +02:00
clawbot dc2d240725 Report handler panics to Sentry when SENTRY_DSN is set (closes #95)
check / check (push) Successful in 3m6s
With SENTRY_DSN set, the server initialises sentry-go with the release
netwatch-server-<version>, adds the sentryhttp middleware with Repanic
as the last router-wide middleware, after the timeout, and flushes
Sentry for 2 seconds on shutdown. A DSN Sentry refuses stops the start
with an error naming SENTRY_DSN. With it empty, nothing is set up.

The metrics middleware stays on the matched routes only, so it runs
inside the Sentry middleware rather than before it.

Model: opus-5-5
2026-10-04 05:57:02 +02:00
clawbot 00c9f8d7d9 Target names, URLs and log lines reach the page as text (closes #29)
check / check (push) Successful in 3m8s
A host row escapes the name and URL it writes into its markup with a new
escapeHTML function, and the debug log builds each line as an element
whose text is set, so neither is read as HTML once targets can be
configured. A unit test builds the row of a target whose name and URL
hold < > " & and ' and checks each comes out escaped; hostRowHTML is
exported for it.

README.md stops calling CONFIG frozen: the interval menu sets
updateInterval, and the timeouts, history span and axis ticks are
computed from it. AppState declares _recoveryProbeId and
_recoveryProbeChecks, the sparkline axis functions drop the parameters
they never used, and HostState's history comment names both entry
shapes.

Model: opus-5-5
2026-10-04 05:35:36 +02:00
clawbot dc11beb6fe Add make add-dependency and make tidy (closes #45)
check / check (push) Successful in 3m29s
No entrypoint could change yarn.lock or go.mod: script/bootstrap
installs with --frozen-lockfile, so adding a package meant running
yarn by hand. make add-dependency PACKAGE=<name>@<version> shims to
the new script/add-dependency: yarn add --dev, then yarn install
--frozen-lockfile. make tidy shims to the new script/tidy, go mod
tidy in backend/; a Go module is added by importing it, or moved by
editing its require line, then make tidy. script/bootstrap is
unchanged.

Model: opus-5-5
2026-10-04 05:17:33 +02:00
clawbot 81d4153e78 Serve Prometheus metrics at /metrics behind basic auth (closes #94)
check / check (push) Successful in 2m53s
With METRICS_USERNAME and METRICS_PASSWORD both set, the backend
records request metrics through go-http-metrics in a registry of its
own, with Go's runtime and process metrics, and serves them at
GET /metrics behind basic auth; nginx passes /metrics to it. With
neither set there is no such route; one alone, or a METRICS_USERNAME
containing ":", stops the start with an error naming the setting.

Only requests that reach the health check or POST /api/v1/reports are
recorded, as the labels are path and method, which clients could
otherwise make up without end; POST /api/v1/reports is registered by
its full path for that.

Deviation: go get and go mod tidy ran directly; no entrypoint added a Go
dependency yet (issue #45).

Model: opus-5-5
2026-10-04 04:58:47 +02:00
clawbot d412815953 Shim make dev and make build, format backend/ markdown (closes #28)
check / check (push) Successful in 4m11s
make dev ran yarn dev inline and there was no make build. Both are now
shims, over script/dev and the new script/build. .prettierignore stops
leaving out backend/, so backend/README.md is formatted and checked;
backend/.golangci.yml is left out by name instead: it is the org
standard file whose sha256 backend/script/lint checks, and prettier
would reindent it. .claude/ stays in .prettierignore, as git does not
ignore it. script/install-precommit and the date on bootstrap's pins go
back to the org model; bootstrap, fmt and fmt-check keep what the Go
backend and eslint need, each with a comment saying so.

Model: opus-5-5
2026-10-04 04:14:59 +02:00
clawbot 9b548da90d Move tailwind off the deprecated module.register() (closes #32)
check / check (push) Successful in 2m40s
Frontend builds on Node 26 or newer, such as make test on a host with
Node 26, printed Node's DEP0205 warning; the build in Dockerfile runs on
Node 22, which never printed it. The trace names @tailwindcss/node,
which @tailwindcss/vite brings in at its own exact version; tailwind
4.3.1 calls module.registerHooks() where Node has it. @tailwindcss/vite
and tailwindcss move to 4.3.3, inside the ranges package.json already
allows, so only yarn.lock changes, every entry still with an integrity
hash. tailwindcss moves too so that one tailwind version is installed.
Deviation: yarn.lock was changed by a raw yarn upgrade; no entrypoint adds a dependency yet (issue #45).

Model: opus-5-5
2026-10-04 03:51:55 +02:00
clawbot 3b1262718d Frontend unit tests cover durations, colours, statistics and health (closes #21)
check / check (push) Successful in 3m38s
New table-driven tests in test/unit/main.test.js: humanDuration, the
figure and sparkline colours either side of each latency boundary, a
target's min, max, average and median over an empty history, an
all-unreachable one, a mixed one, one of only answers and one whose
answers have different numbers of digits, and the four health states
either side of their thresholds, with targets found unreachable counted
as timed out. src/main.js exports the four names they need.

package.json gains a test script, which script/frontend-test runs with
the dot reporter and, if a test fails, again with the spec reporter
before failing. NODE_OPTIONS picks the reporter, since yarn appends its
arguments after the test files.

Model: opus-5-5
2026-10-04 03:25:06 +02:00
clawbot 82bd142429 Log fx through slog, snake_case health check keys (closes #27)
check / check (push) Successful in 4m40s
fx wrote its own steps of starting and stopping as plain text to
stderr. It now logs them with its slog event logger through the
backend's logger, so off a terminal every line the backend's own
logger and fx write is JSON. A malformed config file now makes
config.New return the error instead of panicking. A test runs the
server as a child process and checks that all it writes is JSON, on
a normal start and stop and with such a config file. The backend
logs its name, version and architecture once at start.

The health check's uptime keys are now uptime_seconds and
uptime_human; its type and method take the names
GO_HTTP_SERVER_CONVENTIONS.md gives.
Rules suppressed: revive and tagliatelle on HealthcheckResponse, whose name and keys come from the conventions; gosec where the test starts its own binary.

Model: opus-5-5
2026-10-04 03:11:43 +02:00
clawbot 924527b744 Lint the frontend with eslint in its own Docker stage (closes #47)
check / check (push) Successful in 3m48s
script/lint ran prettier --check, the same check script/fmt-check
runs, so the JavaScript had no linter. eslint now runs with its
recommended rules, set in eslint.config.js, in a new frontend-lint
stage of Dockerfile built from the pinned node image and the lockfile.
The frontend stage copies a file from it, as the builder stage does
from the Go lint stage, so the image cannot build unless eslint passed.
script/lint builds both lint stages with --no-cache and runs no linter
on the host; script/fmt-check keeps prettier on the host, and
script/frontend-check drops its lint step. The viewport harness fixes
the two kinds of finding eslint made. bootstrap wants node 22.13.0, as
eslint 10 does. Also covers item 2 of
#28.

Model: opus-5-5
2026-10-04 01:58:07 +02:00
clawbot 2f0489e3a4 Tap-target check expects a pin button per WAN host row (closes #46)
check / check (push) Successful in 3m50s
The viewport harness's tap-target check required at least 10 visible
pin buttons while 26 render, so pin buttons missing from up to 16 rows
went unnoticed. It now expects one per WAN host row. The host row count
the harness gathers, which the app-rendered check also reads, counts
only the WAN host rows, since the local host rows have no pin button.
Each control's minimum is now worked out from the gathered facts.

TODO.md's harness entry no longer says every check guards itself: the
overflow, viewport-edge and clipped-text checks rely on app-rendered.

Model: opus-5-5
2026-10-04 01:53:02 +02:00
clawbot 91856fa170 Each target's row shows its result as soon as its check ends (closes #91)
check / check (push) Successful in 3m43s
tick drew no row until the round's slowest check ended, up to 24
seconds at a 30-second interval since checks time out at 80% of it.
Each check now pushes its sample and redraws its row as it ends. When
the last check ends, every row is redrawn, so none still reads "paused"
after a pause and resume, and sorting, the summary, the health box and
offline detection run once. A check that ends while paused or after its
round is given up draws nothing, and the first round is still discarded
as a whole. The row is looked up when the check ends, as a pin click can
re-sort the rows mid-round.

Unit tests run tick on the mocked clock against a stand-in page.

Model: opus-5-5
2026-10-04 01:41:04 +02:00
clawbot 39ee6ca839 Entrypoint acts as root on nothing outside /data (closes #80)
check / check (push) Successful in 1m58s
`bin/entrypoint.sh` now runs `netwatch-server prepare-data-dir`, which
refuses a `DATA_DIR` that is not `/data` or a path below it written in
full, then creates `DATA_DIR`, gives `/data` and everything in it to
`netwatch`, and sets mode 750 on `/data` and `DATA_DIR`. Every step goes
through a Go `os.Root` opened on `/data`, and the modes are set on the
opened directories rather than by name, so neither a symbolic link
already there nor one a host process swaps in while the container
starts can make root create or change anything outside `/data`. The
README says which `DATA_DIR` values are accepted.

Model: opus-5-5
2026-10-03 18:24:38 +02:00
clawbot e4df415676 Delete the oldest report files to stay under the size cap (closes #54)
check / check (push) Successful in 1m57s
When a report would take the report files past DATA_DIR_MAX_BYTES,
reportbuf now deletes the oldest report files until it fits, and does
the same at start when files left by an earlier run are already past
it. A file joins the files that may be deleted, at its place by name,
only once it is completely written, so a file still being written is
never deleted. A report is refused with 507 only when the reports
waiting to be written fill the cap on their own, and then no file is
deleted. The reports of a failed write stop counting, and the part of
its file written is removed. A file whose deletion fails keeps
counting; one already deleted by hand counts as freed.

Model: opus-5-5
2026-10-03 17:51:08 +02:00
clawbot 9e4d3fdb54 Lint's .golangci.yml check says which fix applies (closes #34)
check / check (push) Successful in 2m0s
backend/script/lint compared .golangci.yml with its pinned sha256 and,
on any failure, said to restore the file from sneak/prompts. Once the
org standard has legitimately changed, that advice loops: the copied
file is right and GOLANGCI_CONFIG_SHA256 is stale. On a mismatch the
script now says to compare the file with the org standard, restore it
if they differ, and update GOLANGCI_CONFIG_SHA256 if they are the same;
it still prints both hashes. A missing .golangci.yml, and a sha256sum
that is missing or prints no hash, get their own messages instead of
being reported as a mismatch. Each failure still exits 1. .golangci.yml
is unchanged.

Model: opus-5-5
2026-10-03 17:14:25 +02:00
clawbot aa35625d47 Go tests run with -race and -cover under Go's own timeout (closes #88)
check / check (push) Successful in 2m30s
backend/script/test runs go test -timeout 30s -race -cover and, if
that fails, runs it again with -v and fails. The root script/test drops
its one 30-second timeout around both halves: from a cold Go build
cache, compiling the tests with -race used it all up. Each half keeps
its own limit. The race detector needs a C compiler: the Dockerfile
builder stage gains gcc and musl-dev, and script/bootstrap installs gcc,
with the C library headers on apt and apk, when gcc is missing; make
build still sets CGO_ENABLED=0. New tests: the health check's answer, a
valid report's answer, a report file's exact lines, and the flush at the
10 MiB threshold. The handlers TestImport stub is gone.

Model: opus-5-5
2026-10-03 16:26:17 +02:00
55 changed files with 4079 additions and 719 deletions
+2 -1
View File
@@ -1,6 +1,7 @@
backend/
dist/
node_modules/
tmp/
yarn.lock
.claude/
# The org standard file, copied verbatim; backend/script/lint checks its sha256.
backend/.golangci.yml
+27 -9
View File
@@ -1,6 +1,6 @@
# The one image netwatch ships: nginx serves the built frontend and
# passes /api/ and /.well-known/healthcheck to netwatch-server, the Go
# backend, which runs in the same container on loopback only.
# passes /api/, /.well-known/healthcheck and /metrics to netwatch-server,
# the Go backend, which runs in the same container on loopback only.
# bin/entrypoint.sh starts and watches both.
# Lint stage — fast feedback on formatting and lint issues. The
@@ -20,7 +20,10 @@ RUN make lint
# golang:1.25-alpine (2026-02-27)
FROM golang:1.25-alpine@sha256:f6751d823c26342f9506c03797d2527668d095b0a15f1862cddb4d927a7a4ced AS builder
RUN apk add --no-cache git make
# gcc and musl-dev are for make test: its race detector needs cgo, which
# Go turns on by itself once a C compiler is present. make build still
# sets CGO_ENABLED=0, so the binary stays static.
RUN apk add --no-cache gcc git make musl-dev
WORKDIR /src
@@ -56,20 +59,35 @@ RUN version="${VERSION:-$(git --git-dir=/git describe --tags --always)}"; \
esac; \
VERSION="$version" make build
# Frontend lint stage — eslint over the JavaScript, as the lint stage
# above lints the Go. The root make lint builds this stage alone too.
# node:22-alpine as of 2026-02-22
FROM node@sha256:e4bf2a82ad0a4037d28035ae71529873c069b13eb0455466ae0bc13363826e34 AS frontend-lint
WORKDIR /app
COPY package.json yarn.lock ./
RUN yarn install --frozen-lockfile
COPY . .
RUN script/frontend-lint
# Frontend stage
# node:22-alpine as of 2026-02-22
FROM node@sha256:e4bf2a82ad0a4037d28035ae71529873c069b13eb0455466ae0bc13363826e34 AS frontend
WORKDIR /app
# Force BuildKit to run the frontend-lint stage before proceeding, as
# the builder stage does with the lint stage: without this no-op copy an
# eslint failure would not gate the image.
COPY --from=frontend-lint /app/yarn.lock /dev/null
COPY package.json yarn.lock ./
RUN yarn install --frozen-lockfile
RUN apk add --no-cache git make
COPY . .
# make frontend-check is the frontend half of make check (test + lint +
# fmt-check); its test step runs the unit tests, then the production
# yarn build, so this both produces dist/ and gates the image on
# lint/fmt-check/test regressions.
# This node stage has neither Go nor Docker; the lint and builder stages
# above gate the backend half.
# make frontend-check runs the frontend tests and format check; its test
# step runs the unit tests, then the production yarn build, so this both
# produces dist/ and gates the image on test and formatting regressions.
# This node stage has neither Go nor Docker; the frontend-lint, lint and
# builder stages above gate the rest.
RUN make frontend-check
# Runtime stage
+17 -5
View File
@@ -1,5 +1,5 @@
.PHONY: bootstrap setup dev test lint fmt fmt-check check frontend-check \
frontend-viewport-test docker hooks
.PHONY: bootstrap setup dev build test lint fmt fmt-check check \
add-dependency tidy frontend-check frontend-viewport-test docker hooks
# Standard targets are thin shims; the implementations live in script/
# per the scripts-to-rule-them-all pattern (see the Entrypoints section
@@ -13,7 +13,11 @@ setup:
@script/setup
dev:
yarn dev
@script/dev
# The frontend only; backend/Makefile's build target builds the Go server.
build:
@script/build
test:
@script/test
@@ -30,8 +34,16 @@ fmt-check:
check:
@script/check
# The frontend half of check, for Dockerfile's node build stage, which
# has neither Go nor Docker. Use check everywhere else.
# make add-dependency PACKAGE=<name>@<version>. PACKAGE reaches the
# script through the environment, so the shell never reads it as code.
add-dependency:
@script/add-dependency "$$PACKAGE"
tidy:
@script/tidy
# The frontend tests and format check, for Dockerfile's frontend stage,
# which has neither Go nor Docker. Use check everywhere else.
frontend-check:
@script/frontend-check
+157 -75
View File
@@ -1,30 +1,33 @@
NetWatch is an MIT-licensed JavaScript single-page application by
[@sneak](https://sneak.berlin) that provides real-time network latency
monitoring to common internet hosts, displayed with color-coded figures and
sparkline graphs, served from a static bucket or Docker container.
sparkline graphs, served from a static bucket or from its Docker image, where a
small Go backend stores the measurements the page reports.
## Getting Started
```bash
# Install dependencies
yarn install
# Install the dependencies and the git pre-commit hook
make setup
# Development server
yarn dev
# Run the page on the Vite dev server
make dev
# Production build
yarn build
# Run the tests, both linters and the format check
make check
# Preview production build
yarn preview
# Build the page into dist/
make build
# Docker
docker build -t netwatch .
# Build the image and run it
make docker
docker run -p 8080:8080 netwatch
```
`yarn dev` proxies `/api` to `http://127.0.0.1:8080`, so a locally running
`netwatch-server` (see `backend/`) receives the reports the page posts.
`make check` and `make docker` need Docker. `make dev` passes `/api` to
`http://127.0.0.1:8080`, where `make run` in `backend/` starts `netwatch-server`
with its defaults, so the reports the page posts are stored in
`backend/data/reports`.
## Entrypoints
@@ -39,26 +42,45 @@ halves, so the root `make check` fails if either one is broken. We provide:
- `script/bootstrap` — install all dependencies (the pinned node via nvm unless
one new enough for the frontend's dependencies is installed, yarn via
corepack, `yarn install --frozen-lockfile`, the pinned Go unless one at least
as new as `backend/go.mod` asks for is installed, and the Go modules), linking
what it installs itself into `~/.local/bin`, which has to be on `PATH`. It
installs no Go linter and not Docker: `make lint` runs the linter in Docker
as new as `backend/go.mod` asks for is installed, the Go modules, and gcc with
the C library headers unless gcc is installed, for the race detector in
`make test`), linking what it installs itself into `~/.local/bin`, which has
to be on `PATH`. It installs no Go linter and not Docker: `make lint` runs
both linters in Docker
- `script/setup` — make a fresh clone ready for development: bootstrap plus the
git pre-commit hook
- `script/dev` — run the Vite dev server, which proxies `/api` to a locally
running `netwatch-server`
- `script/build` — build the frontend for production into `dist/`;
`backend/script/build` builds the Go server
- `script/projectname` — print the project name (used for the Docker image tag)
- `script/test` — run `script/frontend-test`, then the backend's Go tests, both
within one 30-second timeout
- `script/lint` — run `script/frontend-lint`, then golangci-lint in Docker, by
building the lint stage of `Dockerfile` without the cache
- `script/test` — run `script/frontend-test`, then `backend/script/test`, the
backend's Go tests with the race detector and coverage
- `script/lint` — run eslint, then golangci-lint, both in Docker, by building
the `frontend-lint` and `lint` stages of `Dockerfile` without the cache
- `script/fmt` — format all files (writes): prettier, then gofmt over `backend/`
- `script/fmt-check` — check formatting (read-only): prettier, then gofmt
- `script/check` — run test, lint, and fmt-check
- `script/add-dependency` — add a frontend package, or move one to another
version: `make add-dependency PACKAGE=<name>@<version>` runs `yarn add --dev`,
which changes `package.json` and `yarn.lock` together, then
`yarn install --frozen-lockfile`
- `script/tidy` — run `go mod tidy` in `backend/`: to add a Go module, import it
and run `make tidy`; to move one to another version, edit its `require` line
in `backend/go.mod`, then run `make tidy`
- `script/frontend-test` — run the unit tests in `test/unit/` with Node's
built-in test runner, then the production build
- `script/frontend-lint` — run prettier in check mode
- `script/frontend-fmt` — format everything prettier understands (writes)
built-in test runner, through the `test` script in `package.json`, and if any
fails, run them again listing every test, and fail; then the production build.
Each run has a 30-second timeout
- `script/frontend-lint` — run eslint with the rules in `eslint.config.js`; it
runs inside the `frontend-lint` stage of `Dockerfile`, which `make lint`
builds
- `script/frontend-fmt` — format everything prettier understands (writes), the
markdown in `backend/` included
- `script/frontend-fmt-check` — check prettier formatting (read-only)
- `script/frontend-check` — the frontend half of `script/check`, for
`Dockerfile`, whose node build stage has neither Go nor Docker
- `script/frontend-check` — run `script/frontend-test` and
`script/frontend-fmt-check`, for the frontend stage of `Dockerfile`, which has
neither Go nor Docker
- `script/frontend-viewport-test` — responsive-layout verification of the built
frontend in a containerised headless Chrome (see
[test/viewport/README.md](test/viewport/README.md)). Not part of
@@ -75,25 +97,30 @@ halves, so the root `make check` fails if either one is broken. We provide:
The narrow-viewport layout lives in the `max-width: 768px` media block in
`src/styles.css`. It is verified automatically by `make frontend-viewport-test`,
which drives a digest-pinned headless Chrome against the built `dist/` and
asserts on computed layout at widths derived from that CSS — one pixel either
side of every breakpoint it declares, plus a 320px floor, a desktop baseline and
two landscape sizes. See [test/viewport/README.md](test/viewport/README.md) for
what it covers and what it genuinely cannot.
asserts on computed layout at widths derived from that CSS — on every breakpoint
it declares and one pixel either side of it, plus a 320px floor, a desktop
baseline and two landscape sizes. See
[test/viewport/README.md](test/viewport/README.md) for what it covers and what
it genuinely cannot.
## Rationale
When debugging network issues, it's useful to have a persistent at-a-glance view
of latency and reachability to multiple well-known internet endpoints. NetWatch
provides this as a zero-dependency SPA that can be deployed anywhere static
files are served, with no backend required.
provides this as a single page that does all its measuring in the browser, so it
can be served from anywhere static files are served. The backend in its Docker
image only stores the measurements the page reports; without it, the page works
the same and nothing is stored.
## Design
The application is a single-page app built with Vite and Tailwind CSS v4. All
code lives in `src/main.js` with a class-based architecture:
The page is built with Vite and Tailwind CSS v4. Its code is all in
`src/main.js`, with a class-based architecture:
- **`CONFIG`**: Frozen configuration object (update interval, timeouts, axis
ticks, etc.)
- **`CONFIG`**: Configuration object (update interval, timeouts, axis ticks,
etc.). The interval menu sets `updateInterval`, the one value the page writes
into `CONFIG`; the timeouts, the time the history spans and the x-axis ticks
are computed from it
- **`HostState`**: Per-host state management — history buffer, latency tracking,
status transitions
- **`AppState`**: Top-level state container — WAN hosts, local hosts, pause
@@ -102,9 +129,12 @@ code lives in `src/main.js` with a class-based architecture:
color-coded line segments, error regions, and DPR-aware scaling
- **UI functions**: `buildUI()` constructs the DOM, `updateHostRow()` /
`updateSummary()` / `updateHealthBox()` handle incremental updates
- **`tick()`**: Main loop — measures all hosts in parallel via `Promise.all`,
pushes samples, redraws UI. When paused, pushes blank markers (no probes, no
false outage)
- **`tick()`**: Main loop — measures all hosts in parallel, pushing each host's
sample and redrawing its row as soon as its check ends, then redraws every
row, the summary and the health box once the last check ends. The first round,
after loading or an interval change, is discarded. The rows are sorted when
the last check ends in round 2, the first one kept, and in rounds 11, 21, 31
and so on. When paused, pushes blank markers (no probes, no false outage)
- **`Reporter`**: Posts collected samples to the backend
### Reporting
@@ -118,12 +148,35 @@ delivered report, and while paused nothing is sent. Delivery failure is quiet
one debug-log line per outage, retried at the next interval, never blocking
probing. The report-building step is a pure function of host state.
### Backend
`netwatch-server`, in `backend/`, is a small Go HTTP server that stores the
reports the page posts. It keeps them in memory and writes them to `DATA_DIR` as
zstd-compressed files of JSON lines: every minute, whenever 10 MiB are waiting,
and when it stops. Its routes:
- `POST /api/v1/reports` — takes a report, without credentials; each client
address may send a limited number a minute, and the report files are capped in
size, the oldest deleted first
- `GET /.well-known/healthcheck` — answers 200 with `"status":"ok"`, the
server's version and its uptime
- `GET /metrics` — Prometheus metrics behind basic auth, only when
`METRICS_USERNAME` and `METRICS_PASSWORD` are set; each client address may
make a limited number of requests to it a minute
In the image, the `builder` stage of `Dockerfile` tests it and builds it with
`backend/script/build`, and `bin/entrypoint.sh` runs it as user `netwatch` on
`127.0.0.1:8081`, behind nginx. Outside the image, `make run` in `backend/`
builds it and runs it on port 8080. Its settings, report storage and limits are
in [backend/README.md](backend/README.md).
### Monitoring targets
- **22 WAN hosts**: datavi.be, Anthropic API, OpenAI API, AWS Console, GCP
Console, Azure, Cloudflare, Fastly, Akamai, GitHub, B2, 7 S3 regional
endpoints (Cape Town, London, Bahrain, Tokyo, Sydney, Oregon, São Paulo), 4
GCS locational endpoints (Iowa, Belgium, Singapore, Sydney)
- **26 WAN hosts**: datavi.be (pinned at start), Anthropic API, OpenAI API, AWS
Console, Google Cloud Console, Microsoft Azure, Cloudflare, Fastly CDN,
Akamai, Google, GitHub, B2, 8 S3 regional endpoints (Cape Town, London,
Bahrain, Tokyo, Singapore, Sydney, Oregon, São Paulo) and 6 Hetzner speed test
servers (Nuremberg, Falkenstein, Helsinki, Ashburn, Hillsboro, Singapore)
- **Local CPE**: Cable modem at 192.168.100.1 (always monitored)
- **Local Gateway**: Auto-detected on startup by probing common default gateway
addresses (192.168.1.1, 192.168.0.1, 192.168.8.1, 10.0.0.1); first responder
@@ -135,15 +188,17 @@ Local hosts are tracked separately from WAN stats.
### Latency measurement
HEAD requests with `mode: 'no-cors'` and `cache: 'no-store'`, timed with
`performance.now()`. Each check times out after 80% of the refresh interval (24
seconds at 30 seconds) and is then recorded as a timeout, so a round's checks
have all finished before the next round is due. When no WAN host answers, a
recovery probe checks 4 random WAN hosts every half second, giving up the checks
it started half a second before. As soon as one answers, a new round starts at
once, as it does after an interval change. A round started early gives up the
last round's checks if they are still waiting, and that round records nothing,
so rounds never overlap. IPv4 only.
GET requests with `mode: 'no-cors'`, `cache: 'no-store'` and a cache-busting
query parameter, timed with `performance.now()`. Each check times out after 80%
of the refresh interval (24 seconds at 30 seconds) and is then recorded as a
timeout, so a round's checks have all finished before the next round is due.
When no WAN host answers, a recovery probe checks 4 WAN hosts, picked at random
when it starts, every half second, giving up the checks it started half a second
before. As soon as one answers, a new round starts at once, as it does after an
interval change. A round started early gives up the last round's checks if they
are still waiting, and that round records nothing more, so rounds never overlap.
The browser chooses between IPv4 and IPv6 for each target, as for any request;
the local targets are IPv4 addresses.
### Color coding
@@ -168,27 +223,44 @@ dist/
## Features
- Real-time monitoring with 2s update interval and 300s history sparklines
- Health indicator: green (HEALTHY) or red (DEGRADED) based on WAN reachability
- Summary stats: reachable count, min/max/avg latency across WAN hosts only
- Fixed chart axes: Y-axis 0–1000ms, X-axis 0–300s
- A round of checks every 3 seconds by default; the interval menu sets 1, 2, 3,
5, 10, 15, 30 or 60 seconds and clears the history
- Sparklines of each target's last 100 rounds: 300 seconds at 3 seconds
- The first round after loading or an interval change is discarded, as DNS and
TLS setup inflate its latencies
- Health indicator from the WAN hosts' latest results: OFFLINE (red) when more
than 10 fail and at most 4 answer, otherwise DEGRADED (orange) when more than
4 fail, otherwise SLOW (yellow) when more than 3 take over 1000ms, otherwise
HEALTHY (green)
- Summary stats across WAN hosts only: how many answered, the min, median,
average and max of their latest latencies, the min and max over the whole
history, and the number of rounds run (`Checks`)
- Fixed chart axes: Y-axis 0–1000ms, higher latencies drawn at the top; X-axis
the time the history spans
- Color-coded latency figures and sparkline line segments
- WAN host rows sorted by latest latency, unreachable last; pinned rows stay on
top, in name order
- Play/pause: pause stops probes but history keeps scrolling (blank gaps, no
false outage)
- Debug log panel, behind a checkbox in the footer, with five levels (error,
warning, notice, info, debug) and the last 1000 lines
- Local and UTC clocks
- Clickable service URLs
- A footer link to the commit the page was built from
- Canvas-based sparkline rendering with devicePixelRatio scaling
- Zero runtime dependencies: all resources bundled into build artifacts
## Deployment
After running `yarn build`, deploy the contents of the `dist/` directory to any
static file host (S3, GCS, Cloudflare Pages, Vercel, Netlify, GitHub Pages) or
use the Docker image behind a reverse proxy.
`make build` writes the page to `dist/`, which any static file host (S3, GCS,
Cloudflare Pages, Vercel, Netlify, GitHub Pages) can serve; with no backend
there, its reports fail quietly and nothing is stored. Or run the Docker image
behind a reverse proxy.
The Docker image, built from `Dockerfile`, is the whole service in one
container: nginx serves the built frontend and passes `/api/` and
`/.well-known/healthcheck` to the Go backend, `netwatch-server`, which listens
only inside the container, on `127.0.0.1:8081`. The image:
The Docker image, built from `Dockerfile` by `make docker`, is the whole service
in one container: nginx serves the built frontend and passes `/api/`,
`/.well-known/healthcheck` and `/metrics` to the Go backend, `netwatch-server`,
which listens only inside the container, on `127.0.0.1:8081`. The image:
- Listens on port 8080 by default (override with `PORT` env var)
- Takes the client address from `X-Forwarded-For` only on requests from the
@@ -218,12 +290,14 @@ What the [upaas](https://git.eeqj.de/sneak/upaas) app for netwatch needs:
- `REPORTS_PER_MINUTE`, default `60`: reports each client address may send a
minute
- `DATA_DIR_MAX_BYTES`, default `1073741824` (1 GiB): the most room the
report files may take
report files may take; the oldest are deleted to stay under it
- `CORS_ALLOWED_ORIGINS`, default empty: other origins whose pages may call
the API
- `DEBUG`, default `false`: debug logging
- `DATA_DIR`, default `/data/reports`: leave unset; reports kept outside
`/data` do not survive a redeploy
- `DATA_DIR`, default `/data/reports`: the directory the reports are kept
in: `/data` or a path below it, with no `.` or `..` part and no extra `/`.
The container also stops if the path goes through a symbolic link that
leads out of `/data` or is written as a full path
- `TRUSTED_PROXIES`, default empty: set it to the address the reverse proxy
in front of the container connects from, as an IP address or CIDR; several
are separated by commas. nginx takes the client address from
@@ -234,6 +308,16 @@ What the [upaas](https://git.eeqj.de/sneak/upaas) app for netwatch needs:
connects from one can write its own `X-Forwarded-For`, and through a port
Docker publishes, every client may connect from the Docker network's
gateway, such as `172.17.0.1`.
- `METRICS_USERNAME` and `METRICS_PASSWORD`, default empty: with both set,
the backend records Prometheus metrics of its requests and serves them at
`/metrics` on the container port, to requests with this user name and
password as their basic auth credentials. With neither set, there are no
metrics and `/metrics` is not found. One set without the other, or a user
name containing `:`, stops the container
- `SENTRY_DSN`, default empty: set to a Sentry project's DSN, the backend
sends its errors to that Sentry project: each request whose handling
crashes, which still gets a 500 response. A value Sentry does not accept
stops the container. Empty, the backend sends nothing to Sentry
- **Health check:** the image's `HEALTHCHECK` requests
`/.well-known/healthcheck` through nginx every 30 seconds, so it fails unless
both nginx and the backend answer. upaas reads the container's health 60
@@ -247,21 +331,19 @@ properties.
## Limitations
- **CORS**: Some hosts may block cross-origin HEAD requests. The app uses
`no-cors` mode which allows the request but provides opaque responses. Latency
is still measurable based on request timing.
- **Local gateway**: The 192.168.100.1 endpoint requires the host to be
accessible from your network.
- **CORS**: The checks are cross-origin requests in `no-cors` mode, so the page
cannot read the answer, only time it: any answer counts as reachable, an error
page included.
- **Local targets**: The cable modem at 192.168.100.1 and the detected gateway
answer only on a network that has them, and only when NetWatch is served from
localhost or a private address (see Monitoring targets).
- **Network conditions**: Measurements reflect browser-to-endpoint latency,
which includes your local network, ISP, and internet routing.
## TODO
- Add unit tests
- Add eslint for JS linting (currently lint target runs prettier only)
- Add configurable host list (environment variable or config file)
- Add latency history export (CSV/JSON)
- Add notification/alert when status changes to DEGRADED
The to-do list is [TODO.md](TODO.md): where the work stands, the next step, the
open work, and what has been done.
## License
+199 -22
View File
@@ -1,28 +1,201 @@
# Workflow
- branch (from `main`)
- branch from `next`
- do the work in Next Step
- move Next Step to the top of Completed Steps
- move the top item of Future Steps into Next Step
- commit (`TODO.md` changes in the same commit as the work)
- merge to `main` if the branch is not protected, otherwise open a PR
- push
- push the branch and open a PR against `next`
# Status
pre-1.0. No git tags. `feat/reportbuf-storage` is merged; the backend, the CI
workflow, and the backend repo standard files are all on `main`. Frontend and
backend are both functional. Working toward the 1.0.0 milestone by closing the
remaining repo-compliance issues on the tracker.
pre-1.0. No git tags. `main` is the stable branch and `next` the development
branch, which every PR targets. The frontend and the Go backend ship as one
Docker image, and the Gitea workflow `.gitea/workflows/check.yml` runs
`script/cibuild` on every push. Working toward 1.0.0.
# Next Step
Confirm the `.gitea/workflows/check.yml` run is green (main always green
policy). The workflow file is already on `main`; what is unverified is that its
latest run passes.
Decide whether the repo moves to the layout `REPO_POLICIES.md` gives, with
`backend/` no longer repeating files from the root
([#30](https://git.eeqj.de/sneak/netwatch/issues/30)).
# Completed Steps
- 2026-10-04: the page's footer no longer says "IPv4 only"
([#111](https://git.eeqj.de/sneak/netwatch/issues/111)): each check is a
`fetch`, the browser picks IPv4 or IPv6 for each WAN host, and the local
targets are IPv4 addresses. The rest of the footer is unchanged
- 2026-10-04: `README.md`, `TODO.md` and `test/viewport/README.md` say what the
tree does (issue #24). The README's Getting Started leads with `make` targets;
a new Backend section says what `netwatch-server` stores, its routes and how
the image builds and runs it, and points to `backend/README.md` for its
settings; the checks are GET requests; the 26 WAN hosts, the four health
states, the summary's figures and the features the list lacked are described
as the page has them; and its TODO section points here, as does the one in
`backend/README.md`, whose open items moved to Future Steps. This file's
Workflow branches from `next` and opens the PR against `next`, Status says
where the repo stands, and Next Step and Future Steps hold only open work,
linked to its issue where one exists. The viewport harness README names Node's
test runner, not `vitest`
- 2026-10-04: in `src/main.js` (issue #102), a target's min, max, median and
average latency come from one list of its answers, through the same function
the summary's figures use, so the median is written once. The latency color
limits are one table in `CONFIG`, read by both the figure's and the
sparkline's color. The health thresholds, the debug log's length, the gateway
check's timeout, the recovery probe's number of hosts and interval, how often
the rows are sorted and the delay before the first sparkline resize are
`CONFIG` entries too. A unit test checks the summary's figures. Nothing the
page does or shows changed; the footer's color legend still writes the limits
out as text
- 2026-10-04: password guesses at `/metrics` are rate limited (issue #104): each
client address, resolved through `TRUSTED_PROXIES` as for reports, may make 60
requests to `/metrics` a minute, counted by `go-chi/httprate` apart from its
reports; past that it gets 429 and its basic auth credentials are not checked.
The limit is a constant in `backend/internal/server/routes.go`. A test uses up
one client's allowance on wrong passwords, gets 429 with the right one, and
checks that another client behind the same nginx still gets in
- 2026-10-04: the backend reports errors to Sentry (issue #95). With
`SENTRY_DSN` set, it sets up `sentry-go` with the release `netwatch-server-`
and its version, reports each panic in a handler through `sentryhttp`, the
last of the middleware every request goes through, which panics again so the
request still gets the 500 from the panic recovery, and waits up to 2 seconds
on shutdown for Sentry to finish sending. A DSN Sentry refuses stops the start
with an error naming `SENTRY_DSN`. With it empty, Sentry is not set up and
nothing is sent to it
- 2026-10-04: a target's name and URL and a debug log message show as the
characters they are and are never read as HTML (issue #29): a host row escapes
the name and URL it writes into its markup, and the debug log sets each line
as text. A unit test checks a target whose name and URL hold `<`, `>`, `"`,
`&` and `'`. `README.md` no longer calls `CONFIG` frozen: the interval menu
sets its `updateInterval`, and the values computed from it follow. `AppState`
declares the recovery probe's two properties, the sparkline axis functions
lose the parameters they did not use, and the comment on a target's history
names both kinds of entry it holds. Nothing the page does changed
- 2026-10-04: a dependency can be added without running yarn or go by hand
(issue #45): `make add-dependency PACKAGE=<name>@<version>` shims to the new
`script/add-dependency`, which runs `yarn add --dev`, so `package.json` and
`yarn.lock` change together, then `yarn install --frozen-lockfile`; the same
command moves a package to another version. `make tidy` shims to the new
`script/tidy`, which runs `go mod tidy` in `backend/`: a Go module is added by
importing it, or moved by editing its `require` line, then `make tidy`.
`script/bootstrap` still installs with `--frozen-lockfile`
- 2026-10-04: the backend serves Prometheus metrics (issue #94). With
`METRICS_USERNAME` and `METRICS_PASSWORD` both set, it records request
duration and response size through `go-http-metrics` and serves them, with
Go's runtime and process metrics, at `GET /metrics` behind basic auth with
those credentials; nginx passes `/metrics` to it as it does `/api/`. With
neither set there are no metrics and `/metrics` is 404; one without the other
stops the start with an error naming both, and so does a `METRICS_USERNAME`
containing `:`, with an error naming it. Only requests that reach the health
check or `POST /api/v1/reports` are recorded, not `/metrics` itself and not
every request as `GO_HTTP_SERVER_CONVENTIONS.md` shows, because the labels are
the request's path and method, which clients can make up without end. For
that, `POST /api/v1/reports` is now registered by its full path instead of
inside a `/api/v1` route group; it answers as before
- 2026-10-04: `script/` and `Makefile` follow the org models (issue #28):
`make dev` shims to the new `script/dev`, the Vite dev server, and the new
`make build` to `script/build`, the frontend production build.
`.prettierignore` no longer leaves out `backend/`, so `make fmt` and
`make fmt-check` cover `backend/README.md`; it leaves out
`backend/.golangci.yml` by name, the org standard file whose sha256
`backend/script/lint` checks. `script/install-precommit` and the date on
`script/bootstrap`'s pins are the org model again; `script/bootstrap`,
`script/fmt` and `script/fmt-check` each say in a comment why they differ from
it
- 2026-10-04: a frontend build on Node 26 or newer, such as `make test` on a
host with Node 26, no longer prints Node's warning that `module.register()` is
deprecated (issue #32); the build in `Dockerfile` runs on Node 22, which never
printed it. The call was in `@tailwindcss/node`, which `@tailwindcss/vite`
brings in at its own exact version; tailwind 4.3.1 calls
`module.registerHooks()` instead where Node has it. `yarn.lock` now has
`@tailwindcss/vite` and `tailwindcss` at 4.3.3, inside the ranges
`package.json` already allowed, and tailwind's own dependencies moved with
them. The built CSS changes only in how it is written out, in tailwind's
Firefox focus-ring rule, which no longer applies to iframes (the page has
none, so nothing on it looks different), and in tailwind's default sans-serif
font list, which the page does not use: `body` sets a monospace font
- 2026-10-03: the frontend's unit tests cover what the page computes (issue
#21): `humanDuration`, the latency colours of a figure and of a sparkline
either side of each boundary, a target's min, max, average and median latency
over an empty history, an all-unreachable one and a mixed one, and each of the
four health states either side of its thresholds. `package.json` has a `test`
script, so `yarn run test` and `npm run test` run them, and
`script/frontend-test` runs it: quietly, and if a test fails, again with every
test listed, and then fails. `src/main.js` now exports `humanDuration`,
`HostState`, `latencyClass` and `latencyHex` for the tests; nothing it does
changed
- 2026-10-03: the backend's logs are one stream (issue #27): fx logs its own
steps of starting and stopping through the backend's logger, so off a terminal
every line the backend's own logger and fx write is JSON, where fx used to
write plain text to stderr. A config file that is found but cannot be read now
stops the start with its error, logged as JSON like a bad setting, where it
used to end in a Go panic. A test runs the server as a child process and
checks both. The backend logs its name, version and architecture once at
start. The health check's uptime keys are now `uptime_seconds` and
`uptime_human`; its path, content type, `"status":"ok"` and 200 are unchanged.
`SENTRY_DSN`, `METRICS_USERNAME` and `METRICS_PASSWORD` are still read and
still unused, and left out of `backend/README.md`, until issues #94 and #95
wire them up
- 2026-10-03: the frontend has a real linter (issue #47, and item 2 of issue
#28): `eslint` with its recommended rules, set in `eslint.config.js`, runs in
a new `frontend-lint` stage of `Dockerfile`, which the frontend stage waits
on, as the builder stage waits on the Go `lint` stage. `script/lint` builds
both stages without the cache and runs no linter on the host.
`script/frontend-lint` runs eslint where it used to repeat the prettier check
that `script/fmt-check` runs on the host, and `script/frontend-check`, run by
the frontend stage, is now the tests and the format check. `script/bootstrap`
wants node 22.13.0 or newer, as eslint 10 does
- 2026-10-03: the tap-target check in `make frontend-viewport-test` expects one
visible pin button per WAN host row (issue #46), where it expected at least 10
of the 26, so pin buttons missing from only some rows now fail it. The host
row count the harness gathers, which the `app-rendered` check also reads, now
counts only the WAN host rows: the local host rows have no pin button
- 2026-10-03: each target's row shows its result as soon as its check ends
(issue #91), where every row waited for the round's slowest check, up to 24
seconds at a 30-second interval. Every row is still redrawn, and sorting, the
summary, the health box and offline detection still run, once, when the
round's last check ends, so no row reads "paused" after a pause and resume
during the round. A check that ends after the user pauses or after its round
is given up shows nothing, and the first round is still discarded as a whole
- 2026-10-03: root no longer acts outside `/data` when it prepares `DATA_DIR`
(issue #80): `bin/entrypoint.sh` runs `netwatch-server prepare-data-dir`,
which refuses a `DATA_DIR` that is not `/data` or a path below it written in
full, then creates `DATA_DIR`, gives `/data` and everything in it to
`netwatch` and sets the modes, all through a Go `os.Root` opened on `/data`.
That refuses any path leading out of `/data`, so neither a symbolic link
already there nor one a host process swaps in during the start can make root
create or change anything elsewhere, and `DATA_DIR=/etc` no longer gives
`/etc` to `netwatch`. The `README.md` section "Running under upaas" says which
values are accepted
- 2026-10-03: `DATA_DIR_MAX_BYTES` is now how much of the report files is kept
(issue #54): when a report would take them past it, the oldest report files
are deleted to make room, each deletion logged, and at start files already
past it are deleted the same way. A file still being written is never deleted.
A report is refused with 507 only when the reports waiting to be written fill
the cap on their own, and then no file is deleted. The reports of a failed
write stop counting, and the part of its file written is removed. A file that
cannot be deleted still counts until the next start; one already deleted by
hand counts as freed
- 2026-10-03: `backend/script/lint` says what went wrong with its
`.golangci.yml` check (issue #34). On a hash mismatch it says to compare the
file with the org standard: if they differ, restore the org standard; if they
are the same, the org standard changed, so update `GOLANGCI_CONFIG_SHA256` in
that script. It used to say only to restore the file, which loops once the org
standard itself has moved. A missing `.golangci.yml`, and a `sha256sum` that
is missing or prints no hash, each get their own message instead of being
reported as a mismatch; every one still fails the lint
- 2026-10-03: the Go tests run with the race detector and coverage (issue #88):
`backend/script/test` runs `go test -timeout 30s -race -cover ./...` and, if
that fails, runs it again with `-v` and fails. Go's `-timeout` bounds the
tests, not their compile; the root `script/test` no longer puts one 30-second
timeout around both halves, which a cold Go build cache could use up on
compiling alone. The race detector needs a C compiler: the builder stage of
`Dockerfile` has gcc and musl-dev, and `script/bootstrap` installs gcc, with
the C library headers on apt and apk, when gcc is missing; the binary is still
built with `CGO_ENABLED=0`. New tests cover the health check's answer, a valid
report's answer, a report file's exact contents, and the flush when the buffer
reaches 10 MiB; the handlers' `TestImport` stub is gone
- 2026-10-03: each target check times out after 80% of the refresh interval
(issue #78), 24 seconds at 30 seconds, where it was capped at 3 seconds. A
round started early, after an interval change or when the recovery probe finds
@@ -192,9 +365,11 @@ latest run passes.
- 2026-08-09: automated responsive-layout harness
(`make frontend-viewport-test`): digest-pinned headless Chrome driven over CDP
against the built `dist/`, viewport widths derived from the breakpoints in
`src/styles.css` ([#13](https://git.eeqj.de/sneak/netwatch/issues/13)). Every
check carries a presence guard so none of them can pass against a page it is
not actually measuring. Found two real layout defects, filed as
`src/styles.css` ([#13](https://git.eeqj.de/sneak/netwatch/issues/13)). The
tap-target and host-row checks each fail when they measured nothing; the
overflow, viewport-edge and clipped-text checks have no such guard of their
own and rely on the `app-rendered` check, which fails the run when the app did
not render. Found two real layout defects, filed as
[#42](https://git.eeqj.de/sneak/netwatch/issues/42) and
[#43](https://git.eeqj.de/sneak/netwatch/issues/43)
- 2026-07-07 Adopted scripts-to-rule-them-all: `script/` entrypoints, Makefile
@@ -214,12 +389,14 @@ latest run passes.
# Future Steps
- Wire `script/frontend-viewport-test` into CI as its own step (deliberately not
part of `make check` today; the decision has real CI-runtime cost and is
tracked separately)
- Compliance top-up as one small commit: add .editorconfig and add the hooks
target to the Makefile
- After merge, confirm .gitea/workflows/check.yml is on main and CI is green
(main always green policy)
- Decide what to do with untracked resume.sh: commit it, gitignore it, or delete
it
- Run `make frontend-viewport-test` in CI as its own step; it is not part of
`make check`, as it needs Docker and takes minutes
- A backend test that posts a report to `POST /api/v1/reports` and checks the
compressed file it is written to
- A backend route that decompresses the stored reports and answers queries on
them
- Prometheus metrics for the backend's in-memory buffer: its size, the number of
flushes and the number of reports
- A configurable host list (an environment variable or a config file)
- Export of the latency history (CSV or JSON)
- A notification when the health status changes to DEGRADED
+87 -37
View File
@@ -32,14 +32,16 @@ pattern as the repo root: the targets in `backend/Makefile` are thin shims over
`test`, `fmt` and `fmt-check`:
- `script/build` — compile the static `netwatch-server` binary with its version
stamped in. The version is `VERSION` from the environment;
when that is unset or empty, it falls back to `git describe` inside a git
checkout, then to `dev`
- `script/test` — run the Go tests under a 30-second timeout
stamped in. The version is `VERSION` from the environment; when that is unset
or empty, it falls back to `git describe` inside a git checkout, then to `dev`
- `script/test` — run the Go tests with the race detector and coverage. Go's
`-timeout 30s` bounds the tests, not their compile. If they fail, they run
again with `-v` for the details, and the script fails. The race detector needs
a C compiler
- `script/lint` — check `.golangci.yml` against its pinned sha256, then run
golangci-lint. It runs inside the golangci-lint image of the lint stage of
the root `Dockerfile`; from a checkout, run `make lint` at the repo root,
which builds that stage
golangci-lint. It runs inside the golangci-lint image of the lint stage of the
root `Dockerfile`; from a checkout, run `make lint` at the repo root, which
builds that stage
- `script/fmt` — format the Go sources (writes)
- `script/fmt-check` — check Go formatting (read-only)
- `script/run` — build and run the server locally
@@ -58,8 +60,9 @@ flushes them to compressed files on disk for later analysis.
## Design
The server is structured as an `fx`-wired Go application under `cmd/netwatch-server/`.
Internal packages in `internal/` follow standard Go project layout:
The server is structured as an `fx`-wired Go application under
`cmd/netwatch-server/`. Internal packages in `internal/` follow standard Go
project layout:
- **`config`**: Loads configuration from environment variables and config files
via Viper.
@@ -80,17 +83,21 @@ Internal packages in `internal/` follow standard Go project layout:
| `BIND_ADDRESS` | empty | IP address to listen on; empty listens on every interface |
| `PORT` | `8080` | HTTP listen port |
| `DATA_DIR` | `./data/reports` | Directory for compressed reports |
| `DATA_DIR_MAX_BYTES` | `1073741824` (1 GiB) | Largest total size of the report files in `DATA_DIR`; see [Report limits](#report-limits) |
| `DATA_DIR_MAX_BYTES` | `1073741824` (1 GiB) | Most bytes of report files kept in `DATA_DIR`, oldest deleted first; see [Report limits](#report-limits) |
| `DEBUG` | `false` | Enable debug logging |
| `TRUSTED_PROXIES` | loopback + RFC1918 | Comma-separated CIDRs whose `X-Forwarded-For` / `X-Real-IP` headers are trusted for client IP resolution |
| `REPORTS_PER_MINUTE` | `60` | Reports each client address may send a minute; see [Report limits](#report-limits) |
| `CORS_ALLOWED_ORIGINS` | empty | Comma-separated origins whose pages may call the API; see [CORS](#cors) |
| `METRICS_USERNAME` | empty | Basic auth user name for `/metrics`; see [Metrics](#metrics) |
| `METRICS_PASSWORD` | empty | Basic auth password for `/metrics`; see [Metrics](#metrics) |
| `SENTRY_DSN` | empty | DSN of the Sentry project to send errors to; see [Sentry](#sentry) |
`TRUSTED_PROXIES` defaults to `127.0.0.1/32,::1/128,10.0.0.0/8,172.16.0.0/12,192.168.0.0/16`.
The loopback entries cover a reverse proxy on the same host. A request whose
direct peer is outside this set has its forwarded headers ignored, and the
direct peer is logged and rate-limited instead. The container image does not use
this default; see [Container image](#container-image).
`TRUSTED_PROXIES` defaults to
`127.0.0.1/32,::1/128,10.0.0.0/8,172.16.0.0/12,192.168.0.0/16`. The loopback
entries cover a reverse proxy on the same host. A request whose direct peer is
outside this set has its forwarded headers ignored, and the direct peer is
logged and rate-limited instead. The container image does not use this default;
see [Container image](#container-image).
A variable set to a value the server cannot use, such as `PORT=abc`,
`DEBUG=maybe` or a `BIND_ADDRESS` that is not an IP address, stops it from
@@ -99,15 +106,16 @@ starting, with an error naming the variable. An empty variable counts as unset.
### Container image
The root `Dockerfile` builds one image in which nginx listens on the public port
8080, serves the frontend, and proxies `/api/` and `/.well-known/healthcheck` to
this server. The image's entrypoint, `bin/entrypoint.sh`, starts the server as
user `netwatch` (uid 1000) with `BIND_ADDRESS=127.0.0.1` and `PORT=8081`, so
only nginx reaches it, and with `TRUSTED_PROXIES=127.0.0.1/32`, so it takes the
client address nginx passes on and no other. `DATA_DIR` is `/data/reports`, on
the `/data` volume; the entrypoint creates it and gives it and `/data` to
`netwatch` before starting the server. nginx replaces the security headers
this server sets with those in the root `security-headers.conf`, so those are
what clients of the image see.
8080, serves the frontend, and proxies `/api/`, `/.well-known/healthcheck` and
`/metrics` to this server. The image's entrypoint, `bin/entrypoint.sh`, starts
the server as user `netwatch` (uid 1000) with `BIND_ADDRESS=127.0.0.1` and
`PORT=8081`, so only nginx reaches it, and with `TRUSTED_PROXIES=127.0.0.1/32`,
so it takes the client address nginx passes on and no other. `DATA_DIR` is
`/data/reports`, on the `/data` volume; before starting the server, the
entrypoint creates it and gives it and `/data` to `netwatch` with
`netwatch-server prepare-data-dir`, which acts on nothing outside `/data`. nginx
replaces the security headers this server sets with those in the root
`security-headers.conf`, so those are what clients of the image see.
The container's own `TRUSTED_PROXIES` goes to nginx instead: IP addresses or
CIDRs, separated by commas, of the reverse proxies in front of the container.
@@ -125,10 +133,11 @@ Reports are written as `reports-<timestamp>-<number>.jsonl.zst` files in
`DATA_DIR`. The timestamp is in UTC to the millisecond, so the names sort by
time. The number starts at 1 when the server starts and goes up by one for each
file the server starts to write, so two files written in the same millisecond
still get different names. A failed write uses up its number, leaving a gap in
the numbers if the file could not be created and otherwise a file under that
number that may be incomplete. Each file contains one JSON object per line,
compressed with zstd. Files are created with `O_EXCL` to prevent overwrites.
still get different names. A failed write uses up its number and leaves a gap in
the numbers: its file, if it was created, is removed. The file stays, counted
toward `DATA_DIR_MAX_BYTES` from the next start, only if removing it fails too.
Each file contains one JSON object per line, compressed with zstd. Files are
created with `O_EXCL` to prevent overwrites.
### Report limits
@@ -148,11 +157,16 @@ credentials, so it is bounded instead. Both refusals below answer with the same
`X-RateLimit-Reset` headers.
- **Size cap.** The report files in `DATA_DIR` may total at most
`DATA_DIR_MAX_BYTES`, counting the files already there at start. Reports
waiting in memory count at their uncompressed size until they are written, so
a report that would take the total past the cap is refused with 507, and
nothing of it is stored. Deleting report files frees room only at the next
start, when the files are counted again. The default of 1 GiB is small enough
for any host; set it to the space you can give `DATA_DIR`.
waiting in memory count at their uncompressed size until they are written;
those lost to a failed write stop counting, and the part of its file written
is removed. When a report would take the total past the cap, the oldest report
files are deleted to make room, and each deletion is logged with the file's
name and size; a file still being written is never deleted. A report is
refused with 507, and nothing of it is stored, only when the reports waiting
to be written fill the cap on their own, and then no file is deleted. At
start, report files past the cap, as after lowering it, are deleted the same
way. So the cap is how much of the newest reports is kept: the default of 1
GiB is small enough for any host; set it to the space you can give `DATA_DIR`.
### CORS
@@ -165,12 +179,48 @@ origin, `scheme://host` with an optional `:port`, as browsers send it: no path,
not even a trailing `/`, and no `*`. Any other entry stops the server from
starting, with an error naming `CORS_ALLOWED_ORIGINS`.
### Metrics
With both `METRICS_USERNAME` and `METRICS_PASSWORD` set, the server serves
Prometheus metrics at `GET /metrics` to requests with those as their basic auth
credentials, and answers any other with 401. For each request that reaches the
health check or `POST /api/v1/reports`, those the rate limit refuses included,
the metrics record its duration and response size, labelled with its path,
method and status; they also count those requests in progress, and include Go's
runtime and process metrics. No other request is recorded: not those to
`/metrics` itself, and not those answered before they reach either route, such
as a CORS preflight, or a request refused with 404 for a path no route has, 405
for a method its route does not take, or 413 for declaring a body length over
the 1 MiB limit. A report whose body goes over the limit without declaring its
length reaches the route, is answered 413 there, and is recorded with that
status. Clients can make up any number of paths and methods, and each would add
labels to the metrics for as long as the server runs. With neither set, nothing
is recorded and `/metrics` answers 404. One without the other stops the server
from starting, with an error naming both; so does a `METRICS_USERNAME`
containing `:`, which basic auth cannot carry, with an error naming it.
`/metrics` is rate limited, so that its password cannot be guessed quickly: each
client address, resolved through `TRUSTED_PROXIES`, may make 60 requests to it a
minute, whatever their credentials. Past that it gets 429 with
`Retry-After: 60`, and its credentials are not checked. The minute slides as it
does for reports (see [Report limits](#report-limits)), so a scraper polling
every 2 seconds or less often is never refused. This allowance is apart from the
one for reports.
### Sentry
With `SENTRY_DSN` set, the server sends its errors to that Sentry project: each
panic in a handler is reported there, under the release `netwatch-server-`
followed by the server's version, and the request still gets 500 from the
server's panic recovery. On shutdown the server waits up to 2 seconds for Sentry
to finish sending. A DSN Sentry refuses stops the server from starting, with an
error naming `SENTRY_DSN`. With it empty, Sentry is not set up, and nothing is
sent to it.
## TODO
- Add integration test that POSTs a report and verifies the compressed output
- Add report decompression/query endpoint
- Add metrics (Prometheus) for buffer size, flush count, report count
- Add retention policy to prune old report files
The to-do list, this backend's open work included, is [TODO.md](../TODO.md) at
the repo root.
## License
+34 -1
View File
@@ -4,6 +4,7 @@ package main
import (
"fmt"
"os"
"os/user"
"sneak.berlin/go/netwatch/internal/config"
"sneak.berlin/go/netwatch/internal/globals"
@@ -15,6 +16,7 @@ import (
"sneak.berlin/go/netwatch/internal/server"
"go.uber.org/fx"
"go.uber.org/fx/fxevent"
)
//nolint:gochecknoglobals // set via ldflags at build time
@@ -37,10 +39,36 @@ func main() {
return
}
// "netwatch-server prepare-data-dir DATA_DIR" gets DATA_DIR ready
// for the netwatch user, or exits 1 with the error; see
// reportbuf.PrepareDataDir. bin/entrypoint.sh runs it as root
// before it starts this server as that user.
if len(os.Args) == 3 && os.Args[1] == "prepare-data-dir" {
netwatch, err := user.Lookup("netwatch")
if err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
err = reportbuf.PrepareDataDir("/data", os.Args[2], netwatch)
if err != nil {
fmt.Fprintf(os.Stderr, "DATA_DIR '%s': %v\n", os.Args[2], err)
os.Exit(1)
}
return
}
globals.Appname = Appname
globals.Version = Version
fx.New(
// fx logs each step of starting and stopping through the
// server's own logger, so off a terminal those lines are
// JSON like every other line.
fx.WithLogger(func(log *logger.Logger) fxevent.Logger {
return &fxevent.SlogLogger{Logger: log.Get()}
}),
fx.Provide(
config.New,
globals.New,
@@ -51,6 +79,11 @@ func main() {
reportbuf.New,
server.New,
),
fx.Invoke(func(*server.Server) {}),
fx.Invoke(
// First, so the name and version are logged even when
// a setting stops the start.
func(log *logger.Logger) { log.Identify() },
func(*server.Server) {},
),
).Run()
}
+249
View File
@@ -0,0 +1,249 @@
package main
import (
"bytes"
"context"
"encoding/json"
"net"
"net/http"
"os"
"os/exec"
"os/signal"
"path/filepath"
"strings"
"syscall"
"testing"
"time"
)
// The tests run main() in a child process, this test binary started
// again with runMainEnv set, because main() can exit its process and
// takes its settings from the environment.
const runMainEnv = "NETWATCH_SERVER_RUN_MAIN"
// childTimeout bounds each child's whole run; it is killed after it.
const childTimeout = 10 * time.Second
func TestMain(m *testing.M) {
if os.Getenv(runMainEnv) != "" {
// A SIGTERM that comes before fx catches it is dropped, not fatal.
signal.Notify(make(chan os.Signal, 1), syscall.SIGTERM)
main()
return
}
os.Exit(m.Run())
}
// TestOutputIsJSON: off a terminal, every line the server writes from
// start to stop is JSON, fx's own lines included.
func TestOutputIsJSON(t *testing.T) {
t.Parallel()
ctx, cancel := context.WithTimeout(t.Context(), childTimeout)
defer cancel()
port := freePort(ctx, t)
child, stdout, stderr := startServer(ctx, t, t.TempDir(), port)
waitForHealthcheck(ctx, t, port)
// The child drops a SIGTERM that comes before fx catches it (see
// TestMain), so send one every 100ms until the test ends. ctx
// bounds the wait: when it ends, the child is killed.
stop := make(chan struct{})
defer close(stop)
go func() {
for {
_ = child.Process.Signal(syscall.SIGTERM)
select {
case <-stop:
return
case <-time.After(100 * time.Millisecond):
}
}
}()
err := child.Wait()
if err != nil {
t.Fatalf("server exit = %v, want success", err)
}
requireJSONLines(t, stdout, stderr)
if !strings.Contains(stdout.String(), `"msg":"starting"`) {
t.Fatalf("no startup line in stdout:\n%s", stdout)
}
}
// TestMalformedConfigFileStopsTheStart: a config file the server finds
// but cannot read stops the start, and the error is logged as JSON.
func TestMalformedConfigFileStopsTheStart(t *testing.T) {
t.Parallel()
ctx, cancel := context.WithTimeout(t.Context(), childTimeout)
defer cancel()
home := t.TempDir()
dir := filepath.Join(home, ".config", "netwatch-server")
err := os.MkdirAll(dir, 0o750)
if err != nil {
t.Fatal(err)
}
err = os.WriteFile(filepath.Join(dir, "netwatch-server.yaml"),
[]byte("PORT: [8080\n"), 0o600)
if err != nil {
t.Fatal(err)
}
child, stdout, stderr := startServer(ctx, t, home, freePort(ctx, t))
err = child.Wait()
if child.ProcessState.ExitCode() != 1 {
t.Fatalf("server exit = %v, want exit status 1", err)
}
requireJSONLines(t, stdout, stderr)
if !strings.Contains(stdout.String(), "netwatch-server.yaml") {
t.Fatalf("no error naming the config file in stdout:\n%s", stdout)
}
}
// TestRefusedSentryDSNStopsTheStart: a SENTRY_DSN that Sentry refuses
// stops the start, and the error, naming SENTRY_DSN, is logged as JSON.
func TestRefusedSentryDSNStopsTheStart(t *testing.T) {
t.Setenv("SENTRY_DSN", "not-a-dsn")
ctx, cancel := context.WithTimeout(t.Context(), childTimeout)
defer cancel()
child, stdout, stderr := startServer(ctx, t, t.TempDir(), freePort(ctx, t))
err := child.Wait()
if child.ProcessState.ExitCode() != 1 {
t.Fatalf("server exit = %v, want exit status 1", err)
}
requireJSONLines(t, stdout, stderr)
if !strings.Contains(stdout.String(), "SENTRY_DSN") {
t.Fatalf("no error naming SENTRY_DSN in stdout:\n%s", stdout)
}
}
// startServer runs main() in a child process listening on
// 127.0.0.1:port, with home as its HOME and working directory and its
// data directory in home, so it touches nothing outside home. Its
// stdout and stderr go to the two buffers returned, which hold all of
// it once child.Wait returns. The child is killed when ctx ends, and
// killed and reaped when the test ends if nothing waited for it.
func startServer(
ctx context.Context,
t *testing.T,
home, port string,
) (*exec.Cmd, *bytes.Buffer, *bytes.Buffer) {
t.Helper()
self, err := os.Executable()
if err != nil {
t.Fatal(err)
}
var stdout, stderr bytes.Buffer
child := exec.CommandContext(ctx, self) //nolint:gosec // this test binary
child.Dir = home
child.Env = append(os.Environ(),
runMainEnv+"=1",
"HOME="+home,
"DATA_DIR="+filepath.Join(home, "data"),
"BIND_ADDRESS=127.0.0.1",
"PORT="+port,
)
child.Stdout = &stdout
child.Stderr = &stderr
err = child.Start()
if err != nil {
t.Fatal(err)
}
t.Cleanup(func() {
if child.ProcessState == nil {
_ = child.Process.Kill()
_ = child.Wait()
}
})
return child, &stdout, &stderr
}
// freePort returns a TCP port on 127.0.0.1 that was free a moment ago.
func freePort(ctx context.Context, t *testing.T) string {
t.Helper()
var lc net.ListenConfig
l, err := lc.Listen(ctx, "tcp", "127.0.0.1:0")
if err != nil {
t.Fatal(err)
}
_ = l.Close()
_, port, err := net.SplitHostPort(l.Addr().String())
if err != nil {
t.Fatal(err)
}
return port
}
// waitForHealthcheck returns once the health check on port answers 200.
func waitForHealthcheck(ctx context.Context, t *testing.T, port string) {
t.Helper()
url := "http://127.0.0.1:" + port + "/.well-known/healthcheck"
for {
req, err := http.NewRequestWithContext(ctx, http.MethodGet, url, nil)
if err != nil {
t.Fatal(err)
}
resp, err := http.DefaultClient.Do(req)
if err == nil {
_ = resp.Body.Close()
if resp.StatusCode == http.StatusOK {
return
}
}
select {
case <-ctx.Done():
t.Fatalf("health check never answered: %v", err)
case <-time.After(50 * time.Millisecond):
}
}
}
// requireJSONLines fails the test on each line of outs that is not
// JSON.
func requireJSONLines(t *testing.T, outs ...*bytes.Buffer) {
t.Helper()
for _, out := range outs {
for line := range strings.Lines(out.String()) {
if !json.Valid([]byte(line)) {
t.Errorf("line is not JSON: %s", line)
}
}
}
}
+14 -3
View File
@@ -3,20 +3,30 @@ module sneak.berlin/go/netwatch
go 1.25.5
require (
github.com/99designs/basicauth-go v0.0.0-20230316000542-bf6f9cbbf0f8
github.com/getsentry/sentry-go v0.49.0
github.com/go-chi/chi/v5 v5.2.5
github.com/go-chi/cors v1.2.2
github.com/go-chi/httprate v0.16.0
github.com/joho/godotenv v1.5.1
github.com/klauspost/compress v1.18.4
github.com/klauspost/compress v1.19.1
github.com/prometheus/client_golang v1.24.1
github.com/slok/go-http-metrics v0.13.0
github.com/spf13/viper v1.21.0
go.uber.org/fx v1.24.0
)
require (
github.com/beorn7/perks v1.0.1 // indirect
github.com/cespare/xxhash/v2 v2.3.0 // indirect
github.com/fsnotify/fsnotify v1.9.0 // indirect
github.com/go-viper/mapstructure/v2 v2.4.0 // indirect
github.com/klauspost/cpuid/v2 v2.2.10 // indirect
github.com/munnerz/goautoneg v0.0.0-20191010083416-a7dc8b61c822 // indirect
github.com/pelletier/go-toml/v2 v2.2.4 // indirect
github.com/prometheus/client_model v0.6.2 // indirect
github.com/prometheus/common v0.70.1 // indirect
github.com/prometheus/procfs v0.21.1 // indirect
github.com/sagikazarmark/locafero v0.11.0 // indirect
github.com/sourcegraph/conc v0.3.1-0.20240121214520-5f936abd7ae8 // indirect
github.com/spf13/afero v1.15.0 // indirect
@@ -28,6 +38,7 @@ require (
go.uber.org/multierr v1.10.0 // indirect
go.uber.org/zap v1.26.0 // indirect
go.yaml.in/yaml/v3 v3.0.4 // indirect
golang.org/x/sys v0.30.0 // indirect
golang.org/x/text v0.28.0 // indirect
golang.org/x/sys v0.47.0 // indirect
golang.org/x/text v0.40.0 // indirect
google.golang.org/protobuf v1.36.11 // indirect
)
+52 -18
View File
@@ -1,37 +1,65 @@
github.com/davecgh/go-spew v1.1.1 h1:vj9j/u1bqnvCEfJOwUhtlOARqs3+rkHYY13jYWTU97c=
github.com/davecgh/go-spew v1.1.1/go.mod h1:J7Y8YcW2NihsgmVo/mv3lAwl/skON4iLHjSsI+c5H38=
github.com/99designs/basicauth-go v0.0.0-20230316000542-bf6f9cbbf0f8 h1:nMpu1t4amK3vJWBibQ5X/Nv0aXL+b69TQf2uK5PH7Go=
github.com/99designs/basicauth-go v0.0.0-20230316000542-bf6f9cbbf0f8/go.mod h1:3cARGAK9CfW3HoxCy1a0G4TKrdiKke8ftOMEOHyySYs=
github.com/beorn7/perks v1.0.1 h1:VlbKKnNfV8bJzeqoa4cOKqO6bYr3WgKZxO8Z16+hsOM=
github.com/beorn7/perks v1.0.1/go.mod h1:G2ZrVWU2WbWT9wwq4/hrbKbnv/1ERSJQ0ibhJ6rlkpw=
github.com/cespare/xxhash/v2 v2.3.0 h1:UL815xU9SqsFlibzuggzjXhog7bL6oX9BbNZnL2UFvs=
github.com/cespare/xxhash/v2 v2.3.0/go.mod h1:VGX0DQ3Q6kWi7AoAeZDth3/j3BFtOZR5XLFGgcrjCOs=
github.com/davecgh/go-spew v1.1.2-0.20180830191138-d8f796af33cc h1:U9qPSI2PIWSS1VwoXQT9A3Wy9MM3WgvqSxFWenqJduM=
github.com/davecgh/go-spew v1.1.2-0.20180830191138-d8f796af33cc/go.mod h1:J7Y8YcW2NihsgmVo/mv3lAwl/skON4iLHjSsI+c5H38=
github.com/frankban/quicktest v1.14.6 h1:7Xjx+VpznH+oBnejlPUj8oUpdxnVs4f8XU8WnHkI4W8=
github.com/frankban/quicktest v1.14.6/go.mod h1:4ptaffx2x8+WTWXmUCuVU6aPUX1/Mz7zb5vbUoiM6w0=
github.com/fsnotify/fsnotify v1.9.0 h1:2Ml+OJNzbYCTzsxtv8vKSFD9PbJjmhYF14k/jKC7S9k=
github.com/fsnotify/fsnotify v1.9.0/go.mod h1:8jBTzvmWwFyi3Pb8djgCCO5IBqzKJ/Jwo8TRcHyHii0=
github.com/getsentry/sentry-go v0.49.0 h1:Ehejknu1l023Ub7QoRBVLAI7g3Jnhqku4oWx4B4Sh5s=
github.com/getsentry/sentry-go v0.49.0/go.mod h1:nuMJAoCfe1u0Bts2ocyNI+TW8HT84vRMqwA5Qq/SKUI=
github.com/go-chi/chi/v5 v5.2.5 h1:Eg4myHZBjyvJmAFjFvWgrqDTXFyOzjj7YIm3L3mu6Ug=
github.com/go-chi/chi/v5 v5.2.5/go.mod h1:X7Gx4mteadT3eDOMTsXzmI4/rwUpOwBHLpAfupzFJP0=
github.com/go-chi/cors v1.2.2 h1:Jmey33TE+b+rB7fT8MUy1u0I4L+NARQlK6LhzKPSyQE=
github.com/go-chi/cors v1.2.2/go.mod h1:sSbTewc+6wYHBBCW7ytsFSn836hqM7JxpglAy2Vzc58=
github.com/go-chi/httprate v0.16.0 h1:8V5DH9j6pSK6UQoBsTpvMyFxycqaKEIToyPKzHJjUa8=
github.com/go-chi/httprate v0.16.0/go.mod h1:A8lo+qRhk+s9LiuP5saS7XCGDXRXMcrueq0NfIuCa/I=
github.com/go-errors/errors v1.4.2 h1:J6MZopCL4uSllY1OfXM374weqZFFItUbrImctkmUxIA=
github.com/go-errors/errors v1.4.2/go.mod h1:sIVyrIiJhuEF+Pj9Ebtd6P/rEYROXFi3BopGUQ5a5Og=
github.com/go-viper/mapstructure/v2 v2.4.0 h1:EBsztssimR/CONLSZZ04E8qAkxNYq4Qp9LvH92wZUgs=
github.com/go-viper/mapstructure/v2 v2.4.0/go.mod h1:oJDH3BJKyqBA2TXFhDsKDGDTlndYOZ6rGS0BRZIxGhM=
github.com/google/go-cmp v0.6.0 h1:ofyhxvXcZhMsU5ulbFiLKl/XBFqE1GSq7atu8tAmTRI=
github.com/google/go-cmp v0.6.0/go.mod h1:17dUlkBOakJ0+DkrSSNjCkIjxS6bF9zb3elmeNGIjoY=
github.com/google/go-cmp v0.7.0 h1:wk8382ETsv4JYUZwIsn6YpYiWiBsYLSJiTsyBybVuN8=
github.com/google/go-cmp v0.7.0/go.mod h1:pXiqmnSA92OHEEa9HXL2W4E7lf9JzCmGVUdgjX3N/iU=
github.com/joho/godotenv v1.5.1 h1:7eLL/+HRGLY0ldzfGMeQkb7vMd0as4CfYvUVzLqw0N0=
github.com/joho/godotenv v1.5.1/go.mod h1:f4LDr5Voq0i2e/R5DDNOoa2zzDfwtkZa6DnEwAbqwq4=
github.com/klauspost/compress v1.18.4 h1:RPhnKRAQ4Fh8zU2FY/6ZFDwTVTxgJ/EMydqSTzE9a2c=
github.com/klauspost/compress v1.18.4/go.mod h1:R0h/fSBs8DE4ENlcrlib3PsXS61voFxhIs2DeRhCvJ4=
github.com/klauspost/compress v1.19.1 h1:VsB4HPswih7mmZ8WleSFQ75c/Ui1M4trX5oAsJnhSlk=
github.com/klauspost/compress v1.19.1/go.mod h1:cwPg85FWrGar70rWktvGQj8/hthj3wpl0PGDogxkrSQ=
github.com/klauspost/cpuid/v2 v2.2.10 h1:tBs3QSyvjDyFTq3uoc/9xFpCuOsJQFNPiAhYdw2skhE=
github.com/klauspost/cpuid/v2 v2.2.10/go.mod h1:hqwkgyIinND0mEev00jJYCxPNVRVXFQeu1XKlok6oO0=
github.com/kr/pretty v0.3.1 h1:flRD4NNwYAUpkphVc1HcthR4KEIFJ65n8Mw5qdRn3LE=
github.com/kr/pretty v0.3.1/go.mod h1:hoEshYVHaxMs3cyo3Yncou5ZscifuDolrwPKZanG3xk=
github.com/kr/text v0.2.0 h1:5Nx0Ya0ZqY2ygV366QzturHI13Jq95ApcVaJBhpS+AY=
github.com/kr/text v0.2.0/go.mod h1:eLer722TekiGuMkidMxC/pM04lWEeraHUUmBw8l2grE=
github.com/kylelemons/godebug v1.1.0 h1:RPNrshWIDI6G2gRW9EHilWtl7Z6Sb1BR0xunSBf0SNc=
github.com/kylelemons/godebug v1.1.0/go.mod h1:9/0rRGxNHcop5bhtWyNeEfOS8JIWk580+fNqagV/RAw=
github.com/munnerz/goautoneg v0.0.0-20191010083416-a7dc8b61c822 h1:C3w9PqII01/Oq1c1nUAm88MOHcQC9l5mIlSMApZMrHA=
github.com/munnerz/goautoneg v0.0.0-20191010083416-a7dc8b61c822/go.mod h1:+n7T8mK8HuQTcFwEeznm/DIxMOiR9yIdICNftLE1DvQ=
github.com/pelletier/go-toml/v2 v2.2.4 h1:mye9XuhQ6gvn5h28+VilKrrPoQVanw5PMw/TB0t5Ec4=
github.com/pelletier/go-toml/v2 v2.2.4/go.mod h1:2gIqNv+qfxSVS7cM2xJQKtLSTLUE9V8t9Stt+h56mCY=
github.com/pmezard/go-difflib v1.0.0 h1:4DBwDE0NGyQoBHbLQYPwSUPoCMWR5BEzIk/f1lZbAQM=
github.com/pmezard/go-difflib v1.0.0/go.mod h1:iKH77koFhYxTK1pcRnkKkqfTogsbg7gZNVY4sRDYZ/4=
github.com/rogpeppe/go-internal v1.9.0 h1:73kH8U+JUqXU8lRuOHeVHaa/SZPifC7BkcraZVejAe8=
github.com/rogpeppe/go-internal v1.9.0/go.mod h1:WtVeX8xhTBvf0smdhujwtBcq4Qrzq/fJaraNFVN+nFs=
github.com/pingcap/errors v0.11.4 h1:lFuQV/oaUMGcD2tqt+01ROSmJs75VG1ToEOkZIZ4nE4=
github.com/pingcap/errors v0.11.4/go.mod h1:Oi8TUi2kEtXXLMJk9l1cGmz20kV3TaQ0usTwv5KuLY8=
github.com/pkg/errors v0.9.1 h1:FEBLx1zS214owpjy7qsBeixbURkuhQAwrK5UwLGTwt4=
github.com/pkg/errors v0.9.1/go.mod h1:bwawxfHBFNV+L2hUp1rHADufV3IMtnDRdf1r5NINEl0=
github.com/pmezard/go-difflib v1.0.1-0.20181226105442-5d4384ee4fb2 h1:Jamvg5psRIccs7FGNTlIRMkT8wgtp5eCXdBlqhYGL6U=
github.com/pmezard/go-difflib v1.0.1-0.20181226105442-5d4384ee4fb2/go.mod h1:iKH77koFhYxTK1pcRnkKkqfTogsbg7gZNVY4sRDYZ/4=
github.com/prometheus/client_golang v1.24.1 h1:JnJkREXzWxUdCuPFpIWZiPispT9xVV59uiuyR2bPlnU=
github.com/prometheus/client_golang v1.24.1/go.mod h1:F+oSRECHg4sse5ucfYpYDeIv/hu68Zo0uoHKetWnzcE=
github.com/prometheus/client_model v0.6.2 h1:oBsgwpGs7iVziMvrGhE53c/GrLUsZdHnqNwqPLxwZyk=
github.com/prometheus/client_model v0.6.2/go.mod h1:y3m2F6Gdpfy6Ut/GBsUqTWZqCUvMVzSfMLjcu6wAwpE=
github.com/prometheus/common v0.70.1 h1:1HvjP4D5oL3t8RsPlwxA9onvvStjtIHYE5XuuwOi/PY=
github.com/prometheus/common v0.70.1/go.mod h1:VdFUQDMZK3VLkurFUVhia6uys/0suUp86TJz5qbJRhc=
github.com/prometheus/procfs v0.21.1 h1:GljZCt+zSTS+NZq88cyQ1LjZ+RCHp3uVuabBWA5+OJI=
github.com/prometheus/procfs v0.21.1/go.mod h1:aB55Cww9pdSJVHk0hUf0inxWyyjPogFIjmHKYgMKmtY=
github.com/rogpeppe/go-internal v1.14.1 h1:UQB4HGPB6osV0SQTLymcB4TgvyWu6ZyliaW0tI/otEQ=
github.com/rogpeppe/go-internal v1.14.1/go.mod h1:MaRKkUm5W0goXpeCfT7UZI6fk/L7L7so1lCWt35ZSgc=
github.com/sagikazarmark/locafero v0.11.0 h1:1iurJgmM9G3PA/I+wWYIOw/5SyBtxapeHDcg+AAIFXc=
github.com/sagikazarmark/locafero v0.11.0/go.mod h1:nVIGvgyzw595SUSUE6tvCp3YYTeHs15MvlmU87WwIik=
github.com/slok/go-http-metrics v0.13.0 h1:lQDyJJx9wKhmbliyUsZ2l6peGnXRHjsjoqPt5VYzcP8=
github.com/slok/go-http-metrics v0.13.0/go.mod h1:HIr7t/HbN2sJaunvnt9wKP9xoBBVZFo1/KiHU3b0w+4=
github.com/sourcegraph/conc v0.3.1-0.20240121214520-5f936abd7ae8 h1:+jumHNA0Wrelhe64i8F6HNlS8pkoyMv5sreGx2Ry5Rw=
github.com/sourcegraph/conc v0.3.1-0.20240121214520-5f936abd7ae8/go.mod h1:3n1Cwaq1E1/1lhQhtRK2ts/ZwZEhjcQeJQ1RuC6Q/8U=
github.com/spf13/afero v1.15.0 h1:b/YBCLWAJdFWJTN9cLhiXXcD7mzKn9Dm86dNnfyQw1I=
@@ -42,6 +70,8 @@ github.com/spf13/pflag v1.0.10 h1:4EBh2KAYBwaONj6b2Ye1GiHfwjqyROoF4RwYO+vPwFk=
github.com/spf13/pflag v1.0.10/go.mod h1:McXfInJRrz4CZXVZOBLb0bTZqETkiAhM9Iw0y3An2Bg=
github.com/spf13/viper v1.21.0 h1:x5S+0EU27Lbphp4UKm1C+1oQO+rKx36vfCoaVebLFSU=
github.com/spf13/viper v1.21.0/go.mod h1:P0lhsswPGWD/1lZJ9ny3fYnVqxiegrlNrEmgLjbTCAY=
github.com/stretchr/objx v0.5.2 h1:xuMeJ0Sdp5ZMRXx/aWO6RZxdr3beISkG5/G/aIRr3pY=
github.com/stretchr/objx v0.5.2/go.mod h1:FRsXN1f5AsAjCGJKqEizvkpNtU+EGNCLh3NxZ/8L+MA=
github.com/stretchr/testify v1.11.1 h1:7s2iGBzp5EwR7/aIZr8ao5+dra3wiQyKjjFuvgVKu7U=
github.com/stretchr/testify v1.11.1/go.mod h1:wZwfW3scLgRK+23gO65QZefKpKQRnfz6sD981Nm4B6U=
github.com/subosito/gotenv v1.6.0 h1:9NlTDc1FTs4qu0DDq7AEtTPNw6SVm7uBMsUCUjABIf8=
@@ -54,20 +84,24 @@ go.uber.org/dig v1.19.0 h1:BACLhebsYdpQ7IROQ1AGPjrXcP5dF80U3gKoFzbaq/4=
go.uber.org/dig v1.19.0/go.mod h1:Us0rSJiThwCv2GteUN0Q7OKvU7n5J4dxZ9JKUXozFdE=
go.uber.org/fx v1.24.0 h1:wE8mruvpg2kiiL1Vqd0CC+tr0/24XIB10Iwp2lLWzkg=
go.uber.org/fx v1.24.0/go.mod h1:AmDeGyS+ZARGKM4tlH4FY2Jr63VjbEDJHtqXTGP5hbo=
go.uber.org/goleak v1.2.0 h1:xqgm/S+aQvhWFTtR0XK3Jvg7z8kGV8P4X14IzwN3Eqk=
go.uber.org/goleak v1.2.0/go.mod h1:XJYK+MuIchqpmGmUSAzotztawfKvYLUIgg7guXrwVUo=
go.uber.org/goleak v1.3.0 h1:2K3zAYmnTNqV73imy9J1T3WC+gmCePx2hEGkimedGto=
go.uber.org/goleak v1.3.0/go.mod h1:CoHD4mav9JJNrW/WLlf7HGZPjdw8EucARQHekz1X6bE=
go.uber.org/multierr v1.10.0 h1:S0h4aNzvfcFsC3dRF1jLoaov7oRaKqRGC/pUEJ2yvPQ=
go.uber.org/multierr v1.10.0/go.mod h1:20+QtiLqy0Nd6FdQB9TLXag12DsQkrbs3htMFfDN80Y=
go.uber.org/zap v1.26.0 h1:sI7k6L95XOKS281NhVKOFCUNIvv9e0w4BF8N3u+tCRo=
go.uber.org/zap v1.26.0/go.mod h1:dtElttAiwGvoJ/vj4IwHBS/gXsEu/pZ50mUIRWuG0so=
go.yaml.in/yaml/v2 v2.4.4 h1:tuyd0P+2Ont/d6e2rl3be67goVK4R6deVxCUX5vyPaQ=
go.yaml.in/yaml/v2 v2.4.4/go.mod h1:gMZqIpDtDqOfM0uNfy0SkpRhvUryYH0Z6wdMYcacYXQ=
go.yaml.in/yaml/v3 v3.0.4 h1:tfq32ie2Jv2UxXFdLJdh3jXuOzWiL1fo0bu/FbuKpbc=
go.yaml.in/yaml/v3 v3.0.4/go.mod h1:DhzuOOF2ATzADvBadXxruRBLzYTpT36CKvDb3+aBEFg=
golang.org/x/sys v0.30.0 h1:QjkSwP/36a20jFYWkSue1YwXzLmsV5Gfq7Eiy72C1uc=
golang.org/x/sys v0.30.0/go.mod h1:/VUhepiaJMQUp4+oa/7Zr1D23ma6VTLIYjOOTFZPUcA=
golang.org/x/text v0.28.0 h1:rhazDwis8INMIwQ4tpjLDzUhx6RlXqZNPEM0huQojng=
golang.org/x/text v0.28.0/go.mod h1:U8nCwOR8jO/marOQ0QbDiOngZVEBB7MAiitBuMjXiNU=
golang.org/x/sys v0.47.0 h1:o7XGOvZQCADBQQ4Y7VNq2dRWQR7JmOUW8Kxx4ZsNgWs=
golang.org/x/sys v0.47.0/go.mod h1:4GL1E5IUh+htKOUEOaiffhrAeqysfVGipDYzABqnCmw=
golang.org/x/text v0.40.0 h1:Ub2Z6/xjgF1WrYQz2nuITOEegKFtiIy+rieRJ5lHZKs=
golang.org/x/text v0.40.0/go.mod h1:hpnzDAfGV753zIKo+wk3u1bVKCGPbrnF7+7LBF/UHVY=
google.golang.org/protobuf v1.36.11 h1:fV6ZwhNocDyBLK0dj+fg8ektcVegBBuEolpbTQyBNVE=
google.golang.org/protobuf v1.36.11/go.mod h1:HTf+CrKn2C3g5S8VImy6tdcUvCska2kB7j23XfzDpco=
gopkg.in/check.v1 v0.0.0-20161208181325-20d25e280405/go.mod h1:Co6ibVJAznAaIkqp8huTwlJQCZ016jof/cbN4VW5Yz0=
gopkg.in/check.v1 v1.0.0-20190902080502-41f04d3bba15 h1:YR8cESwS4TdDjEe65xsg0ogRM/Nc3DYOhEAlW+xobZo=
gopkg.in/check.v1 v1.0.0-20190902080502-41f04d3bba15/go.mod h1:Co6ibVJAznAaIkqp8huTwlJQCZ016jof/cbN4VW5Yz0=
gopkg.in/check.v1 v1.0.0-20201130134442-10cb98267c6c h1:Hei/4ADfdWqJk1ZMxUNpqntNwaWcugrBjAiHlqqRiVk=
gopkg.in/check.v1 v1.0.0-20201130134442-10cb98267c6c/go.mod h1:JHkPIbrfpd72SG/EVd6muEfDQjcINNoR0C8j2r3qZ4Q=
gopkg.in/yaml.v3 v3.0.1 h1:fxVm/GzAzEWqLHuvctI91KS9hhNmmWOoWu0XTYJS7CA=
gopkg.in/yaml.v3 v3.0.1/go.mod h1:K4uyk7z7BCEPqu6E+C64Yfv1cQ7kz7rIZviUmN+EgEM=
+25 -3
View File
@@ -44,6 +44,15 @@ var (
errNotPort = errors.New("must be a port number, 1 to 65535")
errNotBool = errors.New("must be true or false")
errNotIP = errors.New("must be an IP address, or empty")
errMetricsCredentials = errors.New(
"METRICS_USERNAME and METRICS_PASSWORD must be set together, " +
"or neither",
)
errMetricsUsernameColon = errors.New(
"METRICS_USERNAME must not contain \":\", " +
"which basic auth cannot carry in a user name",
)
)
// Params defines the dependencies for Config.
@@ -73,7 +82,8 @@ type Config struct {
// New loads configuration from env, .env files, and config
// files, returning a fully resolved Config. It fails, with an error
// naming the setting, on a value the server cannot use.
// naming the setting, on a value the server cannot use, and on a
// config file it finds but cannot read.
func New(
_ fx.Lifecycle,
params Params,
@@ -106,8 +116,8 @@ func New(
if err != nil {
var notFound viper.ConfigFileNotFoundError
if !errors.As(err, &notFound) {
log.Error("config file malformed", "error", err)
panic(err)
return nil, fmt.Errorf("config file %s: %w",
viper.ConfigFileUsed(), err)
}
}
@@ -177,6 +187,18 @@ func (s *Config) check() error {
}
}
// The server records and serves metrics only with both set, so
// one alone is a mistake that would otherwise go unnoticed.
if (s.MetricsUsername == "") != (s.MetricsPassword == "") {
return errMetricsCredentials
}
// Basic auth splits the credentials at the first ":", so with one
// in the user name every request to /metrics would get 401.
if strings.Contains(s.MetricsUsername, ":") {
return errMetricsUsernameColon
}
return checkOrigins(s.CORSAllowedOrigins)
}
+26
View File
@@ -103,6 +103,32 @@ func TestDataDirMaxBytesMustBeANumber(t *testing.T) {
requireConfigError(t, "DATA_DIR_MAX_BYTES")
}
// TestMetricsCredentialsGoTogether: with only one of the two set, the
// server would quietly serve no metrics, so the start fails, naming
// both.
func TestMetricsCredentialsGoTogether(t *testing.T) {
for _, set := range []string{"METRICS_USERNAME", "METRICS_PASSWORD"} {
t.Run(set, func(t *testing.T) {
t.Setenv("METRICS_USERNAME", "")
t.Setenv("METRICS_PASSWORD", "")
t.Setenv(set, "prometheus")
requireConfigError(t, "METRICS_USERNAME")
requireConfigError(t, "METRICS_PASSWORD")
})
}
}
// TestMetricsUsernameMustNotContainColon: basic auth splits the
// credentials at the first ":", so such a user name would get 401 on
// every request to /metrics.
func TestMetricsUsernameMustNotContainColon(t *testing.T) {
t.Setenv("METRICS_USERNAME", "prom:etheus")
t.Setenv("METRICS_PASSWORD", "secret")
requireConfigError(t, "METRICS_USERNAME")
}
// TestCORSAllowedOriginsMustBeOrigins: "*" would let every origin in,
// and an entry that is not a plain origin would match no page.
func TestCORSAllowedOriginsMustBeOrigins(t *testing.T) {
@@ -1,13 +0,0 @@
package handlers_test
import (
"testing"
_ "sneak.berlin/go/netwatch/internal/handlers"
)
func TestImport(t *testing.T) {
t.Parallel()
// Compilation check — verifies the package parses
// and all imports resolve.
}
+1 -1
View File
@@ -6,6 +6,6 @@ import "net/http"
// endpoint.
func (s *Handlers) HandleHealthCheck() http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
s.respondJSON(w, r, s.hc.Check(), http.StatusOK)
s.respondJSON(w, r, s.hc.Healthcheck(), http.StatusOK)
}
}
@@ -0,0 +1,120 @@
package handlers_test
import (
"encoding/json"
"maps"
"net/http"
"net/http/httptest"
"slices"
"testing"
"time"
"sneak.berlin/go/netwatch/internal/globals"
"sneak.berlin/go/netwatch/internal/handlers"
"sneak.berlin/go/netwatch/internal/healthcheck"
"sneak.berlin/go/netwatch/internal/logger"
"go.uber.org/fx/fxtest"
)
// newStartedHandlers builds Handlers with a real health check for the
// server named in g, and starts them, which records the time the
// uptime counts from.
func newStartedHandlers(t *testing.T, g *globals.Globals) *handlers.Handlers {
t.Helper()
lc := fxtest.NewLifecycle(t)
log, err := logger.New(lc, logger.Params{Globals: g})
if err != nil {
t.Fatalf("logger: %v", err)
}
hc, err := healthcheck.New(lc,
healthcheck.Params{Globals: g, Logger: log})
if err != nil {
t.Fatalf("health check: %v", err)
}
h, err := handlers.New(lc,
handlers.Params{Globals: g, Healthcheck: hc, Logger: log})
if err != nil {
t.Fatalf("handlers: %v", err)
}
lc.RequireStart()
t.Cleanup(lc.RequireStop)
return h
}
// TestHandleHealthCheck checks the health check's answer: 200, a JSON
// content type, and a JSON object with exactly the fields of
// healthcheck.HealthcheckResponse, carrying this server's name and
// version and an uptime counted from its start.
func TestHandleHealthCheck(t *testing.T) {
t.Parallel()
g := &globals.Globals{Appname: "netwatch-server", Version: "v1.2.3"}
h := newStartedHandlers(t, g)
rec := httptest.NewRecorder()
req := httptest.NewRequestWithContext(t.Context(),
http.MethodGet, "/.well-known/healthcheck", http.NoBody)
h.HandleHealthCheck().ServeHTTP(rec, req)
if rec.Code != http.StatusOK {
t.Fatalf("status = %d, want %d", rec.Code, http.StatusOK)
}
contentType := rec.Header().Get("Content-Type")
if contentType != "application/json; charset=utf-8" {
t.Errorf("Content-Type = %q, want %q",
contentType, "application/json; charset=utf-8")
}
var body map[string]any
err := json.Unmarshal(rec.Body.Bytes(), &body)
if err != nil {
t.Fatalf("body not a JSON object: %v (%q)", err, rec.Body.String())
}
fields := []string{
"appname", "now", "status", "uptime_human", "uptime_seconds", "version",
}
if got := slices.Sorted(maps.Keys(body)); !slices.Equal(got, fields) {
t.Fatalf("fields = %v, want %v", got, fields)
}
for field, want := range map[string]string{
"appname": g.Appname, "status": "ok", "version": g.Version,
} {
if body[field] != want {
t.Errorf("%s = %v, want %q", field, body[field], want)
}
}
now, _ := body["now"].(string)
at, err := time.Parse(time.RFC3339Nano, now)
if err != nil || time.Since(at).Abs() > time.Minute {
t.Errorf("now = %q, want the current time in RFC 3339 (%v)", now, err)
}
// Started just now, so the uptime is well under a minute.
human, _ := body["uptime_human"].(string)
uptime, err := time.ParseDuration(human)
if err != nil || uptime > time.Minute {
t.Errorf("uptime_human = %q, want a duration under a minute (%v)",
human, err)
}
seconds, ok := body["uptime_seconds"].(float64)
if !ok || seconds < 0 || seconds > time.Minute.Seconds() {
t.Errorf("uptime_seconds = %v, want a number of seconds under a minute",
body["uptime_seconds"])
}
}
+4 -3
View File
@@ -86,11 +86,12 @@ func (s *Handlers) decodeErrorStatus(err error) int {
}
// appendErrorStatus logs a failure to store a report and returns
// the status to send: 507 when the report files are at their size
// cap, otherwise 500.
// the status to send: 507 when the reports waiting to be written fill
// the size cap, otherwise 500.
func (s *Handlers) appendErrorStatus(err error) int {
if errors.Is(err, reportbuf.ErrFull) {
s.log.Warn("report refused: report files at their size cap")
s.log.Warn("report refused: " +
"reports waiting to be written fill the size cap")
return http.StatusInsufficientStorage
}
+60 -3
View File
@@ -46,6 +46,63 @@ func decodeStatus(t *testing.T, body []byte) string {
return resp.Status
}
// TestHandleReportAcceptsValidReports checks the answer to a valid
// report, one with no hosts and one shaped as the frontend sends them:
// 200 and {"status":"ok"} as JSON.
func TestHandleReportAcceptsValidReports(t *testing.T) {
t.Parallel()
tests := []struct {
name string
body string
}{
{
name: "no hosts",
body: `{"clientId":"c1","geo":null,"hosts":[],` +
`"timestamp":"2026-10-03T12:00:00.000Z"}`,
},
{
name: "a host with a latency and an error sample",
body: `{"clientId":"c1","geo":null,"hosts":[{` +
`"name":"Example","url":"https://example.com/",` +
`"status":"error","history":[` +
`{"t":1790000000000,"latency":42,"error":null},` +
`{"t":1790000003000,"latency":null,"error":"timeout"}]}],` +
`"timestamp":"2026-10-03T12:00:00.000Z"}`,
},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
t.Parallel()
h := newTestHandlers(stubAppender{}, io.Discard)
rec := httptest.NewRecorder()
req := httptest.NewRequestWithContext(t.Context(),
http.MethodPost, "/api/v1/reports",
strings.NewReader(tt.body),
)
h.HandleReport().ServeHTTP(rec, req)
if rec.Code != http.StatusOK {
t.Fatalf("status = %d, want %d", rec.Code, http.StatusOK)
}
contentType := rec.Header().Get("Content-Type")
if contentType != "application/json; charset=utf-8" {
t.Errorf("Content-Type = %q, want %q",
contentType, "application/json; charset=utf-8")
}
if got := rec.Body.String(); got != "{\"status\":\"ok\"}\n" {
t.Errorf("body = %q, want %q", got, "{\"status\":\"ok\"}\n")
}
})
}
}
func TestHandleReportStorageFailureIsNon2xx(t *testing.T) {
t.Parallel()
@@ -68,9 +125,9 @@ func TestHandleReportStorageFailureIsNon2xx(t *testing.T) {
}
}
// TestHandleReportFullIs507 checks the answer when the report files
// are at their size cap: 507 and the usual error body, which tells
// the client nothing more.
// TestHandleReportFullIs507 checks the answer when the reports waiting
// to be written fill the size cap: 507 and the usual error body, which
// tells the client nothing more.
func TestHandleReportFullIs507(t *testing.T) {
t.Parallel()
+11 -8
View File
@@ -30,14 +30,17 @@ type Healthcheck struct {
params *Params
}
// Response is the JSON payload returned by the health check
// endpoint.
type Response struct {
// HealthcheckResponse is the JSON payload returned by the health
// check endpoint. Its name and its snake_case keys are the ones
// GO_HTTP_SERVER_CONVENTIONS.md gives.
//
//nolint:revive,tagliatelle // name and keys from the conventions
type HealthcheckResponse struct {
Appname string `json:"appname"`
Now string `json:"now"`
Status string `json:"status"`
UptimeHuman string `json:"uptimeHuman"`
UptimeSeconds int64 `json:"uptimeSeconds"`
UptimeHuman string `json:"uptime_human"`
UptimeSeconds int64 `json:"uptime_seconds"`
Version string `json:"version"`
}
@@ -65,9 +68,9 @@ func New(
return s, nil
}
// Check returns the current health status of the application.
func (s *Healthcheck) Check() *Response {
return &Response{
// Healthcheck returns the current health status of the application.
func (s *Healthcheck) Healthcheck() *HealthcheckResponse {
return &HealthcheckResponse{
Appname: s.params.Globals.Appname,
Now: time.Now().UTC().Format(time.RFC3339Nano),
Status: "ok",
+27
View File
@@ -18,9 +18,14 @@ import (
"sneak.berlin/go/netwatch/internal/globals"
"sneak.berlin/go/netwatch/internal/logger"
basicauth "github.com/99designs/basicauth-go"
"github.com/go-chi/chi/v5/middleware"
"github.com/go-chi/cors"
"github.com/go-chi/httprate"
"github.com/prometheus/client_golang/prometheus"
metrics "github.com/slok/go-http-metrics/metrics/prometheus"
ghmm "github.com/slok/go-http-metrics/middleware"
"github.com/slok/go-http-metrics/middleware/std"
"go.uber.org/fx"
)
@@ -364,3 +369,25 @@ func (s *Middleware) RateLimit(
),
)
}
// Metrics returns middleware that records each request's duration and
// response size, and the requests in progress, in registry. They are
// labelled by the request path.
func (s *Middleware) Metrics(
registry prometheus.Registerer,
) func(http.Handler) http.Handler {
mdlw := ghmm.New(ghmm.Config{
Recorder: metrics.NewRecorder(metrics.Config{Registry: registry}),
})
return std.HandlerProvider("", mdlw)
}
// MetricsAuth returns middleware that lets a request through only with
// METRICS_USERNAME and METRICS_PASSWORD as its basic auth credentials,
// and answers any other with 401.
func (s *Middleware) MetricsAuth() func(http.Handler) http.Handler {
return basicauth.New("metrics", map[string][]string{
s.params.Config.MetricsUsername: {s.params.Config.MetricsPassword},
})
}
+150
View File
@@ -0,0 +1,150 @@
package reportbuf
import (
"errors"
"io/fs"
"os"
"os/user"
"path/filepath"
"strconv"
"syscall"
)
// ErrDataDirOutsideVolume is returned by PrepareDataDir for a DATA_DIR
// that is not the volume or a path below it, written in full.
var ErrDataDirOutsideVolume = errors.New(
"must be /data or a path below it, with no '.', '..' or extra '/'")
// PrepareDataDir gets dir, the server's DATA_DIR, ready for owner, the
// user the server runs as, so that a host directory mounted at volume,
// /data in the image, needs no preparing: it creates dir, gives volume
// and everything in it to owner, and gives volume and dir the mode the
// server gives a directory it creates. dir must be volume or a path
// below it, with no '.', '..', empty part or '/' at the end.
//
// bin/entrypoint.sh runs this as root, which would follow a symbolic
// link anywhere, so every step goes through an os.Root opened on
// volume: it follows a link only when it is written as a relative
// path that stays inside volume, and refuses any other. A process on
// the host can swap a link onto a path in volume at any moment while
// this runs. Even then, the os.Root checks each link as it reaches
// it. MkdirAll creates each directory inside a parent it already has
// open, never following a link at the name it creates, and follows a
// link on the path only as the os.Root allows, so a relative link
// inside volume can lead it to create directories elsewhere inside
// volume. Lchown never changes what a link points to, and the modes
// are set on directories already opened (see chmodDir), so the most
// that process can do is make a step fail or wait, or act on
// something else inside volume.
func PrepareDataDir(volume, dir string, owner *user.User) error {
// rel is dir as a path from volume; IsLocal is false for one that
// leads out of it.
rel, err := filepath.Rel(volume, dir)
if err != nil || dir != filepath.Clean(dir) || !filepath.IsLocal(rel) {
return ErrDataDirOutsideVolume
}
uid, err := strconv.Atoi(owner.Uid)
if err != nil {
return err
}
gid, err := strconv.Atoi(owner.Gid)
if err != nil {
return err
}
root, err := os.OpenRoot(volume)
if err != nil {
return err
}
defer func() { _ = root.Close() }()
err = root.MkdirAll(rel, dirPerms)
if err != nil {
return err
}
err = lchownAll(root, ".", uid, gid)
if err != nil {
return err
}
err = chmodDir(root, ".")
if err != nil {
return err
}
return chmodDir(root, rel)
}
// lchownAll gives name, a directory inside root, and everything in it
// to uid and gid. It reads each directory opened through root, not
// through root.FS(), which refuses a name that is not valid UTF-8, and
// calls Lchown on every entry, which gives a symbolic link itself to
// them, not what it points to. It goes into an entry only when the
// read found a directory there, so it follows no link it finds; one
// swapped in for that directory afterwards is followed only as the
// os.Root allows.
func lchownAll(root *os.Root, name string, uid, gid int) error {
err := root.Lchown(name, uid, gid)
if err != nil {
return err
}
dir, err := root.Open(name)
if err != nil {
return err
}
entries, err := dir.ReadDir(-1)
_ = dir.Close()
if err != nil {
return err
}
for _, entry := range entries {
entryName := filepath.Join(name, entry.Name())
if entry.IsDir() {
err = lchownAll(root, entryName, uid, gid)
} else {
err = root.Lchown(entryName, uid, gid)
}
if err != nil {
return err
}
}
return nil
}
// chmodDir gives name, a directory inside root, the mode the server
// gives a directory it creates. Root.Chmod would not hold: on Linux it
// checks that name is not a symbolic link, then sets the mode by name,
// following a link swapped in between. So chmodDir opens name through
// root and sets the mode on the open directory. It refuses anything
// but a directory: a directory has no second name (hard link), so the
// one opened is inside root, where any other file could be a hard link
// to one outside.
func chmodDir(root *os.Root, name string) error {
dir, err := root.Open(name)
if err != nil {
return err
}
defer func() { _ = dir.Close() }()
info, err := dir.Stat()
if err != nil {
return err
}
if !info.IsDir() {
return &fs.PathError{Op: "chmod", Path: name, Err: syscall.ENOTDIR}
}
return dir.Chmod(dirPerms)
}
+307
View File
@@ -0,0 +1,307 @@
package reportbuf_test
import (
"errors"
"io/fs"
"os"
"os/user"
"path/filepath"
"strconv"
"syscall"
"testing"
"sneak.berlin/go/netwatch/internal/reportbuf"
)
// reports is the last part of DATA_DIR in these tests, as in the
// image's /data/reports.
const reports = "reports"
// currentUser is the user the test runs as, the only owner a test not
// run as root can give files to.
func currentUser() *user.User {
return &user.User{
Uid: strconv.Itoa(os.Getuid()),
Gid: strconv.Itoa(os.Getgid()),
}
}
// tempDirMode700 is a new directory in a t.TempDir with mode 0700, so
// a test can tell that PrepareDataDir left its mode alone.
func tempDirMode700(t *testing.T) string {
t.Helper()
dir := filepath.Join(t.TempDir(), "d")
err := os.Mkdir(dir, 0o700)
if err != nil {
t.Fatal(err)
}
return dir
}
func requireMode(t *testing.T, path string, want fs.FileMode) {
t.Helper()
info, err := os.Stat(path)
if err != nil {
t.Fatal(err)
}
if info.Mode() != want {
t.Errorf("%s: mode %v, want %v", path, info.Mode(), want)
}
}
func requireMissing(t *testing.T, path string) {
t.Helper()
_, err := os.Lstat(path)
if !errors.Is(err, fs.ErrNotExist) {
t.Errorf("%s: created, or Lstat failed: %v", path, err)
}
}
func requireOwner(t *testing.T, path string, uid, gid uint32) {
t.Helper()
info, err := os.Lstat(path)
if err != nil {
t.Fatal(err)
}
stat, _ := info.Sys().(*syscall.Stat_t)
if stat.Uid != uid || stat.Gid != gid {
t.Errorf("%s: owner %d:%d, want %d:%d", path, stat.Uid, stat.Gid,
uid, gid)
}
}
func TestPrepareDataDirCreatesDataDir(t *testing.T) {
t.Parallel()
volume := t.TempDir()
dir := filepath.Join(volume, "a", reports)
err := reportbuf.PrepareDataDir(volume, dir, currentUser())
if err != nil {
t.Fatal(err)
}
requireMode(t, volume, fs.ModeDir|0o750)
requireMode(t, dir, fs.ModeDir|0o750)
}
// TestPrepareDataDirSetsModeOfExistingDataDir: a DATA_DIR already on
// the host with another mode gets the mode too, not only a new one.
func TestPrepareDataDirSetsModeOfExistingDataDir(t *testing.T) {
t.Parallel()
volume := t.TempDir()
dir := filepath.Join(volume, reports)
err := os.Mkdir(dir, 0o700)
if err != nil {
t.Fatal(err)
}
err = reportbuf.PrepareDataDir(volume, dir, currentUser())
if err != nil {
t.Fatal(err)
}
requireMode(t, dir, fs.ModeDir|0o750)
}
func TestPrepareDataDirTakesTheVolumeItself(t *testing.T) {
t.Parallel()
volume := t.TempDir()
err := reportbuf.PrepareDataDir(volume, volume, currentUser())
if err != nil {
t.Fatal(err)
}
requireMode(t, volume, fs.ModeDir|0o750)
}
// TestPrepareDataDirRefusesDataDirOutsideVolume covers a DATA_DIR that
// is relative, outside the volume, or not written in full.
func TestPrepareDataDirRefusesDataDirOutsideVolume(t *testing.T) {
t.Parallel()
volume := tempDirMode700(t)
for _, dir := range []string{
reports, volume + "/../new", volume + "/", volume + "//" + reports,
volume + "/./" + reports, volume + "/" + reports + "/..", volume + "x",
"/etc",
} {
err := reportbuf.PrepareDataDir(volume, dir, currentUser())
if !errors.Is(err, reportbuf.ErrDataDirOutsideVolume) {
t.Errorf("%q: error = %v, want ErrDataDirOutsideVolume", dir, err)
}
}
requireMissing(t, filepath.Join(filepath.Dir(volume), "new"))
requireMode(t, volume, fs.ModeDir|0o700)
}
// TestPrepareDataDirRefusesLinkOutOfVolume puts a symbolic link to a
// directory outside the volume on the path to DATA_DIR, written as a
// full path and as one that climbs out with '..', and as DATA_DIR
// itself, where the mode would be set through it.
func TestPrepareDataDirRefusesLinkOutOfVolume(t *testing.T) {
t.Parallel()
for _, tc := range []struct {
name string
climbsOut bool
link, dir string
}{
{"full path", false, "x", "x/reports"},
{"climbs out", true, "x", "x/reports"},
{"DATA_DIR itself", false, reports, reports},
} {
t.Run(tc.name, func(t *testing.T) {
t.Parallel()
outside := tempDirMode700(t)
volume := t.TempDir()
climbOut, err := filepath.Rel(volume, outside)
if err != nil {
t.Fatal(err)
}
target := outside
if tc.climbsOut {
target = climbOut
}
err = os.Symlink(target, filepath.Join(volume, tc.link))
if err != nil {
t.Fatal(err)
}
err = reportbuf.PrepareDataDir(volume,
filepath.Join(volume, tc.dir), currentUser())
if err == nil {
t.Error("no error")
}
requireMissing(t, filepath.Join(outside, reports))
requireMode(t, outside, fs.ModeDir|0o700)
})
}
}
// TestPrepareDataDirRefusesDanglingLink: DATA_DIR is a symbolic link
// to a name in the volume that does not exist, which is not created.
func TestPrepareDataDirRefusesDanglingLink(t *testing.T) {
t.Parallel()
volume := t.TempDir()
err := os.Symlink("missing", filepath.Join(volume, reports))
if err != nil {
t.Fatal(err)
}
err = reportbuf.PrepareDataDir(volume, filepath.Join(volume, reports),
currentUser())
if err == nil {
t.Error("no error")
}
requireMissing(t, filepath.Join(volume, "missing"))
}
// TestPrepareDataDirTakesDirectoryNamedInLatin1: a host directory can
// hold names that are not valid UTF-8, here "café" written in Latin-1.
// A test not run as root can only check that PrepareDataDir goes into
// such a directory and gives it, and what it holds, to the current
// user.
func TestPrepareDataDirTakesDirectoryNamedInLatin1(t *testing.T) {
t.Parallel()
volume := t.TempDir()
latin1 := filepath.Join(volume, "caf\xe9")
err := os.Mkdir(latin1, 0o700)
if err != nil {
t.Fatal(err)
}
err = os.WriteFile(filepath.Join(latin1, "f"), nil, 0o600)
if err != nil {
t.Fatal(err)
}
owner := currentUser()
err = reportbuf.PrepareDataDir(volume, filepath.Join(volume, reports),
owner)
if err != nil {
t.Fatal(err)
}
uid, _ := strconv.ParseUint(owner.Uid, 10, 32)
gid, _ := strconv.ParseUint(owner.Gid, 10, 32)
requireOwner(t, latin1, uint32(uid), uint32(gid))
requireOwner(t, filepath.Join(latin1, "f"), uint32(uid), uint32(gid))
}
// TestPrepareDataDirGivesVolumeToOwner gives everything in the volume
// to a uid and gid that own nothing, which only root can do. A
// symbolic link in the volume to a directory outside it is given to
// them itself; what it points to is left as it was.
func TestPrepareDataDirGivesVolumeToOwner(t *testing.T) {
t.Parallel()
if os.Geteuid() != 0 {
t.Skip("only root can give files to another uid")
}
outside := t.TempDir()
volume := t.TempDir()
old := filepath.Join(volume, "old")
err := os.WriteFile(filepath.Join(outside, "f"), nil, 0o600)
if err != nil {
t.Fatal(err)
}
err = os.Mkdir(old, 0o700)
if err != nil {
t.Fatal(err)
}
err = os.WriteFile(filepath.Join(old, "f"), nil, 0o600)
if err != nil {
t.Fatal(err)
}
err = os.Symlink(outside, filepath.Join(old, "link"))
if err != nil {
t.Fatal(err)
}
err = reportbuf.PrepareDataDir(volume, filepath.Join(volume, reports),
&user.User{Uid: "4242", Gid: "4343"})
if err != nil {
t.Fatal(err)
}
for _, path := range []string{
volume, filepath.Join(volume, reports), old,
filepath.Join(old, "f"), filepath.Join(old, "link"),
} {
requireOwner(t, path, 4242, 4343)
}
requireOwner(t, outside, 0, 0)
requireOwner(t, filepath.Join(outside, "f"), 0, 0)
}
+15 -1
View File
@@ -1,6 +1,13 @@
package reportbuf
import "time"
import (
"os"
"time"
)
// FlushSizeThreshold exposes the buffer size at which Append starts
// writing a report file to the external tests.
const FlushSizeThreshold = flushSizeThreshold
// Flush writes the buffered reports to a file now, as the periodic
// flush does, so tests need not wait a minute for it.
@@ -13,3 +20,10 @@ func (b *Buffer) Flush() error {
func (b *Buffer) StopClock(at time.Time) {
b.now = func() time.Time { return at }
}
// OnFileCreated makes the buffer call fn with each report file it
// writes from now on, once the file is created and before anything is
// written to it.
func (b *Buffer) OnFileCreated(fn func(f *os.File)) {
b.fileCreated = fn
}
+166 -56
View File
@@ -1,5 +1,6 @@
// Package reportbuf accumulates telemetry reports in memory
// and periodically flushes them to zstd-compressed JSONL files.
// and periodically flushes them to zstd-compressed JSONL files,
// deleting the oldest files to keep them under a size cap.
package reportbuf
import (
@@ -12,6 +13,7 @@ import (
"log/slog"
"os"
"path/filepath"
"slices"
"strings"
"sync"
"sync/atomic"
@@ -37,9 +39,10 @@ const (
fileSuffix = ".jsonl.zst"
)
// ErrFull is returned by Append when storing the report would
// take the report files past the configured maximum size.
var ErrFull = errors.New("report files at their size cap")
// ErrFull is returned by Append when the reports waiting to be
// written leave no room for the report under the configured maximum
// size, however many report files are deleted.
var ErrFull = errors.New("reports waiting to be written fill the size cap")
// Params defines the dependencies for Buffer.
type Params struct {
@@ -49,15 +52,34 @@ type Params struct {
Logger *logger.Logger
}
// reportFile is a report file that may be deleted to make room, with
// the size it counts for in usedBytes.
type reportFile struct {
name string
size int64
}
// Buffer accumulates JSON lines in memory and flushes them
// to zstd-compressed files on disk.
type Buffer struct {
buf bytes.Buffer
dataDir string
done chan struct{}
log *slog.Logger
maxBytes int64
mu sync.Mutex
buf bytes.Buffer
dataDir string
done chan struct{}
// fileCreated is called with each report file once it is
// created, before anything is written to it: it does nothing,
// except in tests that hold the write open or make it fail.
fileCreated func(f *os.File)
// files are the report files that may be deleted to make room,
// in name order, which is oldest first: those in dataDir at
// start, and each one this buffer writes, put in at its place by
// name once it is complete, even when an older file's write
// completes after a newer one's. A file still being written is
// not among them. filesBytes is their total size.
files []reportFile
filesBytes int64
log *slog.Logger
maxBytes int64
mu sync.Mutex
// now is the clock report files are named by: time.Now, except
// in tests that need two flushes to share a timestamp.
now func() time.Time
@@ -66,8 +88,8 @@ type Buffer struct {
seq atomic.Uint64
stopOnce sync.Once
// usedBytes is what Append checks against maxBytes: the size
// of the report files in dataDir, plus the reports not yet
// written to one at their uncompressed size.
// of the report files in dataDir, plus the reports waiting to
// be written to one at their uncompressed size.
usedBytes int64
}
@@ -83,11 +105,12 @@ func New(
}
b := &Buffer{
dataDir: dir,
done: make(chan struct{}),
log: params.Logger.Get(),
maxBytes: params.Config.DataDirMaxBytes,
now: time.Now,
dataDir: dir,
done: make(chan struct{}),
fileCreated: func(*os.File) {},
log: params.Logger.Get(),
maxBytes: params.Config.DataDirMaxBytes,
now: time.Now,
}
lc.Append(fx.Hook{
@@ -97,12 +120,28 @@ func New(
return fmt.Errorf("create data dir: %w", err)
}
// Report files left by earlier runs count too.
b.usedBytes, err = reportFilesSize(b.dataDir)
// Report files left by earlier runs count too, and are
// the first to be deleted to make room.
files, err := reportFiles(b.dataDir)
if err != nil {
return err
}
b.mu.Lock()
b.files = files
for _, f := range files {
b.filesBytes += f.size
}
b.usedBytes = b.filesBytes
// The files may be past the cap, if it was lowered since
// the last run.
b.deleteOldestFiles(0)
b.mu.Unlock()
go b.flushLoop()
return nil
@@ -128,9 +167,10 @@ func New(
}
// Append marshals v as a single JSON line and appends it to
// the buffer. It stores nothing and returns ErrFull if the line
// would take usedBytes past maxBytes. If the buffer reaches the
// size threshold, it is drained and written to disk
// the buffer. If the line would take usedBytes past maxBytes, the
// oldest report files are deleted to make room; it stores nothing
// and returns ErrFull if that cannot make room. If the buffer
// reaches the size threshold, it is drained and written to disk
// asynchronously.
func (b *Buffer) Append(v any) error {
line, err := json.Marshal(v)
@@ -142,6 +182,8 @@ func (b *Buffer) Append(v any) error {
b.mu.Lock()
b.deleteOldestFiles(lineBytes)
if b.usedBytes+lineBytes > b.maxBytes {
b.mu.Unlock()
@@ -171,6 +213,37 @@ func (b *Buffer) Append(v any) error {
return nil
}
// deleteOldestFiles deletes report files, oldest first, until n more
// bytes fit under maxBytes. It deletes none when the reports waiting
// to be written leave no room for n even with every file gone, since
// that would lose the files for nothing. The caller must hold b.mu.
func (b *Buffer) deleteOldestFiles(n int64) {
for b.usedBytes+n > b.maxBytes && len(b.files) > 0 {
if b.usedBytes-b.filesBytes+n > b.maxBytes {
return
}
f := b.files[0]
b.files = b.files[1:]
b.filesBytes -= f.size
// A file already gone, deleted by hand, has freed its room too.
err := os.Remove(filepath.Join(b.dataDir, f.name))
if err != nil && !errors.Is(err, fs.ErrNotExist) {
// The file is still there, so it still counts. It is
// not tried again until the next start.
b.log.Error("delete report file failed",
"file", f.name, "error", err)
continue
}
b.usedBytes -= f.size
b.log.Info("deleted report file to make room",
"file", f.name, "bytes", f.size)
}
}
// flushLoop runs a ticker that periodically flushes buffered
// data to disk until the done channel is closed.
func (b *Buffer) flushLoop() {
@@ -217,15 +290,48 @@ func (b *Buffer) drainBuf() []byte {
return data
}
// writeFile creates a timestamped zstd-compressed JSONL file
// in the data directory.
// writeFile writes data, reports drained from the buffer, to a new
// timestamped zstd-compressed JSONL file in the data directory.
func (b *Buffer) writeFile(data []byte) error {
// The timestamp comes first, so the names sort by time; the number
// after it tells apart files named in the same millisecond.
ts := b.now().UTC().Format("2006-01-02T15-04-05.000Z")
name := fmt.Sprintf("%s%s-%d%s", filePrefix, ts, b.seq.Add(1), fileSuffix)
path := filepath.Join(b.dataDir, name)
size, err := b.createFile(filepath.Join(b.dataDir, name), data)
// The reports no longer wait to be written, so they stop counting
// at their uncompressed size. If the write failed they are lost;
// otherwise they count as the file, which from here on may be
// deleted to make room.
b.mu.Lock()
defer b.mu.Unlock()
b.usedBytes -= int64(len(data))
if err != nil {
return err
}
b.usedBytes += size
b.filesBytes += size
// At its place by name, not at the end: another write, of a newer
// file, may have completed while this one was being written.
i, _ := slices.BinarySearchFunc(b.files, name,
func(f reportFile, target string) int {
return strings.Compare(f.name, target)
})
b.files = slices.Insert(b.files, i, reportFile{name: name, size: size})
return nil
}
// createFile creates the file at path holding data compressed with
// zstd, and returns its size. If the write fails once the file is
// created, it removes the file, so that a failed write leaves nothing
// behind to take room.
func (b *Buffer) createFile(path string, data []byte) (int64, error) {
// path is built from the operator-supplied dataDir plus a
// generated timestamp and number, so it carries no external input.
f, err := os.OpenFile( //nolint:gosec // see comment above
@@ -234,61 +340,65 @@ func (b *Buffer) writeFile(data []byte) error {
filePerms,
)
if err != nil {
return fmt.Errorf("create report file: %w", err)
return 0, fmt.Errorf("create report file: %w", err)
}
// Closes the file on the early returns below. The success
// path closes it explicitly to check the error; closing it
// a second time here is harmless.
defer func() { _ = f.Close() }()
b.fileCreated(f)
size, err := writeCompressed(f, data)
if err != nil {
_ = f.Close()
return 0, errors.Join(err, os.Remove(path))
}
err = f.Close()
if err != nil {
err = fmt.Errorf("close report file: %w", err)
return 0, errors.Join(err, os.Remove(path))
}
return size, nil
}
// writeCompressed writes data to f compressed with zstd, and returns
// the size of f.
func writeCompressed(f *os.File, data []byte) (int64, error) {
enc, err := zstd.NewWriter(f)
if err != nil {
return fmt.Errorf("create zstd encoder: %w", err)
return 0, fmt.Errorf("create zstd encoder: %w", err)
}
_, err = enc.Write(data)
if err != nil {
_ = enc.Close()
return fmt.Errorf("write compressed data: %w", err)
return 0, fmt.Errorf("write compressed data: %w", err)
}
err = enc.Close()
if err != nil {
return fmt.Errorf("close zstd encoder: %w", err)
return 0, fmt.Errorf("close zstd encoder: %w", err)
}
info, err := f.Stat()
if err != nil {
return fmt.Errorf("stat report file: %w", err)
return 0, fmt.Errorf("stat report file: %w", err)
}
err = f.Close()
if err != nil {
return fmt.Errorf("close report file: %w", err)
}
// The reports counted at their uncompressed size while they
// waited; now they count as the file. After a failed write they
// stay counted as they were, which errs toward refusing reports
// early rather than letting the files pass the cap.
b.mu.Lock()
b.usedBytes += info.Size() - int64(len(data))
b.mu.Unlock()
return nil
return info.Size(), nil
}
// reportFilesSize returns the total size of the report files in
// dir.
func reportFilesSize(dir string) (int64, error) {
// reportFiles returns the report files in dir, oldest first:
// os.ReadDir sorts them by name, and the names sort by time.
func reportFiles(dir string) ([]reportFile, error) {
entries, err := os.ReadDir(dir)
if err != nil {
return 0, fmt.Errorf("read data dir: %w", err)
return nil, fmt.Errorf("read data dir: %w", err)
}
var total int64
files := make([]reportFile, 0, len(entries))
for _, entry := range entries {
name := entry.Name()
@@ -299,11 +409,11 @@ func reportFilesSize(dir string) (int64, error) {
info, err := entry.Info()
if err != nil {
return 0, fmt.Errorf("stat report file: %w", err)
return nil, fmt.Errorf("stat report file: %w", err)
}
total += info.Size()
files = append(files, reportFile{name: name, size: info.Size()})
}
return total, nil
return files, nil
}
+522 -44
View File
@@ -139,11 +139,19 @@ func lineBytes(t *testing.T, report any) int {
return len(line) + 1
}
// TestAppendPastCapIsRefused fills the cap with a report not yet
// written. The next report is refused, and the report file already in
// DATA_DIR is kept: it is smaller than a report, so deleting it could
// not make room.
func TestAppendPastCapIsRefused(t *testing.T) {
report := map[string]string{"id": "cap"}
dir := t.TempDir()
earlier := reportFilePath(dir, 1)
t.Setenv("DATA_DIR", t.TempDir())
t.Setenv("DATA_DIR_MAX_BYTES", strconv.Itoa(lineBytes(t, report)))
writeBytes(t, earlier, 1)
t.Setenv("DATA_DIR", dir)
t.Setenv("DATA_DIR_MAX_BYTES", strconv.Itoa(1+lineBytes(t, report)))
buf := startBuffer(t)
@@ -156,24 +164,38 @@ func TestAppendPastCapIsRefused(t *testing.T) {
if !errors.Is(err, reportbuf.ErrFull) {
t.Fatalf("report past the cap: error = %v, want ErrFull", err)
}
if !exists(t, earlier) {
t.Fatal("report file deleted, though that could not make room")
}
}
// TestCapCountsReportFilesAlreadyInDataDir starts on a data
// directory holding a report file from an earlier run, and a file
// that is not a report, which must not count.
func TestCapCountsReportFilesAlreadyInDataDir(t *testing.T) {
const earlierBytes = 100
// TestOldestReportFileDeletedFirst starts on a data directory holding
// report files from an earlier run, and a file that is not a report,
// which neither counts nor is ever deleted. Nothing is deleted while
// there is room; then only the oldest report file is.
func TestOldestReportFileDeletedFirst(t *testing.T) {
const fileBytes = 100
report := map[string]string{"id": "cap"}
report := map[string]string{"id": "oldest"}
dir := t.TempDir()
oldest := reportFilePath(dir, 1)
kept := []string{
reportFilePath(dir, 2),
reportFilePath(dir, 3),
filepath.Join(dir, "notes.txt"),
}
writeBytes(t, filepath.Join(dir, "reports-2026-01-01T00-00-00.000Z.jsonl.zst"),
earlierBytes)
writeBytes(t, filepath.Join(dir, "notes.txt"), 10*earlierBytes)
writeBytes(t, oldest, fileBytes)
for _, path := range kept {
writeBytes(t, path, fileBytes)
}
t.Setenv("DATA_DIR", dir)
// Room for the three report files and one report.
t.Setenv("DATA_DIR_MAX_BYTES",
strconv.Itoa(earlierBytes+lineBytes(t, report)))
strconv.Itoa(3*fileBytes+lineBytes(t, report)))
buf := startBuffer(t)
@@ -182,9 +204,55 @@ func TestCapCountsReportFilesAlreadyInDataDir(t *testing.T) {
t.Fatalf("report that fills the cap exactly: %v", err)
}
if !exists(t, oldest) {
t.Fatal("oldest report file deleted while there was room")
}
err = buf.Append(report)
if !errors.Is(err, reportbuf.ErrFull) {
t.Fatalf("report past the cap: error = %v, want ErrFull", err)
if err != nil {
t.Fatalf("report past the cap: %v", err)
}
if exists(t, oldest) {
t.Fatal("oldest report file kept when room was needed")
}
for _, path := range kept {
if !exists(t, path) {
t.Fatalf("%s deleted; only the oldest report file should be", path)
}
}
}
// TestStartDeletesFilesPastCap starts on report files past the cap,
// as after the cap is lowered: the oldest are deleted until the rest
// fit.
func TestStartDeletesFilesPastCap(t *testing.T) {
const fileBytes = 100
dir := t.TempDir()
oldest := reportFilePath(dir, 1)
kept := []string{reportFilePath(dir, 2), reportFilePath(dir, 3)}
writeBytes(t, oldest, fileBytes)
for _, path := range kept {
writeBytes(t, path, fileBytes)
}
t.Setenv("DATA_DIR", dir)
t.Setenv("DATA_DIR_MAX_BYTES", strconv.Itoa(len(kept)*fileBytes))
startBuffer(t)
if exists(t, oldest) {
t.Fatal("oldest report file kept, though the files were past the cap")
}
for _, path := range kept {
if !exists(t, path) {
t.Fatalf("%s deleted, though the rest fit without it", path)
}
}
}
@@ -195,8 +263,9 @@ func TestWrittenReportsCountAtFileSize(t *testing.T) {
// Repetitive, so its file is far smaller than its JSON.
report := map[string]string{"id": strings.Repeat("a", 1000)}
size := lineBytes(t, report)
dir := t.TempDir()
t.Setenv("DATA_DIR", t.TempDir())
t.Setenv("DATA_DIR", dir)
// Room for the report twice over only if the first one counts
// at its file's size by the time the second arrives.
t.Setenv("DATA_DIR_MAX_BYTES", strconv.Itoa(2*size-1))
@@ -217,13 +286,22 @@ func TestWrittenReportsCountAtFileSize(t *testing.T) {
if err != nil {
t.Fatalf("second report, after the first was written: %v", err)
}
if !hasReportFile(t, dir) {
t.Fatal("first report's file deleted, though the second fit beside it")
}
}
// TestWrittenReportsKeepCounting writes one report file after another
// under a small cap: each report must be taken while the files on disk
// leave room for it, and refused once they do not.
func TestWrittenReportsKeepCounting(t *testing.T) {
const maxBytes = 200
// TestWrittenFilesDeletedToMakeRoom writes one report file after
// another under a small cap. Every report must be taken; files are
// deleted only when the report would not fit beside them, and the
// files kept leave room for it.
func TestWrittenFilesDeletedToMakeRoom(t *testing.T) {
const (
maxBytes = 200
// One file each, which take far more than maxBytes together.
reports = 50
)
report := map[string]string{"id": "written"}
size := int64(lineBytes(t, report))
@@ -234,23 +312,24 @@ func TestWrittenReportsKeepCounting(t *testing.T) {
buf := startBuffer(t)
// Every file takes at least a byte, so they fill the cap within
// maxBytes rounds.
for range maxBytes {
used := reportFilesBytes(t, dir)
for range reports {
before := reportFilesBytes(t, dir)
err := buf.Append(report)
if used+size > maxBytes {
if !errors.Is(err, reportbuf.ErrFull) {
t.Fatalf("with %d bytes of report files: error = %v, "+
"want ErrFull", used, err)
}
return
if err != nil {
t.Fatalf("with %d bytes of report files: %v", before, err)
}
if err != nil {
t.Fatalf("with %d bytes of report files: %v", used, err)
after := reportFilesBytes(t, dir)
if before+size <= maxBytes && after != before {
t.Fatalf("files deleted, though the report fit beside "+
"their %d bytes", before)
}
if after+size > maxBytes {
t.Fatalf("%d bytes of report files kept, leaving no room "+
"for the report", after)
}
err = buf.Flush()
@@ -258,8 +337,217 @@ func TestWrittenReportsKeepCounting(t *testing.T) {
t.Fatalf("flush: %v", err)
}
}
}
t.Fatal("the report files never filled the cap")
// TestFailedDeletionStillCounts makes deleting the oldest report file
// fail. It is still there, so it still takes room, and the next oldest
// is deleted in its place.
func TestFailedDeletionStillCounts(t *testing.T) {
const fileBytes = 100
report := map[string]string{"id": "stuck"}
dir := t.TempDir()
stuck := reportFilePath(dir, 1)
next := reportFilePath(dir, 2)
newest := reportFilePath(dir, 3)
for _, path := range []string{stuck, next, newest} {
writeBytes(t, path, fileBytes)
}
t.Setenv("DATA_DIR", dir)
t.Setenv("DATA_DIR_MAX_BYTES", strconv.Itoa(3*fileBytes))
buf := startBuffer(t)
// A directory that is not empty cannot be deleted, even by root,
// which the tests run as in the backend image.
err := os.Remove(stuck)
if err != nil {
t.Fatalf("remove %s: %v", stuck, err)
}
err = os.Mkdir(stuck, 0o750)
if err != nil {
t.Fatalf("make directory %s: %v", stuck, err)
}
writeBytes(t, filepath.Join(stuck, "file"), 1)
err = buf.Append(report)
if err != nil {
t.Fatalf("report past the cap: %v", err)
}
if exists(t, next) {
t.Fatal("next oldest report file kept: the failed deletion " +
"counted as making room")
}
if !exists(t, newest) {
t.Fatal("newest report file deleted, though deleting one made room")
}
}
// TestFileDeletedByHandFreesRoom deletes the oldest report file by
// hand after start. When room is needed, its room counts as freed, so
// no other file is deleted.
func TestFileDeletedByHandFreesRoom(t *testing.T) {
const fileBytes = 100
report := map[string]string{"id": "by-hand"}
dir := t.TempDir()
gone := reportFilePath(dir, 1)
kept := reportFilePath(dir, 2)
writeBytes(t, gone, fileBytes)
writeBytes(t, kept, fileBytes)
t.Setenv("DATA_DIR", dir)
t.Setenv("DATA_DIR_MAX_BYTES", strconv.Itoa(2*fileBytes))
buf := startBuffer(t)
err := os.Remove(gone)
if err != nil {
t.Fatalf("remove %s: %v", gone, err)
}
err = buf.Append(report)
if err != nil {
t.Fatalf("report past the cap: %v", err)
}
if !exists(t, kept) {
t.Fatal("report file deleted, though the one deleted by hand " +
"had made room")
}
}
// TestFileBeingWrittenIsNeverDeleted holds the write of one report file
// open while a second write completes, then sends a report that needs
// room. Deleting either file would make it, and the one being written is
// the older, but only the complete one may be deleted. Once the first
// write is complete, its file is deleted when room is needed.
func TestFileBeingWrittenIsNeverDeleted(t *testing.T) {
report := map[string]string{"id": "writing"}
t.Setenv("DATA_DIR", t.TempDir())
// Room for two reports waiting to be written, but not for two
// beside a report file.
t.Setenv("DATA_DIR_MAX_BYTES", strconv.Itoa(2*lineBytes(t, report)))
buf := startBuffer(t)
created := make(chan string)
release := make(chan struct{})
buf.OnFileCreated(func(f *os.File) {
created <- f.Name()
<-release
})
err := buf.Append(report)
if err != nil {
t.Fatalf("first report: %v", err)
}
flushed := make(chan error)
go func() { flushed <- buf.Flush() }()
writing := <-created
// Only the first write is held; the second goes through, and
// so do the writes after it, the final one at stop included.
var complete string
buf.OnFileCreated(func(f *os.File) { complete = f.Name() })
// Errorf, not Fatalf, until the first write is released, so that a
// failure here does not leave it held.
err = buf.Append(report)
if err != nil {
t.Errorf("second report: %v", err)
}
err = buf.Flush()
if err != nil {
t.Errorf("flush of the second report: %v", err)
}
err = buf.Append(report)
if err != nil {
t.Errorf("report that needs room: %v", err)
}
if !exists(t, writing) {
t.Error("report file deleted while it was being written")
}
if exists(t, complete) {
t.Error("complete report file kept, though room was needed")
}
close(release)
err = <-flushed
if err != nil {
t.Fatalf("flush of the first report: %v", err)
}
err = buf.Append(report)
if err != nil {
t.Fatalf("report after the first write was complete: %v", err)
}
if exists(t, writing) {
t.Fatal("complete report file kept when room was needed")
}
}
// TestFailedWriteStopsCounting makes a write fail once its file is
// created. Its reports are lost, so they stop counting, and the part of
// the file written is removed, so it takes no room.
func TestFailedWriteStopsCounting(t *testing.T) {
report := map[string]string{"id": "failed"}
t.Setenv("DATA_DIR", t.TempDir())
t.Setenv("DATA_DIR_MAX_BYTES", strconv.Itoa(lineBytes(t, report)))
buf := startBuffer(t)
var failed string
// Closing the file under the write makes the write fail.
buf.OnFileCreated(func(f *os.File) {
failed = f.Name()
_ = f.Close()
})
err := buf.Append(report)
if err != nil {
t.Fatalf("report that fills the cap exactly: %v", err)
}
err = buf.Flush()
if err == nil {
t.Fatal("flush succeeded, though its file was closed under it")
}
if exists(t, failed) {
t.Fatal("file of the failed write kept")
}
// Writes from here on, the final one at stop included, succeed.
buf.OnFileCreated(func(*os.File) {})
err = buf.Append(report)
if err != nil {
t.Fatalf("report after the failed write: %v", err)
}
}
// TestConcurrentAppendsStopAtCap appends from many goroutines at once
@@ -334,7 +622,11 @@ func TestTwoFlushesInOneMillisecond(t *testing.T) {
}
}
files := readReportFiles(t, dir)
files, err := readReportFiles(dir)
if err != nil {
t.Fatalf("read report files: %v", err)
}
if len(files) != flushes {
t.Fatalf("%d report files after %d flushes", len(files), flushes)
}
@@ -347,6 +639,88 @@ func TestTwoFlushesInOneMillisecond(t *testing.T) {
}
}
// TestReportFileHoldsTheLinesAppended flushes three reports and reads
// their file back: it must decompress to exactly their JSON lines, in
// the order they were appended.
func TestReportFileHoldsTheLinesAppended(t *testing.T) {
dir := t.TempDir()
t.Setenv("DATA_DIR", dir)
buf := startBuffer(t)
for _, id := range []int{1, 2, 3} {
err := buf.Append(map[string]int{"id": id})
if err != nil {
t.Fatalf("append report %d: %v", id, err)
}
}
err := buf.Flush()
if err != nil {
t.Fatalf("flush: %v", err)
}
files, err := readReportFiles(dir)
if err != nil {
t.Fatalf("read report files: %v", err)
}
want := `{"id":1}` + "\n" + `{"id":2}` + "\n" + `{"id":3}` + "\n"
if len(files) != 1 || files[0] != want {
t.Fatalf("report files = %q, want one holding %q", files, want)
}
}
// TestFlushAtSizeThreshold appends reports until the buffer holds
// FlushSizeThreshold bytes. The append that gets it there must write
// them all to one report file, with no call to Flush and the periodic
// flush a minute away, and no earlier append may write one.
func TestFlushAtSizeThreshold(t *testing.T) {
dir := t.TempDir()
t.Setenv("DATA_DIR", dir)
buf := startBuffer(t)
// Large reports, so the threshold takes a few hundred appends.
pad := strings.Repeat("a", 64<<10)
var appended strings.Builder
for id := 0; appended.Len() < reportbuf.FlushSizeThreshold; id++ {
report := map[string]any{"id": id, "pad": pad}
err := buf.Append(report)
if err != nil {
t.Fatalf("append report %d: %v", id, err)
}
line, err := json.Marshal(report)
if err != nil {
t.Fatalf("marshal report %d: %v", id, err)
}
appended.Write(line)
appended.WriteByte('\n')
}
// Append writes the file in the background, so wait for it.
deadline := time.Now().Add(10 * time.Second)
for {
files, err := readReportFiles(dir)
if err == nil && len(files) == 1 && files[0] == appended.String() {
return
}
if time.Now().After(deadline) {
t.Fatalf("%d report files (error: %v), want one holding the "+
"%d bytes appended", len(files), err, appended.Len())
}
time.Sleep(10 * time.Millisecond)
}
}
// reportFilesBytes returns the total size of the report files in dir.
func reportFilesBytes(t *testing.T, dir string) int64 {
t.Helper()
@@ -371,20 +745,19 @@ func reportFilesBytes(t *testing.T, dir string) int64 {
}
// readReportFiles returns the decompressed contents of each report
// file in dir.
func readReportFiles(t *testing.T, dir string) []string {
t.Helper()
// file in dir. A file still being written does not decompress, so it
// gives an error.
func readReportFiles(dir string) ([]string, error) {
files := os.DirFS(dir)
names, err := fs.Glob(files, "reports-*.jsonl.zst")
if err != nil {
t.Fatalf("list report files: %v", err)
return nil, fmt.Errorf("list report files: %w", err)
}
dec, err := zstd.NewReader(nil)
if err != nil {
t.Fatalf("create zstd decoder: %v", err)
return nil, fmt.Errorf("create zstd decoder: %w", err)
}
defer dec.Close()
@@ -393,18 +766,26 @@ func readReportFiles(t *testing.T, dir string) []string {
for _, name := range names {
compressed, readErr := fs.ReadFile(files, name)
if readErr != nil {
t.Fatalf("read %s: %v", name, readErr)
return nil, fmt.Errorf("read %s: %w", name, readErr)
}
data, decErr := dec.DecodeAll(compressed, nil)
if decErr != nil {
t.Fatalf("decompress %s: %v", name, decErr)
return nil, fmt.Errorf("decompress %s: %w", name, decErr)
}
contents = append(contents, string(data))
}
return contents
return contents, nil
}
// reportFilePath returns the path in dir of a report file named as
// written on the given day of January 2026, so that a lower day sorts
// as older.
func reportFilePath(dir string, day int) string {
return filepath.Join(dir,
fmt.Sprintf("reports-2026-01-%02dT00-00-00.000Z-1.jsonl.zst", day))
}
func writeBytes(t *testing.T, path string, n int) {
@@ -416,6 +797,21 @@ func writeBytes(t *testing.T, path string, n int) {
}
}
func exists(t *testing.T, path string) bool {
t.Helper()
_, err := os.Stat(path)
if errors.Is(err, fs.ErrNotExist) {
return false
}
if err != nil {
t.Fatalf("stat %s: %v", path, err)
}
return true
}
func hasReportFile(t *testing.T, dir string) bool {
t.Helper()
@@ -439,3 +835,85 @@ func hasReportFile(t *testing.T, dir string) bool {
return false
}
// TestFilesDeletedOldestFirstWhenWritesOverlap holds the write of an
// older report file open until a newer one's write completes, then
// releases it. When room is needed, the older file is deleted first,
// though its write was the last to complete.
func TestFilesDeletedOldestFirstWhenWritesOverlap(t *testing.T) {
const maxBytes = 1000
report := map[string]string{"id": "overlap"}
t.Setenv("DATA_DIR", t.TempDir())
t.Setenv("DATA_DIR_MAX_BYTES", strconv.Itoa(maxBytes))
buf := startBuffer(t)
created := make(chan string)
release := make(chan struct{})
buf.OnFileCreated(func(f *os.File) {
created <- f.Name()
<-release
})
err := buf.Append(report)
if err != nil {
t.Fatalf("older report: %v", err)
}
flushed := make(chan error)
go func() { flushed <- buf.Flush() }()
older := <-created
// Only the older write is held; the newer one goes through, and
// so do the writes after it, the final one at stop included.
var newer string
buf.OnFileCreated(func(f *os.File) { newer = f.Name() })
// Errorf, not Fatalf, until the older write is released, so that a
// failure here does not leave it held.
err = buf.Append(report)
if err != nil {
t.Errorf("newer report: %v", err)
}
err = buf.Flush()
if err != nil {
t.Errorf("flush of the newer report: %v", err)
}
close(release)
err = <-flushed
if err != nil {
t.Fatalf("flush of the older report: %v", err)
}
info, err := os.Stat(newer)
if err != nil {
t.Fatalf("stat %s: %v", newer, err)
}
// A report that fits beside the newer file alone, so deleting the
// older one makes exactly the room it needs.
pad := maxBytes - int(info.Size()) - lineBytes(t, map[string]string{"id": ""})
err = buf.Append(map[string]string{"id": strings.Repeat("a", pad)})
if err != nil {
t.Fatalf("report that needs room: %v", err)
}
if exists(t, older) {
t.Fatal("older report file kept when room was needed")
}
if !exists(t, newer) {
t.Fatal("newer report file deleted before the older one")
}
}
+12
View File
@@ -1,9 +1,21 @@
package server
import "github.com/go-chi/chi/v5"
// Router exposes the router to the external tests, which add routes
// of their own to it after SetupRoutes.
func (s *Server) Router() *chi.Mux {
return s.router
}
// MaxRequestBodyBytes exposes the router-wide body limit to the
// external tests.
const MaxRequestBodyBytes = maxRequestBodyBytes
// MetricsRequestsPerMinute exposes the /metrics rate limit to the
// external tests.
const MetricsRequestsPerMinute = metricsRequestsPerMinute
// ListenAddr exposes the address the server listens on to the
// external tests.
func (s *Server) ListenAddr() string {
+49 -5
View File
@@ -3,8 +3,12 @@ package server
import (
"time"
sentryhttp "github.com/getsentry/sentry-go/http"
"github.com/go-chi/chi/v5"
"github.com/go-chi/chi/v5/middleware"
"github.com/prometheus/client_golang/prometheus"
"github.com/prometheus/client_golang/prometheus/collectors"
"github.com/prometheus/client_golang/prometheus/promhttp"
)
const (
@@ -14,6 +18,12 @@ const (
// can mount s.mw.MaxBodyBytes with a smaller value to lower
// its bound, but cannot raise it: this cap runs first.
maxRequestBodyBytes int64 = 1 << 20 // 1 MiB
// metricsRequestsPerMinute is how many requests to /metrics each
// client address may make a minute, whatever their credentials. A
// scraper polling every 2 seconds sends half of it, which httprate
// never refuses.
metricsRequestsPerMinute = 60
)
// SetupRoutes configures the chi router with middleware and
@@ -29,13 +39,47 @@ func (s *Server) SetupRoutes() {
s.router.Use(s.mw.MaxBodyBytes(maxRequestBodyBytes))
s.router.Use(middleware.Timeout(requestTimeout))
s.router.Get(
"/.well-known/healthcheck",
s.h.HandleHealthCheck(),
// Sentry reports a panic, then panics again, so that s.mw.Recoverer
// still answers 500.
if s.params.Config.SentryDSN != "" {
s.router.Use(sentryhttp.New(sentryhttp.Options{Repanic: true}).Handle)
}
// The metrics go in a registry of this server's own, not in
// Prometheus' default one, which takes them only once per process.
registry := prometheus.NewRegistry()
registry.MustRegister(
collectors.NewGoCollector(),
collectors.NewProcessCollector(collectors.ProcessCollectorOpts{}),
)
s.router.Route("/api/v1", func(r chi.Router) {
// Requests are measured only once chi has matched them to one of
// these routes, by path and method. The metrics are labelled with
// both, which any client can make up, so measuring every request
// would let clients add labels without bound. A Route here would
// be matched by its path prefix alone, so each path is given in
// full.
s.router.Group(func(r chi.Router) {
// config.New refuses one of the two credentials without the
// other.
if s.params.Config.MetricsUsername != "" {
r.Use(s.mw.Metrics(registry))
}
r.Get("/.well-known/healthcheck", s.h.HandleHealthCheck())
r.With(s.mw.RateLimit(s.params.Config.ReportsPerMinute)).
Post("/reports", s.h.HandleReport())
Post("/api/v1/reports", s.h.HandleReport())
})
// The rate limit comes before the basic auth, so a client past it
// gets 429 and its password is not checked.
if s.params.Config.MetricsUsername != "" {
s.router.With(
s.mw.RateLimit(metricsRequestsPerMinute),
s.mw.MetricsAuth(),
).Get("/metrics", promhttp.HandlerFor(
registry, promhttp.HandlerOpts{},
).ServeHTTP)
}
}
+200
View File
@@ -1,10 +1,12 @@
package server_test
import (
"io"
"net/http"
"net/http/httptest"
"strings"
"testing"
"time"
"sneak.berlin/go/netwatch/internal/config"
"sneak.berlin/go/netwatch/internal/globals"
@@ -15,6 +17,7 @@ import (
"sneak.berlin/go/netwatch/internal/reportbuf"
"sneak.berlin/go/netwatch/internal/server"
"github.com/getsentry/sentry-go"
"go.uber.org/fx"
"go.uber.org/fx/fxtest"
)
@@ -107,6 +110,203 @@ func TestCORSAllowedOriginsReachTheRouter(t *testing.T) {
}
}
// TestNoMetricsWithoutCredentials: with neither metrics setting set,
// there is no /metrics.
func TestNoMetricsWithoutCredentials(t *testing.T) {
t.Setenv("METRICS_USERNAME", "")
t.Setenv("METRICS_PASSWORD", "")
srv := newServer(t)
srv.SetupRoutes()
rec := httptest.NewRecorder()
req := httptest.NewRequestWithContext(t.Context(),
http.MethodGet, "/metrics", http.NoBody)
srv.ServeHTTP(rec, req)
if rec.Code != http.StatusNotFound {
t.Fatalf("status = %d, want %d", rec.Code, http.StatusNotFound)
}
}
// TestMetricsBehindBasicAuth: with both metrics settings set, /metrics
// answers only with them as basic auth credentials, and shows a
// request to a route but not one to a path no route has.
func TestMetricsBehindBasicAuth(t *testing.T) {
t.Setenv("METRICS_USERNAME", "prometheus")
t.Setenv("METRICS_PASSWORD", "right")
srv := newServer(t)
srv.SetupRoutes()
get := func(path, username, password string) *httptest.ResponseRecorder {
rec := httptest.NewRecorder()
req := httptest.NewRequestWithContext(t.Context(),
http.MethodGet, path, http.NoBody)
if username != "" {
req.SetBasicAuth(username, password)
}
srv.ServeHTTP(rec, req)
return rec
}
get("/.well-known/healthcheck", "", "")
get("/api/v1/no-such-route", "", "")
for _, creds := range [][2]string{
{"", ""},
{"prometheus", "wrong"},
{"someone", "right"},
} {
rec := get("/metrics", creds[0], creds[1])
if rec.Code != http.StatusUnauthorized {
t.Errorf("credentials %q: status = %d, want %d",
creds, rec.Code, http.StatusUnauthorized)
}
}
rec := get("/metrics", "prometheus", "right")
if rec.Code != http.StatusOK {
t.Fatalf("right credentials: status = %d, want %d",
rec.Code, http.StatusOK)
}
body := rec.Body.String()
if !strings.Contains(body, `handler="/.well-known/healthcheck"`) {
t.Errorf("metrics show no health check request:\n%s", body)
}
if strings.Contains(body, "no-such-route") {
t.Errorf("metrics show a request to a path no route has:\n%s", body)
}
if !strings.Contains(body, "go_goroutines") {
t.Errorf("metrics show no Go runtime metrics:\n%s", body)
}
}
// TestMetricsAreRateLimited: a client that has used up its /metrics
// allowance on wrong passwords gets 429 even with the right one, which
// is then not checked, while another client behind the same nginx
// still gets in.
func TestMetricsAreRateLimited(t *testing.T) {
t.Setenv("METRICS_USERNAME", "prometheus")
t.Setenv("METRICS_PASSWORD", "right")
// As in the container: nginx connects from loopback and names the
// client in X-Forwarded-For.
t.Setenv("TRUSTED_PROXIES", "127.0.0.1/32")
srv := newServer(t)
srv.SetupRoutes()
get := func(client, password string) int {
rec := httptest.NewRecorder()
req := httptest.NewRequestWithContext(t.Context(),
http.MethodGet, "/metrics", http.NoBody)
req.RemoteAddr = "127.0.0.1:40000"
req.Header.Set("X-Forwarded-For", client)
req.SetBasicAuth("prometheus", password)
srv.ServeHTTP(rec, req)
return rec.Code
}
for i := range server.MetricsRequestsPerMinute {
if code := get("203.0.113.7", "wrong"); code != http.StatusUnauthorized {
t.Fatalf("guess %d: status = %d, want %d",
i+1, code, http.StatusUnauthorized)
}
}
if code := get("203.0.113.7", "right"); code != http.StatusTooManyRequests {
t.Fatalf("right password past the limit: status = %d, want %d",
code, http.StatusTooManyRequests)
}
if code := get("203.0.113.8", "right"); code != http.StatusOK {
t.Fatalf("another client: status = %d, want %d",
code, http.StatusOK)
}
}
// TestMetricsInTwoServers: two servers in one process can both have
// metrics on.
func TestMetricsInTwoServers(t *testing.T) {
t.Setenv("METRICS_USERNAME", "prometheus")
t.Setenv("METRICS_PASSWORD", "right")
for range 2 {
newServer(t).SetupRoutes()
}
}
// TestSentry: with SENTRY_DSN empty there is no Sentry client. With it
// pointing at a local server standing in for Sentry, a panic in a
// handler reaches that server, and the request still gets the 500 from
// the panic recovery.
func TestSentry(t *testing.T) {
const panicMessage = "handler panic for TestSentry"
// sentry.Init sets the client for the whole process; take it away
// again so that no other test reports to Sentry.
t.Cleanup(func() { sentry.CurrentHub().BindClient(nil) })
t.Setenv("SENTRY_DSN", "")
newServer(t)
if sentry.CurrentHub().Client() != nil {
t.Fatal("a Sentry client exists with SENTRY_DSN empty")
}
// The body of the first request the stand-in for Sentry receives.
received := make(chan string, 1)
sentryServer := httptest.NewServer(http.HandlerFunc(
func(_ http.ResponseWriter, r *http.Request) {
body, _ := io.ReadAll(r.Body)
select {
case received <- string(body):
default:
}
},
))
defer sentryServer.Close()
t.Setenv("SENTRY_DSN",
"http://key@"+sentryServer.Listener.Addr().String()+"/1")
srv := newServer(t)
srv.SetupRoutes()
srv.Router().Get("/panic", func(http.ResponseWriter, *http.Request) {
panic(panicMessage)
})
rec := httptest.NewRecorder()
req := httptest.NewRequestWithContext(t.Context(),
http.MethodGet, "/panic", http.NoBody)
srv.ServeHTTP(rec, req)
if rec.Code != http.StatusInternalServerError {
t.Fatalf("status = %d, want %d",
rec.Code, http.StatusInternalServerError)
}
// Sentry sends from a goroutine of its own.
select {
case body := <-received:
if !strings.Contains(body, panicMessage) {
t.Fatalf("the Sentry server received no report of the panic:\n%s",
body)
}
case <-time.After(5 * time.Second):
t.Fatal("nothing reached the Sentry server")
}
}
// TestHealthCheckRejectsOversizeBody sends the health check, which
// never reads its body, a body one byte over the limit. Only the
// router-wide body limit can reject it.
+40 -1
View File
@@ -7,8 +7,10 @@ package server
import (
"context"
"fmt"
"log/slog"
"net/http"
"time"
"sneak.berlin/go/netwatch/internal/config"
"sneak.berlin/go/netwatch/internal/globals"
@@ -16,10 +18,15 @@ import (
"sneak.berlin/go/netwatch/internal/logger"
"sneak.berlin/go/netwatch/internal/middleware"
"github.com/getsentry/sentry-go"
"github.com/go-chi/chi/v5"
"go.uber.org/fx"
)
// sentryFlushTimeout is how long shutdown waits for Sentry to send
// what it still holds.
const sentryFlushTimeout = 2 * time.Second
// Params defines the dependencies for Server.
type Params struct {
fx.In
@@ -56,6 +63,11 @@ func New(
s.log = params.Logger.Get()
s.shutdowner = params.Shutdowner
err := s.enableSentry()
if err != nil {
return nil, err
}
lc.Append(fx.Hook{
OnStart: func(_ context.Context) error {
// Build the router and http.Server synchronously
@@ -87,10 +99,37 @@ func (s *Server) ServeHTTP(
s.router.ServeHTTP(w, r)
}
// enableSentry sets Sentry up when SENTRY_DSN is set, so that
// SetupRoutes can report panics to it. With SENTRY_DSN empty it does
// nothing. A DSN Sentry refuses stops the start.
func (s *Server) enableSentry() error {
if s.params.Config.SentryDSN == "" {
return nil
}
err := sentry.Init(sentry.ClientOptions{
Dsn: s.params.Config.SentryDSN,
Release: s.params.Globals.Appname + "-" + s.params.Globals.Version,
})
if err != nil {
return fmt.Errorf("SENTRY_DSN: %w", err)
}
s.log.Info("sentry error reporting activated")
return nil
}
// shutdown gracefully stops the HTTP server within the
// deadline of the context fx provides for OnStop.
// deadline of the context fx provides for OnStop, then gives
// Sentry, if set up, time to send what it still holds.
func (s *Server) shutdown(ctx context.Context) error {
err := s.httpServer.Shutdown(ctx)
if s.params.Config.SentryDSN != "" {
sentry.Flush(sentryFlushTimeout)
}
if err != nil {
s.log.Error("server clean shutdown failed", "error", err)
+20 -2
View File
@@ -14,16 +14,34 @@ set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
# The sha256 of the org standard .golangci.yml. When that file changes in
# sneak/prompts and is copied here again, this changes with it.
GOLANGCI_CONFIG_SHA256="a79b63a254602a5318db5d0e9a06bc71b84bf0c1d896305229d8bfed1d1b1776"
main() {
cd "$ROOT"
if [ ! -f .golangci.yml ]; then
echo "backend/.golangci.yml is missing. Copy the org standard verbatim" >&2
echo "from https://git.eeqj.de/sneak/prompts/raw/branch/main/.golangci.yml" >&2
exit 1
fi
actual="$(sha256sum .golangci.yml | cut -d' ' -f1)"
if [ -z "$actual" ]; then
echo "sha256sum is missing or printed no hash, so" >&2
echo "backend/.golangci.yml could not be checked." >&2
exit 1
fi
if [ "$actual" != "$GOLANGCI_CONFIG_SHA256" ]; then
echo ".golangci.yml has drifted from the org standard." >&2
echo "backend/.golangci.yml does not match GOLANGCI_CONFIG_SHA256" >&2
echo "in backend/script/lint." >&2
echo " expected $GOLANGCI_CONFIG_SHA256" >&2
echo " actual $actual" >&2
echo "Restore it verbatim from sneak/prompts; do not edit it." >&2
echo "Compare it with the org standard," >&2
echo "https://git.eeqj.de/sneak/prompts/raw/branch/main/.golangci.yml" >&2
echo "- If they differ, it was edited here: restore the org standard" >&2
echo " verbatim. Do not edit it." >&2
echo "- If they are the same, the org standard changed: set" >&2
echo " GOLANGCI_CONFIG_SHA256 in backend/script/lint to the actual hash." >&2
exit 1
fi
golangci-lint run ./...
+10 -2
View File
@@ -1,12 +1,20 @@
#!/bin/sh
# script/test: run the backend test suite.
# script/test: run the backend test suite with the race detector and
# coverage. Go's own -timeout bounds the tests and not their compile,
# so a cold build cache cannot fail it. The race detector needs cgo,
# and so a C compiler. If the tests fail, they run again with -v for
# the details, and the script fails even if that run passes.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
main() {
cd "$ROOT"
timeout 30 go test ./...
go test -timeout 30s -race -cover ./... || {
echo "--- Rerunning with -v for details ---"
go test -timeout 30s -race -v ./...
exit 1
}
}
main "$@"
+6 -19
View File
@@ -63,26 +63,13 @@ done > /etc/nginx/trusted-proxies.conf
# netwatch-server keeps its report files in DATA_DIR, on the /data
# volume, which may be a host directory owned by root or by another
# uid. Both are given to the netwatch user here, with the mode the
# server gives a directory it creates, so the host directory needs no
# preparing.
#
# chown and chmod, run as root, change whatever a symbolic link on the
# path points to, anywhere in the container, and the netwatch user can
# put one in /data. So the start stops unless readlink -f, which
# follows every link on a path, gives /data and DATA_DIR back as they
# are. It also writes a path in full, so a DATA_DIR with '.', '..' or
# an extra '/' in it is refused too.
# uid. Here, as root, netwatch-server prepare-data-dir creates DATA_DIR
# and gives /data and everything in it to the netwatch user, so the
# host directory needs no preparing. It stops the start, naming
# DATA_DIR, unless DATA_DIR is /data or a path below it, and it acts on
# nothing outside /data, whatever symbolic links it meets there.
export DATA_DIR="${DATA_DIR:-/data/reports}"
mkdir -p "$DATA_DIR" || exit 1
if [ "$(readlink -f /data)" != /data ] ||
[ "$(readlink -f "$DATA_DIR")" != "$DATA_DIR" ]; then
echo "entrypoint: DATA_DIR must be a full path with no '.', '..'," \
"extra '/' or symbolic link on it or on /data, not '$DATA_DIR'" >&2
exit 1
fi
chown -R netwatch:netwatch /data "$DATA_DIR" || exit 1
chmod 750 /data "$DATA_DIR" || exit 1
netwatch-server prepare-data-dir "$DATA_DIR" || exit 1
# A stop signal is only noted here; the loop below acts on it.
stop_requested=""
+27
View File
@@ -0,0 +1,27 @@
// eslint's recommended rules over every JavaScript file in the repo.
// script/frontend-lint runs it, in the frontend-lint stage of Dockerfile.
import js from "@eslint/js";
import globals from "globals";
import { defineConfig } from "eslint/config";
export default defineConfig([
js.configs.recommended,
{
// The page, and facts.js, which the viewport test runs inside it.
files: ["src/**/*.js", "test/viewport/facts.js"],
languageOptions: {
globals: {
...globals.browser,
// vite.config.js defines these when it builds the page.
__COMMIT_HASH__: "readonly",
__COMMIT_FULL__: "readonly",
},
},
},
{
// Run by node: the tests, the viewport harness and these configs.
files: ["test/**/*.js", "*.js"],
ignores: ["test/viewport/facts.js"],
languageOptions: { globals: globals.node },
},
]);
+6
View File
@@ -68,4 +68,10 @@ server {
location = /.well-known/healthcheck {
proxy_pass http://127.0.0.1:8081;
}
# The backend's Prometheus metrics, behind its own basic auth. Unless
# METRICS_USERNAME and METRICS_PASSWORD are set, it answers 404.
location = /metrics {
proxy_pass http://127.0.0.1:8081;
}
}
+5 -1
View File
@@ -6,12 +6,16 @@
"scripts": {
"dev": "vite",
"build": "vite build",
"preview": "vite preview"
"preview": "vite preview",
"test": "node --test test/unit/*.test.js"
},
"license": "MIT",
"devDependencies": {
"@eslint/js": "^10.0.1",
"@tailwindcss/vite": "^4.1.18",
"autoprefixer": "^10.4.23",
"eslint": "^10.12.0",
"globals": "^17.13.0",
"postcss": "^8.5.6",
"prettier": "^3.8.1",
"puppeteer-core": "25.5.0",
+30
View File
@@ -0,0 +1,30 @@
#!/bin/sh
# script/add-dependency: add a frontend package, or move one to another
# version, with yarn add, which changes package.json and yarn.lock
# together; then install from yarn.lock with --frozen-lockfile, as
# script/bootstrap does, to show it installs as written. --dev because
# no frontend package is needed when the page runs: it ships as the
# built dist/.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
usage() {
echo "usage: make add-dependency PACKAGE=<name>@<version>" >&2
exit 2
}
main() {
# Exactly one package. A value beginning with - would reach yarn as
# an option; yarn would quietly drop all but the first of several
# packages given in one value.
[ "$#" -eq 1 ] || usage
case "$1" in
"" | -* | *[[:space:]]*) usage ;;
esac
cd "$ROOT"
yarn add --dev "$1"
yarn install --frozen-lockfile
}
main "$@"
+17 -8
View File
@@ -17,16 +17,18 @@
#
# golangci-lint is not installed: make lint runs it in Docker, which
# this script does not install either.
#
# Unlike the org model: Go and gcc for backend/, a newer node for eslint.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
# Pinned versions, 2026-07-07
# Pinned versions, 2026-07-06
NODE_VERSION="22.17.0"
# The oldest node the frontend's dependencies accept: the "engines"
# field of puppeteer-core 25.5.0, the most demanding of them, asks for
# 22.12.0 or newer, 2026-09-29. An older installed node is not used.
NODE_MIN_VERSION="22.12.0"
# field of eslint 10.12.0, the most demanding of them, asks for 22.13.0
# or newer, 2026-10-03. An older installed node is not used.
NODE_MIN_VERSION="22.13.0"
NVM_VERSION="0.40.3"
# sha256 of https://github.com/nvm-sh/nvm/archive/refs/tags/v0.40.3.tar.gz
NVM_SHA256="5f4d6aaa04a177dc93c985e31dbc411ab6b8c6e1e21d8015dbc1372625fcd1d0"
@@ -64,7 +66,8 @@ detect_pkgmgr() {
fi
}
# pkg_install <nix-attr> <apt-pkg> <brew-formula> <apk-pkg>
# pkg_install <nix-attr> <apt-pkgs> <brew-formula> <apk-pkgs>: the apt
# and apk arguments may each list several packages, separated by spaces.
pkg_install() {
detect_pkgmgr
case "$PKGMGR" in
@@ -74,10 +77,10 @@ pkg_install() {
$SUDO env DEBIAN_FRONTEND=noninteractive apt-get update
APT_UPDATED=1
fi
$SUDO env DEBIAN_FRONTEND=noninteractive apt-get install -y "$2"
$SUDO env DEBIAN_FRONTEND=noninteractive apt-get install -y $2
;;
brew) brew install "$3" ;;
apk) apk add --no-cache "$4" ;;
apk) apk add --no-cache $4 ;;
esac
}
@@ -253,6 +256,12 @@ main() {
if missing make; then pkg_install gnumake make make make; fi
if missing git; then pkg_install git git git git; fi
# The race detector in make test needs cgo, which Go turns on only
# when it finds its C compiler, gcc on Linux. apt and apk ship the C
# library headers apart from gcc.
if missing gcc; then
pkg_install gcc "gcc libc6-dev" gcc "gcc musl-dev"
fi
ensure_node
ensure_yarn
@@ -263,7 +272,7 @@ main() {
if missing docker; then
echo "bootstrap: docker not found; make lint, and so make check" >&2
echo " and the pre-commit hook, need it to run the Go linter" >&2
echo " and the pre-commit hook, need it to run the linters" >&2
fi
if [ -n "$path_hint" ] && [ -d "$BIN_DIR" ]; then
echo "bootstrap: add $BIN_DIR to the front of your PATH, e.g." >&2
Executable
+13
View File
@@ -0,0 +1,13 @@
#!/bin/sh
# script/build: build the frontend for production into dist/. The Go
# backend is built by backend/script/build.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
main() {
cd "$ROOT"
yarn build
}
main "$@"
Executable
+13
View File
@@ -0,0 +1,13 @@
#!/bin/sh
# script/dev: run the frontend's Vite dev server. It proxies /api to a
# netwatch-server running locally (see vite.config.js).
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
main() {
cd "$ROOT"
yarn dev
}
main "$@"
+1
View File
@@ -1,6 +1,7 @@
#!/bin/sh
# script/fmt: format the whole repo (writes): prettier over everything
# it understands, then gofmt over the Go backend.
# The org model formats only markdown; this repo also has JS and Go.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
+1
View File
@@ -1,6 +1,7 @@
#!/bin/sh
# script/fmt-check: check formatting across the whole repo (read-only).
# Same scope as script/fmt, but fails instead of writing.
# The org model checks only markdown; this repo also has JS and Go.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
+6 -6
View File
@@ -1,9 +1,10 @@
#!/bin/sh
# script/frontend-check: run the frontend half of the checks only (test,
# lint, fmt-check). This exists for the frontend stage of Dockerfile, a
# node image with neither Go nor Docker; the Dockerfile's lint and
# backend build stages gate the backend half. Everywhere else, use
# script/check, which covers the whole repo. Must not modify any files.
# script/frontend-check: run the frontend tests and format check only.
# This exists for the frontend stage of Dockerfile, a node image with
# neither Go nor Docker; the Dockerfile's frontend-lint stage runs the
# frontend linter, and its lint and builder stages gate the backend.
# Everywhere else, use script/check, which covers the whole repo. Must
# not modify any files.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
@@ -11,7 +12,6 @@ ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
main() {
cd "$ROOT"
"$ROOT/script/frontend-test"
"$ROOT/script/frontend-lint"
"$ROOT/script/frontend-fmt-check"
}
+2 -2
View File
@@ -1,7 +1,7 @@
#!/bin/sh
# script/frontend-fmt: format the frontend and every other file prettier
# understands, repo-wide (writes). backend/ is in .prettierignore; Go
# sources are formatted by backend/script/fmt.
# understands, repo-wide (writes), the markdown in backend/ included.
# Prettier does not read Go; backend/script/fmt formats the Go sources.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
+5 -2
View File
@@ -1,12 +1,15 @@
#!/bin/sh
# script/frontend-lint: run the frontend linter (prettier in check mode).
# script/frontend-lint: run eslint over the frontend. This runs inside
# the frontend-lint stage of Dockerfile, the digest-pinned node image
# with the packages from yarn.lock. From a checkout, run `make lint`,
# which builds that stage.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
main() {
cd "$ROOT"
yarn prettier --check .
yarn eslint .
}
main "$@"
+11 -3
View File
@@ -1,14 +1,22 @@
#!/bin/sh
# script/frontend-test: run the frontend test suite: the unit tests in
# test/unit/ with Node's built-in test runner, then the production
# build, which fails on broken code.
# test/unit/, through the test script in package.json, then the
# production build, which fails on broken code. The tests print a dot
# each; if any fails, they run again with every test listed, and the
# script fails even if that run passes. NODE_OPTIONS chooses the
# reporter because yarn adds its arguments after the test files, where
# node would take a reporter option for one more file.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
main() {
cd "$ROOT"
timeout 30 node --test test/unit/*.test.js
NODE_OPTIONS=--test-reporter=dot timeout 30 yarn --silent run test || {
echo "--- Rerunning with every test listed for details ---"
NODE_OPTIONS=--test-reporter=spec timeout 30 yarn --silent run test
exit 1
}
timeout 30 yarn build
}
+2 -2
View File
@@ -8,8 +8,8 @@ ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
main() {
cd "$ROOT"
hook=".git/hooks/pre-commit"
printf '#!/bin/sh\nset -e\nscript/precommit\n' > "$hook"
chmod +x "$hook"
printf '#!/bin/sh\nset -e\nscript/precommit\n' > .git/hooks/pre-commit
chmod +x .git/hooks/pre-commit
echo "pre-commit hook installed: runs script/precommit"
}
+10 -8
View File
@@ -1,19 +1,21 @@
#!/bin/sh
# script/lint: lint the whole repo: prettier over the frontend, then the
# Go linter over backend/.
# script/lint: lint the whole repo: eslint over the frontend, then the Go
# linter over backend/.
#
# The Go linter runs only in Docker: this builds the lint stage of
# Dockerfile, the digest-pinned golangci-lint image, which runs the
# backend's fmt-check and lint targets. --no-cache makes the linter
# really run every time rather than reuse an earlier result, and the
# stage is built for its checks alone, so no image is kept.
# No linter runs on the host: this builds the frontend-lint and lint
# stages of Dockerfile, the digest-pinned node and golangci-lint images.
# The first runs eslint; the second runs the backend's fmt-check and
# lint targets. --no-cache makes each linter really run every time
# rather than reuse an earlier result, and each stage is built for its
# checks alone, so no image is kept.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
main() {
cd "$ROOT"
"$ROOT/script/frontend-lint"
timeout 300 docker build --no-cache --target frontend-lint \
--output type=cacheonly .
timeout 300 docker build --no-cache --target lint \
--output type=cacheonly .
}
+6 -3
View File
@@ -1,14 +1,17 @@
#!/bin/sh
# script/test: run the test suite for the whole repo: the frontend at
# the repo root, then the Go backend in backend/. Both halves together
# get 30 seconds; each also keeps its own limit for the Dockerfiles.
# the repo root, then the Go backend in backend/. Each half has its own
# 30-second limit, and there is none around both: from a cold Go build
# cache, compiling the backend's tests with the race detector can take
# 30 seconds on its own, and Go's -timeout leaves the compile out.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
main() {
cd "$ROOT"
timeout 30 sh -c 'script/frontend-test && backend/script/test'
script/frontend-test
backend/script/test
}
main "$@"
Executable
+13
View File
@@ -0,0 +1,13 @@
#!/bin/sh
# script/tidy: run go mod tidy in backend/, which adds the modules the
# Go sources import, drops those they no longer do, and updates go.sum.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
main() {
cd "$ROOT/backend"
go mod tidy
}
main "$@"
+161 -117
View File
@@ -8,6 +8,8 @@
// display their real value in the latency figure. The history buffer holds
// maxHistoryPoints samples (historyDuration / updateInterval).
// reportInterval is how often collected samples are POSTed to the backend.
// The interval menu changes updateInterval while the page runs; the
// getters compute their values from it each time they are read.
export const CONFIG = {
updateInterval: 3000,
maxHistoryPoints: 100,
@@ -28,6 +30,39 @@ export const CONFIG = {
return [0, 1, 2, 3, 4, 5].map((i) => Math.round((d * i) / 5));
},
canvasHeight: 96,
// A latency figure and its sparkline take the color of the first entry
// whose limit, in ms, the latency is below.
latencyColors: [
{ below: 50, hex: "#22c55e", className: "text-green-500" },
{ below: 100, hex: "#84cc16", className: "text-lime-500" },
{ below: 200, hex: "#eab308", className: "text-yellow-500" },
{ below: 500, hex: "#f97316", className: "text-orange-500" },
{ below: Infinity, hex: "#ef4444", className: "text-red-500" },
],
// The health is offline when more than offlineTimeouts WAN hosts timed
// out or were unreachable and at most offlineReachable answered;
// otherwise degraded when more than degradedTimeouts timed out or were
// unreachable; otherwise slow when more than slowHosts answered after
// more than slowLatency ms.
offlineTimeouts: 10,
offlineReachable: 4,
degradedTimeouts: 4,
slowHosts: 3,
slowLatency: 1000,
// The debug log keeps its last maxLogEntries lines.
maxLogEntries: 1000,
// A gateway candidate that has not answered after gatewayTimeout ms is
// passed over.
gatewayTimeout: 1500,
// When no WAN host answers, the recovery probe checks recoveryProbeHosts
// random ones every recoveryProbeInterval ms.
recoveryProbeHosts: 4,
recoveryProbeInterval: 500,
// The rows are sorted after the first round that is not discarded, then
// every roundsPerSort rounds.
roundsPerSort: 10,
// The sparklines are sized and drawn again resizeDelay ms after start.
resizeDelay: 100,
};
// WAN endpoints to monitor. These are used for the aggregate health/stats
@@ -112,7 +147,8 @@ const debugLog = [];
const log = (() => {
function append(level, message) {
debugLog.push({ timestamp: new Date(), level, message });
if (debugLog.length > 1000) debugLog.splice(0, debugLog.length - 1000);
if (debugLog.length > CONFIG.maxLogEntries)
debugLog.splice(0, debugLog.length - CONFIG.maxLogEntries);
const panel = document.getElementById("debug-panel");
if (panel && !panel.classList.contains("hidden")) renderDebugLog();
}
@@ -151,7 +187,7 @@ function formatUTCTimestamp(date) {
// --- Duration Formatting -----------------------------------------------------
function humanDuration(seconds) {
export function humanDuration(seconds) {
const h = Math.floor(seconds / 3600);
const m = Math.floor((seconds % 3600) / 60);
const s = seconds % 60;
@@ -172,7 +208,10 @@ async function detectGateway() {
const result = await Promise.any(
GATEWAY_CANDIDATES.map(async (url) => {
const controller = new AbortController();
const timeoutId = setTimeout(() => controller.abort(), 1500);
const timeoutId = setTimeout(
() => controller.abort(),
CONFIG.gatewayTimeout,
);
try {
await fetch(url, {
method: "GET",
@@ -197,11 +236,35 @@ async function detectGateway() {
// --- App State ---------------------------------------------------------------
class HostState {
// The min, max, median and average of latencies, a list of numbers, or all
// null when it is empty. The median of an even count is the mean of the
// middle two; it and the average are rounded.
function latencyStats(latencies) {
if (latencies.length === 0)
return { min: null, max: null, med: null, avg: null };
const sorted = [...latencies].sort((a, b) => a - b);
const mid = Math.floor(sorted.length / 2);
return {
min: sorted[0],
max: sorted[sorted.length - 1],
med:
sorted.length % 2
? sorted[mid]
: Math.round((sorted[mid - 1] + sorted[mid]) / 2),
avg: Math.round(
latencies.reduce((a, b) => a + b, 0) / latencies.length,
),
};
}
export class HostState {
constructor(host, pinned = false) {
this.name = host.name;
this.url = host.url;
this.history = []; // { timestamp, latency, paused }
// Each entry is either a check's result, { timestamp, latency,
// error }, or a round skipped while paused, { timestamp,
// latency: null, paused: true }.
this.history = [];
this.lastLatency = null;
this.status = "pending"; // 'online' | 'offline' | 'error' | 'pending'
this.pinned = pinned;
@@ -225,45 +288,23 @@ class HostState {
this._trim();
}
averageLatency() {
const valid = this.history.filter((p) => p.latency !== null);
if (valid.length === 0) return null;
return Math.round(
valid.reduce((s, p) => s + p.latency, 0) / valid.length,
// The min, max, median and average latency of the checks in the history
// that got an answer.
historyStats() {
return latencyStats(
this.history
.filter((p) => p.latency !== null)
.map((p) => p.latency),
);
}
minLatency() {
const valid = this.history.filter((p) => p.latency !== null);
if (valid.length === 0) return null;
return Math.min(...valid.map((p) => p.latency));
}
maxLatency() {
const valid = this.history.filter((p) => p.latency !== null);
if (valid.length === 0) return null;
return Math.max(...valid.map((p) => p.latency));
}
medianLatency() {
const sorted = this.history
.filter((p) => p.latency !== null)
.map((p) => p.latency)
.sort((a, b) => a - b);
if (sorted.length === 0) return null;
const mid = Math.floor(sorted.length / 2);
return sorted.length % 2
? sorted[mid]
: Math.round((sorted[mid - 1] + sorted[mid]) / 2);
}
_trim() {
while (this.history.length > CONFIG.maxHistoryPoints)
this.history.shift();
}
}
class AppState {
export class AppState {
constructor(localHosts) {
this.wan = WAN_HOSTS.map(
(h) => new HostState(h, h.name === "datavi.be"),
@@ -271,6 +312,10 @@ class AppState {
this.local = localHosts.map((h) => new HostState(h));
this.paused = false;
this.tickCount = 0;
// The recovery probe's timer, null while it is not running, and the
// checks it started last.
this._recoveryProbeId = null;
this._recoveryProbeChecks = null;
}
get allHosts() {
@@ -279,33 +324,13 @@ class AppState {
/** WAN-only stats from latest sample (excludes local) */
wanStats() {
const reachable = this.wan.filter((h) => h.lastLatency !== null);
const latencies = reachable.map((h) => h.lastLatency);
const total = this.wan.length;
if (latencies.length === 0)
return {
reachable: 0,
total,
min: null,
max: null,
med: null,
avg: null,
};
const sorted = [...latencies].sort((a, b) => a - b);
const mid = Math.floor(sorted.length / 2);
const med =
sorted.length % 2
? sorted[mid]
: Math.round((sorted[mid - 1] + sorted[mid]) / 2);
const latencies = this.wan
.filter((h) => h.lastLatency !== null)
.map((h) => h.lastLatency);
return {
reachable: latencies.length,
total,
min: Math.min(...latencies),
max: Math.max(...latencies),
med,
avg: Math.round(
latencies.reduce((a, b) => a + b, 0) / latencies.length,
),
total: this.wan.length,
...latencyStats(latencies),
};
}
@@ -331,12 +356,16 @@ class AppState {
const timeouts = this.wan.filter(
(h) => h.status === "error" || h.status === "offline",
).length;
if (timeouts > 10 && reachable <= 4) return "offline";
if (timeouts > 4) return "degraded";
if (
timeouts > CONFIG.offlineTimeouts &&
reachable <= CONFIG.offlineReachable
)
return "offline";
if (timeouts > CONFIG.degradedTimeouts) return "degraded";
const slow = this.wan.filter(
(h) => h.lastLatency !== null && h.lastLatency > 1000,
(h) => h.lastLatency !== null && h.lastLatency > CONFIG.slowLatency,
).length;
if (slow > 3) return "slow";
if (slow > CONFIG.slowHosts) return "slow";
return "healthy";
}
@@ -546,23 +575,15 @@ export async function measureLatency(url, signal) {
// --- Color Helpers -----------------------------------------------------------
function latencyHex(latency) {
export function latencyHex(latency) {
if (latency === null) return "#6b7280";
if (latency < 50) return "#22c55e";
if (latency < 100) return "#84cc16";
if (latency < 200) return "#eab308";
if (latency < 500) return "#f97316";
return "#ef4444";
return CONFIG.latencyColors.find((c) => latency < c.below).hex;
}
function latencyClass(latency, status) {
export function latencyClass(latency, status) {
if (status === "offline" || status === "error" || latency === null)
return "text-gray-500";
if (latency < 50) return "text-green-500";
if (latency < 100) return "text-lime-500";
if (latency < 200) return "text-yellow-500";
if (latency < 500) return "text-orange-500";
return "text-red-500";
return CONFIG.latencyColors.find((c) => latency < c.below).className;
}
// --- Sparkline Renderer ------------------------------------------------------
@@ -580,8 +601,8 @@ class SparklineRenderer {
const ch = h - m.top - m.bottom;
ctx.clearRect(0, 0, w, h);
SparklineRenderer._drawYAxis(ctx, w, h, m, ch);
SparklineRenderer._drawXAxis(ctx, w, h, m, cw);
SparklineRenderer._drawYAxis(ctx, w, m, ch);
SparklineRenderer._drawXAxis(ctx, h, m, cw);
const len = history.length;
const pw = cw / (CONFIG.maxHistoryPoints - 1);
@@ -597,7 +618,7 @@ class SparklineRenderer {
SparklineRenderer._drawTip(ctx, history, getX, getY);
}
static _drawYAxis(ctx, w, h, m, ch) {
static _drawYAxis(ctx, w, m, ch) {
ctx.font = "300 12px monospace";
ctx.textAlign = "right";
ctx.textBaseline = "middle";
@@ -614,7 +635,7 @@ class SparklineRenderer {
}
}
static _drawXAxis(ctx, w, h, m, cw) {
static _drawXAxis(ctx, h, m, cw) {
ctx.textAlign = "center";
ctx.textBaseline = "top";
for (const tick of CONFIG.xAxisTicks) {
@@ -702,7 +723,18 @@ class SparklineRenderer {
// horizontally.
const STATUS_TEXT_CLASS = "status-text text-xs text-right col-span-2 mt-5";
function hostRowHTML(host, index, showPin = true) {
// Escapes text for HTML, so it shows as written inside an element or a
// quoted attribute and is never read as markup.
function escapeHTML(text) {
return text
.replaceAll("&", "&amp;")
.replaceAll("<", "&lt;")
.replaceAll(">", "&gt;")
.replaceAll('"', "&quot;")
.replaceAll("'", "&#39;");
}
export function hostRowHTML(host, index, showPin = true) {
const pinColor = host.pinned
? "text-blue-500"
: "text-gray-600 hover:text-gray-400";
@@ -721,12 +753,12 @@ function hostRowHTML(host, index, showPin = true) {
<div class="w-[420px] flex-shrink-0 grid grid-cols-[minmax(0,1fr)_auto] items-center">
<div class="flex items-center gap-2 min-w-[200px]">
<div class="w-3 h-3 rounded-full flex-shrink-0 bg-[#6b7280]"></div>
<span class="font-medium text-white truncate">${host.name}</span>
<span class="font-medium text-white truncate">${escapeHTML(host.name)}</span>
</div>
<div class="latency-value text-4xl font-bold tabular-nums text-right mt-3" data-host="${index}">
<span class="text-gray-500">---</span>
</div>
<a href="${host.url}" target="_blank" rel="noopener" class="text-xs text-gray-500 truncate block col-span-2 -mt-2">${host.url}</a>
<a href="${escapeHTML(host.url)}" target="_blank" rel="noopener" class="text-xs text-gray-500 truncate block col-span-2 -mt-2">${escapeHTML(host.url)}</a>
<div class="${STATUS_TEXT_CLASS} text-gray-500" data-host="${index}">waiting...</div>
</div>
<div class="flex-grow sparkline-container rounded overflow-hidden border border-gray-700/30">
@@ -825,7 +857,7 @@ function buildUI(state) {
</div>
<footer class="mt-8 text-center text-gray-600 text-xs">
<p>Latency measured via GET requests | IPv4 only | CORS restrictions may affect some measurements</p>
<p>Latency measured via GET requests | CORS restrictions may affect some measurements</p>
<p class="mt-2">
<span class="inline-block w-3 h-3 rounded-full bg-green-500 mr-1 align-middle"></span>&lt;50ms
<span class="inline-block w-3 h-3 rounded-full bg-lime-500 mr-1 ml-3 align-middle"></span>&lt;100ms
@@ -892,10 +924,7 @@ function updateHostRow(host, index) {
latencyEl.innerHTML = `<span class="text-gray-500">---</span>`;
}
const avg = host.averageLatency();
const med = host.medianLatency();
const min = host.minLatency();
const max = host.maxLatency();
const { min, med, avg, max } = host.historyStats();
if (host.status === "online" && avg !== null) {
statusEl.innerHTML = statusStatsHTML([
["min", min],
@@ -1042,14 +1071,16 @@ function renderDebugLog() {
info: "text-gray-300",
debug: "text-gray-500",
};
el.innerHTML = debugLog
.map((entry) => {
el.replaceChildren(
...debugLog.map((entry) => {
const ts = formatUTCTimestamp(entry.timestamp);
const cls = levelColors[entry.level] || "text-gray-400";
const lvl = entry.level.toUpperCase().padEnd(7);
return `<div class="${cls}">${ts} ${lvl} ${entry.message}</div>`;
})
.join("");
const line = document.createElement("div");
line.className = levelColors[entry.level] || "text-gray-400";
line.textContent = `${ts} ${lvl} ${entry.message}`;
return line;
}),
);
el.scrollTop = el.scrollHeight;
}
@@ -1105,7 +1136,7 @@ function sortAndRebuildWAN(state) {
// --- Main Loop ---------------------------------------------------------------
async function tick(state, signal, onOffline) {
export async function tick(state, signal, onOffline) {
const ts = Date.now();
if (state.paused) {
@@ -1126,12 +1157,26 @@ async function tick(state, signal, onOffline) {
log.debug(`Tick #${state.tickCount + 1} started`);
const results = await Promise.all(
state.allHosts.map((h) => measureLatency(h.url, signal)),
// Each host's row shows its result as soon as its check ends. The
// result is discarded if by then the user has paused or the next round
// has given up this one's checks, and in the first tick (tickCount is
// still 0), which is discarded as a whole below. The row is looked up
// when the check ends, as a pin click may have re-sorted the rows since
// the round started.
await Promise.all(
state.allHosts.map(async (host) => {
const r = await measureLatency(host.url, signal);
if (state.paused || signal.aborted || state.tickCount === 0) {
return;
}
host.pushSample(ts, r);
updateHostRow(host, state.allHosts.indexOf(host));
log.debug(`${host.name}: ${r.error ? r.error : r.latency + "ms"}`);
}),
);
// User may have paused, or the next round may have given up this
// one's checks, while awaiting results — discard them
// one's checks, while awaiting results — skip the rest of the round
if (state.paused || signal.aborted) return;
state.tickCount++;
@@ -1142,15 +1187,13 @@ async function tick(state, signal, onOffline) {
return;
}
state.allHosts.forEach((host, i) => {
const r = results[i];
host.pushSample(ts, r);
updateHostRow(host, i);
log.debug(`${host.name}: ${r.error ? r.error : r.latency + "ms"}`);
});
// Redraw every row: if the user paused and resumed during this round,
// rows whose check ended before the resume still read "paused"
state.allHosts.forEach((host, i) => updateHostRow(host, i));
// Sort after the first real check, then every 10 ticks thereafter
if (state.tickCount === 2 || state.tickCount % 10 === 1) {
// Sort after the first real check, then every CONFIG.roundsPerSort
// ticks thereafter
if (state.tickCount === 2 || state.tickCount % CONFIG.roundsPerSort === 1) {
sortAndRebuildWAN(state);
}
@@ -1173,9 +1216,10 @@ async function tick(state, signal, onOffline) {
// --- Recovery Probe ----------------------------------------------------------
// When offline, check 4 random WAN hosts every 500ms, giving up the checks
// started 500ms before, so at most 4 are ever waiting. As soon as one
// answers, stop probing and start a new round at once.
// When offline, check CONFIG.recoveryProbeHosts random WAN hosts every
// CONFIG.recoveryProbeInterval ms, giving up the checks started one interval
// before, so at most that many are ever waiting. As soon as one answers,
// stop probing and start a new round at once.
function startRecoveryProbe(state, startRounds) {
if (state._recoveryProbeId) return; // already running
const candidates = [...state.wan];
@@ -1183,7 +1227,7 @@ function startRecoveryProbe(state, startRounds) {
const j = Math.floor(Math.random() * (i + 1));
[candidates[i], candidates[j]] = [candidates[j], candidates[i]];
}
const canaries = candidates.slice(0, 4);
const canaries = candidates.slice(0, CONFIG.recoveryProbeHosts);
log.notice(
`Recovery probe started (${canaries.map((h) => h.name).join(", ")})`,
);
@@ -1200,7 +1244,7 @@ function startRecoveryProbe(state, startRounds) {
startRounds();
});
}
}, 500);
}, CONFIG.recoveryProbeInterval);
}
function stopRecoveryProbe(state) {
@@ -1213,7 +1257,7 @@ function stopRecoveryProbe(state) {
// --- Pause / Resume ----------------------------------------------------------
function greyOutUI(state) {
export function greyOutUI(state) {
// Grey out all host rows
state.allHosts.forEach((host, i) => {
const latencyEl = document.querySelector(
@@ -1385,10 +1429,10 @@ async function init() {
setInterval(updateClocks, 1000);
// Rounds never overlap: a round first gives up the last round's checks
// if they are still waiting, and the last round then records nothing.
// At a steady interval they never are, as they time out at 80% of it;
// they can be when a round starts early, after an interval change or
// when the recovery probe finds a target answering.
// if they are still waiting, and the last round then records nothing
// more. At a steady interval they never are, as they time out at 80% of
// it; they can be when a round starts early, after an interval change
// or when the recovery probe finds a target answering.
let roundChecks = new AbortController();
function doTick() {
roundChecks.abort();
@@ -1464,7 +1508,7 @@ async function init() {
});
window.addEventListener("resize", () => handleResize(state));
setTimeout(() => handleResize(state), 100);
setTimeout(() => handleResize(state), CONFIG.resizeDelay);
}
// Bootstrap only when loaded as the page: a real DOM containing the #app
+380 -14
View File
@@ -1,19 +1,76 @@
// Unit tests for src/main.js, run by script/frontend-test with Node's
// built-in test runner. Importing the module does not start the page.
// Unit tests for src/main.js, run with Node's built-in test runner by the
// test script in package.json. Importing the module does not start the
// page.
import { test } from "node:test";
import { beforeEach, test } from "node:test";
import assert from "node:assert/strict";
import { CONFIG, measureLatency } from "../../src/main.js";
import {
AppState,
CONFIG,
greyOutUI,
hostRowHTML,
HostState,
humanDuration,
latencyClass,
latencyHex,
measureLatency,
tick,
} from "../../src/main.js";
// measureLatency writes timeouts to the debug log, which looks for its
// panel in the page. There is no page here.
globalThis.document = { getElementById: () => null };
// There is no page here, so the tests stand in for it. The debug log looks
// for its panel by id and finds none. Each element of a host's row that
// tick or greyOutUI draws into is a plain object, made the first time a
// test looks it up and kept in elements under its selector until the next
// test starts. As on a page, writing its text replaces its markup; the
// status dot greyOutUI looks for in it is not there. Drawing a sparkline
// does nothing; it looks for the pixel ratio on window and finds none.
let elements;
beforeEach(() => {
elements = {};
});
const doNothing = () => {};
const canvasContext = {
clearRect: doNothing,
beginPath: doNothing,
moveTo: doNothing,
lineTo: doNothing,
stroke: doNothing,
fill: doNothing,
fillRect: doNothing,
fillText: doNothing,
arc: doNothing,
};
globalThis.window = {};
globalThis.document = {
getElementById: () => null,
querySelector: (selector) =>
(elements[selector] ??= {
getContext: () => canvasContext,
querySelector: () => null,
set textContent(text) {
this.innerHTML = text;
},
}),
};
// What tick last wrote into the latency figure in host's row, or undefined
// if it has written nothing there.
function latencyFigure(state, host) {
const index = state.allHosts.indexOf(host);
return elements[`.latency-value[data-host="${index}"]`]?.innerHTML;
}
// What was last written into the status text in host's row.
function statusText(state, host) {
const index = state.allHosts.indexOf(host);
return elements[`.status-text[data-host="${index}"]`]?.innerHTML;
}
// Mocks the clock for test t, so that a check lasting seconds takes no real
// time, and replaces fetch with a target that answers after answerAfter
// milliseconds of that clock, or never when answerAfter is Infinity. Both
// are restored when the test ends.
function mockTarget(t, answerAfter) {
// time, and replaces fetch with targets that each answer after
// answerAfter(url) milliseconds of that clock, or never when that is
// Infinity. Both are restored when the test ends.
function mockTargets(t, answerAfter) {
t.mock.timers.enable({ apis: ["setTimeout", "Date"] });
t.mock.method(performance, "now", () => Date.now());
t.mock.method(
@@ -21,7 +78,9 @@ function mockTarget(t, answerAfter) {
"fetch",
(url, { signal }) =>
new Promise((resolve, reject) => {
if (answerAfter !== Infinity) setTimeout(resolve, answerAfter);
if (answerAfter(url) !== Infinity) {
setTimeout(resolve, answerAfter(url));
}
signal.addEventListener("abort", () => reject(signal.reason));
}),
);
@@ -42,7 +101,7 @@ for (const interval of [10000, 30000]) {
test(`at a ${interval}ms interval, an answer after ${slowAnswer}ms is recorded with its real time`, async (t) => {
CONFIG.updateInterval = interval;
mockTarget(t, slowAnswer);
mockTargets(t, () => slowAnswer);
const check = measureLatency("https://target.test");
t.mock.timers.tick(slowAnswer);
assert.deepEqual(await settled(check), {
@@ -53,7 +112,7 @@ for (const interval of [10000, 30000]) {
test(`at a ${interval}ms interval, a target that never answers is recorded as a timeout after ${timeout}ms`, async (t) => {
CONFIG.updateInterval = interval;
mockTarget(t, Infinity);
mockTargets(t, () => Infinity);
const check = measureLatency("https://target.test");
t.mock.timers.tick(timeout - 1);
assert.equal(await settled(check), "still waiting");
@@ -64,3 +123,310 @@ for (const interval of [10000, 30000]) {
});
});
}
test("at a 30000ms interval, a target answering after 1000ms shows in its row while another target's check is still waiting", async (t) => {
CONFIG.updateInterval = 30000;
const state = new AppState([
{ name: "Answering", url: "https://answering.test" },
]);
const answering = state.local[0];
const waiting = state.wan[0];
// No target but the answering one ever answers.
mockTargets(t, (url) => (url.startsWith(answering.url) ? 1000 : Infinity));
// The third tick: the first is discarded as a whole, and the second ends
// by sorting the rows, which rebuilds a page that is not here.
state.tickCount = 2;
const round = tick(state, new AbortController().signal);
t.mock.timers.tick(1000);
assert.equal(await settled(round), "still waiting");
assert.match(latencyFigure(state, answering), />1000</);
assert.equal(latencyFigure(state, waiting), undefined);
assert.equal(state.tickCount, 2);
// The round ends, once, when the last check times out.
t.mock.timers.tick(CONFIG.requestTimeout - 1000);
assert.notEqual(await settled(round), "still waiting");
assert.equal(state.tickCount, 3);
});
// In the next three tests, the answering target's check is still waiting
// when something happens after which its result must not show.
test("at a 30000ms interval, a check still waiting when its round is given up does not show in its row", async (t) => {
CONFIG.updateInterval = 30000;
const state = new AppState([
{ name: "Answering", url: "https://answering.test" },
]);
const answering = state.local[0];
mockTargets(t, (url) => (url.startsWith(answering.url) ? 1000 : Infinity));
state.tickCount = 2;
const roundChecks = new AbortController();
const round = tick(state, roundChecks.signal);
t.mock.timers.tick(500);
// As a round started early does to the last round's checks.
roundChecks.abort();
assert.notEqual(await settled(round), "still waiting");
assert.equal(latencyFigure(state, answering), undefined);
});
test("at a 30000ms interval, a check still waiting when the user pauses does not show in its row", async (t) => {
CONFIG.updateInterval = 30000;
const state = new AppState([
{ name: "Answering", url: "https://answering.test" },
]);
const answering = state.local[0];
mockTargets(t, (url) => (url.startsWith(answering.url) ? 1000 : Infinity));
state.tickCount = 2;
const round = tick(state, new AbortController().signal);
t.mock.timers.tick(500);
state.paused = true;
t.mock.timers.tick(500);
assert.equal(await settled(round), "still waiting");
assert.equal(latencyFigure(state, answering), undefined);
});
test("at a 30000ms interval, a check in the first round does not show in its row", async (t) => {
CONFIG.updateInterval = 30000;
const state = new AppState([
{ name: "Answering", url: "https://answering.test" },
]);
const answering = state.local[0];
mockTargets(t, (url) => (url.startsWith(answering.url) ? 1000 : Infinity));
const round = tick(state, new AbortController().signal);
t.mock.timers.tick(1000);
assert.equal(await settled(round), "still waiting");
assert.equal(latencyFigure(state, answering), undefined);
});
test("at a 30000ms interval, after the user pauses and resumes during a round, no row reads paused once its last check ends", async (t) => {
CONFIG.updateInterval = 30000;
const state = new AppState([
{ name: "Answering", url: "https://answering.test" },
]);
const answering = state.local[0];
mockTargets(t, (url) => (url.startsWith(answering.url) ? 1000 : Infinity));
state.tickCount = 2;
const round = tick(state, new AbortController().signal);
t.mock.timers.tick(1000);
assert.equal(await settled(round), "still waiting");
// The user pauses, which greys out every row, and resumes, which leaves
// the rows as they are, as togglePause does. The answering target's
// check has already ended, so only the redraw of every row at the end
// of the round can take "paused" out of its row.
state.paused = true;
greyOutUI(state);
state.paused = false;
assert.equal(statusText(state, answering), "paused");
// The round ends when the last check times out.
t.mock.timers.tick(CONFIG.requestTimeout - 1000);
assert.notEqual(await settled(round), "still waiting");
for (const host of state.allHosts) {
assert.notEqual(statusText(state, host), "paused", host.name);
}
});
// The page shows &lt; &gt; &quot; &amp; and &#39; in a row's markup as
// < > " & and '.
test(`a target whose name and URL hold < > " & and ' shows those characters in its row`, () => {
const host = new HostState({
name: `<b>"x" & 'y'</b>`,
url: `https://x.test/<b>?a="x"&b='y'`,
});
const row = hostRowHTML(host, 0);
assert.doesNotMatch(row, /<b>/);
assert.ok(
row.includes(
">&lt;b&gt;&quot;x&quot; &amp; &#39;y&#39;&lt;/b&gt;</span>",
),
);
assert.ok(
row.includes(
'href="https://x.test/&lt;b&gt;?a=&quot;x&quot;&amp;b=&#39;y&#39;"',
),
);
assert.ok(
row.includes(
">https://x.test/&lt;b&gt;?a=&quot;x&quot;&amp;b=&#39;y&#39;</a>",
),
);
});
for (const [seconds, text] of [
[0, "0s"],
[1, "1s"],
[59, "59s"],
[60, "1m"],
[61, "1m1s"],
[3599, "59m59s"],
[3600, "1h"],
[3601, "1h1s"],
[3660, "1h1m"],
[3661, "1h1m1s"],
]) {
test(`humanDuration writes ${seconds} seconds as ${text}`, () => {
assert.equal(humanDuration(seconds), text);
});
}
// The colour of a target's latency figure, as a class, and of its sparkline
// for an answer after latency ms, as the colour coding in README.md gives
// them. The rows sit either side of each boundary.
for (const [latency, figure, sparkline] of [
[0, "text-green-500", "#22c55e"],
[49, "text-green-500", "#22c55e"],
[50, "text-lime-500", "#84cc16"],
[99, "text-lime-500", "#84cc16"],
[100, "text-yellow-500", "#eab308"],
[199, "text-yellow-500", "#eab308"],
[200, "text-orange-500", "#f97316"],
[499, "text-orange-500", "#f97316"],
[500, "text-red-500", "#ef4444"],
]) {
test(`an answer after ${latency}ms has a ${figure} figure and a ${sparkline} sparkline`, () => {
assert.equal(latencyClass(latency, "online"), figure);
assert.equal(latencyHex(latency), sparkline);
});
}
test("a check that timed out or found its target unreachable has a grey figure and sparkline", () => {
assert.equal(latencyClass(null, "error"), "text-gray-500");
assert.equal(latencyClass(null, "offline"), "text-gray-500");
assert.equal(latencyHex(null), "#6b7280");
});
// A target whose checks, in turn, answered after each of latencies ms, or,
// for null, found it unreachable.
function hostAfter(latencies) {
const host = new HostState({ name: "Target", url: "https://target.test" });
for (const latency of latencies) {
host.pushSample(
Date.now(),
latency === null
? { latency: null, error: "unreachable" }
: { latency, error: null },
);
}
return host;
}
for (const { history, latencies, statistics } of [
{
history: "no checks",
latencies: [],
statistics: { min: null, max: null, average: null, median: null },
},
{
history: "only unreachable checks",
latencies: [null, null, null],
statistics: { min: null, max: null, average: null, median: null },
},
{
history: "three answers and an unreachable check",
latencies: [30, null, 10, 20],
statistics: { min: 10, max: 30, average: 20, median: 20 },
},
{
// The median of an even number of answers is the mean of the
// middle two. It and the average, 23.75, are rounded.
history: "four answers",
latencies: [10, 40, 20, 25],
statistics: { min: 10, max: 40, average: 24, median: 23 },
},
{
// Sorted as text rather than as numbers, these answers would put
// 100 in the middle. The average, 39.67, is rounded.
history: "answers with different numbers of digits",
latencies: [100, 9, 10],
statistics: { min: 9, max: 100, average: 40, median: 10 },
},
]) {
test(`a target's min, max, average and median latency over ${history}`, () => {
const { min, max, avg, med } = hostAfter(latencies).historyStats();
assert.deepEqual({ min, max, average: avg, median: med }, statistics);
});
}
// The summary's figures come from each WAN target's last check, by the same
// rules as a target's own: here four answered, one was found unreachable
// and the rest have not been checked yet. The median, 22.5, and the
// average, 21.25, are rounded.
test("the summary's min, max, median and average latency over the WAN targets' last checks", () => {
const state = new AppState([]);
[30, 10, null, 25, 20].forEach((latency, i) =>
state.wan[i].pushSample(
Date.now(),
latency === null
? { latency: null, error: "unreachable" }
: { latency, error: null },
),
);
assert.deepEqual(state.wanStats(), {
reachable: 4,
total: state.wan.length,
min: 10,
max: 30,
med: 23,
avg: 21,
});
});
// An app state in which, of the WAN targets, the first timedOut timed out,
// the next unreachable were found unreachable, the next answered answered
// after latency ms, and the rest have not been checked yet.
function stateAfter({ timedOut, unreachable, answered, latency }) {
const state = new AppState([]);
const results = [
...Array(timedOut).fill({ latency: null, error: "timeout" }),
...Array(unreachable).fill({ latency: null, error: "unreachable" }),
...Array(answered).fill({ latency, error: null }),
];
results.forEach((result, i) => state.wan[i].pushSample(Date.now(), result));
return state;
}
// For each number the health is decided by, the rows put it one under, at
// and one over its threshold.
for (const { timedOut, unreachable = 0, answered, latency, health } of [
// Offline: more than 10 timed out and at most 4 answered.
{ timedOut: 9, answered: 4, latency: 30, health: "degraded" },
{ timedOut: 10, answered: 4, latency: 30, health: "degraded" },
{ timedOut: 11, answered: 4, latency: 30, health: "offline" },
{ timedOut: 11, answered: 3, latency: 30, health: "offline" },
{ timedOut: 11, answered: 5, latency: 30, health: "degraded" },
// Otherwise degraded: more than 4 timed out.
{ timedOut: 3, answered: 10, latency: 30, health: "healthy" },
{ timedOut: 4, answered: 10, latency: 30, health: "healthy" },
{ timedOut: 5, answered: 10, latency: 30, health: "degraded" },
// Otherwise slow: more than 3 answered after more than 1000ms.
{ timedOut: 0, answered: 4, latency: 999, health: "healthy" },
{ timedOut: 0, answered: 4, latency: 1000, health: "healthy" },
{ timedOut: 0, answered: 4, latency: 1001, health: "slow" },
{ timedOut: 0, answered: 2, latency: 1001, health: "healthy" },
{ timedOut: 0, answered: 3, latency: 1001, health: "healthy" },
// A target found unreachable counts as one that timed out.
{
timedOut: 5,
unreachable: 6,
answered: 4,
latency: 30,
health: "offline",
},
{
timedOut: 0,
unreachable: 5,
answered: 10,
latency: 30,
health: "degraded",
},
]) {
test(`with ${timedOut} WAN targets timed out, ${unreachable} found unreachable and ${answered} answering after ${latency}ms, the health is ${health}`, () => {
const state = stateAfter({ timedOut, unreachable, answered, latency });
assert.equal(state.healthStatus(), health);
});
}
+13 -11
View File
@@ -49,10 +49,11 @@ sizes straddling the breakpoint for the rotation case.
excluded: it is a design choice, not breakage.
- **tap-targets-44px** — every interactive control is at least 44x44 CSS px on
touch viewports, _and_ each selector in the control list matched at least the
number of visible elements it declares. The second half is what stops the
check passing vacuously: with size alone, a renamed class would take its
controls out of the measured set and the check would report "all 0 controls
are at least 44x44" and pass. See below.
number of visible elements it declares: one of each single control, and one
pin button per WAN host row. The second half is what stops the check passing
vacuously: with size alone, a renamed class would take its controls out of the
measured set and the check would report "all 0 controls are at least 44x44"
and pass. See below.
- **host-rows-stacked / host-rows-side-by-side** — the rows genuinely reflow.
Computed `flex-direction` _and_ the actual geometry are checked, and in the
narrow layout the info block and the sparkline must each occupy essentially
@@ -76,7 +77,7 @@ the internet, so the app's latency probes cannot reach anything real. The
harness answers them itself from a fixed delay table, with a deterministic
fraction failed outright, so the rows render a realistic spread of one-, two-
and three-digit latencies plus some unreachable rows. That spread is what the
layout has to survive; 24 identical `---` placeholders would not exercise it.
layout has to survive; a `---` placeholder in every row would not exercise it.
## What this cannot verify
@@ -103,10 +104,11 @@ Everything else this issue was actually about — does the layout reflow, does
anything overflow, is content clipped, are the controls big enough — is a
function of viewport width and CSS, and is covered above.
## Relation to the unit test framework (#21)
## Relation to the unit tests
Complementary layers, not two stacks. `vitest` (#21) will exercise module-level
logic in-process with no browser. This harness exercises rendered layout in a
real engine and is the only thing here that can see a media query. Neither
replaces the other; assertions about computed styles and element geometry belong
here, assertions about functions belong in `vitest`.
Complementary layers, not two stacks. The unit tests in `test/unit/`, which
`make test` runs with Node's built-in test runner, exercise the functions
`src/main.js` exports in-process with no browser. This harness exercises
rendered layout in a real engine and is the only thing here that can see a media
query. Neither replaces the other; assertions about computed styles and element
geometry belong here, assertions about functions belong in `test/unit/`.
+14 -13
View File
@@ -15,7 +15,8 @@ export const MIN_TAP_TARGET_PX = 44;
// The controls named in the definition of done, plus the pause button.
// Each carries the smallest number of *visible* instances the page has to
// contain for the tap-target oracle to be measuring anything at all.
// contain, worked out from the facts gathered from that page, for the
// tap-target oracle to be measuring every control it should.
//
// Without those floors the check is inert: `undersized` is empty both when
// every control is large enough and when the selectors have gone stale and
@@ -24,13 +25,12 @@ export const MIN_TAP_TARGET_PX = 44;
// for all three singleton controls vanishing at once — so the floor is per
// selector, and one stale selector out of four fails the check.
export const INTERACTIVE_CONTROLS = [
{ selector: "#pause-btn", minCount: 1 },
{ selector: "#interval-select", minCount: 1 },
// One per pinnable host row. `app-rendered` already requires at least
// 10 host rows, so a count below that means the pin buttons stopped
// being rendered per row rather than that there were fewer hosts.
{ selector: ".pin-btn", minCount: 10 },
{ selector: "#debug-toggle", minCount: 1 },
{ selector: "#pause-btn", minCount: () => 1 },
{ selector: "#interval-select", minCount: () => 1 },
// One per WAN host row, so pin buttons missing from even one row fail
// the check rather than only a drop below some fixed number.
{ selector: ".pin-btn", minCount: (facts) => facts.wanRowCount },
{ selector: "#debug-toggle", minCount: () => 1 },
];
export const INTERACTIVE_SELECTORS = INTERACTIVE_CONTROLS.map(
@@ -105,8 +105,8 @@ export function evaluateChecks(facts, viewport, probes) {
// never rendered. Everything below is only meaningful if this holds.
check(
"app-rendered",
facts.rowCount >= 10 && facts.numericLatencies >= 5,
`${facts.rowCount} host rows, ${facts.numericLatencies} showing a numeric latency`,
facts.wanRowCount >= 10 && facts.numericLatencies >= 5,
`${facts.wanRowCount} WAN host rows, ${facts.numericLatencies} showing a numeric latency`,
);
const viewportWidth = Math.min(facts.innerWidth, facts.documentClientWidth);
@@ -162,7 +162,8 @@ export function evaluateChecks(facts, viewport, probes) {
seen.set(target.selector, (seen.get(target.selector) ?? 0) + 1);
}
const missing = INTERACTIVE_CONTROLS.filter(
(control) => (seen.get(control.selector) ?? 0) < control.minCount,
(control) =>
(seen.get(control.selector) ?? 0) < control.minCount(facts),
);
const undersized = facts.tapTargets.filter(
@@ -184,11 +185,11 @@ export function evaluateChecks(facts, viewport, probes) {
const detail = [];
if (missing.length > 0) {
detail.push(
"oracle is not measuring the page: " +
"oracle is not measuring every control: " +
summarise(
missing,
(c) =>
`${c.selector} matched ${seen.get(c.selector) ?? 0} visible element(s), expected at least ${c.minCount}`,
`${c.selector} matched ${seen.get(c.selector) ?? 0} visible element(s), expected at least ${c.minCount(facts)}`,
4,
),
);
+3 -1
View File
@@ -173,7 +173,9 @@ export function collectLayoutFacts(options) {
clipped,
tapTargets,
rows,
rowCount: document.querySelectorAll(".host-row").length,
// WAN host rows only: each has a pin button, and the tap-target
// check expects one per row. The local host rows have none.
wanRowCount: document.querySelectorAll("#wan-hosts .host-row").length,
numericLatencies: Array.from(
document.querySelectorAll(".latency-value"),
).filter((el) => /\d/.test(el.textContent)).length,
+7 -3
View File
@@ -16,6 +16,10 @@ import { collectLayoutFacts } from "./facts.js";
import { evaluateChecks, INTERACTIVE_SELECTORS } from "./checks.js";
import { deriveViewports } from "./viewports.js";
// page.waitForFunction and page.evaluate run the functions given them in
// the page, where these are defined.
/* global document, requestAnimationFrame */
function required(name) {
const value = process.env[name];
if (!value) {
@@ -62,7 +66,6 @@ function stableHash(text) {
async function connectBrowser() {
const deadline = Date.now() + BROWSER_TIMEOUT_MS;
let lastError;
for (;;) {
try {
const response = await fetch(`${CDP_URL}/json/version`);
@@ -77,9 +80,10 @@ async function connectBrowser() {
});
return { browser, version: info.Browser };
} catch (error) {
lastError = error;
if (Date.now() > deadline) {
throw new Error(`browser never came up: ${lastError}`);
throw new Error(`browser never came up: ${error}`, {
cause: error,
});
}
await sleep(250);
}
+724 -193
View File
File diff suppressed because it is too large Load Diff