20 Commits
Author SHA1 Message Date
clawbot b1eefe0c29 Re-vendor the shared files from sneak/prompts at dd4027b (closes #113)
check / check (push) Waiting to run
The shared files are the sneak/prompts copies at dd4027b, plus this
repository's own entries. make lint and make test each build one
Dockerfile phase without the cache, both covering the frontend; the
builder stage waits on both and takes its version from git describe
unless VERSION is given. The test phase keeps Go's module and build
caches in memory, out of the image make test tags. golangci-lint moves
to v2.14.0 with the new .golangci.yml; one test spells X-Request-ID as
canonicalheader asks. prettier formats only JavaScript, CSS, HTML and
Markdown, so .golangci.yml stays as fetched. script/fmt and
script/fmt-check put ~/.local/bin on PATH. script/bootstrap keeps a Go
only if it is exactly GO_VERSION, and re-checks the go on PATH after
installing.

Model: opus-5-5
2026-10-07 13:41:50 +02:00
clawbot dcdee6bfab Hetzner checks answer again; failed checks go to the console (closes #114)
check / check (push) Waiting to run
The Hetzner speed-test servers close the connection without an answer
when the URL has a query string, and every check added ?_cb= and the
time, so all six showed unreachable. Checks now fetch each target's URL
as written; cache: "no-store" still keeps the browser's cache out of the
measurement.

Each recorded check that fails writes one console.error line, and the
same line to the debug log: the target's name and URL, the time, what
failed and how long the request took. A target that answers after
failed checks writes one console.info line. Checks the page does not
record write nothing, so the debug log no longer lists failures in the
first round or the recovery probe.

Model: opus-5-5
2026-10-07 12:24:51 +02:00
clawbot 186f932eb8 Footer no longer says IPv4 only (closes #111)
check / check (push) Successful in 2m48s
Each check is a fetch: the browser picks IPv4 or IPv6 for each WAN
host, and the local targets are IPv4 addresses, so "IPv4 only" was
wrong for the page as a whole. The footer's first line drops it and
its separator; the rest of the footer is unchanged. TODO.md moves this
to Completed Steps and the layout decision from Future Steps into Next
Step.

Model: opus-5-5
2026-10-04 07:38:06 +02:00
clawbot 64e142c17f README, TODO and the viewport README say what the tree does (closes #24)
check / check (push) Successful in 2m57s
README.md: Getting Started leads with make targets; a Backend section
gives netwatch-server's routes and how the image builds and runs it;
the checks are GET requests; the WAN host list, health states, summary
figures, sorting and missing features match src/main.js; the TODO
section points to TODO.md, which holds the one to-do list.

backend/README.md: its TODO section points to TODO.md too, whose
Future Steps take its three open items.

TODO.md: Workflow branches from next and opens the PR against next;
Status, Next Step and Future Steps describe the open work, linked to
its issue where one exists.

test/viewport/README.md: the unit tests run on Node's test runner,
not vitest.

Model: opus-5-5
2026-10-04 07:03:05 +02:00
clawbot 161f955ae2 Latency statistics written once, thresholds read from CONFIG (closes #102)
check / check (push) Successful in 2m50s
A target's min, max, median and average latency now come from one list
of its answers, through latencyStats(), which the summary's figures use
too, so the median is written once. latencyHex() and latencyClass() read
one table of color limits in CONFIG. The health thresholds, the debug
log's length, the gateway check's timeout, the recovery probe's number
of hosts and interval, how often the rows are sorted and the delay
before the first sparkline resize are CONFIG entries, read where the
numbers were. HostState's minLatency(), maxLatency(), averageLatency()
and medianLatency() are gone; the statistics test reads historyStats().
Nothing the page does or shows changes. A unit test now covers the
summary's figures.

Model: opus-5-5
2026-10-04 06:51:24 +02:00
clawbot d40e67d4ab Rate limit password attempts on /metrics (closes #104)
check / check (push) Successful in 3m32s
Each client address may make 60 requests to /metrics a minute,
through the same httprate middleware and TRUSTED_PROXIES
resolution the report route uses, with an allowance of its own.
The limit runs before the basic auth, so past it the answer is 429
and the password is not checked. backend/README.md says so; a test
uses up one client's allowance on wrong passwords, gets 429 with
the right one, and checks that another client behind the same
nginx still gets in.

Model: opus-5-5
2026-10-04 06:19:06 +02:00
clawbot dc2d240725 Report handler panics to Sentry when SENTRY_DSN is set (closes #95)
check / check (push) Successful in 3m6s
With SENTRY_DSN set, the server initialises sentry-go with the release
netwatch-server-<version>, adds the sentryhttp middleware with Repanic
as the last router-wide middleware, after the timeout, and flushes
Sentry for 2 seconds on shutdown. A DSN Sentry refuses stops the start
with an error naming SENTRY_DSN. With it empty, nothing is set up.

The metrics middleware stays on the matched routes only, so it runs
inside the Sentry middleware rather than before it.

Model: opus-5-5
2026-10-04 05:57:02 +02:00
clawbot 00c9f8d7d9 Target names, URLs and log lines reach the page as text (closes #29)
check / check (push) Successful in 3m8s
A host row escapes the name and URL it writes into its markup with a new
escapeHTML function, and the debug log builds each line as an element
whose text is set, so neither is read as HTML once targets can be
configured. A unit test builds the row of a target whose name and URL
hold < > " & and ' and checks each comes out escaped; hostRowHTML is
exported for it.

README.md stops calling CONFIG frozen: the interval menu sets
updateInterval, and the timeouts, history span and axis ticks are
computed from it. AppState declares _recoveryProbeId and
_recoveryProbeChecks, the sparkline axis functions drop the parameters
they never used, and HostState's history comment names both entry
shapes.

Model: opus-5-5
2026-10-04 05:35:36 +02:00
clawbot dc11beb6fe Add make add-dependency and make tidy (closes #45)
check / check (push) Successful in 3m29s
No entrypoint could change yarn.lock or go.mod: script/bootstrap
installs with --frozen-lockfile, so adding a package meant running
yarn by hand. make add-dependency PACKAGE=<name>@<version> shims to
the new script/add-dependency: yarn add --dev, then yarn install
--frozen-lockfile. make tidy shims to the new script/tidy, go mod
tidy in backend/; a Go module is added by importing it, or moved by
editing its require line, then make tidy. script/bootstrap is
unchanged.

Model: opus-5-5
2026-10-04 05:17:33 +02:00
clawbot 81d4153e78 Serve Prometheus metrics at /metrics behind basic auth (closes #94)
check / check (push) Successful in 2m53s
With METRICS_USERNAME and METRICS_PASSWORD both set, the backend
records request metrics through go-http-metrics in a registry of its
own, with Go's runtime and process metrics, and serves them at
GET /metrics behind basic auth; nginx passes /metrics to it. With
neither set there is no such route; one alone, or a METRICS_USERNAME
containing ":", stops the start with an error naming the setting.

Only requests that reach the health check or POST /api/v1/reports are
recorded, as the labels are path and method, which clients could
otherwise make up without end; POST /api/v1/reports is registered by
its full path for that.

Deviation: go get and go mod tidy ran directly; no entrypoint added a Go
dependency yet (issue #45).

Model: opus-5-5
2026-10-04 04:58:47 +02:00
clawbot d412815953 Shim make dev and make build, format backend/ markdown (closes #28)
check / check (push) Successful in 4m11s
make dev ran yarn dev inline and there was no make build. Both are now
shims, over script/dev and the new script/build. .prettierignore stops
leaving out backend/, so backend/README.md is formatted and checked;
backend/.golangci.yml is left out by name instead: it is the org
standard file whose sha256 backend/script/lint checks, and prettier
would reindent it. .claude/ stays in .prettierignore, as git does not
ignore it. script/install-precommit and the date on bootstrap's pins go
back to the org model; bootstrap, fmt and fmt-check keep what the Go
backend and eslint need, each with a comment saying so.

Model: opus-5-5
2026-10-04 04:14:59 +02:00
clawbot 9b548da90d Move tailwind off the deprecated module.register() (closes #32)
check / check (push) Successful in 2m40s
Frontend builds on Node 26 or newer, such as make test on a host with
Node 26, printed Node's DEP0205 warning; the build in Dockerfile runs on
Node 22, which never printed it. The trace names @tailwindcss/node,
which @tailwindcss/vite brings in at its own exact version; tailwind
4.3.1 calls module.registerHooks() where Node has it. @tailwindcss/vite
and tailwindcss move to 4.3.3, inside the ranges package.json already
allows, so only yarn.lock changes, every entry still with an integrity
hash. tailwindcss moves too so that one tailwind version is installed.
Deviation: yarn.lock was changed by a raw yarn upgrade; no entrypoint adds a dependency yet (issue #45).

Model: opus-5-5
2026-10-04 03:51:55 +02:00
clawbot 3b1262718d Frontend unit tests cover durations, colours, statistics and health (closes #21)
check / check (push) Successful in 3m38s
New table-driven tests in test/unit/main.test.js: humanDuration, the
figure and sparkline colours either side of each latency boundary, a
target's min, max, average and median over an empty history, an
all-unreachable one, a mixed one, one of only answers and one whose
answers have different numbers of digits, and the four health states
either side of their thresholds, with targets found unreachable counted
as timed out. src/main.js exports the four names they need.

package.json gains a test script, which script/frontend-test runs with
the dot reporter and, if a test fails, again with the spec reporter
before failing. NODE_OPTIONS picks the reporter, since yarn appends its
arguments after the test files.

Model: opus-5-5
2026-10-04 03:25:06 +02:00
clawbot 82bd142429 Log fx through slog, snake_case health check keys (closes #27)
check / check (push) Successful in 4m40s
fx wrote its own steps of starting and stopping as plain text to
stderr. It now logs them with its slog event logger through the
backend's logger, so off a terminal every line the backend's own
logger and fx write is JSON. A malformed config file now makes
config.New return the error instead of panicking. A test runs the
server as a child process and checks that all it writes is JSON, on
a normal start and stop and with such a config file. The backend
logs its name, version and architecture once at start.

The health check's uptime keys are now uptime_seconds and
uptime_human; its type and method take the names
GO_HTTP_SERVER_CONVENTIONS.md gives.
Rules suppressed: revive and tagliatelle on HealthcheckResponse, whose name and keys come from the conventions; gosec where the test starts its own binary.

Model: opus-5-5
2026-10-04 03:11:43 +02:00
clawbot 924527b744 Lint the frontend with eslint in its own Docker stage (closes #47)
check / check (push) Successful in 3m48s
script/lint ran prettier --check, the same check script/fmt-check
runs, so the JavaScript had no linter. eslint now runs with its
recommended rules, set in eslint.config.js, in a new frontend-lint
stage of Dockerfile built from the pinned node image and the lockfile.
The frontend stage copies a file from it, as the builder stage does
from the Go lint stage, so the image cannot build unless eslint passed.
script/lint builds both lint stages with --no-cache and runs no linter
on the host; script/fmt-check keeps prettier on the host, and
script/frontend-check drops its lint step. The viewport harness fixes
the two kinds of finding eslint made. bootstrap wants node 22.13.0, as
eslint 10 does. Also covers item 2 of
#28.

Model: opus-5-5
2026-10-04 01:58:07 +02:00
clawbot 2f0489e3a4 Tap-target check expects a pin button per WAN host row (closes #46)
check / check (push) Successful in 3m50s
The viewport harness's tap-target check required at least 10 visible
pin buttons while 26 render, so pin buttons missing from up to 16 rows
went unnoticed. It now expects one per WAN host row. The host row count
the harness gathers, which the app-rendered check also reads, counts
only the WAN host rows, since the local host rows have no pin button.
Each control's minimum is now worked out from the gathered facts.

TODO.md's harness entry no longer says every check guards itself: the
overflow, viewport-edge and clipped-text checks rely on app-rendered.

Model: opus-5-5
2026-10-04 01:53:02 +02:00
clawbot 91856fa170 Each target's row shows its result as soon as its check ends (closes #91)
check / check (push) Successful in 3m43s
tick drew no row until the round's slowest check ended, up to 24
seconds at a 30-second interval since checks time out at 80% of it.
Each check now pushes its sample and redraws its row as it ends. When
the last check ends, every row is redrawn, so none still reads "paused"
after a pause and resume, and sorting, the summary, the health box and
offline detection run once. A check that ends while paused or after its
round is given up draws nothing, and the first round is still discarded
as a whole. The row is looked up when the check ends, as a pin click can
re-sort the rows mid-round.

Unit tests run tick on the mocked clock against a stand-in page.

Model: opus-5-5
2026-10-04 01:41:04 +02:00
clawbot 39ee6ca839 Entrypoint acts as root on nothing outside /data (closes #80)
check / check (push) Successful in 1m58s
`bin/entrypoint.sh` now runs `netwatch-server prepare-data-dir`, which
refuses a `DATA_DIR` that is not `/data` or a path below it written in
full, then creates `DATA_DIR`, gives `/data` and everything in it to
`netwatch`, and sets mode 750 on `/data` and `DATA_DIR`. Every step goes
through a Go `os.Root` opened on `/data`, and the modes are set on the
opened directories rather than by name, so neither a symbolic link
already there nor one a host process swaps in while the container
starts can make root create or change anything outside `/data`. The
README says which `DATA_DIR` values are accepted.

Model: opus-5-5
2026-10-03 18:24:38 +02:00
clawbot e4df415676 Delete the oldest report files to stay under the size cap (closes #54)
check / check (push) Successful in 1m57s
When a report would take the report files past DATA_DIR_MAX_BYTES,
reportbuf now deletes the oldest report files until it fits, and does
the same at start when files left by an earlier run are already past
it. A file joins the files that may be deleted, at its place by name,
only once it is completely written, so a file still being written is
never deleted. A report is refused with 507 only when the reports
waiting to be written fill the cap on their own, and then no file is
deleted. The reports of a failed write stop counting, and the part of
its file written is removed. A file whose deletion fails keeps
counting; one already deleted by hand counts as freed.

Model: opus-5-5
2026-10-03 17:51:08 +02:00
clawbot 9e4d3fdb54 Lint's .golangci.yml check says which fix applies (closes #34)
check / check (push) Successful in 2m0s
backend/script/lint compared .golangci.yml with its pinned sha256 and,
on any failure, said to restore the file from sneak/prompts. Once the
org standard has legitimately changed, that advice loops: the copied
file is right and GOLANGCI_CONFIG_SHA256 is stale. On a mismatch the
script now says to compare the file with the org standard, restore it
if they differ, and update GOLANGCI_CONFIG_SHA256 if they are the same;
it still prints both hashes. A missing .golangci.yml, and a sha256sum
that is missing or prints no hash, get their own messages instead of
being reported as a mismatch. Each failure still exits 1. .golangci.yml
is unchanged.

Model: opus-5-5
2026-10-03 17:14:25 +02:00
63 changed files with 4540 additions and 962 deletions
+79 -7
View File
@@ -1,9 +1,81 @@
node_modules
dist
tmp
.DS_Store
*.log
# .dockerignore does NOT use .gitignore semantics. Docker matches with
# moby/patternmatcher: filepath.Match plus `**`, so `*` does not cross
# `/` and an unprefixed pattern is anchored at the context root. Every
# depth-independent pattern therefore needs `**/`, or `config/.env` and
# `certs/server.key` still ship while this file reads as solved. Only
# genuinely root-anchored entries go unprefixed. Never transplant these
# into .gitignore, where `**/` is wrong.
#
# Matching is case-sensitive, so secrets use character ranges rather
# than an ALL-CAPS twin, which would still miss `Server.Key`.
#
# Extend with this repo's own host-built artifacts, written anchored:
# `/myapp`, never `**/myapp`, which also matches `cmd/myapp/` and
# deletes the package directory from the context.
# .git is sent without its config. Without a VERSION build argument the
# stage that compiles runs `git describe --tags --always` on .git, which
# does not need .git/config; that file can hold a credential, such as a
# password in a remote URL or the token the CI checkout step stores there.
# Each submodule keeps a config with the same exposure in its git directory
# under .git/modules/, nested again for a submodule's own submodules, or in
# its own .git directory when it keeps one.
# KNOWN GAP: a submodule whose name has a `config` segment (`config`,
# `deploy/config`, `config/lib`) loses its whole git directory, because
# `**/.git/modules/**/config` also matches that segment's directory
# under .git/modules/. Go's version stamping then fails the build;
# nothing leaks. Name such a submodule without that segment:
# `git submodule add --name`.
**/.git/config
**/.git/modules/**/config
# Agent scratch: one full checkout of the repo per in-flight agent.
# Anchored because it occurs once where agents run at the repo root.
# KNOWN GAP: a repo running agents in subdirectories still ships
# `services/api/.claude/` and must add its own anchored entry.
.claude
# .git is sent so the build can stamp the version, without its config.
.git/config
# Environment files. `*.env` covers bare `.env` and the `prod.env`
# convention. Re-include a committed template with a negation if the
# build needs one: `!docs/example.env`.
**/*.[eE][nN][vV]
**/.[eE][nN][vV].*
**/.[eE][nN][vV][rR][cC]
# Private keys and the bundles carrying them. Public certificates
# (*.crt, *.cer) are deliberately absent: they are legitimate inputs.
**/*.[pP][eE][mM]
**/*.[kK][eE][yY]
**/*.[pP]12
**/*.[pP][fF][xX]
**/[iI][dD]_[rR][sS][aA]
**/[iI][dD]_[dD][sS][aA]
**/[iI][dD]_[eE][cC][dD][sS][aA]
**/[iI][dD]_[eE][cC][dD][sS][aA]_[sS][kK]
**/[iI][dD]_[eE][dD]25519
**/[iI][dD]_[eE][dD]25519_[sS][kK]
# Dependencies: restored inside the image, never copied in.
**/node_modules
# OS metadata.
**/.DS_Store
**/Thumbs.db
# Editor state: never a build input, and it churns COPY.
**/*.swp
**/*.swo
**/*~
**/*.bak
**/.idea
**/.vscode
**/*.sublime-*
# This repository's own entries, after the shared content above: the
# frontend build, the viewport test's output, what `make build` and
# `make run` in backend/ leave behind, and log files at the root.
/dist
/tmp
/backend/netwatch-server
/backend/data
/*.log
+5
View File
@@ -10,3 +10,8 @@ insert_final_newline = true
[Makefile]
indent_style = tab
# This repository's own entries, after the shared content above.
[*.go]
indent_style = tab
+1 -4
View File
@@ -6,7 +6,4 @@ jobs:
steps:
# actions/checkout v4.2.2, 2026-02-22
- uses: actions/checkout@11bd71901bbe5b1630ceea73d27597364c9af683
# script/cibuild bootstraps, runs every check and builds the
# image. script/bootstrap links what it installs into
# ~/.local/bin, so that has to be on PATH for the rest.
- run: PATH="$HOME/.local/bin:$PATH" script/cibuild
- run: script/cibuild
+38 -5
View File
@@ -11,18 +11,51 @@ Thumbs.db
.vscode/
*.sublime-*
# Agent scratch (worktrees of this repo, created and destroyed by
# in-flight tooling). Unanchored: .gitignore patterns already match at
# every depth, so no prefix is wanted here. This is not a .dockerignore
# entry and must not be given a `**/` prefix on the way into one.
.claude/
# Node
node_modules/
# Environment / secrets
.env
.env.*
*.pem
*.key
# Secrets. Unanchored like every entry above, so each matches at every
# depth. Matching is case-sensitive on Linux, so names use character
# ranges rather than a lowercase form that misses `Server.Key`.
# Environment files. `*.env` covers bare `.env` and the `prod.env`
# convention. Only the templates `example.env` and `sample.env` are
# re-included below. A repository that commits any other template adds
# its own negation after these lines, for example `!.env.example`.
*.[eE][nN][vV]
.[eE][nN][vV].*
.[eE][nN][vV][rR][cC]
!example.env
!sample.env
# Private keys and the bundles carrying them.
*.[pP][eE][mM]
*.[kK][eE][yY]
*.[pP]12
*.[pP][fF][xX]
[iI][dD]_[rR][sS][aA]
[iI][dD]_[dD][sS][aA]
[iI][dD]_[eE][cC][dD][sS][aA]
[iI][dD]_[eE][cC][dD][sS][aA]_[sS][kK]
[iI][dD]_[eE][dD]25519
[iI][dD]_[eE][dD]25519_[sS][kK]
# This repository's own entries, after the shared content above.
# Build output
dist/
tmp/
/backend/netwatch-server
# Go test binaries and coverage output
*.test
*.out
# Logs
*.log
-4
View File
@@ -1,6 +1,2 @@
backend/
dist/
node_modules/
tmp/
yarn.lock
.claude/
+108 -68
View File
@@ -1,81 +1,121 @@
# The one image netwatch ships: nginx serves the built frontend and
# passes /api/ and /.well-known/healthcheck to netwatch-server, the Go
# backend, which runs in the same container on loopback only.
# passes /api/, /.well-known/healthcheck and /metrics to netwatch-server,
# the Go backend, which runs in the same container on loopback only.
# bin/entrypoint.sh starts and watches both.
# Lint stage — fast feedback on formatting and lint issues. The
# golangci/golangci-lint image ships Go, gofmt, make and the linter, so
# nothing is installed here. The root make lint builds this stage alone.
# golangci/golangci-lint:v2.12.2 (2026-08-10)
FROM golangci/golangci-lint@sha256:5cceeef04e53efe1470638d4b4b4f5ceefd574955ab3941b2d9a68a8c9ad5240 AS lint
WORKDIR /src
COPY backend/go.mod backend/go.sum ./
RUN go mod download
COPY backend/ .
RUN make fmt-check
RUN make lint
# Backend build stage
# golang:1.25-alpine (2026-02-27)
FROM golang:1.25-alpine@sha256:f6751d823c26342f9506c03797d2527668d095b0a15f1862cddb4d927a7a4ced AS builder
# gcc and musl-dev are for make test: its race detector needs cgo, which
# Go turns on by itself once a C compiler is present. make build still
# sets CGO_ENABLED=0, so the binary stays static.
RUN apk add --no-cache gcc git make musl-dev
WORKDIR /src
# Force BuildKit to run the lint stage before proceeding. BuildKit runs
# stages in parallel by default; without this no-op copy a lint failure
# would not gate compilation.
COPY --from=lint /src/go.sum /dev/null
COPY backend/go.mod backend/go.sum ./
RUN go mod download
COPY backend/ .
RUN make test
# make build is a shim around backend/script/build, the one definition
# of the build command:
# CGO_ENABLED=0 go build -trimpath -ldflags "-s -w -X main.Version=..."
# That script reads VERSION from the environment, so it is handed over
# there rather than as a make variable.
#
# The version is the VERSION build argument when one is given, otherwise
# `git describe --tags --always` of the repo's .git: the tag on a tagged
# commit, tag-N-gHASH on a commit after one, the short commit when no
# tag is reachable. A version that still comes out empty, dev or unknown
# fails the build. .git goes to /git, not /src/.git, where go build would
# find it and record VCS details of a work tree holding only backend/.
COPY .git /git
ARG VERSION
RUN version="${VERSION:-$(git --git-dir=/git describe --tags --always)}"; \
case "$version" in ""|dev|unknown) \
echo "version is '$version' although .git is present" >&2; \
exit 1 ;; \
esac; \
VERSION="$version" make build
# The lint and test phases are the gates: `make lint` (script/lint)
# builds the lint stage alone and `make test` (script/test) the test
# stage alone, and the builder stage depends on both, so the image
# cannot be built unless they pass. Each covers the frontend as well,
# through a copy from a node stage. Inside them each tool is invoked
# directly, never through make or script/, whose lint and test are
# themselves docker builds.
# Frontend stage
# Frontend lint stage: eslint with the rules in eslint.config.js. The
# lint phase below runs it.
# node:22-alpine as of 2026-02-22
FROM node@sha256:e4bf2a82ad0a4037d28035ae71529873c069b13eb0455466ae0bc13363826e34 AS frontend-lint
WORKDIR /app
COPY package.json yarn.lock ./
RUN yarn install --frozen-lockfile
COPY . .
RUN yarn eslint .
# Lint phase: golangci-lint over the backend with backend/.golangci.yml,
# and eslint through the copy from frontend-lint at the end. The
# golangci/golangci-lint image ships Go and the linter.
# golangci/golangci-lint:v2.14.0, 2026-09-24
FROM golangci/golangci-lint@sha256:ad862ba6b3798cbe0fd9fd7408d498fd74fbd2623a92406b2fd3898faf0bf98f AS lint
WORKDIR /src
COPY backend/go.mod backend/go.sum ./
RUN go mod download
COPY backend/ .
RUN golangci-lint run --config .golangci.yml ./...
# Nothing is wanted from frontend-lint; the copy is what makes this
# phase run it.
COPY --from=frontend-lint /app/yarn.lock /dev/null
# Frontend stage: the unit tests in test/unit/, then the production
# build into dist/, which the runtime stage serves. The test phase below
# runs it. The tests print a dot each; if any fails, they run again with
# every test listed, and the step fails even if that run passes.
# NODE_OPTIONS chooses the reporter because yarn adds its arguments
# after the test files, where node would take a reporter option for one
# more file.
# node:22-alpine as of 2026-02-22
FROM node@sha256:e4bf2a82ad0a4037d28035ae71529873c069b13eb0455466ae0bc13363826e34 AS frontend
WORKDIR /app
COPY package.json yarn.lock ./
RUN yarn install --frozen-lockfile
RUN apk add --no-cache git make
# vite.config.js reads the commit for the page's footer with git.
RUN apk add --no-cache git
COPY . .
# make frontend-check is the frontend half of make check (test + lint +
# fmt-check); its test step runs the unit tests, then the production
# yarn build, so this both produces dist/ and gates the image on
# lint/fmt-check/test regressions.
# This node stage has neither Go nor Docker; the lint and builder stages
# above gate the backend half.
RUN make frontend-check
RUN NODE_OPTIONS=--test-reporter=dot timeout 90 yarn --silent run test || \
{ echo "--- Rerunning with every test listed for details ---"; \
NODE_OPTIONS=--test-reporter=spec timeout 90 yarn --silent run test; \
exit 1; }
RUN yarn build
# Runtime stage
# Test phase: the backend's tests with the race detector and coverage,
# and the frontend's through the copy from the frontend stage at the
# end. -race needs cgo and so a C compiler, which the Debian Go image
# ships and the alpine one does not. -timeout 90s is a backstop above
# the 60-second cap on the suite. The rerun with -v only shows details:
# the step fails however it ends, because the first run already failed.
# Go's module and build caches are in memory (tmpfs) for the go test
# step alone, so the image make test tags does not carry them and the
# build spends no time writing them into it. They start empty on every
# build, so no stored test result can stand in for a run.
# golang:1.25.7-trixie, 2026-10-07
FROM golang@sha256:2b174ffcf56c7ad0c47d30d2630693265639ddf2a5141149c2da34db921791b4 AS test
WORKDIR /src
COPY backend/ .
RUN --mount=type=tmpfs,target=/go/pkg/mod \
--mount=type=tmpfs,target=/root/.cache/go-build \
go test -timeout 90s -race -cover ./... || \
{ echo "--- Rerunning with -v for details ---"; \
go test -timeout 90s -race -v ./...; exit 1; }
# Nothing is wanted from the frontend stage; the copy is what makes
# this phase run its tests.
COPY --from=frontend /app/yarn.lock /dev/null
# Backend build stage. Nothing is wanted from the two phases; the copies
# are what make BuildKit build them first, so this stage cannot run
# unless lint and test passed.
# golang:1.25-alpine (2026-02-27)
FROM golang@sha256:f6751d823c26342f9506c03797d2527668d095b0a15f1862cddb4d927a7a4ced AS builder
COPY --from=lint /src/go.sum /dev/null
COPY --from=test /src/go.sum /dev/null
RUN apk add --no-cache git
# A tar-stream context keeps the sender's file owners, which git refuses.
RUN git config --system --add safe.directory /src
WORKDIR /src
COPY backend/go.mod backend/go.sum backend/
RUN cd backend && go mod download
COPY . .
# backend/script/build is the one definition of the build command:
# CGO_ENABLED=0 go build -trimpath -ldflags "-s -w -X main.Version=..."
# It reads VERSION from the environment.
#
# The version is the VERSION build argument when one is given, otherwise
# `git describe --tags --always` on the .git in the build context: the
# tag on a tagged commit, tag-N-gHASH on a commit after one, the short
# commit when no tag is reachable. With .git present, a version that is
# still empty, dev or unknown fails the build: git is missing or could
# not read the checkout.
ARG VERSION
RUN version="${VERSION:-$(git describe --tags --always)}"; \
if [ -e .git ]; then \
case "$version" in ""|dev|unknown) \
echo "version is '$version' although .git is present" >&2; \
exit 1 ;; \
esac; \
fi; \
VERSION="$version" backend/script/build
# Runtime stage, and the last one: a plain `docker build .` builds it
# and the stages it copies from, the two phases included.
# nginx:stable-alpine as of 2026-02-22
FROM nginx@sha256:15e96e59aa3b0aada3a121296e3bce117721f42d88f5f64217ef4b18f458c6ab
@@ -91,7 +131,7 @@ RUN rm /etc/nginx/conf.d/default.conf
COPY nginx.conf /etc/nginx/templates/netwatch.conf.template
COPY security-headers.conf /etc/nginx/security-headers.conf
COPY --from=frontend /app/dist /usr/share/nginx/html
COPY --from=builder /src/netwatch-server /usr/local/bin/netwatch-server
COPY --from=builder /src/backend/netwatch-server /usr/local/bin/netwatch-server
COPY bin/entrypoint.sh /usr/local/bin/entrypoint.sh
# bin/entrypoint.sh creates DATA_DIR at start and gives it and /data to
+15 -9
View File
@@ -1,5 +1,5 @@
.PHONY: bootstrap setup dev test lint fmt fmt-check check frontend-check \
frontend-viewport-test docker hooks
.PHONY: bootstrap setup dev build test lint fmt fmt-check check \
add-dependency tidy frontend-viewport-test docker hooks
# Standard targets are thin shims; the implementations live in script/
# per the scripts-to-rule-them-all pattern (see the Entrypoints section
@@ -13,7 +13,11 @@ setup:
@script/setup
dev:
yarn dev
@script/dev
# The frontend only; backend/Makefile's build target builds the Go server.
build:
@script/build
test:
@script/test
@@ -30,14 +34,16 @@ fmt-check:
check:
@script/check
# The frontend half of check, for Dockerfile's node build stage, which
# has neither Go nor Docker. Use check everywhere else.
frontend-check:
@script/frontend-check
# make add-dependency PACKAGE=<name>@<version>. PACKAGE reaches the
# script through the environment, so the shell never reads it as code.
add-dependency:
@script/add-dependency "$$PACKAGE"
tidy:
@script/tidy
# Responsive-layout verification in a containerised browser. Kept out of
# check: it needs Docker and takes minutes, where make test has to stay
# under 20 seconds.
# check: it takes minutes.
frontend-viewport-test:
@script/frontend-viewport-test
+169 -81
View File
@@ -1,30 +1,33 @@
NetWatch is an MIT-licensed JavaScript single-page application by
[@sneak](https://sneak.berlin) that provides real-time network latency
monitoring to common internet hosts, displayed with color-coded figures and
sparkline graphs, served from a static bucket or Docker container.
sparkline graphs, served from a static bucket or from its Docker image, where a
small Go backend stores the measurements the page reports.
## Getting Started
```bash
# Install dependencies
yarn install
# Install the dependencies and the git pre-commit hook
make setup
# Development server
yarn dev
# Run the page on the Vite dev server
make dev
# Production build
yarn build
# Run the tests, both linters and the format check
make check
# Preview production build
yarn preview
# Build the page into dist/
make build
# Docker
docker build -t netwatch .
# Build the image and run it
make docker
docker run -p 8080:8080 netwatch
```
`yarn dev` proxies `/api` to `http://127.0.0.1:8080`, so a locally running
`netwatch-server` (see `backend/`) receives the reports the page posts.
`make check` and `make docker` need Docker. `make dev` passes `/api` to
`http://127.0.0.1:8080`, where `make run` in `backend/` starts `netwatch-server`
with its defaults, so the reports the page posts are stored in
`backend/data/reports`.
## Entrypoints
@@ -38,33 +41,48 @@ halves, so the root `make check` fails if either one is broken. We provide:
- `script/bootstrap` — install all dependencies (the pinned node via nvm unless
one new enough for the frontend's dependencies is installed, yarn via
corepack, `yarn install --frozen-lockfile`, the pinned Go unless one at least
as new as `backend/go.mod` asks for is installed, the Go modules, and gcc with
the C library headers unless gcc is installed, for the race detector in
`make test`), linking what it installs itself into `~/.local/bin`, which has
to be on `PATH`. It installs no Go linter and not Docker: `make lint` runs the
linter in Docker
corepack, `yarn install --frozen-lockfile`, Go `GO_VERSION` unless the `go` on
`PATH` is exactly that version, the Go modules, and gcc with the C library
headers unless gcc is installed, for the race detector in `make test` in
`backend/`), linking what it installs itself into `~/.local/bin`, which has to
be on `PATH`. It installs no Go linter and not Docker: `make test` and
`make lint` run in Docker
- `script/setup` — make a fresh clone ready for development: bootstrap plus the
git pre-commit hook
- `script/dev` — run the Vite dev server, which proxies `/api` to a locally
running `netwatch-server`
- `script/build` — build the frontend for production into `dist/`;
`backend/script/build` builds the Go server
- `script/projectname` — print the project name (used for the Docker image tag)
- `script/test` — run `script/frontend-test`, then the backend's Go tests, each
under its own 30-second timeout
- `script/lint` — run `script/frontend-lint`, then golangci-lint in Docker, by
building the lint stage of `Dockerfile` without the cache
- `script/fmt` — format all files (writes): prettier, then gofmt over `backend/`
- `script/test` — build the `test` phase of `Dockerfile` without the cache: the
frontend's unit tests and production build in its `frontend` stage, and the
backend's Go tests with the race detector and coverage
- `script/lint` — build the `lint` phase of `Dockerfile` without the cache:
eslint in its `frontend-lint` stage, and golangci-lint over `backend/`
- `script/fmt` — format all files (writes): prettier over the JavaScript, CSS,
HTML and Markdown, then gofmt over `backend/`. It runs on the host, as
`script/fmt-check` does, with `~/.local/bin` put on `PATH`
- `script/fmt-check` — check formatting (read-only): prettier, then gofmt
- `script/check` — run test, lint, and fmt-check
- `script/frontend-test` — run the unit tests in `test/unit/` with Node's
built-in test runner, then the production build
- `script/frontend-lint` — run prettier in check mode
- `script/frontend-fmt` — format everything prettier understands (writes)
- `script/add-dependency` — add a frontend package, or move one to another
version: `make add-dependency PACKAGE=<name>@<version>` runs `yarn add --dev`,
which changes `package.json` and `yarn.lock` together, then
`yarn install --frozen-lockfile`
- `script/tidy` — run `go mod tidy` in `backend/`: to add a Go module, import it
and run `make tidy`; to move one to another version, edit its `require` line
in `backend/go.mod`, then run `make tidy`
- `script/frontend-test` — run the unit tests in `test/unit/` on the host with
Node's built-in test runner, through the `test` script in `package.json`, and
if any fails, run them again listing every test, and fail; then the production
build. Each test run has a 90-second timeout. `make test` runs the same in
Docker
- `script/frontend-fmt` — format the JavaScript, CSS, HTML and Markdown with
prettier (writes), the markdown in `backend/` included
- `script/frontend-fmt-check` — check prettier formatting (read-only)
- `script/frontend-check` — the frontend half of `script/check`, for
`Dockerfile`, whose node build stage has neither Go nor Docker
- `script/frontend-viewport-test` — responsive-layout verification of the built
frontend in a containerised headless Chrome (see
[test/viewport/README.md](test/viewport/README.md)). Not part of
`script/check`: it needs Docker and takes minutes.
`script/check`: it takes minutes.
- `script/docker` — build the image from `Dockerfile` without the build cache,
tagged `netwatch` via `script/projectname`
- `script/cibuild` — CI entrypoint: runs `script/bootstrap` and `script/check`,
@@ -77,25 +95,30 @@ halves, so the root `make check` fails if either one is broken. We provide:
The narrow-viewport layout lives in the `max-width: 768px` media block in
`src/styles.css`. It is verified automatically by `make frontend-viewport-test`,
which drives a digest-pinned headless Chrome against the built `dist/` and
asserts on computed layout at widths derived from that CSS — one pixel either
side of every breakpoint it declares, plus a 320px floor, a desktop baseline and
two landscape sizes. See [test/viewport/README.md](test/viewport/README.md) for
what it covers and what it genuinely cannot.
asserts on computed layout at widths derived from that CSS — on every breakpoint
it declares and one pixel either side of it, plus a 320px floor, a desktop
baseline and two landscape sizes. See
[test/viewport/README.md](test/viewport/README.md) for what it covers and what
it genuinely cannot.
## Rationale
When debugging network issues, it's useful to have a persistent at-a-glance view
of latency and reachability to multiple well-known internet endpoints. NetWatch
provides this as a zero-dependency SPA that can be deployed anywhere static
files are served, with no backend required.
provides this as a single page that does all its measuring in the browser, so it
can be served from anywhere static files are served. The backend in its Docker
image only stores the measurements the page reports; without it, the page works
the same and nothing is stored.
## Design
The application is a single-page app built with Vite and Tailwind CSS v4. All
code lives in `src/main.js` with a class-based architecture:
The page is built with Vite and Tailwind CSS v4. Its code is all in
`src/main.js`, with a class-based architecture:
- **`CONFIG`**: Frozen configuration object (update interval, timeouts, axis
ticks, etc.)
- **`CONFIG`**: Configuration object (update interval, timeouts, axis ticks,
etc.). The interval menu sets `updateInterval`, the one value the page writes
into `CONFIG`; the timeouts, the time the history spans and the x-axis ticks
are computed from it
- **`HostState`**: Per-host state management — history buffer, latency tracking,
status transitions
- **`AppState`**: Top-level state container — WAN hosts, local hosts, pause
@@ -105,9 +128,11 @@ code lives in `src/main.js` with a class-based architecture:
- **UI functions**: `buildUI()` constructs the DOM, `updateHostRow()` /
`updateSummary()` / `updateHealthBox()` handle incremental updates
- **`tick()`**: Main loop — measures all hosts in parallel, pushing each host's
sample and redrawing its row as soon as its check ends, then sorts and redraws
the summary and health box once the last check ends. When paused, pushes blank
markers (no probes, no false outage)
sample and redrawing its row as soon as its check ends, then redraws every
row, the summary and the health box once the last check ends. The first round,
after loading or an interval change, is discarded. The rows are sorted when
the last check ends in round 2, the first one kept, and in rounds 11, 21, 31
and so on. When paused, pushes blank markers (no probes, no false outage)
- **`Reporter`**: Posts collected samples to the backend
### Reporting
@@ -121,12 +146,35 @@ delivered report, and while paused nothing is sent. Delivery failure is quiet
one debug-log line per outage, retried at the next interval, never blocking
probing. The report-building step is a pure function of host state.
### Backend
`netwatch-server`, in `backend/`, is a small Go HTTP server that stores the
reports the page posts. It keeps them in memory and writes them to `DATA_DIR` as
zstd-compressed files of JSON lines: every minute, whenever 10 MiB are waiting,
and when it stops. Its routes:
- `POST /api/v1/reports` — takes a report, without credentials; each client
address may send a limited number a minute, and the report files are capped in
size, the oldest deleted first
- `GET /.well-known/healthcheck` — answers 200 with `"status":"ok"`, the
server's version and its uptime
- `GET /metrics` — Prometheus metrics behind basic auth, only when
`METRICS_USERNAME` and `METRICS_PASSWORD` are set; each client address may
make a limited number of requests to it a minute
In the image, the `test` phase of `Dockerfile` tests it, the `builder` stage
builds it with `backend/script/build`, and `bin/entrypoint.sh` runs it as user
`netwatch` on `127.0.0.1:8081`, behind nginx. Outside the image, `make run` in
`backend/` builds it and runs it on port 8080. Its settings, report storage and
limits are in [backend/README.md](backend/README.md).
### Monitoring targets
- **22 WAN hosts**: datavi.be, Anthropic API, OpenAI API, AWS Console, GCP
Console, Azure, Cloudflare, Fastly, Akamai, GitHub, B2, 7 S3 regional
endpoints (Cape Town, London, Bahrain, Tokyo, Sydney, Oregon, São Paulo), 4
GCS locational endpoints (Iowa, Belgium, Singapore, Sydney)
- **26 WAN hosts**: datavi.be (pinned at start), Anthropic API, OpenAI API, AWS
Console, Google Cloud Console, Microsoft Azure, Cloudflare, Fastly CDN,
Akamai, Google, GitHub, B2, 8 S3 regional endpoints (Cape Town, London,
Bahrain, Tokyo, Singapore, Sydney, Oregon, São Paulo) and 6 Hetzner speed test
servers (Nuremberg, Falkenstein, Helsinki, Ashburn, Hillsboro, Singapore)
- **Local CPE**: Cable modem at 192.168.100.1 (always monitored)
- **Local Gateway**: Auto-detected on startup by probing common default gateway
addresses (192.168.1.1, 192.168.0.1, 192.168.8.1, 10.0.0.1); first responder
@@ -138,15 +186,28 @@ Local hosts are tracked separately from WAN stats.
### Latency measurement
HEAD requests with `mode: 'no-cors'` and `cache: 'no-store'`, timed with
`performance.now()`. Each check times out after 80% of the refresh interval (24
seconds at 30 seconds) and is then recorded as a timeout, so a round's checks
have all finished before the next round is due. When no WAN host answers, a
recovery probe checks 4 random WAN hosts every half second, giving up the checks
it started half a second before. As soon as one answers, a new round starts at
once, as it does after an interval change. A round started early gives up the
last round's checks if they are still waiting, and that round records nothing
more, so rounds never overlap. IPv4 only.
GET requests to each target's URL as written, with `mode: 'no-cors'` and
`cache: 'no-store'`, timed with `performance.now()`. No query string is added:
the Hetzner speed-test servers close the connection without an answer when the
URL has one, and `no-store` keeps the browser's cache out of the measurement.
Each check times out after 80% of the refresh interval (24 seconds at 30
seconds) and is then recorded as a timeout, so a round's checks have all
finished before the next round is due. When no WAN host answers, a recovery
probe checks 4 WAN hosts, picked at random when it starts, every half second,
giving up the checks it started half a second before. As soon as one answers, a
new round starts at once, as it does after an interval change. A round started
early gives up the last round's checks if they are still waiting, and that round
records nothing more, so rounds never overlap. The browser chooses between IPv4
and IPv6 for each target, as for any request; the local targets are IPv4
addresses.
Each recorded check that fails writes one line to the browser console with
`console.error`, and the same line to the debug log: the target's name and URL,
the time, whether it timed out, answered over the time limit or hit a network
error (with the error the browser gives the page, such as
`TypeError: Failed to fetch`, which does not say why), and how long the request
took. A target that answers after failed checks writes one `console.info` line
with its latency and how many checks in a row had failed.
### Color coding
@@ -171,27 +232,44 @@ dist/
## Features
- Real-time monitoring with 2s update interval and 300s history sparklines
- Health indicator: green (HEALTHY) or red (DEGRADED) based on WAN reachability
- Summary stats: reachable count, min/max/avg latency across WAN hosts only
- Fixed chart axes: Y-axis 0–1000ms, X-axis 0–300s
- A round of checks every 3 seconds by default; the interval menu sets 1, 2, 3,
5, 10, 15, 30 or 60 seconds and clears the history
- Sparklines of each target's last 100 rounds: 300 seconds at 3 seconds
- The first round after loading or an interval change is discarded, as DNS and
TLS setup inflate its latencies
- Health indicator from the WAN hosts' latest results: OFFLINE (red) when more
than 10 fail and at most 4 answer, otherwise DEGRADED (orange) when more than
4 fail, otherwise SLOW (yellow) when more than 3 take over 1000ms, otherwise
HEALTHY (green)
- Summary stats across WAN hosts only: how many answered, the min, median,
average and max of their latest latencies, the min and max over the whole
history, and the number of rounds run (`Checks`)
- Fixed chart axes: Y-axis 0–1000ms, higher latencies drawn at the top; X-axis
the time the history spans
- Color-coded latency figures and sparkline line segments
- WAN host rows sorted by latest latency, unreachable last; pinned rows stay on
top, in name order
- Play/pause: pause stops probes but history keeps scrolling (blank gaps, no
false outage)
- Debug log panel, behind a checkbox in the footer, with five levels (error,
warning, notice, info, debug) and the last 1000 lines
- Local and UTC clocks
- Clickable service URLs
- A footer link to the commit the page was built from
- Canvas-based sparkline rendering with devicePixelRatio scaling
- Zero runtime dependencies: all resources bundled into build artifacts
## Deployment
After running `yarn build`, deploy the contents of the `dist/` directory to any
static file host (S3, GCS, Cloudflare Pages, Vercel, Netlify, GitHub Pages) or
use the Docker image behind a reverse proxy.
`make build` writes the page to `dist/`, which any static file host (S3, GCS,
Cloudflare Pages, Vercel, Netlify, GitHub Pages) can serve; with no backend
there, its reports fail quietly and nothing is stored. Or run the Docker image
behind a reverse proxy.
The Docker image, built from `Dockerfile`, is the whole service in one
container: nginx serves the built frontend and passes `/api/` and
`/.well-known/healthcheck` to the Go backend, `netwatch-server`, which listens
only inside the container, on `127.0.0.1:8081`. The image:
The Docker image, built from `Dockerfile` by `make docker`, is the whole service
in one container: nginx serves the built frontend and passes `/api/`,
`/.well-known/healthcheck` and `/metrics` to the Go backend, `netwatch-server`,
which listens only inside the container, on `127.0.0.1:8081`. The image:
- Listens on port 8080 by default (override with `PORT` env var)
- Takes the client address from `X-Forwarded-For` only on requests from the
@@ -221,12 +299,14 @@ What the [upaas](https://git.eeqj.de/sneak/upaas) app for netwatch needs:
- `REPORTS_PER_MINUTE`, default `60`: reports each client address may send a
minute
- `DATA_DIR_MAX_BYTES`, default `1073741824` (1 GiB): the most room the
report files may take
report files may take; the oldest are deleted to stay under it
- `CORS_ALLOWED_ORIGINS`, default empty: other origins whose pages may call
the API
- `DEBUG`, default `false`: debug logging
- `DATA_DIR`, default `/data/reports`: leave unset; reports kept outside
`/data` do not survive a redeploy
- `DATA_DIR`, default `/data/reports`: the directory the reports are kept
in: `/data` or a path below it, with no `.` or `..` part and no extra `/`.
The container also stops if the path goes through a symbolic link that
leads out of `/data` or is written as a full path
- `TRUSTED_PROXIES`, default empty: set it to the address the reverse proxy
in front of the container connects from, as an IP address or CIDR; several
are separated by commas. nginx takes the client address from
@@ -237,6 +317,16 @@ What the [upaas](https://git.eeqj.de/sneak/upaas) app for netwatch needs:
connects from one can write its own `X-Forwarded-For`, and through a port
Docker publishes, every client may connect from the Docker network's
gateway, such as `172.17.0.1`.
- `METRICS_USERNAME` and `METRICS_PASSWORD`, default empty: with both set,
the backend records Prometheus metrics of its requests and serves them at
`/metrics` on the container port, to requests with this user name and
password as their basic auth credentials. With neither set, there are no
metrics and `/metrics` is not found. One set without the other, or a user
name containing `:`, stops the container
- `SENTRY_DSN`, default empty: set to a Sentry project's DSN, the backend
sends its errors to that Sentry project: each request whose handling
crashes, which still gets a 500 response. A value Sentry does not accept
stops the container. Empty, the backend sends nothing to Sentry
- **Health check:** the image's `HEALTHCHECK` requests
`/.well-known/healthcheck` through nginx every 30 seconds, so it fails unless
both nginx and the backend answer. upaas reads the container's health 60
@@ -250,21 +340,19 @@ properties.
## Limitations
- **CORS**: Some hosts may block cross-origin HEAD requests. The app uses
`no-cors` mode which allows the request but provides opaque responses. Latency
is still measurable based on request timing.
- **Local gateway**: The 192.168.100.1 endpoint requires the host to be
accessible from your network.
- **CORS**: The checks are cross-origin requests in `no-cors` mode, so the page
cannot read the answer, only time it: any answer counts as reachable, an error
page included.
- **Local targets**: The cable modem at 192.168.100.1 and the detected gateway
answer only on a network that has them, and only when NetWatch is served from
localhost or a private address (see Monitoring targets).
- **Network conditions**: Measurements reflect browser-to-endpoint latency,
which includes your local network, ISP, and internet routing.
## TODO
- Add unit tests
- Add eslint for JS linting (currently lint target runs prettier only)
- Add configurable host list (environment variable or config file)
- Add latency history export (CSV/JSON)
- Add notification/alert when status changes to DEGRADED
The to-do list is [TODO.md](TODO.md): where the work stands, the next step, the
open work, and what has been done.
## License
+355 -84
View File
@@ -1,6 +1,6 @@
---
title: Repository Policies
last_modified: 2026-07-06
last_modified: 2026-10-04
---
This document covers repository structure, tooling, and workflow standards. Code
@@ -60,17 +60,28 @@ style conventions are in separate documents:
prerequisite since nvm requires bash. yarn is then pinned via
`corepack prepare yarn@<version> --activate`. Never install "latest" or "lts";
always exact versions. `script/cibuild` runs the CI build: it changes to the
repo root and runs `docker build .`; the Gitea workflow calls it. Four further
scripts are our own extensions to the standard: `script/check` runs
`script/test`, `script/lint`, and `script/fmt-check`; `script/precommit` is
what the git pre-commit hook runs, and it calls `script/check`;
`script/install-precommit` installs the git pre-commit hook (the `make hooks`
target shims to it); and `script/projectname` (literally that filename) simply
outputs the project's name. Scripts that need the name call
`script/projectname` — e.g. `script/docker` assembles its image tag from it —
so those scripts stay byte-identical across all repos. Repo-type-specific
pre-commit extras (e.g. `go mod tidy` verification in Go repos) belong in
`script/precommit`, not in the hook itself. Model scripts are at
repo root, runs `script/bootstrap`, runs `script/check`, and builds the image
with the version; the Gitea workflow calls it. **`script/cibuild` runs
`script/bootstrap` first**, because the workflow checks out the repo and runs
nothing else, while `script/fmt-check` runs the formatter on the host: on a
pristine checkout with nothing installed the run dies there, after the
containerised gates have passed. **The bootstrap alone is not enough**:
`script/bootstrap` installs node and yarn under nvm and leaves neither on the
`PATH` of the shell that called it, so a bare `yarn` still exits 127. The host
entrypoints that need yarn — `script/fmt` and `script/fmt-check` — therefore
source nvm for the pinned node version before invoking it, exactly as
`script/bootstrap`'s own install step does. A runner carrying nothing but
docker and git then gets through `script/check`. Four further scripts are our
own extensions to the standard: `script/check` runs `script/test`,
`script/lint` and `script/fmt-check`; `script/precommit` is what the git
pre-commit hook runs, and it calls `script/check`; `script/install-precommit`
installs the git pre-commit hook (the `make hooks` target shims to it); and
`script/projectname` (literally that filename) simply outputs the project's
name. Scripts that need the name call `script/projectname` — e.g.
`script/docker` assembles its image tag from it — so those scripts stay
byte-identical across all repos. Repo-type-specific pre-commit extras (e.g.
`go mod tidy` verification in Go repos) belong in `script/precommit`, not in
the hook itself. Model scripts are at
`https://git.eeqj.de/sneak/prompts/raw/branch/main/script/<name>`. The README
must document the provided scripts in an **Entrypoints** section (see the
README requirements below).
@@ -89,87 +100,198 @@ style conventions are in separate documents:
contributor should be able to understand the entire development workflow by
reading the Makefile.
- Every repo should have a `Dockerfile`. All Dockerfiles must run `make check`
as a build step so the build fails if the branch is not green. For non-server
repos, the Dockerfile should bring up a development environment and run
`make check`. For server repos, `make check` should run as an early build
stage before the final image is assembled. Dockerfiles install development
prerequisites by running `script/bootstrap` rather than duplicating installs
inline; COPY `script/` and the dependency manifests (`package.json` +
`yarn.lock`, `go.mod` + `go.sum`, etc.) before running it so the bootstrap
layer stays cached until dependencies change.
- Every repo should have a `Dockerfile`, and it carries the repo's gates: a
`lint` phase and a `test` phase, with the final stage depending on both so the
image cannot be built unless they pass. For non-server repos the final stage
brings up a development environment; for server repos it is the runtime image.
The gate phases and the build stage start from their pinned base images and
install what those images lack either inline, as the canonical Go `Dockerfile`
below does for `git`, or by running `script/bootstrap`, as the `prompts`
repo's own `Dockerfile` does for its yarn packages. The development
environment stage installs development prerequisites by running
`script/bootstrap` rather than duplicating its installs inline. A stage that
runs `script/bootstrap` COPYs `script/` and the dependency manifests
(`package.json` + `yarn.lock`, `go.mod` + `go.sum`, etc.) before running it.
- **Dockerfiles must use a separate lint stage for fail-fast feedback.** Go
repos use a multistage build where linting runs in an independent stage based
on the `golangci/golangci-lint` image (pinned by hash). This stage runs
`make fmt-check` and `make lint` before the full build begins. The build stage
then declares an explicit dependency on the lint stage via
`COPY --from=lint /src/go.sum /dev/null`, which forces BuildKit to complete
linting before proceeding to compilation and tests. This ensures lint failures
surface in seconds rather than minutes, without blocking on dependency
download or compilation in the build stage.
- **Linting and testing run in Docker, as phases of the `Dockerfile`.** There is
no separate lint file. `script/lint` and `script/test` each build one phase
and nothing else:
The standard pattern for a Go repo Dockerfile is:
```sh
docker build --no-cache --target lint -t "$(script/projectname)-lint" .
docker build --no-cache --target test -t "$(script/projectname)-test" .
```
**A stage that is not the last one in the file is built only when the final
stage's chain depends on it, or when `--target` names it.** That is why the
two gates are always invoked by name here, and why the final stage carries a
`COPY --from=` of a harmless file from each of them: without that edge a
plain `docker build .` builds the last stage alone and exits 0 having linted
and tested nothing.
**Every `docker build` in `script/` is tagged**, here and in
`script/cibuild` and `script/docker`. An untagged build leaves a dangling
image behind on every invocation, on every developer host and every CI
runner; a tagged one replaces the previous image.
Inside a phase the tool is invoked directly — `golangci-lint`, `go test`,
`eslint`, `prettier` — never through `make lint` or `script/test`, which are
themselves a `docker build` and would recurse into a daemon that does not
exist in a build step. Formatting is the exception and stays on the host:
`script/fmt` writes the working tree, and `script/fmt-check` is its
read-only twin.
**No lint verdict may come from a host invocation of the linter.** On a
shared host golangci-lint reads a result cache keyed on file content rather
than location, so a second checkout of the same content is served the first
one's findings, and a host-global lock in `$TMPDIR` makes concurrent runs
exit non-zero with `parallel golangci-lint is running` — a status a caller
cannot tell from real findings. Both have produced wrong verdicts in this
org, in both directions. A container has its own cache, its own `TMPDIR` and
a digest-pinned binary, so neither is reachable.
- **Any build that runs checks is built with `--no-cache`.** Docker invalidates
a `COPY` layer only when the copied content changes, so on an unchanged tree
the check `RUN` is served from cache, nothing executes, and the build still
exits 0. Every `docker build` in `script/` therefore passes `--no-cache`:
`script/lint`, `script/test`, `script/cibuild` and `script/docker` are the
four, and there is no fifth — `script/check` runs the two gate phases and
`script/fmt-check`, and builds no image of its own. A bare `docker build .` is
not evidence that anything ran: a sub-second build reporting success is a
cache hit, not a result. Never invalidate by pruning — `docker builder prune`
and friends destroy a build cache shared with every other build on the host.
When a check is added or changed, prove it works by planting a defect it must
catch and watching the run fail on it, then revert the defect. A green run
alone shows neither that the check ran nor that it covers what it should.
- **The gate phases are separate stages, and the build stage depends on both.**
The lint phase is based on the `golangci/golangci-lint` image (pinned by
hash), so lint failures surface in seconds rather than after a full compile,
and the test phase is based on the Debian Go image. The canonical Go repo
`Dockerfile`:
```dockerfile
# Lint stage — fast feedback on formatting and lint issues
# Lint phase
# golangci/golangci-lint:v2.x.x, YYYY-MM-DD
FROM golangci/golangci-lint@sha256:... AS lint
WORKDIR /src
COPY go.mod go.sum ./
RUN go mod download
COPY . .
RUN make fmt-check
RUN make lint
RUN golangci-lint run --config .golangci.yml ./...
# Build stage
# golang:1.x-alpine, YYYY-MM-DD
FROM golang@sha256:... AS builder
# Test phase. -race needs cgo and so a C compiler, which the Debian Go
# image ships and the alpine one does not.
# golang:1.x, YYYY-MM-DD
FROM golang@sha256:... AS test
WORKDIR /src
# Force BuildKit to run the lint stage before proceeding
COPY --from=lint /src/go.sum /dev/null
COPY go.mod go.sum ./
RUN go mod download
COPY . .
RUN make test
RUN go test -timeout 90s -race -cover ./... || \
{ echo "--- Rerunning with -v for details ---"; \
go test -timeout 90s -race -v ./...; exit 1; }
ARG VERSION=dev
RUN CGO_ENABLED=0 go build -trimpath \
# Build stage. Nothing is wanted from either phase above; the copies
# are what make BuildKit build them first, so this stage cannot run
# unless lint and test passed.
# golang:1.x-alpine, YYYY-MM-DD
FROM golang@sha256:... AS builder
COPY --from=lint /src/go.sum /dev/null
COPY --from=test /src/go.sum /dev/null
RUN apk add --no-cache git
# A tar-stream context keeps the sender's file owners, which git refuses.
RUN git config --system --add safe.directory /src
WORKDIR /src
COPY go.mod go.sum ./
RUN go mod download
COPY . .
# The VERSION build arg when one is given, otherwise
# `git describe --tags --always` on the .git in the build context. With
# .git present, a version that is still empty, dev or unknown fails the
# build: git is missing or could not read the checkout.
ARG VERSION
RUN VERSION="${VERSION:-$(git describe --tags --always)}"; \
if [ -e .git ]; then \
case "$VERSION" in ""|dev|unknown) \
echo "version is '$VERSION' although .git is present" >&2; \
exit 1 ;; \
esac; \
fi; \
CGO_ENABLED=0 go build -trimpath \
-ldflags="-s -w -X main.Version=${VERSION}" \
-o /app ./cmd/app/
# Runtime stage
# Runtime stage, and the last one
FROM alpine@sha256:...
COPY --from=builder /app /usr/local/bin/app
ENTRYPOINT ["app"]
```
Key points:
- The lint stage uses the `golangci/golangci-lint` image directly (it
includes both Go and the linter), so there is no need to install the
linter separately.
- `COPY --from=lint /src/go.sum /dev/null` is a no-op file copy that creates
a stage dependency. BuildKit runs stages in parallel by default; without
this line, the build stage would not wait for lint to finish and a lint
failure might not fail the overall build.
- The lint phase uses the `golangci/golangci-lint` image directly (it has
both Go and the linter), so nothing needs installing.
- `COPY --from=<phase> /src/go.sum /dev/null` is a no-op copy whose only
purpose is the ordering edge. BuildKit runs stages in parallel by default,
and a stage nothing depends on is not built at all, so without these two
lines a red gate would not fail the build.
- Keep the runtime stage last, and if you add a stage after it, give it the
same two copies. A plain `docker build .` builds the last stage's chain
and nothing else.
- If the project uses `//go:embed` directives that reference build artifacts
(e.g. a web frontend compiled in a separate stage), the lint stage must
(e.g. a web frontend compiled in a separate stage), the lint phase must
create placeholder files so the embed directives resolve. Example:
`RUN mkdir -p web/dist && touch web/dist/index.html web/dist/style.css`.
The lint stage should not depend on the actual build output — it exists to
fail fast.
- If the project requires CGO or system libraries for linting (e.g.
`vips-dev`), install them in the lint stage with `apk add`.
- The build stage runs `make test` after compilation setup. Tests run in the
build stage, not the lint stage, because they may require compiled
artifacts or heavier dependencies.
- If the project requires CGO or system libraries for linting, install them
in the lint phase. The `golangci/golangci-lint` image is Debian-based and
has no `apk`, so install with `apt-get` under the Debian package name
(`libvips-dev`, where alpine says `vips-dev`), and delete the package
lists in the same `RUN`, so the layer does not keep them:
```dockerfile
RUN apt-get update \
&& apt-get install -y --no-install-recommends libvips-dev \
&& rm -rf /var/lib/apt/lists/*
```
- `.dockerignore` lets `.git` into the build context. It keeps out every git
`config` at any depth (`**/.git/config`, `**/.git/modules/**/config`): the
repository's own, each submodule's under `.git/modules/`, and that of a
submodule keeping its own `.git` directory. `git describe` does not need
them, and each can hold a credential: a password in a remote URL, or the
token the CI checkout step stores there. A submodule whose name has a
`config` segment (`config`, `deploy/config`, `config/lib`) loses its whole
git directory to `**/.git/modules/**/config`, and Go's version stamping
then fails the build: give it a name without that segment
(`git submodule add --name`). The stage that compiles has `git` (the
Debian Go image has it; an alpine one needs `apk add --no-cache git`) and
takes the version from the `VERSION` build argument when one is given,
otherwise from `git describe --tags --always`. That gives the tag on a
tagged commit; on a later commit, the tag, the number of commits since it
and the short commit (`v1.2.3-4-gabc1234`); and the short commit when no
tag is reachable. The stage that compiles also marks its working directory
safe for git (`git config --system --add safe.directory /src`): a context
sent as a tar stream keeps the sender's file owners, and git refuses a
checkout owned by another user, so the version would come out empty.
`ARG VERSION` has no default, and the build fails if the context carries
`.git` and the version still comes out empty, `dev` or `unknown`. A plain
`docker build .` with no build arguments must succeed; a Dockerfile that
refuses an empty build argument drops that refusal and keeps the argument.
- Every repo should have a Gitea Actions workflow (`.gitea/workflows/`) that
runs `script/cibuild` (which runs `docker build .`) on push. Since the
Dockerfile already runs `make check`, a successful build implies all checks
pass.
runs `script/cibuild` on push, and checks out the repo as its only other step.
That script bootstraps, runs the gate phases, and then builds the image, so a
successful run means every check passed; a bare `docker build .` does not
carry the same guarantee, because its gate phases may come from the cache. The
image build is uncached and so runs the gate phases a second time. That is the
price of the rule above, and it is worth paying: the image that ships is built
from a run of its own gates rather than from a cache entry. A separate
workflow limited to `main` by a `branches` list under `on: push` cannot be
checked by review: to try a change to it, add the feature branch to that list
and push, then remove the branch from the list again before merging. Keep any
job in it that publishes behind `if: github.ref_name == 'main'`, so the run
from the feature branch publishes nothing.
- Use platform-standard formatters: `black` for Python, `prettier` for
JS/CSS/Markdown/HTML, `go fmt` for Go. Always use default configuration with
@@ -189,14 +311,21 @@ style conventions are in separate documents:
module under test to verify it compiles/parses. There is no excuse for
`make test` to be a no-op.
- `make test` must complete in under 20 seconds. Add a 30-second timeout in the
Makefile.
- `make test` must complete in under 60 seconds. That is the hard cap, and a
suite that exceeds it fails. Under 20 seconds is the target. A suite between
20 and 60 seconds is still green, but the overage must be filed as an
improvement bug against that repo. Add a 90-second timeout to the test
invocation (`go test -timeout 90s`). The backstop deliberately sits above the
hard cap so that it catches a genuinely hung test rather than a merely slow
one.
- **`make test` should use the conditional verbose rerun pattern.** Run tests
without `-v` (verbose) first. If tests fail, automatically rerun with `-v` to
show full output. This keeps CI logs and `docker build` output clean on
success (just package/suite summaries) while providing full diagnostic detail
on failure (every test case, every assertion). The general shell pattern:
- **The test command should use the conditional verbose rerun pattern.** Run
tests without `-v` (verbose) first. If tests fail, automatically rerun with
`-v` to show full output. This keeps CI logs and `docker build` output clean
on success (just package/suite summaries) while providing full diagnostic
detail on failure (every test case, every assertion). The command lives in the
`test` phase of the `Dockerfile`, since `script/test` builds that phase; the
Makefile form below is the same pattern for any repo-local invocation:
```makefile
test:
@@ -209,11 +338,26 @@ style conventions are in separate documents:
```makefile
test:
@go test -timeout 30s -race -cover ./... || \
@go test -count=1 -timeout 90s -race -cover ./... || \
{ echo "--- Rerunning with -v for details ---"; \
go test -timeout 30s -race -v ./...; exit 1; }
go test -count=1 -timeout 90s -race -v ./...; exit 1; }
```
`-count=1` is required on both invocations: it defeats Go's test _result_
cache, so neither run can report a stored pass in place of running the
tests. It leaves the build cache alone, so it costs the runtime of the suite
and no recompilation.
That cache is Go's own, separate from Docker's layer cache. Go stores a
passing result in its cache directory (`GOCACHE`), and when the same tests
run again on unchanged code it prints that result, marked `(cached)`,
without running them. That matters on a developer's machine, where this
target runs and the directory lasts from one run to the next. The `test`
phase of the `Dockerfile` needs no `-count=1`: its base image holds no
result for this repo's tests and nothing before its `go test` step runs a
test, so there is nothing to replay. `--no-cache` (above) is what makes that
step run on an unchanged tree.
Python example:
```makefile
@@ -239,10 +383,84 @@ style conventions are in separate documents:
must be in `.gitignore`. No exceptions.
- `.gitignore` should be comprehensive from the start: OS files (`.DS_Store`),
editor files (`.swp`, `*~`), language build artifacts, and `node_modules/`.
Fetch the standard `.gitignore` from
`https://git.eeqj.de/sneak/prompts/raw/branch/main/.gitignore` when setting up
a new repo.
editor files (`.swp`, `*~`), in-repo agent scratch directories (`.claude/`),
language build artifacts, and `node_modules/`. Fetch the standard `.gitignore`
from `https://git.eeqj.de/sneak/prompts/raw/branch/main/.gitignore` when
setting up a new repo. These patterns are written to `.gitignore`'s own
semantics, in which an unanchored pattern already matches at every depth; they
are not a `.dockerignore` and must not be transplanted into one unmodified.
- **`.dockerignore` does not use `.gitignore` semantics, and copying patterns
across unmodified leaves secrets in the build context.** Docker matches with
`moby/patternmatcher`: `filepath.Match` semantics plus a `**` extension, so
`*` does not cross `/` and a pattern without a leading `**/` is anchored at
the build-context root. A `.dockerignore` listing `.env`, `*.pem` and `*.key`
therefore excludes only the copies at the repository root, while `config/.env`
and `certs/server.key` still reach the context and can land in an image layer
— which is more dangerous than a short file with no secret patterns at all,
because it reads as solved and stops anyone looking. Give every
depth-independent pattern the `**/` prefix and leave only genuinely
root-anchored entries unprefixed: `.claude`, and the repo's own host-built
binary, written `/myapp` and never `**/myapp`, which would also match
`cmd/myapp/` and delete the package directory from the context. Matching is
case-sensitive, and an ALL-CAPS twin per pattern still misses `Server.Key`, so
secret names use character ranges — `**/*.[kK][eE][yY]`, `**/*.[pP][eE][mM]`,
and likewise for `.envrc` and the extensionless SSH keys. Where such a pattern
also catches something the build needs, re-include it with a negation
(`!docs/example.env`); deleting the pattern reopens the exposure for every
other file it covers. Fetch the standard `.dockerignore` from
`https://git.eeqj.de/sneak/prompts/raw/branch/main/.dockerignore` and extend
it with the repo's own artifacts.
- **In-repo agent scratch belongs in both files, written to each file's own
semantics.** `.claude/` holds one worktree per in-flight agent — an entire
additional checkout of the repo — so under `COPY . .` the build context
inflates by a multiple of the repo and another session's unreviewed work can
be copied into an image layer. In `.gitignore` the entry is `.claude/`,
unanchored. In `.dockerignore` it is `.claude`, anchored and with **no** `**/`
prefix, because the prefixed form would also delete any nested directory of
that name from the build. Anchoring carries a known gap that the canonical
`.dockerignore` states in its own comment, since consuming repos receive the
file and not the tracker: the directory is created in the agent's working
directory, so a repo running agents in subdirectories still ships
`services/api/.claude/` and must add its own anchored entry there.
- **A plain `docker build .` of a clone stamps the version that
`git describe --tags --always` gives**, derived from the `.git` in the build
context as the canonical `Dockerfile` above shows. Without its failure check,
a missing `git` or an unreadable checkout would leave `-X main.Version=` empty
and the build would still exit 0. `script/docker` and `script/cibuild` pass
the version they compute on the host; it takes precedence. They do this
byte-identically across repos:
```sh
# Own line: a failing command substitution inside an argument does not
# trip `set -e`, so the inline form degrades to an empty constant.
version="$(git describe --tags --always --dirty 2>/dev/null || true)"
[ -n "$version" ] || version="unknown"
docker build --no-cache \
--build-arg VERSION="$version" \
-t "$(script/projectname)" .
```
`--always` makes an untagged repo yield an abbreviated commit hash rather
than failing, and the `[ -n "$version" ]` line is the single place the
fallback is applied — a live check that fires on a build from an export with
no `.git` and on a repository with no commits yet. Do not fold it into the
substitution as `|| echo unknown`, which makes the guard unreachable. The
Dockerfile's side is `ARG VERSION` in the stage that compiles, declared
there because `ARG` is stage-scoped; passing `VERSION` to a repo whose
Dockerfile declares no such `ARG` is ignored and costs nothing, which is why
the scripts stay byte-identical. One consequence for CI: the standard
checkout action clones shallow and fetches no tags, so a repo that embeds a
tag-derived version must set `fetch-depth: 0` on its checkout step.
- **Verify `.dockerignore` by enumerating the image, not by reading the
patterns.** Plant files at the root _and_ at least two directories deep, build
a probe image that does `COPY . .`, and list what actually landed
(`docker run --rm --entrypoint find IMAGE /app`). The `transferring context`
size is not a substitute: a nested secret is a few bytes, and BuildKit
transfers only the delta from the previous build.
- **No build artifacts in version control.** Code-derived data (compiled
bundles, minified output, generated assets) must never be committed to the
@@ -258,9 +476,56 @@ style conventions are in separate documents:
- Make all changes on a feature branch. You can do whatever you want on a
feature branch.
- `.golangci.yml` is standardized and must _NEVER_ be modified by an agent, only
manually by the user. Fetch from
`https://git.eeqj.de/sneak/prompts/raw/branch/main/.golangci.yml`.
- `.golangci.yml` is standardized. The vendored copy in a consuming repo must
_NEVER_ be modified by an agent: fetch it from
`https://git.eeqj.de/sneak/prompts/raw/branch/main/.golangci.yml` and keep it
byte-identical, so that no repo can quietly loosen its own linting. Linter
configuration changes are made to the canonical copy in the `prompts` repo and
reach consuming repos by re-vendoring; an agent may open a PR against
canonical, which only the user merges. One list is exempt from byte-identity,
because it cannot be written once for every repo: the `deny` list of the
`test-support` depguard rule, where a repo names its own test-support packages
by full import path. A repo adds entries there and changes nothing else, and a
re-vendor carries its entries forward. The canonical golangci-lint version is
v2.14.0 (released 2026-09-24), pinned as the digest of the lint phase's base
image
(`golangci/golangci-lint@sha256:ad862ba6b3798cbe0fd9fd7408d498fd74fbd2623a92406b2fd3898faf0bf98f`,
which reports `2.14.0 built with go1.27.0 from 114493f9`). A module's `go`
directive must not name a newer Go minor version than the one golangci-lint
was built with, or golangci-lint refuses to lint it: this release lints
`go 1.27.1` but not `go 1.28`. That digest is the only pin, since no repo
installs golangci-lint on the host. A repo sets the lint phase digest to the
one named here and re-vendors `.golangci.yml` in the same commit, whichever of
the two prompted the change: the canonical copy can name linters that an older
golangci-lint rejects, and a newer golangci-lint can add linters that
`default: all` switches on until the canonical copy disables them.
- **`script/bootstrap` installs a pinned tool by comparing versions, never by
testing presence.** An `if ! command -v <tool>; then install; fi` guard tests
`PATH` only, so on an already-provisioned machine the pin is inert and a
version bump is a silent no-op — while the Dockerfile, installing into a clean
image, gets the pinned version, so a local `make check` and `make docker` can
disagree about what the tool even is. The canonical form:
- compares the installed version against the pin over the **whole** version
token; a parser that stops at the first `-` reports `2.12.2` for a host
running `2.12.2-rc1` and skips the install;
- treats absent, non-zero, empty or unrecognised `--version` output as a
mismatch, so the failure direction is a redundant install and never a
skipped one;
- after installing, re-resolves the binary the way callers do — `hash -r`,
then through `PATH`, not through the directory the installer wrote to —
and fails naming the resolved path, since an install that a shadowing
binary hides succeeds while changing nothing any caller sees;
- is actually called, and prints the version on both success paths: a
function defined and never invoked has the same exit status and the same
empty output as one that worked.
Keep it POSIX sh: no arrays, no `[[`, no `grep -P`.
A Go tool a repo needs on the host is installed with `go install` pinned to
a commit hash (`go install <package>@<commit hash>`). It is never tracked as
a `go.mod` tool dependency or through a `tools.go` file, either of which
pulls the tool's own dependencies into the repo's `go.mod` and `go.sum`.
- When pinning images or packages by hash, add a comment above the reference
with the version and date (YYYY-MM-DD).
@@ -374,12 +639,14 @@ style conventions are in separate documents:
settings.
- Avoid putting files in the repo root unless necessary. Root should contain
only project-level config files (`README.md`, `Makefile`, `Dockerfile`,
`LICENSE`, `.gitignore`, `.editorconfig`, `REPO_POLICIES.md`, and
language-specific config). Everything else goes in a subdirectory. Canonical
subdirectory names:
only project-level config files (`README.md`, `AGENTS.md`, `Makefile`,
`Dockerfile`, `LICENSE`, `.gitignore`, `.editorconfig`, `REPO_POLICIES.md`,
and language-specific config). Everything else goes in a subdirectory.
Canonical subdirectory names:
- `bin/` — executable scripts and tools
- `cmd/` — Go command entrypoints
- `cmd/` — Go command entrypoints; thin only: one `main.go` per binary whose
body is a single call into `internal/` or `pkg/`, no project logic in
`cmd/`
- `configs/` — configuration templates and examples
- `deploy/` — deployment manifests (k8s, compose, terraform)
- `docs/` — documentation and markdown (README.md stays in root)
@@ -406,3 +673,7 @@ style conventions are in separate documents:
- Go: `go.mod`, `go.sum`, `.golangci.yml`
- JS: `package.json`, `yarn.lock`, `.prettierrc`, `.prettierignore`
- Python: `pyproject.toml`
- Guidance for coding agents lives in one `AGENTS.md` at the repository root. It
is never committed under a file or directory named after one agent tool, such
as `CLAUDE.md` or `.claude/`, and never split into separate memory files.
+213 -26
View File
@@ -1,34 +1,217 @@
# Workflow
- branch (from `main`)
- branch from `next`
- do the work in Next Step
- move Next Step to the top of Completed Steps
- move the top item of Future Steps into Next Step
- commit (`TODO.md` changes in the same commit as the work)
- merge to `main` if the branch is not protected, otherwise open a PR
- push
- push the branch and open a PR against `next`
# Status
pre-1.0. No git tags. `feat/reportbuf-storage` is merged; the backend, the CI
workflow, and the backend repo standard files are all on `main`. Frontend and
backend are both functional. Working toward the 1.0.0 milestone by closing the
remaining repo-compliance issues on the tracker.
pre-1.0. No git tags. `main` is the stable branch and `next` the development
branch, which every PR targets. The frontend and the Go backend ship as one
Docker image, and the Gitea workflow `.gitea/workflows/check.yml` runs
`script/cibuild` on every push. Working toward 1.0.0.
# Next Step
Confirm the `.gitea/workflows/check.yml` run is green (main always green
policy). The workflow file is already on `main`; what is unverified is that its
latest run passes.
Decide whether the repo moves to the layout `REPO_POLICIES.md` gives, with
`backend/` no longer repeating files from the root
([#30](https://git.eeqj.de/sneak/netwatch/issues/30)).
# Completed Steps
- 2026-10-07: the files shared from `sneak/prompts` are its copies at commit
`dd4027b` ([#113](https://git.eeqj.de/sneak/netwatch/issues/113)), with this
repository's own entries after the shared content in `.gitignore`,
`.dockerignore` and `.editorconfig`. `make lint` and `make test` each build
one phase of `Dockerfile` without the cache: `lint` runs golangci-lint v2.14.0
over `backend/` and, through the `frontend-lint` stage, eslint; `test` runs
the Go tests on the Debian Go image and, through the `frontend` stage, the
frontend's unit tests and build. The builder stage waits on both, and without
a `VERSION` build argument takes the version from `git describe` on the `.git`
in the build context. Neither linter nor the tests run on the host for
`make check`; `backend/script/lint` runs the root one, and
`backend/.golangci.yml` is no longer checked against a sha256. prettier
formats only the JavaScript, CSS, HTML and Markdown, so the shared
`.golangci.yml` stays as fetched. `script/fmt` and `script/fmt-check` put
`~/.local/bin` on `PATH`, which the shared workflow no longer does.
`script/frontend-lint` and `script/frontend-check` are gone
- 2026-10-07: the six Hetzner targets answer again, and failed checks are
written to the browser console
([#114](https://git.eeqj.de/sneak/netwatch/issues/114)). The Hetzner
speed-test servers close the connection without an answer when the URL has a
query string, and every check added `?_cb=` and the time; checks now fetch
each target's URL as written, with `cache: 'no-store'` as before. Each
recorded check that fails writes one `console.error` line, and the same line
to the debug log, with the target's name and URL, the time, what failed and
how long the request took; a target that answers after failed checks writes
one `console.info` line. Checks the page does not record write none, so the
debug log no longer lists failures in the first round or in the recovery probe
- 2026-10-04: the page's footer no longer says "IPv4 only"
([#111](https://git.eeqj.de/sneak/netwatch/issues/111)): each check is a
`fetch`, the browser picks IPv4 or IPv6 for each WAN host, and the local
targets are IPv4 addresses. The rest of the footer is unchanged
- 2026-10-04: `README.md`, `TODO.md` and `test/viewport/README.md` say what the
tree does (issue #24). The README's Getting Started leads with `make` targets;
a new Backend section says what `netwatch-server` stores, its routes and how
the image builds and runs it, and points to `backend/README.md` for its
settings; the checks are GET requests; the 26 WAN hosts, the four health
states, the summary's figures and the features the list lacked are described
as the page has them; and its TODO section points here, as does the one in
`backend/README.md`, whose open items moved to Future Steps. This file's
Workflow branches from `next` and opens the PR against `next`, Status says
where the repo stands, and Next Step and Future Steps hold only open work,
linked to its issue where one exists. The viewport harness README names Node's
test runner, not `vitest`
- 2026-10-04: in `src/main.js` (issue #102), a target's min, max, median and
average latency come from one list of its answers, through the same function
the summary's figures use, so the median is written once. The latency color
limits are one table in `CONFIG`, read by both the figure's and the
sparkline's color. The health thresholds, the debug log's length, the gateway
check's timeout, the recovery probe's number of hosts and interval, how often
the rows are sorted and the delay before the first sparkline resize are
`CONFIG` entries too. A unit test checks the summary's figures. Nothing the
page does or shows changed; the footer's color legend still writes the limits
out as text
- 2026-10-04: password guesses at `/metrics` are rate limited (issue #104): each
client address, resolved through `TRUSTED_PROXIES` as for reports, may make 60
requests to `/metrics` a minute, counted by `go-chi/httprate` apart from its
reports; past that it gets 429 and its basic auth credentials are not checked.
The limit is a constant in `backend/internal/server/routes.go`. A test uses up
one client's allowance on wrong passwords, gets 429 with the right one, and
checks that another client behind the same nginx still gets in
- 2026-10-04: the backend reports errors to Sentry (issue #95). With
`SENTRY_DSN` set, it sets up `sentry-go` with the release `netwatch-server-`
and its version, reports each panic in a handler through `sentryhttp`, the
last of the middleware every request goes through, which panics again so the
request still gets the 500 from the panic recovery, and waits up to 2 seconds
on shutdown for Sentry to finish sending. A DSN Sentry refuses stops the start
with an error naming `SENTRY_DSN`. With it empty, Sentry is not set up and
nothing is sent to it
- 2026-10-04: a target's name and URL and a debug log message show as the
characters they are and are never read as HTML (issue #29): a host row escapes
the name and URL it writes into its markup, and the debug log sets each line
as text. A unit test checks a target whose name and URL hold `<`, `>`, `"`,
`&` and `'`. `README.md` no longer calls `CONFIG` frozen: the interval menu
sets its `updateInterval`, and the values computed from it follow. `AppState`
declares the recovery probe's two properties, the sparkline axis functions
lose the parameters they did not use, and the comment on a target's history
names both kinds of entry it holds. Nothing the page does changed
- 2026-10-04: a dependency can be added without running yarn or go by hand
(issue #45): `make add-dependency PACKAGE=<name>@<version>` shims to the new
`script/add-dependency`, which runs `yarn add --dev`, so `package.json` and
`yarn.lock` change together, then `yarn install --frozen-lockfile`; the same
command moves a package to another version. `make tidy` shims to the new
`script/tidy`, which runs `go mod tidy` in `backend/`: a Go module is added by
importing it, or moved by editing its `require` line, then `make tidy`.
`script/bootstrap` still installs with `--frozen-lockfile`
- 2026-10-04: the backend serves Prometheus metrics (issue #94). With
`METRICS_USERNAME` and `METRICS_PASSWORD` both set, it records request
duration and response size through `go-http-metrics` and serves them, with
Go's runtime and process metrics, at `GET /metrics` behind basic auth with
those credentials; nginx passes `/metrics` to it as it does `/api/`. With
neither set there are no metrics and `/metrics` is 404; one without the other
stops the start with an error naming both, and so does a `METRICS_USERNAME`
containing `:`, with an error naming it. Only requests that reach the health
check or `POST /api/v1/reports` are recorded, not `/metrics` itself and not
every request as `GO_HTTP_SERVER_CONVENTIONS.md` shows, because the labels are
the request's path and method, which clients can make up without end. For
that, `POST /api/v1/reports` is now registered by its full path instead of
inside a `/api/v1` route group; it answers as before
- 2026-10-04: `script/` and `Makefile` follow the org models (issue #28):
`make dev` shims to the new `script/dev`, the Vite dev server, and the new
`make build` to `script/build`, the frontend production build.
`.prettierignore` no longer leaves out `backend/`, so `make fmt` and
`make fmt-check` cover `backend/README.md`; it leaves out
`backend/.golangci.yml` by name, the org standard file whose sha256
`backend/script/lint` checks. `script/install-precommit` and the date on
`script/bootstrap`'s pins are the org model again; `script/bootstrap`,
`script/fmt` and `script/fmt-check` each say in a comment why they differ from
it
- 2026-10-04: a frontend build on Node 26 or newer, such as `make test` on a
host with Node 26, no longer prints Node's warning that `module.register()` is
deprecated (issue #32); the build in `Dockerfile` runs on Node 22, which never
printed it. The call was in `@tailwindcss/node`, which `@tailwindcss/vite`
brings in at its own exact version; tailwind 4.3.1 calls
`module.registerHooks()` instead where Node has it. `yarn.lock` now has
`@tailwindcss/vite` and `tailwindcss` at 4.3.3, inside the ranges
`package.json` already allowed, and tailwind's own dependencies moved with
them. The built CSS changes only in how it is written out, in tailwind's
Firefox focus-ring rule, which no longer applies to iframes (the page has
none, so nothing on it looks different), and in tailwind's default sans-serif
font list, which the page does not use: `body` sets a monospace font
- 2026-10-03: the frontend's unit tests cover what the page computes (issue
#21): `humanDuration`, the latency colours of a figure and of a sparkline
either side of each boundary, a target's min, max, average and median latency
over an empty history, an all-unreachable one and a mixed one, and each of the
four health states either side of its thresholds. `package.json` has a `test`
script, so `yarn run test` and `npm run test` run them, and
`script/frontend-test` runs it: quietly, and if a test fails, again with every
test listed, and then fails. `src/main.js` now exports `humanDuration`,
`HostState`, `latencyClass` and `latencyHex` for the tests; nothing it does
changed
- 2026-10-03: the backend's logs are one stream (issue #27): fx logs its own
steps of starting and stopping through the backend's logger, so off a terminal
every line the backend's own logger and fx write is JSON, where fx used to
write plain text to stderr. A config file that is found but cannot be read now
stops the start with its error, logged as JSON like a bad setting, where it
used to end in a Go panic. A test runs the server as a child process and
checks both. The backend logs its name, version and architecture once at
start. The health check's uptime keys are now `uptime_seconds` and
`uptime_human`; its path, content type, `"status":"ok"` and 200 are unchanged.
`SENTRY_DSN`, `METRICS_USERNAME` and `METRICS_PASSWORD` are still read and
still unused, and left out of `backend/README.md`, until issues #94 and #95
wire them up
- 2026-10-03: the frontend has a real linter (issue #47, and item 2 of issue
#28): `eslint` with its recommended rules, set in `eslint.config.js`, runs in
a new `frontend-lint` stage of `Dockerfile`, which the frontend stage waits
on, as the builder stage waits on the Go `lint` stage. `script/lint` builds
both stages without the cache and runs no linter on the host.
`script/frontend-lint` runs eslint where it used to repeat the prettier check
that `script/fmt-check` runs on the host, and `script/frontend-check`, run by
the frontend stage, is now the tests and the format check. `script/bootstrap`
wants node 22.13.0 or newer, as eslint 10 does
- 2026-10-03: the tap-target check in `make frontend-viewport-test` expects one
visible pin button per WAN host row (issue #46), where it expected at least 10
of the 26, so pin buttons missing from only some rows now fail it. The host
row count the harness gathers, which the `app-rendered` check also reads, now
counts only the WAN host rows: the local host rows have no pin button
- 2026-10-03: each target's row shows its result as soon as its check ends
(issue #91), where every row waited for the round's slowest check, up to 24
seconds at a 30-second interval. Sorting, the summary, the health box and
offline detection still run once, when the round's last check ends. A check
that ends after the user pauses or after its round is given up shows nothing,
and the first round is still discarded as a whole
seconds at a 30-second interval. Every row is still redrawn, and sorting, the
summary, the health box and offline detection still run, once, when the
round's last check ends, so no row reads "paused" after a pause and resume
during the round. A check that ends after the user pauses or after its round
is given up shows nothing, and the first round is still discarded as a whole
- 2026-10-03: root no longer acts outside `/data` when it prepares `DATA_DIR`
(issue #80): `bin/entrypoint.sh` runs `netwatch-server prepare-data-dir`,
which refuses a `DATA_DIR` that is not `/data` or a path below it written in
full, then creates `DATA_DIR`, gives `/data` and everything in it to
`netwatch` and sets the modes, all through a Go `os.Root` opened on `/data`.
That refuses any path leading out of `/data`, so neither a symbolic link
already there nor one a host process swaps in during the start can make root
create or change anything elsewhere, and `DATA_DIR=/etc` no longer gives
`/etc` to `netwatch`. The `README.md` section "Running under upaas" says which
values are accepted
- 2026-10-03: `DATA_DIR_MAX_BYTES` is now how much of the report files is kept
(issue #54): when a report would take them past it, the oldest report files
are deleted to make room, each deletion logged, and at start files already
past it are deleted the same way. A file still being written is never deleted.
A report is refused with 507 only when the reports waiting to be written fill
the cap on their own, and then no file is deleted. The reports of a failed
write stop counting, and the part of its file written is removed. A file that
cannot be deleted still counts until the next start; one already deleted by
hand counts as freed
- 2026-10-03: `backend/script/lint` says what went wrong with its
`.golangci.yml` check (issue #34). On a hash mismatch it says to compare the
file with the org standard: if they differ, restore the org standard; if they
are the same, the org standard changed, so update `GOLANGCI_CONFIG_SHA256` in
that script. It used to say only to restore the file, which loops once the org
standard itself has moved. A missing `.golangci.yml`, and a `sha256sum` that
is missing or prints no hash, each get their own message instead of being
reported as a mismatch; every one still fails the lint
- 2026-10-03: the Go tests run with the race detector and coverage (issue #88):
`backend/script/test` runs `go test -timeout 30s -race -cover ./...` and, if
that fails, runs it again with `-v` and fails. Go's `-timeout` bounds the
@@ -209,9 +392,11 @@ latest run passes.
- 2026-08-09: automated responsive-layout harness
(`make frontend-viewport-test`): digest-pinned headless Chrome driven over CDP
against the built `dist/`, viewport widths derived from the breakpoints in
`src/styles.css` ([#13](https://git.eeqj.de/sneak/netwatch/issues/13)). Every
check carries a presence guard so none of them can pass against a page it is
not actually measuring. Found two real layout defects, filed as
`src/styles.css` ([#13](https://git.eeqj.de/sneak/netwatch/issues/13)). The
tap-target and host-row checks each fail when they measured nothing; the
overflow, viewport-edge and clipped-text checks have no such guard of their
own and rely on the `app-rendered` check, which fails the run when the app did
not render. Found two real layout defects, filed as
[#42](https://git.eeqj.de/sneak/netwatch/issues/42) and
[#43](https://git.eeqj.de/sneak/netwatch/issues/43)
- 2026-07-07 Adopted scripts-to-rule-them-all: `script/` entrypoints, Makefile
@@ -231,12 +416,14 @@ latest run passes.
# Future Steps
- Wire `script/frontend-viewport-test` into CI as its own step (deliberately not
part of `make check` today; the decision has real CI-runtime cost and is
tracked separately)
- Compliance top-up as one small commit: add .editorconfig and add the hooks
target to the Makefile
- After merge, confirm .gitea/workflows/check.yml is on main and CI is green
(main always green policy)
- Decide what to do with untracked resume.sh: commit it, gitignore it, or delete
it
- Run `make frontend-viewport-test` in CI as its own step; it is not part of
`make check`, as it needs Docker and takes minutes
- A backend test that posts a report to `POST /api/v1/reports` and checks the
compressed file it is written to
- A backend route that decompresses the stored reports and answers queries on
them
- Prometheus metrics for the backend's in-memory buffer: its size, the number of
flushes and the number of reports
- A configurable host list (an environment variable or a config file)
- Export of the latency history (CSV or JSON)
- A notification when the health status changes to DEGRADED
+1
View File
@@ -17,6 +17,7 @@ linters:
disable:
# Genuinely incompatible with project patterns
- exhaustruct # Requires all struct fields
- exhaustruct_v5 # Requires all struct fields (successor to exhaustruct)
- godot # Requires comments to end with periods
- wrapcheck # Too verbose for internal packages
- varnamelen # Short names like db, id are idiomatic Go
+91 -43
View File
@@ -28,21 +28,21 @@ docker run -p 8080:8080 netwatch
This directory follows the same
[Scripts to Rule Them All](https://github.com/github/scripts-to-rule-them-all)
pattern as the repo root: the targets in `backend/Makefile` are thin shims over
`backend/script/`. The root `Dockerfile` runs them, and the root scripts call
`test`, `fmt` and `fmt-check`:
`backend/script/`. The root `Dockerfile` runs `build`, and the root scripts call
`fmt` and `fmt-check`:
- `script/build` — compile the static `netwatch-server` binary with its version
stamped in. The version is `VERSION` from the environment;
when that is unset or empty, it falls back to `git describe` inside a git
checkout, then to `dev`
- `script/test` — run the Go tests with the race detector and coverage. Go's
`-timeout 30s` bounds the tests, not their compile. If they fail, they run
again with `-v` for the details, and the script fails. The race detector needs
a C compiler
- `script/lint` — check `.golangci.yml` against its pinned sha256, then run
golangci-lint. It runs inside the golangci-lint image of the lint stage of
the root `Dockerfile`; from a checkout, run `make lint` at the repo root,
which builds that stage
stamped in. The version is `VERSION` from the environment; when that is unset
or empty, it falls back to `git describe` inside a git checkout, then to `dev`
- `script/test` — run the Go tests on the host with the race detector and
coverage; the root `make test` runs them in the `test` phase of the root
`Dockerfile`. Go's `-timeout 90s` bounds the tests, not their compile, and
`-count=1` keeps Go from reporting a stored pass. If they fail, they run again
with `-v` for the details, and the script fails. The race detector needs a C
compiler
- `script/lint` — run the root `script/lint`, which builds the `lint` phase of
the root `Dockerfile`: golangci-lint over this directory, and eslint over the
frontend. golangci-lint never runs on the host
- `script/fmt` — format the Go sources (writes)
- `script/fmt-check` — check Go formatting (read-only)
- `script/run` — build and run the server locally
@@ -61,8 +61,9 @@ flushes them to compressed files on disk for later analysis.
## Design
The server is structured as an `fx`-wired Go application under `cmd/netwatch-server/`.
Internal packages in `internal/` follow standard Go project layout:
The server is structured as an `fx`-wired Go application under
`cmd/netwatch-server/`. Internal packages in `internal/` follow standard Go
project layout:
- **`config`**: Loads configuration from environment variables and config files
via Viper.
@@ -83,17 +84,21 @@ Internal packages in `internal/` follow standard Go project layout:
| `BIND_ADDRESS` | empty | IP address to listen on; empty listens on every interface |
| `PORT` | `8080` | HTTP listen port |
| `DATA_DIR` | `./data/reports` | Directory for compressed reports |
| `DATA_DIR_MAX_BYTES` | `1073741824` (1 GiB) | Largest total size of the report files in `DATA_DIR`; see [Report limits](#report-limits) |
| `DATA_DIR_MAX_BYTES` | `1073741824` (1 GiB) | Most bytes of report files kept in `DATA_DIR`, oldest deleted first; see [Report limits](#report-limits) |
| `DEBUG` | `false` | Enable debug logging |
| `TRUSTED_PROXIES` | loopback + RFC1918 | Comma-separated CIDRs whose `X-Forwarded-For` / `X-Real-IP` headers are trusted for client IP resolution |
| `REPORTS_PER_MINUTE` | `60` | Reports each client address may send a minute; see [Report limits](#report-limits) |
| `CORS_ALLOWED_ORIGINS` | empty | Comma-separated origins whose pages may call the API; see [CORS](#cors) |
| `METRICS_USERNAME` | empty | Basic auth user name for `/metrics`; see [Metrics](#metrics) |
| `METRICS_PASSWORD` | empty | Basic auth password for `/metrics`; see [Metrics](#metrics) |
| `SENTRY_DSN` | empty | DSN of the Sentry project to send errors to; see [Sentry](#sentry) |
`TRUSTED_PROXIES` defaults to `127.0.0.1/32,::1/128,10.0.0.0/8,172.16.0.0/12,192.168.0.0/16`.
The loopback entries cover a reverse proxy on the same host. A request whose
direct peer is outside this set has its forwarded headers ignored, and the
direct peer is logged and rate-limited instead. The container image does not use
this default; see [Container image](#container-image).
`TRUSTED_PROXIES` defaults to
`127.0.0.1/32,::1/128,10.0.0.0/8,172.16.0.0/12,192.168.0.0/16`. The loopback
entries cover a reverse proxy on the same host. A request whose direct peer is
outside this set has its forwarded headers ignored, and the direct peer is
logged and rate-limited instead. The container image does not use this default;
see [Container image](#container-image).
A variable set to a value the server cannot use, such as `PORT=abc`,
`DEBUG=maybe` or a `BIND_ADDRESS` that is not an IP address, stops it from
@@ -102,15 +107,16 @@ starting, with an error naming the variable. An empty variable counts as unset.
### Container image
The root `Dockerfile` builds one image in which nginx listens on the public port
8080, serves the frontend, and proxies `/api/` and `/.well-known/healthcheck` to
this server. The image's entrypoint, `bin/entrypoint.sh`, starts the server as
user `netwatch` (uid 1000) with `BIND_ADDRESS=127.0.0.1` and `PORT=8081`, so
only nginx reaches it, and with `TRUSTED_PROXIES=127.0.0.1/32`, so it takes the
client address nginx passes on and no other. `DATA_DIR` is `/data/reports`, on
the `/data` volume; the entrypoint creates it and gives it and `/data` to
`netwatch` before starting the server. nginx replaces the security headers
this server sets with those in the root `security-headers.conf`, so those are
what clients of the image see.
8080, serves the frontend, and proxies `/api/`, `/.well-known/healthcheck` and
`/metrics` to this server. The image's entrypoint, `bin/entrypoint.sh`, starts
the server as user `netwatch` (uid 1000) with `BIND_ADDRESS=127.0.0.1` and
`PORT=8081`, so only nginx reaches it, and with `TRUSTED_PROXIES=127.0.0.1/32`,
so it takes the client address nginx passes on and no other. `DATA_DIR` is
`/data/reports`, on the `/data` volume; before starting the server, the
entrypoint creates it and gives it and `/data` to `netwatch` with
`netwatch-server prepare-data-dir`, which acts on nothing outside `/data`. nginx
replaces the security headers this server sets with those in the root
`security-headers.conf`, so those are what clients of the image see.
The container's own `TRUSTED_PROXIES` goes to nginx instead: IP addresses or
CIDRs, separated by commas, of the reverse proxies in front of the container.
@@ -128,10 +134,11 @@ Reports are written as `reports-<timestamp>-<number>.jsonl.zst` files in
`DATA_DIR`. The timestamp is in UTC to the millisecond, so the names sort by
time. The number starts at 1 when the server starts and goes up by one for each
file the server starts to write, so two files written in the same millisecond
still get different names. A failed write uses up its number, leaving a gap in
the numbers if the file could not be created and otherwise a file under that
number that may be incomplete. Each file contains one JSON object per line,
compressed with zstd. Files are created with `O_EXCL` to prevent overwrites.
still get different names. A failed write uses up its number and leaves a gap in
the numbers: its file, if it was created, is removed. The file stays, counted
toward `DATA_DIR_MAX_BYTES` from the next start, only if removing it fails too.
Each file contains one JSON object per line, compressed with zstd. Files are
created with `O_EXCL` to prevent overwrites.
### Report limits
@@ -151,11 +158,16 @@ credentials, so it is bounded instead. Both refusals below answer with the same
`X-RateLimit-Reset` headers.
- **Size cap.** The report files in `DATA_DIR` may total at most
`DATA_DIR_MAX_BYTES`, counting the files already there at start. Reports
waiting in memory count at their uncompressed size until they are written, so
a report that would take the total past the cap is refused with 507, and
nothing of it is stored. Deleting report files frees room only at the next
start, when the files are counted again. The default of 1 GiB is small enough
for any host; set it to the space you can give `DATA_DIR`.
waiting in memory count at their uncompressed size until they are written;
those lost to a failed write stop counting, and the part of its file written
is removed. When a report would take the total past the cap, the oldest report
files are deleted to make room, and each deletion is logged with the file's
name and size; a file still being written is never deleted. A report is
refused with 507, and nothing of it is stored, only when the reports waiting
to be written fill the cap on their own, and then no file is deleted. At
start, report files past the cap, as after lowering it, are deleted the same
way. So the cap is how much of the newest reports is kept: the default of 1
GiB is small enough for any host; set it to the space you can give `DATA_DIR`.
### CORS
@@ -168,12 +180,48 @@ origin, `scheme://host` with an optional `:port`, as browsers send it: no path,
not even a trailing `/`, and no `*`. Any other entry stops the server from
starting, with an error naming `CORS_ALLOWED_ORIGINS`.
### Metrics
With both `METRICS_USERNAME` and `METRICS_PASSWORD` set, the server serves
Prometheus metrics at `GET /metrics` to requests with those as their basic auth
credentials, and answers any other with 401. For each request that reaches the
health check or `POST /api/v1/reports`, those the rate limit refuses included,
the metrics record its duration and response size, labelled with its path,
method and status; they also count those requests in progress, and include Go's
runtime and process metrics. No other request is recorded: not those to
`/metrics` itself, and not those answered before they reach either route, such
as a CORS preflight, or a request refused with 404 for a path no route has, 405
for a method its route does not take, or 413 for declaring a body length over
the 1 MiB limit. A report whose body goes over the limit without declaring its
length reaches the route, is answered 413 there, and is recorded with that
status. Clients can make up any number of paths and methods, and each would add
labels to the metrics for as long as the server runs. With neither set, nothing
is recorded and `/metrics` answers 404. One without the other stops the server
from starting, with an error naming both; so does a `METRICS_USERNAME`
containing `:`, which basic auth cannot carry, with an error naming it.
`/metrics` is rate limited, so that its password cannot be guessed quickly: each
client address, resolved through `TRUSTED_PROXIES`, may make 60 requests to it a
minute, whatever their credentials. Past that it gets 429 with
`Retry-After: 60`, and its credentials are not checked. The minute slides as it
does for reports (see [Report limits](#report-limits)), so a scraper polling
every 2 seconds or less often is never refused. This allowance is apart from the
one for reports.
### Sentry
With `SENTRY_DSN` set, the server sends its errors to that Sentry project: each
panic in a handler is reported there, under the release `netwatch-server-`
followed by the server's version, and the request still gets 500 from the
server's panic recovery. On shutdown the server waits up to 2 seconds for Sentry
to finish sending. A DSN Sentry refuses stops the server from starting, with an
error naming `SENTRY_DSN`. With it empty, Sentry is not set up, and nothing is
sent to it.
## TODO
- Add integration test that POSTs a report and verifies the compressed output
- Add report decompression/query endpoint
- Add metrics (Prometheus) for buffer size, flush count, report count
- Add retention policy to prune old report files
The to-do list, this backend's open work included, is [TODO.md](../TODO.md) at
the repo root.
## License
+34 -1
View File
@@ -4,6 +4,7 @@ package main
import (
"fmt"
"os"
"os/user"
"sneak.berlin/go/netwatch/internal/config"
"sneak.berlin/go/netwatch/internal/globals"
@@ -15,6 +16,7 @@ import (
"sneak.berlin/go/netwatch/internal/server"
"go.uber.org/fx"
"go.uber.org/fx/fxevent"
)
//nolint:gochecknoglobals // set via ldflags at build time
@@ -37,10 +39,36 @@ func main() {
return
}
// "netwatch-server prepare-data-dir DATA_DIR" gets DATA_DIR ready
// for the netwatch user, or exits 1 with the error; see
// reportbuf.PrepareDataDir. bin/entrypoint.sh runs it as root
// before it starts this server as that user.
if len(os.Args) == 3 && os.Args[1] == "prepare-data-dir" {
netwatch, err := user.Lookup("netwatch")
if err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
err = reportbuf.PrepareDataDir("/data", os.Args[2], netwatch)
if err != nil {
fmt.Fprintf(os.Stderr, "DATA_DIR '%s': %v\n", os.Args[2], err)
os.Exit(1)
}
return
}
globals.Appname = Appname
globals.Version = Version
fx.New(
// fx logs each step of starting and stopping through the
// server's own logger, so off a terminal those lines are
// JSON like every other line.
fx.WithLogger(func(log *logger.Logger) fxevent.Logger {
return &fxevent.SlogLogger{Logger: log.Get()}
}),
fx.Provide(
config.New,
globals.New,
@@ -51,6 +79,11 @@ func main() {
reportbuf.New,
server.New,
),
fx.Invoke(func(*server.Server) {}),
fx.Invoke(
// First, so the name and version are logged even when
// a setting stops the start.
func(log *logger.Logger) { log.Identify() },
func(*server.Server) {},
),
).Run()
}
+249
View File
@@ -0,0 +1,249 @@
package main
import (
"bytes"
"context"
"encoding/json"
"net"
"net/http"
"os"
"os/exec"
"os/signal"
"path/filepath"
"strings"
"syscall"
"testing"
"time"
)
// The tests run main() in a child process, this test binary started
// again with runMainEnv set, because main() can exit its process and
// takes its settings from the environment.
const runMainEnv = "NETWATCH_SERVER_RUN_MAIN"
// childTimeout bounds each child's whole run; it is killed after it.
const childTimeout = 10 * time.Second
func TestMain(m *testing.M) {
if os.Getenv(runMainEnv) != "" {
// A SIGTERM that comes before fx catches it is dropped, not fatal.
signal.Notify(make(chan os.Signal, 1), syscall.SIGTERM)
main()
return
}
os.Exit(m.Run())
}
// TestOutputIsJSON: off a terminal, every line the server writes from
// start to stop is JSON, fx's own lines included.
func TestOutputIsJSON(t *testing.T) {
t.Parallel()
ctx, cancel := context.WithTimeout(t.Context(), childTimeout)
defer cancel()
port := freePort(ctx, t)
child, stdout, stderr := startServer(ctx, t, t.TempDir(), port)
waitForHealthcheck(ctx, t, port)
// The child drops a SIGTERM that comes before fx catches it (see
// TestMain), so send one every 100ms until the test ends. ctx
// bounds the wait: when it ends, the child is killed.
stop := make(chan struct{})
defer close(stop)
go func() {
for {
_ = child.Process.Signal(syscall.SIGTERM)
select {
case <-stop:
return
case <-time.After(100 * time.Millisecond):
}
}
}()
err := child.Wait()
if err != nil {
t.Fatalf("server exit = %v, want success", err)
}
requireJSONLines(t, stdout, stderr)
if !strings.Contains(stdout.String(), `"msg":"starting"`) {
t.Fatalf("no startup line in stdout:\n%s", stdout)
}
}
// TestMalformedConfigFileStopsTheStart: a config file the server finds
// but cannot read stops the start, and the error is logged as JSON.
func TestMalformedConfigFileStopsTheStart(t *testing.T) {
t.Parallel()
ctx, cancel := context.WithTimeout(t.Context(), childTimeout)
defer cancel()
home := t.TempDir()
dir := filepath.Join(home, ".config", "netwatch-server")
err := os.MkdirAll(dir, 0o750)
if err != nil {
t.Fatal(err)
}
err = os.WriteFile(filepath.Join(dir, "netwatch-server.yaml"),
[]byte("PORT: [8080\n"), 0o600)
if err != nil {
t.Fatal(err)
}
child, stdout, stderr := startServer(ctx, t, home, freePort(ctx, t))
err = child.Wait()
if child.ProcessState.ExitCode() != 1 {
t.Fatalf("server exit = %v, want exit status 1", err)
}
requireJSONLines(t, stdout, stderr)
if !strings.Contains(stdout.String(), "netwatch-server.yaml") {
t.Fatalf("no error naming the config file in stdout:\n%s", stdout)
}
}
// TestRefusedSentryDSNStopsTheStart: a SENTRY_DSN that Sentry refuses
// stops the start, and the error, naming SENTRY_DSN, is logged as JSON.
func TestRefusedSentryDSNStopsTheStart(t *testing.T) {
t.Setenv("SENTRY_DSN", "not-a-dsn")
ctx, cancel := context.WithTimeout(t.Context(), childTimeout)
defer cancel()
child, stdout, stderr := startServer(ctx, t, t.TempDir(), freePort(ctx, t))
err := child.Wait()
if child.ProcessState.ExitCode() != 1 {
t.Fatalf("server exit = %v, want exit status 1", err)
}
requireJSONLines(t, stdout, stderr)
if !strings.Contains(stdout.String(), "SENTRY_DSN") {
t.Fatalf("no error naming SENTRY_DSN in stdout:\n%s", stdout)
}
}
// startServer runs main() in a child process listening on
// 127.0.0.1:port, with home as its HOME and working directory and its
// data directory in home, so it touches nothing outside home. Its
// stdout and stderr go to the two buffers returned, which hold all of
// it once child.Wait returns. The child is killed when ctx ends, and
// killed and reaped when the test ends if nothing waited for it.
func startServer(
ctx context.Context,
t *testing.T,
home, port string,
) (*exec.Cmd, *bytes.Buffer, *bytes.Buffer) {
t.Helper()
self, err := os.Executable()
if err != nil {
t.Fatal(err)
}
var stdout, stderr bytes.Buffer
child := exec.CommandContext(ctx, self) //nolint:gosec // this test binary
child.Dir = home
child.Env = append(os.Environ(),
runMainEnv+"=1",
"HOME="+home,
"DATA_DIR="+filepath.Join(home, "data"),
"BIND_ADDRESS=127.0.0.1",
"PORT="+port,
)
child.Stdout = &stdout
child.Stderr = &stderr
err = child.Start()
if err != nil {
t.Fatal(err)
}
t.Cleanup(func() {
if child.ProcessState == nil {
_ = child.Process.Kill()
_ = child.Wait()
}
})
return child, &stdout, &stderr
}
// freePort returns a TCP port on 127.0.0.1 that was free a moment ago.
func freePort(ctx context.Context, t *testing.T) string {
t.Helper()
var lc net.ListenConfig
l, err := lc.Listen(ctx, "tcp", "127.0.0.1:0")
if err != nil {
t.Fatal(err)
}
_ = l.Close()
_, port, err := net.SplitHostPort(l.Addr().String())
if err != nil {
t.Fatal(err)
}
return port
}
// waitForHealthcheck returns once the health check on port answers 200.
func waitForHealthcheck(ctx context.Context, t *testing.T, port string) {
t.Helper()
url := "http://127.0.0.1:" + port + "/.well-known/healthcheck"
for {
req, err := http.NewRequestWithContext(ctx, http.MethodGet, url, nil)
if err != nil {
t.Fatal(err)
}
resp, err := http.DefaultClient.Do(req)
if err == nil {
_ = resp.Body.Close()
if resp.StatusCode == http.StatusOK {
return
}
}
select {
case <-ctx.Done():
t.Fatalf("health check never answered: %v", err)
case <-time.After(50 * time.Millisecond):
}
}
}
// requireJSONLines fails the test on each line of outs that is not
// JSON.
func requireJSONLines(t *testing.T, outs ...*bytes.Buffer) {
t.Helper()
for _, out := range outs {
for line := range strings.Lines(out.String()) {
if !json.Valid([]byte(line)) {
t.Errorf("line is not JSON: %s", line)
}
}
}
}
+14 -3
View File
@@ -3,20 +3,30 @@ module sneak.berlin/go/netwatch
go 1.25.5
require (
github.com/99designs/basicauth-go v0.0.0-20230316000542-bf6f9cbbf0f8
github.com/getsentry/sentry-go v0.49.0
github.com/go-chi/chi/v5 v5.2.5
github.com/go-chi/cors v1.2.2
github.com/go-chi/httprate v0.16.0
github.com/joho/godotenv v1.5.1
github.com/klauspost/compress v1.18.4
github.com/klauspost/compress v1.19.1
github.com/prometheus/client_golang v1.24.1
github.com/slok/go-http-metrics v0.13.0
github.com/spf13/viper v1.21.0
go.uber.org/fx v1.24.0
)
require (
github.com/beorn7/perks v1.0.1 // indirect
github.com/cespare/xxhash/v2 v2.3.0 // indirect
github.com/fsnotify/fsnotify v1.9.0 // indirect
github.com/go-viper/mapstructure/v2 v2.4.0 // indirect
github.com/klauspost/cpuid/v2 v2.2.10 // indirect
github.com/munnerz/goautoneg v0.0.0-20191010083416-a7dc8b61c822 // indirect
github.com/pelletier/go-toml/v2 v2.2.4 // indirect
github.com/prometheus/client_model v0.6.2 // indirect
github.com/prometheus/common v0.70.1 // indirect
github.com/prometheus/procfs v0.21.1 // indirect
github.com/sagikazarmark/locafero v0.11.0 // indirect
github.com/sourcegraph/conc v0.3.1-0.20240121214520-5f936abd7ae8 // indirect
github.com/spf13/afero v1.15.0 // indirect
@@ -28,6 +38,7 @@ require (
go.uber.org/multierr v1.10.0 // indirect
go.uber.org/zap v1.26.0 // indirect
go.yaml.in/yaml/v3 v3.0.4 // indirect
golang.org/x/sys v0.30.0 // indirect
golang.org/x/text v0.28.0 // indirect
golang.org/x/sys v0.47.0 // indirect
golang.org/x/text v0.40.0 // indirect
google.golang.org/protobuf v1.36.11 // indirect
)
+52 -18
View File
@@ -1,37 +1,65 @@
github.com/davecgh/go-spew v1.1.1 h1:vj9j/u1bqnvCEfJOwUhtlOARqs3+rkHYY13jYWTU97c=
github.com/davecgh/go-spew v1.1.1/go.mod h1:J7Y8YcW2NihsgmVo/mv3lAwl/skON4iLHjSsI+c5H38=
github.com/99designs/basicauth-go v0.0.0-20230316000542-bf6f9cbbf0f8 h1:nMpu1t4amK3vJWBibQ5X/Nv0aXL+b69TQf2uK5PH7Go=
github.com/99designs/basicauth-go v0.0.0-20230316000542-bf6f9cbbf0f8/go.mod h1:3cARGAK9CfW3HoxCy1a0G4TKrdiKke8ftOMEOHyySYs=
github.com/beorn7/perks v1.0.1 h1:VlbKKnNfV8bJzeqoa4cOKqO6bYr3WgKZxO8Z16+hsOM=
github.com/beorn7/perks v1.0.1/go.mod h1:G2ZrVWU2WbWT9wwq4/hrbKbnv/1ERSJQ0ibhJ6rlkpw=
github.com/cespare/xxhash/v2 v2.3.0 h1:UL815xU9SqsFlibzuggzjXhog7bL6oX9BbNZnL2UFvs=
github.com/cespare/xxhash/v2 v2.3.0/go.mod h1:VGX0DQ3Q6kWi7AoAeZDth3/j3BFtOZR5XLFGgcrjCOs=
github.com/davecgh/go-spew v1.1.2-0.20180830191138-d8f796af33cc h1:U9qPSI2PIWSS1VwoXQT9A3Wy9MM3WgvqSxFWenqJduM=
github.com/davecgh/go-spew v1.1.2-0.20180830191138-d8f796af33cc/go.mod h1:J7Y8YcW2NihsgmVo/mv3lAwl/skON4iLHjSsI+c5H38=
github.com/frankban/quicktest v1.14.6 h1:7Xjx+VpznH+oBnejlPUj8oUpdxnVs4f8XU8WnHkI4W8=
github.com/frankban/quicktest v1.14.6/go.mod h1:4ptaffx2x8+WTWXmUCuVU6aPUX1/Mz7zb5vbUoiM6w0=
github.com/fsnotify/fsnotify v1.9.0 h1:2Ml+OJNzbYCTzsxtv8vKSFD9PbJjmhYF14k/jKC7S9k=
github.com/fsnotify/fsnotify v1.9.0/go.mod h1:8jBTzvmWwFyi3Pb8djgCCO5IBqzKJ/Jwo8TRcHyHii0=
github.com/getsentry/sentry-go v0.49.0 h1:Ehejknu1l023Ub7QoRBVLAI7g3Jnhqku4oWx4B4Sh5s=
github.com/getsentry/sentry-go v0.49.0/go.mod h1:nuMJAoCfe1u0Bts2ocyNI+TW8HT84vRMqwA5Qq/SKUI=
github.com/go-chi/chi/v5 v5.2.5 h1:Eg4myHZBjyvJmAFjFvWgrqDTXFyOzjj7YIm3L3mu6Ug=
github.com/go-chi/chi/v5 v5.2.5/go.mod h1:X7Gx4mteadT3eDOMTsXzmI4/rwUpOwBHLpAfupzFJP0=
github.com/go-chi/cors v1.2.2 h1:Jmey33TE+b+rB7fT8MUy1u0I4L+NARQlK6LhzKPSyQE=
github.com/go-chi/cors v1.2.2/go.mod h1:sSbTewc+6wYHBBCW7ytsFSn836hqM7JxpglAy2Vzc58=
github.com/go-chi/httprate v0.16.0 h1:8V5DH9j6pSK6UQoBsTpvMyFxycqaKEIToyPKzHJjUa8=
github.com/go-chi/httprate v0.16.0/go.mod h1:A8lo+qRhk+s9LiuP5saS7XCGDXRXMcrueq0NfIuCa/I=
github.com/go-errors/errors v1.4.2 h1:J6MZopCL4uSllY1OfXM374weqZFFItUbrImctkmUxIA=
github.com/go-errors/errors v1.4.2/go.mod h1:sIVyrIiJhuEF+Pj9Ebtd6P/rEYROXFi3BopGUQ5a5Og=
github.com/go-viper/mapstructure/v2 v2.4.0 h1:EBsztssimR/CONLSZZ04E8qAkxNYq4Qp9LvH92wZUgs=
github.com/go-viper/mapstructure/v2 v2.4.0/go.mod h1:oJDH3BJKyqBA2TXFhDsKDGDTlndYOZ6rGS0BRZIxGhM=
github.com/google/go-cmp v0.6.0 h1:ofyhxvXcZhMsU5ulbFiLKl/XBFqE1GSq7atu8tAmTRI=
github.com/google/go-cmp v0.6.0/go.mod h1:17dUlkBOakJ0+DkrSSNjCkIjxS6bF9zb3elmeNGIjoY=
github.com/google/go-cmp v0.7.0 h1:wk8382ETsv4JYUZwIsn6YpYiWiBsYLSJiTsyBybVuN8=
github.com/google/go-cmp v0.7.0/go.mod h1:pXiqmnSA92OHEEa9HXL2W4E7lf9JzCmGVUdgjX3N/iU=
github.com/joho/godotenv v1.5.1 h1:7eLL/+HRGLY0ldzfGMeQkb7vMd0as4CfYvUVzLqw0N0=
github.com/joho/godotenv v1.5.1/go.mod h1:f4LDr5Voq0i2e/R5DDNOoa2zzDfwtkZa6DnEwAbqwq4=
github.com/klauspost/compress v1.18.4 h1:RPhnKRAQ4Fh8zU2FY/6ZFDwTVTxgJ/EMydqSTzE9a2c=
github.com/klauspost/compress v1.18.4/go.mod h1:R0h/fSBs8DE4ENlcrlib3PsXS61voFxhIs2DeRhCvJ4=
github.com/klauspost/compress v1.19.1 h1:VsB4HPswih7mmZ8WleSFQ75c/Ui1M4trX5oAsJnhSlk=
github.com/klauspost/compress v1.19.1/go.mod h1:cwPg85FWrGar70rWktvGQj8/hthj3wpl0PGDogxkrSQ=
github.com/klauspost/cpuid/v2 v2.2.10 h1:tBs3QSyvjDyFTq3uoc/9xFpCuOsJQFNPiAhYdw2skhE=
github.com/klauspost/cpuid/v2 v2.2.10/go.mod h1:hqwkgyIinND0mEev00jJYCxPNVRVXFQeu1XKlok6oO0=
github.com/kr/pretty v0.3.1 h1:flRD4NNwYAUpkphVc1HcthR4KEIFJ65n8Mw5qdRn3LE=
github.com/kr/pretty v0.3.1/go.mod h1:hoEshYVHaxMs3cyo3Yncou5ZscifuDolrwPKZanG3xk=
github.com/kr/text v0.2.0 h1:5Nx0Ya0ZqY2ygV366QzturHI13Jq95ApcVaJBhpS+AY=
github.com/kr/text v0.2.0/go.mod h1:eLer722TekiGuMkidMxC/pM04lWEeraHUUmBw8l2grE=
github.com/kylelemons/godebug v1.1.0 h1:RPNrshWIDI6G2gRW9EHilWtl7Z6Sb1BR0xunSBf0SNc=
github.com/kylelemons/godebug v1.1.0/go.mod h1:9/0rRGxNHcop5bhtWyNeEfOS8JIWk580+fNqagV/RAw=
github.com/munnerz/goautoneg v0.0.0-20191010083416-a7dc8b61c822 h1:C3w9PqII01/Oq1c1nUAm88MOHcQC9l5mIlSMApZMrHA=
github.com/munnerz/goautoneg v0.0.0-20191010083416-a7dc8b61c822/go.mod h1:+n7T8mK8HuQTcFwEeznm/DIxMOiR9yIdICNftLE1DvQ=
github.com/pelletier/go-toml/v2 v2.2.4 h1:mye9XuhQ6gvn5h28+VilKrrPoQVanw5PMw/TB0t5Ec4=
github.com/pelletier/go-toml/v2 v2.2.4/go.mod h1:2gIqNv+qfxSVS7cM2xJQKtLSTLUE9V8t9Stt+h56mCY=
github.com/pmezard/go-difflib v1.0.0 h1:4DBwDE0NGyQoBHbLQYPwSUPoCMWR5BEzIk/f1lZbAQM=
github.com/pmezard/go-difflib v1.0.0/go.mod h1:iKH77koFhYxTK1pcRnkKkqfTogsbg7gZNVY4sRDYZ/4=
github.com/rogpeppe/go-internal v1.9.0 h1:73kH8U+JUqXU8lRuOHeVHaa/SZPifC7BkcraZVejAe8=
github.com/rogpeppe/go-internal v1.9.0/go.mod h1:WtVeX8xhTBvf0smdhujwtBcq4Qrzq/fJaraNFVN+nFs=
github.com/pingcap/errors v0.11.4 h1:lFuQV/oaUMGcD2tqt+01ROSmJs75VG1ToEOkZIZ4nE4=
github.com/pingcap/errors v0.11.4/go.mod h1:Oi8TUi2kEtXXLMJk9l1cGmz20kV3TaQ0usTwv5KuLY8=
github.com/pkg/errors v0.9.1 h1:FEBLx1zS214owpjy7qsBeixbURkuhQAwrK5UwLGTwt4=
github.com/pkg/errors v0.9.1/go.mod h1:bwawxfHBFNV+L2hUp1rHADufV3IMtnDRdf1r5NINEl0=
github.com/pmezard/go-difflib v1.0.1-0.20181226105442-5d4384ee4fb2 h1:Jamvg5psRIccs7FGNTlIRMkT8wgtp5eCXdBlqhYGL6U=
github.com/pmezard/go-difflib v1.0.1-0.20181226105442-5d4384ee4fb2/go.mod h1:iKH77koFhYxTK1pcRnkKkqfTogsbg7gZNVY4sRDYZ/4=
github.com/prometheus/client_golang v1.24.1 h1:JnJkREXzWxUdCuPFpIWZiPispT9xVV59uiuyR2bPlnU=
github.com/prometheus/client_golang v1.24.1/go.mod h1:F+oSRECHg4sse5ucfYpYDeIv/hu68Zo0uoHKetWnzcE=
github.com/prometheus/client_model v0.6.2 h1:oBsgwpGs7iVziMvrGhE53c/GrLUsZdHnqNwqPLxwZyk=
github.com/prometheus/client_model v0.6.2/go.mod h1:y3m2F6Gdpfy6Ut/GBsUqTWZqCUvMVzSfMLjcu6wAwpE=
github.com/prometheus/common v0.70.1 h1:1HvjP4D5oL3t8RsPlwxA9onvvStjtIHYE5XuuwOi/PY=
github.com/prometheus/common v0.70.1/go.mod h1:VdFUQDMZK3VLkurFUVhia6uys/0suUp86TJz5qbJRhc=
github.com/prometheus/procfs v0.21.1 h1:GljZCt+zSTS+NZq88cyQ1LjZ+RCHp3uVuabBWA5+OJI=
github.com/prometheus/procfs v0.21.1/go.mod h1:aB55Cww9pdSJVHk0hUf0inxWyyjPogFIjmHKYgMKmtY=
github.com/rogpeppe/go-internal v1.14.1 h1:UQB4HGPB6osV0SQTLymcB4TgvyWu6ZyliaW0tI/otEQ=
github.com/rogpeppe/go-internal v1.14.1/go.mod h1:MaRKkUm5W0goXpeCfT7UZI6fk/L7L7so1lCWt35ZSgc=
github.com/sagikazarmark/locafero v0.11.0 h1:1iurJgmM9G3PA/I+wWYIOw/5SyBtxapeHDcg+AAIFXc=
github.com/sagikazarmark/locafero v0.11.0/go.mod h1:nVIGvgyzw595SUSUE6tvCp3YYTeHs15MvlmU87WwIik=
github.com/slok/go-http-metrics v0.13.0 h1:lQDyJJx9wKhmbliyUsZ2l6peGnXRHjsjoqPt5VYzcP8=
github.com/slok/go-http-metrics v0.13.0/go.mod h1:HIr7t/HbN2sJaunvnt9wKP9xoBBVZFo1/KiHU3b0w+4=
github.com/sourcegraph/conc v0.3.1-0.20240121214520-5f936abd7ae8 h1:+jumHNA0Wrelhe64i8F6HNlS8pkoyMv5sreGx2Ry5Rw=
github.com/sourcegraph/conc v0.3.1-0.20240121214520-5f936abd7ae8/go.mod h1:3n1Cwaq1E1/1lhQhtRK2ts/ZwZEhjcQeJQ1RuC6Q/8U=
github.com/spf13/afero v1.15.0 h1:b/YBCLWAJdFWJTN9cLhiXXcD7mzKn9Dm86dNnfyQw1I=
@@ -42,6 +70,8 @@ github.com/spf13/pflag v1.0.10 h1:4EBh2KAYBwaONj6b2Ye1GiHfwjqyROoF4RwYO+vPwFk=
github.com/spf13/pflag v1.0.10/go.mod h1:McXfInJRrz4CZXVZOBLb0bTZqETkiAhM9Iw0y3An2Bg=
github.com/spf13/viper v1.21.0 h1:x5S+0EU27Lbphp4UKm1C+1oQO+rKx36vfCoaVebLFSU=
github.com/spf13/viper v1.21.0/go.mod h1:P0lhsswPGWD/1lZJ9ny3fYnVqxiegrlNrEmgLjbTCAY=
github.com/stretchr/objx v0.5.2 h1:xuMeJ0Sdp5ZMRXx/aWO6RZxdr3beISkG5/G/aIRr3pY=
github.com/stretchr/objx v0.5.2/go.mod h1:FRsXN1f5AsAjCGJKqEizvkpNtU+EGNCLh3NxZ/8L+MA=
github.com/stretchr/testify v1.11.1 h1:7s2iGBzp5EwR7/aIZr8ao5+dra3wiQyKjjFuvgVKu7U=
github.com/stretchr/testify v1.11.1/go.mod h1:wZwfW3scLgRK+23gO65QZefKpKQRnfz6sD981Nm4B6U=
github.com/subosito/gotenv v1.6.0 h1:9NlTDc1FTs4qu0DDq7AEtTPNw6SVm7uBMsUCUjABIf8=
@@ -54,20 +84,24 @@ go.uber.org/dig v1.19.0 h1:BACLhebsYdpQ7IROQ1AGPjrXcP5dF80U3gKoFzbaq/4=
go.uber.org/dig v1.19.0/go.mod h1:Us0rSJiThwCv2GteUN0Q7OKvU7n5J4dxZ9JKUXozFdE=
go.uber.org/fx v1.24.0 h1:wE8mruvpg2kiiL1Vqd0CC+tr0/24XIB10Iwp2lLWzkg=
go.uber.org/fx v1.24.0/go.mod h1:AmDeGyS+ZARGKM4tlH4FY2Jr63VjbEDJHtqXTGP5hbo=
go.uber.org/goleak v1.2.0 h1:xqgm/S+aQvhWFTtR0XK3Jvg7z8kGV8P4X14IzwN3Eqk=
go.uber.org/goleak v1.2.0/go.mod h1:XJYK+MuIchqpmGmUSAzotztawfKvYLUIgg7guXrwVUo=
go.uber.org/goleak v1.3.0 h1:2K3zAYmnTNqV73imy9J1T3WC+gmCePx2hEGkimedGto=
go.uber.org/goleak v1.3.0/go.mod h1:CoHD4mav9JJNrW/WLlf7HGZPjdw8EucARQHekz1X6bE=
go.uber.org/multierr v1.10.0 h1:S0h4aNzvfcFsC3dRF1jLoaov7oRaKqRGC/pUEJ2yvPQ=
go.uber.org/multierr v1.10.0/go.mod h1:20+QtiLqy0Nd6FdQB9TLXag12DsQkrbs3htMFfDN80Y=
go.uber.org/zap v1.26.0 h1:sI7k6L95XOKS281NhVKOFCUNIvv9e0w4BF8N3u+tCRo=
go.uber.org/zap v1.26.0/go.mod h1:dtElttAiwGvoJ/vj4IwHBS/gXsEu/pZ50mUIRWuG0so=
go.yaml.in/yaml/v2 v2.4.4 h1:tuyd0P+2Ont/d6e2rl3be67goVK4R6deVxCUX5vyPaQ=
go.yaml.in/yaml/v2 v2.4.4/go.mod h1:gMZqIpDtDqOfM0uNfy0SkpRhvUryYH0Z6wdMYcacYXQ=
go.yaml.in/yaml/v3 v3.0.4 h1:tfq32ie2Jv2UxXFdLJdh3jXuOzWiL1fo0bu/FbuKpbc=
go.yaml.in/yaml/v3 v3.0.4/go.mod h1:DhzuOOF2ATzADvBadXxruRBLzYTpT36CKvDb3+aBEFg=
golang.org/x/sys v0.30.0 h1:QjkSwP/36a20jFYWkSue1YwXzLmsV5Gfq7Eiy72C1uc=
golang.org/x/sys v0.30.0/go.mod h1:/VUhepiaJMQUp4+oa/7Zr1D23ma6VTLIYjOOTFZPUcA=
golang.org/x/text v0.28.0 h1:rhazDwis8INMIwQ4tpjLDzUhx6RlXqZNPEM0huQojng=
golang.org/x/text v0.28.0/go.mod h1:U8nCwOR8jO/marOQ0QbDiOngZVEBB7MAiitBuMjXiNU=
golang.org/x/sys v0.47.0 h1:o7XGOvZQCADBQQ4Y7VNq2dRWQR7JmOUW8Kxx4ZsNgWs=
golang.org/x/sys v0.47.0/go.mod h1:4GL1E5IUh+htKOUEOaiffhrAeqysfVGipDYzABqnCmw=
golang.org/x/text v0.40.0 h1:Ub2Z6/xjgF1WrYQz2nuITOEegKFtiIy+rieRJ5lHZKs=
golang.org/x/text v0.40.0/go.mod h1:hpnzDAfGV753zIKo+wk3u1bVKCGPbrnF7+7LBF/UHVY=
google.golang.org/protobuf v1.36.11 h1:fV6ZwhNocDyBLK0dj+fg8ektcVegBBuEolpbTQyBNVE=
google.golang.org/protobuf v1.36.11/go.mod h1:HTf+CrKn2C3g5S8VImy6tdcUvCska2kB7j23XfzDpco=
gopkg.in/check.v1 v0.0.0-20161208181325-20d25e280405/go.mod h1:Co6ibVJAznAaIkqp8huTwlJQCZ016jof/cbN4VW5Yz0=
gopkg.in/check.v1 v1.0.0-20190902080502-41f04d3bba15 h1:YR8cESwS4TdDjEe65xsg0ogRM/Nc3DYOhEAlW+xobZo=
gopkg.in/check.v1 v1.0.0-20190902080502-41f04d3bba15/go.mod h1:Co6ibVJAznAaIkqp8huTwlJQCZ016jof/cbN4VW5Yz0=
gopkg.in/check.v1 v1.0.0-20201130134442-10cb98267c6c h1:Hei/4ADfdWqJk1ZMxUNpqntNwaWcugrBjAiHlqqRiVk=
gopkg.in/check.v1 v1.0.0-20201130134442-10cb98267c6c/go.mod h1:JHkPIbrfpd72SG/EVd6muEfDQjcINNoR0C8j2r3qZ4Q=
gopkg.in/yaml.v3 v3.0.1 h1:fxVm/GzAzEWqLHuvctI91KS9hhNmmWOoWu0XTYJS7CA=
gopkg.in/yaml.v3 v3.0.1/go.mod h1:K4uyk7z7BCEPqu6E+C64Yfv1cQ7kz7rIZviUmN+EgEM=
+25 -3
View File
@@ -44,6 +44,15 @@ var (
errNotPort = errors.New("must be a port number, 1 to 65535")
errNotBool = errors.New("must be true or false")
errNotIP = errors.New("must be an IP address, or empty")
errMetricsCredentials = errors.New(
"METRICS_USERNAME and METRICS_PASSWORD must be set together, " +
"or neither",
)
errMetricsUsernameColon = errors.New(
"METRICS_USERNAME must not contain \":\", " +
"which basic auth cannot carry in a user name",
)
)
// Params defines the dependencies for Config.
@@ -73,7 +82,8 @@ type Config struct {
// New loads configuration from env, .env files, and config
// files, returning a fully resolved Config. It fails, with an error
// naming the setting, on a value the server cannot use.
// naming the setting, on a value the server cannot use, and on a
// config file it finds but cannot read.
func New(
_ fx.Lifecycle,
params Params,
@@ -106,8 +116,8 @@ func New(
if err != nil {
var notFound viper.ConfigFileNotFoundError
if !errors.As(err, &notFound) {
log.Error("config file malformed", "error", err)
panic(err)
return nil, fmt.Errorf("config file %s: %w",
viper.ConfigFileUsed(), err)
}
}
@@ -177,6 +187,18 @@ func (s *Config) check() error {
}
}
// The server records and serves metrics only with both set, so
// one alone is a mistake that would otherwise go unnoticed.
if (s.MetricsUsername == "") != (s.MetricsPassword == "") {
return errMetricsCredentials
}
// Basic auth splits the credentials at the first ":", so with one
// in the user name every request to /metrics would get 401.
if strings.Contains(s.MetricsUsername, ":") {
return errMetricsUsernameColon
}
return checkOrigins(s.CORSAllowedOrigins)
}
+26
View File
@@ -103,6 +103,32 @@ func TestDataDirMaxBytesMustBeANumber(t *testing.T) {
requireConfigError(t, "DATA_DIR_MAX_BYTES")
}
// TestMetricsCredentialsGoTogether: with only one of the two set, the
// server would quietly serve no metrics, so the start fails, naming
// both.
func TestMetricsCredentialsGoTogether(t *testing.T) {
for _, set := range []string{"METRICS_USERNAME", "METRICS_PASSWORD"} {
t.Run(set, func(t *testing.T) {
t.Setenv("METRICS_USERNAME", "")
t.Setenv("METRICS_PASSWORD", "")
t.Setenv(set, "prometheus")
requireConfigError(t, "METRICS_USERNAME")
requireConfigError(t, "METRICS_PASSWORD")
})
}
}
// TestMetricsUsernameMustNotContainColon: basic auth splits the
// credentials at the first ":", so such a user name would get 401 on
// every request to /metrics.
func TestMetricsUsernameMustNotContainColon(t *testing.T) {
t.Setenv("METRICS_USERNAME", "prom:etheus")
t.Setenv("METRICS_PASSWORD", "secret")
requireConfigError(t, "METRICS_USERNAME")
}
// TestCORSAllowedOriginsMustBeOrigins: "*" would let every origin in,
// and an entry that is not a plain origin would match no page.
func TestCORSAllowedOriginsMustBeOrigins(t *testing.T) {
+1 -1
View File
@@ -6,6 +6,6 @@ import "net/http"
// endpoint.
func (s *Handlers) HandleHealthCheck() http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
s.respondJSON(w, r, s.hc.Check(), http.StatusOK)
s.respondJSON(w, r, s.hc.Healthcheck(), http.StatusOK)
}
}
@@ -50,8 +50,8 @@ func newStartedHandlers(t *testing.T, g *globals.Globals) *handlers.Handlers {
// TestHandleHealthCheck checks the health check's answer: 200, a JSON
// content type, and a JSON object with exactly the fields of
// healthcheck.Response, carrying this server's name and version and
// an uptime counted from its start.
// healthcheck.HealthcheckResponse, carrying this server's name and
// version and an uptime counted from its start.
func TestHandleHealthCheck(t *testing.T) {
t.Parallel()
@@ -82,7 +82,7 @@ func TestHandleHealthCheck(t *testing.T) {
}
fields := []string{
"appname", "now", "status", "uptimeHuman", "uptimeSeconds", "version",
"appname", "now", "status", "uptime_human", "uptime_seconds", "version",
}
if got := slices.Sorted(maps.Keys(body)); !slices.Equal(got, fields) {
t.Fatalf("fields = %v, want %v", got, fields)
@@ -104,17 +104,17 @@ func TestHandleHealthCheck(t *testing.T) {
}
// Started just now, so the uptime is well under a minute.
human, _ := body["uptimeHuman"].(string)
human, _ := body["uptime_human"].(string)
uptime, err := time.ParseDuration(human)
if err != nil || uptime > time.Minute {
t.Errorf("uptimeHuman = %q, want a duration under a minute (%v)",
t.Errorf("uptime_human = %q, want a duration under a minute (%v)",
human, err)
}
seconds, ok := body["uptimeSeconds"].(float64)
seconds, ok := body["uptime_seconds"].(float64)
if !ok || seconds < 0 || seconds > time.Minute.Seconds() {
t.Errorf("uptimeSeconds = %v, want a number of seconds under a minute",
body["uptimeSeconds"])
t.Errorf("uptime_seconds = %v, want a number of seconds under a minute",
body["uptime_seconds"])
}
}
+4 -3
View File
@@ -86,11 +86,12 @@ func (s *Handlers) decodeErrorStatus(err error) int {
}
// appendErrorStatus logs a failure to store a report and returns
// the status to send: 507 when the report files are at their size
// cap, otherwise 500.
// the status to send: 507 when the reports waiting to be written fill
// the size cap, otherwise 500.
func (s *Handlers) appendErrorStatus(err error) int {
if errors.Is(err, reportbuf.ErrFull) {
s.log.Warn("report refused: report files at their size cap")
s.log.Warn("report refused: " +
"reports waiting to be written fill the size cap")
return http.StatusInsufficientStorage
}
+3 -3
View File
@@ -125,9 +125,9 @@ func TestHandleReportStorageFailureIsNon2xx(t *testing.T) {
}
}
// TestHandleReportFullIs507 checks the answer when the report files
// are at their size cap: 507 and the usual error body, which tells
// the client nothing more.
// TestHandleReportFullIs507 checks the answer when the reports waiting
// to be written fill the size cap: 507 and the usual error body, which
// tells the client nothing more.
func TestHandleReportFullIs507(t *testing.T) {
t.Parallel()
+11 -8
View File
@@ -30,14 +30,17 @@ type Healthcheck struct {
params *Params
}
// Response is the JSON payload returned by the health check
// endpoint.
type Response struct {
// HealthcheckResponse is the JSON payload returned by the health
// check endpoint. Its name and its snake_case keys are the ones
// GO_HTTP_SERVER_CONVENTIONS.md gives.
//
//nolint:revive,tagliatelle // name and keys from the conventions
type HealthcheckResponse struct {
Appname string `json:"appname"`
Now string `json:"now"`
Status string `json:"status"`
UptimeHuman string `json:"uptimeHuman"`
UptimeSeconds int64 `json:"uptimeSeconds"`
UptimeHuman string `json:"uptime_human"`
UptimeSeconds int64 `json:"uptime_seconds"`
Version string `json:"version"`
}
@@ -65,9 +68,9 @@ func New(
return s, nil
}
// Check returns the current health status of the application.
func (s *Healthcheck) Check() *Response {
return &Response{
// Healthcheck returns the current health status of the application.
func (s *Healthcheck) Healthcheck() *HealthcheckResponse {
return &HealthcheckResponse{
Appname: s.params.Globals.Appname,
Now: time.Now().UTC().Format(time.RFC3339Nano),
Status: "ok",
+27
View File
@@ -18,9 +18,14 @@ import (
"sneak.berlin/go/netwatch/internal/globals"
"sneak.berlin/go/netwatch/internal/logger"
basicauth "github.com/99designs/basicauth-go"
"github.com/go-chi/chi/v5/middleware"
"github.com/go-chi/cors"
"github.com/go-chi/httprate"
"github.com/prometheus/client_golang/prometheus"
metrics "github.com/slok/go-http-metrics/metrics/prometheus"
ghmm "github.com/slok/go-http-metrics/middleware"
"github.com/slok/go-http-metrics/middleware/std"
"go.uber.org/fx"
)
@@ -364,3 +369,25 @@ func (s *Middleware) RateLimit(
),
)
}
// Metrics returns middleware that records each request's duration and
// response size, and the requests in progress, in registry. They are
// labelled by the request path.
func (s *Middleware) Metrics(
registry prometheus.Registerer,
) func(http.Handler) http.Handler {
mdlw := ghmm.New(ghmm.Config{
Recorder: metrics.NewRecorder(metrics.Config{Registry: registry}),
})
return std.HandlerProvider("", mdlw)
}
// MetricsAuth returns middleware that lets a request through only with
// METRICS_USERNAME and METRICS_PASSWORD as its basic auth credentials,
// and answers any other with 401.
func (s *Middleware) MetricsAuth() func(http.Handler) http.Handler {
return basicauth.New("metrics", map[string][]string{
s.params.Config.MetricsUsername: {s.params.Config.MetricsPassword},
})
}
@@ -343,7 +343,7 @@ func TestLoggingCutsRequestStringsToBound(t *testing.T) {
http.MethodGet, "/"+long, http.NoBody)
req.Header.Set("User-Agent", long)
req.Header.Set("Referer", long)
req.Header.Set("X-Request-Id", long)
req.Header.Set("X-Request-ID", long)
handler.ServeHTTP(httptest.NewRecorder(), req)
+150
View File
@@ -0,0 +1,150 @@
package reportbuf
import (
"errors"
"io/fs"
"os"
"os/user"
"path/filepath"
"strconv"
"syscall"
)
// ErrDataDirOutsideVolume is returned by PrepareDataDir for a DATA_DIR
// that is not the volume or a path below it, written in full.
var ErrDataDirOutsideVolume = errors.New(
"must be /data or a path below it, with no '.', '..' or extra '/'")
// PrepareDataDir gets dir, the server's DATA_DIR, ready for owner, the
// user the server runs as, so that a host directory mounted at volume,
// /data in the image, needs no preparing: it creates dir, gives volume
// and everything in it to owner, and gives volume and dir the mode the
// server gives a directory it creates. dir must be volume or a path
// below it, with no '.', '..', empty part or '/' at the end.
//
// bin/entrypoint.sh runs this as root, which would follow a symbolic
// link anywhere, so every step goes through an os.Root opened on
// volume: it follows a link only when it is written as a relative
// path that stays inside volume, and refuses any other. A process on
// the host can swap a link onto a path in volume at any moment while
// this runs. Even then, the os.Root checks each link as it reaches
// it. MkdirAll creates each directory inside a parent it already has
// open, never following a link at the name it creates, and follows a
// link on the path only as the os.Root allows, so a relative link
// inside volume can lead it to create directories elsewhere inside
// volume. Lchown never changes what a link points to, and the modes
// are set on directories already opened (see chmodDir), so the most
// that process can do is make a step fail or wait, or act on
// something else inside volume.
func PrepareDataDir(volume, dir string, owner *user.User) error {
// rel is dir as a path from volume; IsLocal is false for one that
// leads out of it.
rel, err := filepath.Rel(volume, dir)
if err != nil || dir != filepath.Clean(dir) || !filepath.IsLocal(rel) {
return ErrDataDirOutsideVolume
}
uid, err := strconv.Atoi(owner.Uid)
if err != nil {
return err
}
gid, err := strconv.Atoi(owner.Gid)
if err != nil {
return err
}
root, err := os.OpenRoot(volume)
if err != nil {
return err
}
defer func() { _ = root.Close() }()
err = root.MkdirAll(rel, dirPerms)
if err != nil {
return err
}
err = lchownAll(root, ".", uid, gid)
if err != nil {
return err
}
err = chmodDir(root, ".")
if err != nil {
return err
}
return chmodDir(root, rel)
}
// lchownAll gives name, a directory inside root, and everything in it
// to uid and gid. It reads each directory opened through root, not
// through root.FS(), which refuses a name that is not valid UTF-8, and
// calls Lchown on every entry, which gives a symbolic link itself to
// them, not what it points to. It goes into an entry only when the
// read found a directory there, so it follows no link it finds; one
// swapped in for that directory afterwards is followed only as the
// os.Root allows.
func lchownAll(root *os.Root, name string, uid, gid int) error {
err := root.Lchown(name, uid, gid)
if err != nil {
return err
}
dir, err := root.Open(name)
if err != nil {
return err
}
entries, err := dir.ReadDir(-1)
_ = dir.Close()
if err != nil {
return err
}
for _, entry := range entries {
entryName := filepath.Join(name, entry.Name())
if entry.IsDir() {
err = lchownAll(root, entryName, uid, gid)
} else {
err = root.Lchown(entryName, uid, gid)
}
if err != nil {
return err
}
}
return nil
}
// chmodDir gives name, a directory inside root, the mode the server
// gives a directory it creates. Root.Chmod would not hold: on Linux it
// checks that name is not a symbolic link, then sets the mode by name,
// following a link swapped in between. So chmodDir opens name through
// root and sets the mode on the open directory. It refuses anything
// but a directory: a directory has no second name (hard link), so the
// one opened is inside root, where any other file could be a hard link
// to one outside.
func chmodDir(root *os.Root, name string) error {
dir, err := root.Open(name)
if err != nil {
return err
}
defer func() { _ = dir.Close() }()
info, err := dir.Stat()
if err != nil {
return err
}
if !info.IsDir() {
return &fs.PathError{Op: "chmod", Path: name, Err: syscall.ENOTDIR}
}
return dir.Chmod(dirPerms)
}
+307
View File
@@ -0,0 +1,307 @@
package reportbuf_test
import (
"errors"
"io/fs"
"os"
"os/user"
"path/filepath"
"strconv"
"syscall"
"testing"
"sneak.berlin/go/netwatch/internal/reportbuf"
)
// reports is the last part of DATA_DIR in these tests, as in the
// image's /data/reports.
const reports = "reports"
// currentUser is the user the test runs as, the only owner a test not
// run as root can give files to.
func currentUser() *user.User {
return &user.User{
Uid: strconv.Itoa(os.Getuid()),
Gid: strconv.Itoa(os.Getgid()),
}
}
// tempDirMode700 is a new directory in a t.TempDir with mode 0700, so
// a test can tell that PrepareDataDir left its mode alone.
func tempDirMode700(t *testing.T) string {
t.Helper()
dir := filepath.Join(t.TempDir(), "d")
err := os.Mkdir(dir, 0o700)
if err != nil {
t.Fatal(err)
}
return dir
}
func requireMode(t *testing.T, path string, want fs.FileMode) {
t.Helper()
info, err := os.Stat(path)
if err != nil {
t.Fatal(err)
}
if info.Mode() != want {
t.Errorf("%s: mode %v, want %v", path, info.Mode(), want)
}
}
func requireMissing(t *testing.T, path string) {
t.Helper()
_, err := os.Lstat(path)
if !errors.Is(err, fs.ErrNotExist) {
t.Errorf("%s: created, or Lstat failed: %v", path, err)
}
}
func requireOwner(t *testing.T, path string, uid, gid uint32) {
t.Helper()
info, err := os.Lstat(path)
if err != nil {
t.Fatal(err)
}
stat, _ := info.Sys().(*syscall.Stat_t)
if stat.Uid != uid || stat.Gid != gid {
t.Errorf("%s: owner %d:%d, want %d:%d", path, stat.Uid, stat.Gid,
uid, gid)
}
}
func TestPrepareDataDirCreatesDataDir(t *testing.T) {
t.Parallel()
volume := t.TempDir()
dir := filepath.Join(volume, "a", reports)
err := reportbuf.PrepareDataDir(volume, dir, currentUser())
if err != nil {
t.Fatal(err)
}
requireMode(t, volume, fs.ModeDir|0o750)
requireMode(t, dir, fs.ModeDir|0o750)
}
// TestPrepareDataDirSetsModeOfExistingDataDir: a DATA_DIR already on
// the host with another mode gets the mode too, not only a new one.
func TestPrepareDataDirSetsModeOfExistingDataDir(t *testing.T) {
t.Parallel()
volume := t.TempDir()
dir := filepath.Join(volume, reports)
err := os.Mkdir(dir, 0o700)
if err != nil {
t.Fatal(err)
}
err = reportbuf.PrepareDataDir(volume, dir, currentUser())
if err != nil {
t.Fatal(err)
}
requireMode(t, dir, fs.ModeDir|0o750)
}
func TestPrepareDataDirTakesTheVolumeItself(t *testing.T) {
t.Parallel()
volume := t.TempDir()
err := reportbuf.PrepareDataDir(volume, volume, currentUser())
if err != nil {
t.Fatal(err)
}
requireMode(t, volume, fs.ModeDir|0o750)
}
// TestPrepareDataDirRefusesDataDirOutsideVolume covers a DATA_DIR that
// is relative, outside the volume, or not written in full.
func TestPrepareDataDirRefusesDataDirOutsideVolume(t *testing.T) {
t.Parallel()
volume := tempDirMode700(t)
for _, dir := range []string{
reports, volume + "/../new", volume + "/", volume + "//" + reports,
volume + "/./" + reports, volume + "/" + reports + "/..", volume + "x",
"/etc",
} {
err := reportbuf.PrepareDataDir(volume, dir, currentUser())
if !errors.Is(err, reportbuf.ErrDataDirOutsideVolume) {
t.Errorf("%q: error = %v, want ErrDataDirOutsideVolume", dir, err)
}
}
requireMissing(t, filepath.Join(filepath.Dir(volume), "new"))
requireMode(t, volume, fs.ModeDir|0o700)
}
// TestPrepareDataDirRefusesLinkOutOfVolume puts a symbolic link to a
// directory outside the volume on the path to DATA_DIR, written as a
// full path and as one that climbs out with '..', and as DATA_DIR
// itself, where the mode would be set through it.
func TestPrepareDataDirRefusesLinkOutOfVolume(t *testing.T) {
t.Parallel()
for _, tc := range []struct {
name string
climbsOut bool
link, dir string
}{
{"full path", false, "x", "x/reports"},
{"climbs out", true, "x", "x/reports"},
{"DATA_DIR itself", false, reports, reports},
} {
t.Run(tc.name, func(t *testing.T) {
t.Parallel()
outside := tempDirMode700(t)
volume := t.TempDir()
climbOut, err := filepath.Rel(volume, outside)
if err != nil {
t.Fatal(err)
}
target := outside
if tc.climbsOut {
target = climbOut
}
err = os.Symlink(target, filepath.Join(volume, tc.link))
if err != nil {
t.Fatal(err)
}
err = reportbuf.PrepareDataDir(volume,
filepath.Join(volume, tc.dir), currentUser())
if err == nil {
t.Error("no error")
}
requireMissing(t, filepath.Join(outside, reports))
requireMode(t, outside, fs.ModeDir|0o700)
})
}
}
// TestPrepareDataDirRefusesDanglingLink: DATA_DIR is a symbolic link
// to a name in the volume that does not exist, which is not created.
func TestPrepareDataDirRefusesDanglingLink(t *testing.T) {
t.Parallel()
volume := t.TempDir()
err := os.Symlink("missing", filepath.Join(volume, reports))
if err != nil {
t.Fatal(err)
}
err = reportbuf.PrepareDataDir(volume, filepath.Join(volume, reports),
currentUser())
if err == nil {
t.Error("no error")
}
requireMissing(t, filepath.Join(volume, "missing"))
}
// TestPrepareDataDirTakesDirectoryNamedInLatin1: a host directory can
// hold names that are not valid UTF-8, here "café" written in Latin-1.
// A test not run as root can only check that PrepareDataDir goes into
// such a directory and gives it, and what it holds, to the current
// user.
func TestPrepareDataDirTakesDirectoryNamedInLatin1(t *testing.T) {
t.Parallel()
volume := t.TempDir()
latin1 := filepath.Join(volume, "caf\xe9")
err := os.Mkdir(latin1, 0o700)
if err != nil {
t.Fatal(err)
}
err = os.WriteFile(filepath.Join(latin1, "f"), nil, 0o600)
if err != nil {
t.Fatal(err)
}
owner := currentUser()
err = reportbuf.PrepareDataDir(volume, filepath.Join(volume, reports),
owner)
if err != nil {
t.Fatal(err)
}
uid, _ := strconv.ParseUint(owner.Uid, 10, 32)
gid, _ := strconv.ParseUint(owner.Gid, 10, 32)
requireOwner(t, latin1, uint32(uid), uint32(gid))
requireOwner(t, filepath.Join(latin1, "f"), uint32(uid), uint32(gid))
}
// TestPrepareDataDirGivesVolumeToOwner gives everything in the volume
// to a uid and gid that own nothing, which only root can do. A
// symbolic link in the volume to a directory outside it is given to
// them itself; what it points to is left as it was.
func TestPrepareDataDirGivesVolumeToOwner(t *testing.T) {
t.Parallel()
if os.Geteuid() != 0 {
t.Skip("only root can give files to another uid")
}
outside := t.TempDir()
volume := t.TempDir()
old := filepath.Join(volume, "old")
err := os.WriteFile(filepath.Join(outside, "f"), nil, 0o600)
if err != nil {
t.Fatal(err)
}
err = os.Mkdir(old, 0o700)
if err != nil {
t.Fatal(err)
}
err = os.WriteFile(filepath.Join(old, "f"), nil, 0o600)
if err != nil {
t.Fatal(err)
}
err = os.Symlink(outside, filepath.Join(old, "link"))
if err != nil {
t.Fatal(err)
}
err = reportbuf.PrepareDataDir(volume, filepath.Join(volume, reports),
&user.User{Uid: "4242", Gid: "4343"})
if err != nil {
t.Fatal(err)
}
for _, path := range []string{
volume, filepath.Join(volume, reports), old,
filepath.Join(old, "f"), filepath.Join(old, "link"),
} {
requireOwner(t, path, 4242, 4343)
}
requireOwner(t, outside, 0, 0)
requireOwner(t, filepath.Join(outside, "f"), 0, 0)
}
+11 -1
View File
@@ -1,6 +1,9 @@
package reportbuf
import "time"
import (
"os"
"time"
)
// FlushSizeThreshold exposes the buffer size at which Append starts
// writing a report file to the external tests.
@@ -17,3 +20,10 @@ func (b *Buffer) Flush() error {
func (b *Buffer) StopClock(at time.Time) {
b.now = func() time.Time { return at }
}
// OnFileCreated makes the buffer call fn with each report file it
// writes from now on, once the file is created and before anything is
// written to it.
func (b *Buffer) OnFileCreated(fn func(f *os.File)) {
b.fileCreated = fn
}
+155 -45
View File
@@ -1,5 +1,6 @@
// Package reportbuf accumulates telemetry reports in memory
// and periodically flushes them to zstd-compressed JSONL files.
// and periodically flushes them to zstd-compressed JSONL files,
// deleting the oldest files to keep them under a size cap.
package reportbuf
import (
@@ -12,6 +13,7 @@ import (
"log/slog"
"os"
"path/filepath"
"slices"
"strings"
"sync"
"sync/atomic"
@@ -37,9 +39,10 @@ const (
fileSuffix = ".jsonl.zst"
)
// ErrFull is returned by Append when storing the report would
// take the report files past the configured maximum size.
var ErrFull = errors.New("report files at their size cap")
// ErrFull is returned by Append when the reports waiting to be
// written leave no room for the report under the configured maximum
// size, however many report files are deleted.
var ErrFull = errors.New("reports waiting to be written fill the size cap")
// Params defines the dependencies for Buffer.
type Params struct {
@@ -49,12 +52,31 @@ type Params struct {
Logger *logger.Logger
}
// reportFile is a report file that may be deleted to make room, with
// the size it counts for in usedBytes.
type reportFile struct {
name string
size int64
}
// Buffer accumulates JSON lines in memory and flushes them
// to zstd-compressed files on disk.
type Buffer struct {
buf bytes.Buffer
dataDir string
done chan struct{}
// fileCreated is called with each report file once it is
// created, before anything is written to it: it does nothing,
// except in tests that hold the write open or make it fail.
fileCreated func(f *os.File)
// files are the report files that may be deleted to make room,
// in name order, which is oldest first: those in dataDir at
// start, and each one this buffer writes, put in at its place by
// name once it is complete, even when an older file's write
// completes after a newer one's. A file still being written is
// not among them. filesBytes is their total size.
files []reportFile
filesBytes int64
log *slog.Logger
maxBytes int64
mu sync.Mutex
@@ -66,8 +88,8 @@ type Buffer struct {
seq atomic.Uint64
stopOnce sync.Once
// usedBytes is what Append checks against maxBytes: the size
// of the report files in dataDir, plus the reports not yet
// written to one at their uncompressed size.
// of the report files in dataDir, plus the reports waiting to
// be written to one at their uncompressed size.
usedBytes int64
}
@@ -85,6 +107,7 @@ func New(
b := &Buffer{
dataDir: dir,
done: make(chan struct{}),
fileCreated: func(*os.File) {},
log: params.Logger.Get(),
maxBytes: params.Config.DataDirMaxBytes,
now: time.Now,
@@ -97,12 +120,28 @@ func New(
return fmt.Errorf("create data dir: %w", err)
}
// Report files left by earlier runs count too.
b.usedBytes, err = reportFilesSize(b.dataDir)
// Report files left by earlier runs count too, and are
// the first to be deleted to make room.
files, err := reportFiles(b.dataDir)
if err != nil {
return err
}
b.mu.Lock()
b.files = files
for _, f := range files {
b.filesBytes += f.size
}
b.usedBytes = b.filesBytes
// The files may be past the cap, if it was lowered since
// the last run.
b.deleteOldestFiles(0)
b.mu.Unlock()
go b.flushLoop()
return nil
@@ -128,9 +167,10 @@ func New(
}
// Append marshals v as a single JSON line and appends it to
// the buffer. It stores nothing and returns ErrFull if the line
// would take usedBytes past maxBytes. If the buffer reaches the
// size threshold, it is drained and written to disk
// the buffer. If the line would take usedBytes past maxBytes, the
// oldest report files are deleted to make room; it stores nothing
// and returns ErrFull if that cannot make room. If the buffer
// reaches the size threshold, it is drained and written to disk
// asynchronously.
func (b *Buffer) Append(v any) error {
line, err := json.Marshal(v)
@@ -142,6 +182,8 @@ func (b *Buffer) Append(v any) error {
b.mu.Lock()
b.deleteOldestFiles(lineBytes)
if b.usedBytes+lineBytes > b.maxBytes {
b.mu.Unlock()
@@ -171,6 +213,37 @@ func (b *Buffer) Append(v any) error {
return nil
}
// deleteOldestFiles deletes report files, oldest first, until n more
// bytes fit under maxBytes. It deletes none when the reports waiting
// to be written leave no room for n even with every file gone, since
// that would lose the files for nothing. The caller must hold b.mu.
func (b *Buffer) deleteOldestFiles(n int64) {
for b.usedBytes+n > b.maxBytes && len(b.files) > 0 {
if b.usedBytes-b.filesBytes+n > b.maxBytes {
return
}
f := b.files[0]
b.files = b.files[1:]
b.filesBytes -= f.size
// A file already gone, deleted by hand, has freed its room too.
err := os.Remove(filepath.Join(b.dataDir, f.name))
if err != nil && !errors.Is(err, fs.ErrNotExist) {
// The file is still there, so it still counts. It is
// not tried again until the next start.
b.log.Error("delete report file failed",
"file", f.name, "error", err)
continue
}
b.usedBytes -= f.size
b.log.Info("deleted report file to make room",
"file", f.name, "bytes", f.size)
}
}
// flushLoop runs a ticker that periodically flushes buffered
// data to disk until the done channel is closed.
func (b *Buffer) flushLoop() {
@@ -217,15 +290,48 @@ func (b *Buffer) drainBuf() []byte {
return data
}
// writeFile creates a timestamped zstd-compressed JSONL file
// in the data directory.
// writeFile writes data, reports drained from the buffer, to a new
// timestamped zstd-compressed JSONL file in the data directory.
func (b *Buffer) writeFile(data []byte) error {
// The timestamp comes first, so the names sort by time; the number
// after it tells apart files named in the same millisecond.
ts := b.now().UTC().Format("2006-01-02T15-04-05.000Z")
name := fmt.Sprintf("%s%s-%d%s", filePrefix, ts, b.seq.Add(1), fileSuffix)
path := filepath.Join(b.dataDir, name)
size, err := b.createFile(filepath.Join(b.dataDir, name), data)
// The reports no longer wait to be written, so they stop counting
// at their uncompressed size. If the write failed they are lost;
// otherwise they count as the file, which from here on may be
// deleted to make room.
b.mu.Lock()
defer b.mu.Unlock()
b.usedBytes -= int64(len(data))
if err != nil {
return err
}
b.usedBytes += size
b.filesBytes += size
// At its place by name, not at the end: another write, of a newer
// file, may have completed while this one was being written.
i, _ := slices.BinarySearchFunc(b.files, name,
func(f reportFile, target string) int {
return strings.Compare(f.name, target)
})
b.files = slices.Insert(b.files, i, reportFile{name: name, size: size})
return nil
}
// createFile creates the file at path holding data compressed with
// zstd, and returns its size. If the write fails once the file is
// created, it removes the file, so that a failed write leaves nothing
// behind to take room.
func (b *Buffer) createFile(path string, data []byte) (int64, error) {
// path is built from the operator-supplied dataDir plus a
// generated timestamp and number, so it carries no external input.
f, err := os.OpenFile( //nolint:gosec // see comment above
@@ -234,61 +340,65 @@ func (b *Buffer) writeFile(data []byte) error {
filePerms,
)
if err != nil {
return fmt.Errorf("create report file: %w", err)
return 0, fmt.Errorf("create report file: %w", err)
}
// Closes the file on the early returns below. The success
// path closes it explicitly to check the error; closing it
// a second time here is harmless.
defer func() { _ = f.Close() }()
b.fileCreated(f)
size, err := writeCompressed(f, data)
if err != nil {
_ = f.Close()
return 0, errors.Join(err, os.Remove(path))
}
err = f.Close()
if err != nil {
err = fmt.Errorf("close report file: %w", err)
return 0, errors.Join(err, os.Remove(path))
}
return size, nil
}
// writeCompressed writes data to f compressed with zstd, and returns
// the size of f.
func writeCompressed(f *os.File, data []byte) (int64, error) {
enc, err := zstd.NewWriter(f)
if err != nil {
return fmt.Errorf("create zstd encoder: %w", err)
return 0, fmt.Errorf("create zstd encoder: %w", err)
}
_, err = enc.Write(data)
if err != nil {
_ = enc.Close()
return fmt.Errorf("write compressed data: %w", err)
return 0, fmt.Errorf("write compressed data: %w", err)
}
err = enc.Close()
if err != nil {
return fmt.Errorf("close zstd encoder: %w", err)
return 0, fmt.Errorf("close zstd encoder: %w", err)
}
info, err := f.Stat()
if err != nil {
return fmt.Errorf("stat report file: %w", err)
return 0, fmt.Errorf("stat report file: %w", err)
}
err = f.Close()
if err != nil {
return fmt.Errorf("close report file: %w", err)
return info.Size(), nil
}
// The reports counted at their uncompressed size while they
// waited; now they count as the file. After a failed write they
// stay counted as they were, which errs toward refusing reports
// early rather than letting the files pass the cap.
b.mu.Lock()
b.usedBytes += info.Size() - int64(len(data))
b.mu.Unlock()
return nil
}
// reportFilesSize returns the total size of the report files in
// dir.
func reportFilesSize(dir string) (int64, error) {
// reportFiles returns the report files in dir, oldest first:
// os.ReadDir sorts them by name, and the names sort by time.
func reportFiles(dir string) ([]reportFile, error) {
entries, err := os.ReadDir(dir)
if err != nil {
return 0, fmt.Errorf("read data dir: %w", err)
return nil, fmt.Errorf("read data dir: %w", err)
}
var total int64
files := make([]reportFile, 0, len(entries))
for _, entry := range entries {
name := entry.Name()
@@ -299,11 +409,11 @@ func reportFilesSize(dir string) (int64, error) {
info, err := entry.Info()
if err != nil {
return 0, fmt.Errorf("stat report file: %w", err)
return nil, fmt.Errorf("stat report file: %w", err)
}
total += info.Size()
files = append(files, reportFile{name: name, size: info.Size()})
}
return total, nil
return files, nil
}
+443 -50
View File
@@ -139,41 +139,19 @@ func lineBytes(t *testing.T, report any) int {
return len(line) + 1
}
// TestAppendPastCapIsRefused fills the cap with a report not yet
// written. The next report is refused, and the report file already in
// DATA_DIR is kept: it is smaller than a report, so deleting it could
// not make room.
func TestAppendPastCapIsRefused(t *testing.T) {
report := map[string]string{"id": "cap"}
t.Setenv("DATA_DIR", t.TempDir())
t.Setenv("DATA_DIR_MAX_BYTES", strconv.Itoa(lineBytes(t, report)))
buf := startBuffer(t)
err := buf.Append(report)
if err != nil {
t.Fatalf("report that fills the cap exactly: %v", err)
}
err = buf.Append(report)
if !errors.Is(err, reportbuf.ErrFull) {
t.Fatalf("report past the cap: error = %v, want ErrFull", err)
}
}
// TestCapCountsReportFilesAlreadyInDataDir starts on a data
// directory holding a report file from an earlier run, and a file
// that is not a report, which must not count.
func TestCapCountsReportFilesAlreadyInDataDir(t *testing.T) {
const earlierBytes = 100
report := map[string]string{"id": "cap"}
dir := t.TempDir()
earlier := reportFilePath(dir, 1)
writeBytes(t, filepath.Join(dir, "reports-2026-01-01T00-00-00.000Z.jsonl.zst"),
earlierBytes)
writeBytes(t, filepath.Join(dir, "notes.txt"), 10*earlierBytes)
writeBytes(t, earlier, 1)
t.Setenv("DATA_DIR", dir)
t.Setenv("DATA_DIR_MAX_BYTES",
strconv.Itoa(earlierBytes+lineBytes(t, report)))
t.Setenv("DATA_DIR_MAX_BYTES", strconv.Itoa(1+lineBytes(t, report)))
buf := startBuffer(t)
@@ -186,6 +164,96 @@ func TestCapCountsReportFilesAlreadyInDataDir(t *testing.T) {
if !errors.Is(err, reportbuf.ErrFull) {
t.Fatalf("report past the cap: error = %v, want ErrFull", err)
}
if !exists(t, earlier) {
t.Fatal("report file deleted, though that could not make room")
}
}
// TestOldestReportFileDeletedFirst starts on a data directory holding
// report files from an earlier run, and a file that is not a report,
// which neither counts nor is ever deleted. Nothing is deleted while
// there is room; then only the oldest report file is.
func TestOldestReportFileDeletedFirst(t *testing.T) {
const fileBytes = 100
report := map[string]string{"id": "oldest"}
dir := t.TempDir()
oldest := reportFilePath(dir, 1)
kept := []string{
reportFilePath(dir, 2),
reportFilePath(dir, 3),
filepath.Join(dir, "notes.txt"),
}
writeBytes(t, oldest, fileBytes)
for _, path := range kept {
writeBytes(t, path, fileBytes)
}
t.Setenv("DATA_DIR", dir)
// Room for the three report files and one report.
t.Setenv("DATA_DIR_MAX_BYTES",
strconv.Itoa(3*fileBytes+lineBytes(t, report)))
buf := startBuffer(t)
err := buf.Append(report)
if err != nil {
t.Fatalf("report that fills the cap exactly: %v", err)
}
if !exists(t, oldest) {
t.Fatal("oldest report file deleted while there was room")
}
err = buf.Append(report)
if err != nil {
t.Fatalf("report past the cap: %v", err)
}
if exists(t, oldest) {
t.Fatal("oldest report file kept when room was needed")
}
for _, path := range kept {
if !exists(t, path) {
t.Fatalf("%s deleted; only the oldest report file should be", path)
}
}
}
// TestStartDeletesFilesPastCap starts on report files past the cap,
// as after the cap is lowered: the oldest are deleted until the rest
// fit.
func TestStartDeletesFilesPastCap(t *testing.T) {
const fileBytes = 100
dir := t.TempDir()
oldest := reportFilePath(dir, 1)
kept := []string{reportFilePath(dir, 2), reportFilePath(dir, 3)}
writeBytes(t, oldest, fileBytes)
for _, path := range kept {
writeBytes(t, path, fileBytes)
}
t.Setenv("DATA_DIR", dir)
t.Setenv("DATA_DIR_MAX_BYTES", strconv.Itoa(len(kept)*fileBytes))
startBuffer(t)
if exists(t, oldest) {
t.Fatal("oldest report file kept, though the files were past the cap")
}
for _, path := range kept {
if !exists(t, path) {
t.Fatalf("%s deleted, though the rest fit without it", path)
}
}
}
// TestWrittenReportsCountAtFileSize checks that once reports are
@@ -195,8 +263,9 @@ func TestWrittenReportsCountAtFileSize(t *testing.T) {
// Repetitive, so its file is far smaller than its JSON.
report := map[string]string{"id": strings.Repeat("a", 1000)}
size := lineBytes(t, report)
dir := t.TempDir()
t.Setenv("DATA_DIR", t.TempDir())
t.Setenv("DATA_DIR", dir)
// Room for the report twice over only if the first one counts
// at its file's size by the time the second arrives.
t.Setenv("DATA_DIR_MAX_BYTES", strconv.Itoa(2*size-1))
@@ -217,13 +286,22 @@ func TestWrittenReportsCountAtFileSize(t *testing.T) {
if err != nil {
t.Fatalf("second report, after the first was written: %v", err)
}
if !hasReportFile(t, dir) {
t.Fatal("first report's file deleted, though the second fit beside it")
}
}
// TestWrittenReportsKeepCounting writes one report file after another
// under a small cap: each report must be taken while the files on disk
// leave room for it, and refused once they do not.
func TestWrittenReportsKeepCounting(t *testing.T) {
const maxBytes = 200
// TestWrittenFilesDeletedToMakeRoom writes one report file after
// another under a small cap. Every report must be taken; files are
// deleted only when the report would not fit beside them, and the
// files kept leave room for it.
func TestWrittenFilesDeletedToMakeRoom(t *testing.T) {
const (
maxBytes = 200
// One file each, which take far more than maxBytes together.
reports = 50
)
report := map[string]string{"id": "written"}
size := int64(lineBytes(t, report))
@@ -234,23 +312,24 @@ func TestWrittenReportsKeepCounting(t *testing.T) {
buf := startBuffer(t)
// Every file takes at least a byte, so they fill the cap within
// maxBytes rounds.
for range maxBytes {
used := reportFilesBytes(t, dir)
for range reports {
before := reportFilesBytes(t, dir)
err := buf.Append(report)
if used+size > maxBytes {
if !errors.Is(err, reportbuf.ErrFull) {
t.Fatalf("with %d bytes of report files: error = %v, "+
"want ErrFull", used, err)
}
return
}
if err != nil {
t.Fatalf("with %d bytes of report files: %v", used, err)
t.Fatalf("with %d bytes of report files: %v", before, err)
}
after := reportFilesBytes(t, dir)
if before+size <= maxBytes && after != before {
t.Fatalf("files deleted, though the report fit beside "+
"their %d bytes", before)
}
if after+size > maxBytes {
t.Fatalf("%d bytes of report files kept, leaving no room "+
"for the report", after)
}
err = buf.Flush()
@@ -258,8 +337,217 @@ func TestWrittenReportsKeepCounting(t *testing.T) {
t.Fatalf("flush: %v", err)
}
}
}
t.Fatal("the report files never filled the cap")
// TestFailedDeletionStillCounts makes deleting the oldest report file
// fail. It is still there, so it still takes room, and the next oldest
// is deleted in its place.
func TestFailedDeletionStillCounts(t *testing.T) {
const fileBytes = 100
report := map[string]string{"id": "stuck"}
dir := t.TempDir()
stuck := reportFilePath(dir, 1)
next := reportFilePath(dir, 2)
newest := reportFilePath(dir, 3)
for _, path := range []string{stuck, next, newest} {
writeBytes(t, path, fileBytes)
}
t.Setenv("DATA_DIR", dir)
t.Setenv("DATA_DIR_MAX_BYTES", strconv.Itoa(3*fileBytes))
buf := startBuffer(t)
// A directory that is not empty cannot be deleted, even by root,
// which the tests run as in the backend image.
err := os.Remove(stuck)
if err != nil {
t.Fatalf("remove %s: %v", stuck, err)
}
err = os.Mkdir(stuck, 0o750)
if err != nil {
t.Fatalf("make directory %s: %v", stuck, err)
}
writeBytes(t, filepath.Join(stuck, "file"), 1)
err = buf.Append(report)
if err != nil {
t.Fatalf("report past the cap: %v", err)
}
if exists(t, next) {
t.Fatal("next oldest report file kept: the failed deletion " +
"counted as making room")
}
if !exists(t, newest) {
t.Fatal("newest report file deleted, though deleting one made room")
}
}
// TestFileDeletedByHandFreesRoom deletes the oldest report file by
// hand after start. When room is needed, its room counts as freed, so
// no other file is deleted.
func TestFileDeletedByHandFreesRoom(t *testing.T) {
const fileBytes = 100
report := map[string]string{"id": "by-hand"}
dir := t.TempDir()
gone := reportFilePath(dir, 1)
kept := reportFilePath(dir, 2)
writeBytes(t, gone, fileBytes)
writeBytes(t, kept, fileBytes)
t.Setenv("DATA_DIR", dir)
t.Setenv("DATA_DIR_MAX_BYTES", strconv.Itoa(2*fileBytes))
buf := startBuffer(t)
err := os.Remove(gone)
if err != nil {
t.Fatalf("remove %s: %v", gone, err)
}
err = buf.Append(report)
if err != nil {
t.Fatalf("report past the cap: %v", err)
}
if !exists(t, kept) {
t.Fatal("report file deleted, though the one deleted by hand " +
"had made room")
}
}
// TestFileBeingWrittenIsNeverDeleted holds the write of one report file
// open while a second write completes, then sends a report that needs
// room. Deleting either file would make it, and the one being written is
// the older, but only the complete one may be deleted. Once the first
// write is complete, its file is deleted when room is needed.
func TestFileBeingWrittenIsNeverDeleted(t *testing.T) {
report := map[string]string{"id": "writing"}
t.Setenv("DATA_DIR", t.TempDir())
// Room for two reports waiting to be written, but not for two
// beside a report file.
t.Setenv("DATA_DIR_MAX_BYTES", strconv.Itoa(2*lineBytes(t, report)))
buf := startBuffer(t)
created := make(chan string)
release := make(chan struct{})
buf.OnFileCreated(func(f *os.File) {
created <- f.Name()
<-release
})
err := buf.Append(report)
if err != nil {
t.Fatalf("first report: %v", err)
}
flushed := make(chan error)
go func() { flushed <- buf.Flush() }()
writing := <-created
// Only the first write is held; the second goes through, and
// so do the writes after it, the final one at stop included.
var complete string
buf.OnFileCreated(func(f *os.File) { complete = f.Name() })
// Errorf, not Fatalf, until the first write is released, so that a
// failure here does not leave it held.
err = buf.Append(report)
if err != nil {
t.Errorf("second report: %v", err)
}
err = buf.Flush()
if err != nil {
t.Errorf("flush of the second report: %v", err)
}
err = buf.Append(report)
if err != nil {
t.Errorf("report that needs room: %v", err)
}
if !exists(t, writing) {
t.Error("report file deleted while it was being written")
}
if exists(t, complete) {
t.Error("complete report file kept, though room was needed")
}
close(release)
err = <-flushed
if err != nil {
t.Fatalf("flush of the first report: %v", err)
}
err = buf.Append(report)
if err != nil {
t.Fatalf("report after the first write was complete: %v", err)
}
if exists(t, writing) {
t.Fatal("complete report file kept when room was needed")
}
}
// TestFailedWriteStopsCounting makes a write fail once its file is
// created. Its reports are lost, so they stop counting, and the part of
// the file written is removed, so it takes no room.
func TestFailedWriteStopsCounting(t *testing.T) {
report := map[string]string{"id": "failed"}
t.Setenv("DATA_DIR", t.TempDir())
t.Setenv("DATA_DIR_MAX_BYTES", strconv.Itoa(lineBytes(t, report)))
buf := startBuffer(t)
var failed string
// Closing the file under the write makes the write fail.
buf.OnFileCreated(func(f *os.File) {
failed = f.Name()
_ = f.Close()
})
err := buf.Append(report)
if err != nil {
t.Fatalf("report that fills the cap exactly: %v", err)
}
err = buf.Flush()
if err == nil {
t.Fatal("flush succeeded, though its file was closed under it")
}
if exists(t, failed) {
t.Fatal("file of the failed write kept")
}
// Writes from here on, the final one at stop included, succeed.
buf.OnFileCreated(func(*os.File) {})
err = buf.Append(report)
if err != nil {
t.Fatalf("report after the failed write: %v", err)
}
}
// TestConcurrentAppendsStopAtCap appends from many goroutines at once
@@ -492,6 +780,14 @@ func readReportFiles(dir string) ([]string, error) {
return contents, nil
}
// reportFilePath returns the path in dir of a report file named as
// written on the given day of January 2026, so that a lower day sorts
// as older.
func reportFilePath(dir string, day int) string {
return filepath.Join(dir,
fmt.Sprintf("reports-2026-01-%02dT00-00-00.000Z-1.jsonl.zst", day))
}
func writeBytes(t *testing.T, path string, n int) {
t.Helper()
@@ -501,6 +797,21 @@ func writeBytes(t *testing.T, path string, n int) {
}
}
func exists(t *testing.T, path string) bool {
t.Helper()
_, err := os.Stat(path)
if errors.Is(err, fs.ErrNotExist) {
return false
}
if err != nil {
t.Fatalf("stat %s: %v", path, err)
}
return true
}
func hasReportFile(t *testing.T, dir string) bool {
t.Helper()
@@ -524,3 +835,85 @@ func hasReportFile(t *testing.T, dir string) bool {
return false
}
// TestFilesDeletedOldestFirstWhenWritesOverlap holds the write of an
// older report file open until a newer one's write completes, then
// releases it. When room is needed, the older file is deleted first,
// though its write was the last to complete.
func TestFilesDeletedOldestFirstWhenWritesOverlap(t *testing.T) {
const maxBytes = 1000
report := map[string]string{"id": "overlap"}
t.Setenv("DATA_DIR", t.TempDir())
t.Setenv("DATA_DIR_MAX_BYTES", strconv.Itoa(maxBytes))
buf := startBuffer(t)
created := make(chan string)
release := make(chan struct{})
buf.OnFileCreated(func(f *os.File) {
created <- f.Name()
<-release
})
err := buf.Append(report)
if err != nil {
t.Fatalf("older report: %v", err)
}
flushed := make(chan error)
go func() { flushed <- buf.Flush() }()
older := <-created
// Only the older write is held; the newer one goes through, and
// so do the writes after it, the final one at stop included.
var newer string
buf.OnFileCreated(func(f *os.File) { newer = f.Name() })
// Errorf, not Fatalf, until the older write is released, so that a
// failure here does not leave it held.
err = buf.Append(report)
if err != nil {
t.Errorf("newer report: %v", err)
}
err = buf.Flush()
if err != nil {
t.Errorf("flush of the newer report: %v", err)
}
close(release)
err = <-flushed
if err != nil {
t.Fatalf("flush of the older report: %v", err)
}
info, err := os.Stat(newer)
if err != nil {
t.Fatalf("stat %s: %v", newer, err)
}
// A report that fits beside the newer file alone, so deleting the
// older one makes exactly the room it needs.
pad := maxBytes - int(info.Size()) - lineBytes(t, map[string]string{"id": ""})
err = buf.Append(map[string]string{"id": strings.Repeat("a", pad)})
if err != nil {
t.Fatalf("report that needs room: %v", err)
}
if exists(t, older) {
t.Fatal("older report file kept when room was needed")
}
if !exists(t, newer) {
t.Fatal("newer report file deleted before the older one")
}
}
+12
View File
@@ -1,9 +1,21 @@
package server
import "github.com/go-chi/chi/v5"
// Router exposes the router to the external tests, which add routes
// of their own to it after SetupRoutes.
func (s *Server) Router() *chi.Mux {
return s.router
}
// MaxRequestBodyBytes exposes the router-wide body limit to the
// external tests.
const MaxRequestBodyBytes = maxRequestBodyBytes
// MetricsRequestsPerMinute exposes the /metrics rate limit to the
// external tests.
const MetricsRequestsPerMinute = metricsRequestsPerMinute
// ListenAddr exposes the address the server listens on to the
// external tests.
func (s *Server) ListenAddr() string {
+49 -5
View File
@@ -3,8 +3,12 @@ package server
import (
"time"
sentryhttp "github.com/getsentry/sentry-go/http"
"github.com/go-chi/chi/v5"
"github.com/go-chi/chi/v5/middleware"
"github.com/prometheus/client_golang/prometheus"
"github.com/prometheus/client_golang/prometheus/collectors"
"github.com/prometheus/client_golang/prometheus/promhttp"
)
const (
@@ -14,6 +18,12 @@ const (
// can mount s.mw.MaxBodyBytes with a smaller value to lower
// its bound, but cannot raise it: this cap runs first.
maxRequestBodyBytes int64 = 1 << 20 // 1 MiB
// metricsRequestsPerMinute is how many requests to /metrics each
// client address may make a minute, whatever their credentials. A
// scraper polling every 2 seconds sends half of it, which httprate
// never refuses.
metricsRequestsPerMinute = 60
)
// SetupRoutes configures the chi router with middleware and
@@ -29,13 +39,47 @@ func (s *Server) SetupRoutes() {
s.router.Use(s.mw.MaxBodyBytes(maxRequestBodyBytes))
s.router.Use(middleware.Timeout(requestTimeout))
s.router.Get(
"/.well-known/healthcheck",
s.h.HandleHealthCheck(),
// Sentry reports a panic, then panics again, so that s.mw.Recoverer
// still answers 500.
if s.params.Config.SentryDSN != "" {
s.router.Use(sentryhttp.New(sentryhttp.Options{Repanic: true}).Handle)
}
// The metrics go in a registry of this server's own, not in
// Prometheus' default one, which takes them only once per process.
registry := prometheus.NewRegistry()
registry.MustRegister(
collectors.NewGoCollector(),
collectors.NewProcessCollector(collectors.ProcessCollectorOpts{}),
)
s.router.Route("/api/v1", func(r chi.Router) {
// Requests are measured only once chi has matched them to one of
// these routes, by path and method. The metrics are labelled with
// both, which any client can make up, so measuring every request
// would let clients add labels without bound. A Route here would
// be matched by its path prefix alone, so each path is given in
// full.
s.router.Group(func(r chi.Router) {
// config.New refuses one of the two credentials without the
// other.
if s.params.Config.MetricsUsername != "" {
r.Use(s.mw.Metrics(registry))
}
r.Get("/.well-known/healthcheck", s.h.HandleHealthCheck())
r.With(s.mw.RateLimit(s.params.Config.ReportsPerMinute)).
Post("/reports", s.h.HandleReport())
Post("/api/v1/reports", s.h.HandleReport())
})
// The rate limit comes before the basic auth, so a client past it
// gets 429 and its password is not checked.
if s.params.Config.MetricsUsername != "" {
s.router.With(
s.mw.RateLimit(metricsRequestsPerMinute),
s.mw.MetricsAuth(),
).Get("/metrics", promhttp.HandlerFor(
registry, promhttp.HandlerOpts{},
).ServeHTTP)
}
}
+200
View File
@@ -1,10 +1,12 @@
package server_test
import (
"io"
"net/http"
"net/http/httptest"
"strings"
"testing"
"time"
"sneak.berlin/go/netwatch/internal/config"
"sneak.berlin/go/netwatch/internal/globals"
@@ -15,6 +17,7 @@ import (
"sneak.berlin/go/netwatch/internal/reportbuf"
"sneak.berlin/go/netwatch/internal/server"
"github.com/getsentry/sentry-go"
"go.uber.org/fx"
"go.uber.org/fx/fxtest"
)
@@ -107,6 +110,203 @@ func TestCORSAllowedOriginsReachTheRouter(t *testing.T) {
}
}
// TestNoMetricsWithoutCredentials: with neither metrics setting set,
// there is no /metrics.
func TestNoMetricsWithoutCredentials(t *testing.T) {
t.Setenv("METRICS_USERNAME", "")
t.Setenv("METRICS_PASSWORD", "")
srv := newServer(t)
srv.SetupRoutes()
rec := httptest.NewRecorder()
req := httptest.NewRequestWithContext(t.Context(),
http.MethodGet, "/metrics", http.NoBody)
srv.ServeHTTP(rec, req)
if rec.Code != http.StatusNotFound {
t.Fatalf("status = %d, want %d", rec.Code, http.StatusNotFound)
}
}
// TestMetricsBehindBasicAuth: with both metrics settings set, /metrics
// answers only with them as basic auth credentials, and shows a
// request to a route but not one to a path no route has.
func TestMetricsBehindBasicAuth(t *testing.T) {
t.Setenv("METRICS_USERNAME", "prometheus")
t.Setenv("METRICS_PASSWORD", "right")
srv := newServer(t)
srv.SetupRoutes()
get := func(path, username, password string) *httptest.ResponseRecorder {
rec := httptest.NewRecorder()
req := httptest.NewRequestWithContext(t.Context(),
http.MethodGet, path, http.NoBody)
if username != "" {
req.SetBasicAuth(username, password)
}
srv.ServeHTTP(rec, req)
return rec
}
get("/.well-known/healthcheck", "", "")
get("/api/v1/no-such-route", "", "")
for _, creds := range [][2]string{
{"", ""},
{"prometheus", "wrong"},
{"someone", "right"},
} {
rec := get("/metrics", creds[0], creds[1])
if rec.Code != http.StatusUnauthorized {
t.Errorf("credentials %q: status = %d, want %d",
creds, rec.Code, http.StatusUnauthorized)
}
}
rec := get("/metrics", "prometheus", "right")
if rec.Code != http.StatusOK {
t.Fatalf("right credentials: status = %d, want %d",
rec.Code, http.StatusOK)
}
body := rec.Body.String()
if !strings.Contains(body, `handler="/.well-known/healthcheck"`) {
t.Errorf("metrics show no health check request:\n%s", body)
}
if strings.Contains(body, "no-such-route") {
t.Errorf("metrics show a request to a path no route has:\n%s", body)
}
if !strings.Contains(body, "go_goroutines") {
t.Errorf("metrics show no Go runtime metrics:\n%s", body)
}
}
// TestMetricsAreRateLimited: a client that has used up its /metrics
// allowance on wrong passwords gets 429 even with the right one, which
// is then not checked, while another client behind the same nginx
// still gets in.
func TestMetricsAreRateLimited(t *testing.T) {
t.Setenv("METRICS_USERNAME", "prometheus")
t.Setenv("METRICS_PASSWORD", "right")
// As in the container: nginx connects from loopback and names the
// client in X-Forwarded-For.
t.Setenv("TRUSTED_PROXIES", "127.0.0.1/32")
srv := newServer(t)
srv.SetupRoutes()
get := func(client, password string) int {
rec := httptest.NewRecorder()
req := httptest.NewRequestWithContext(t.Context(),
http.MethodGet, "/metrics", http.NoBody)
req.RemoteAddr = "127.0.0.1:40000"
req.Header.Set("X-Forwarded-For", client)
req.SetBasicAuth("prometheus", password)
srv.ServeHTTP(rec, req)
return rec.Code
}
for i := range server.MetricsRequestsPerMinute {
if code := get("203.0.113.7", "wrong"); code != http.StatusUnauthorized {
t.Fatalf("guess %d: status = %d, want %d",
i+1, code, http.StatusUnauthorized)
}
}
if code := get("203.0.113.7", "right"); code != http.StatusTooManyRequests {
t.Fatalf("right password past the limit: status = %d, want %d",
code, http.StatusTooManyRequests)
}
if code := get("203.0.113.8", "right"); code != http.StatusOK {
t.Fatalf("another client: status = %d, want %d",
code, http.StatusOK)
}
}
// TestMetricsInTwoServers: two servers in one process can both have
// metrics on.
func TestMetricsInTwoServers(t *testing.T) {
t.Setenv("METRICS_USERNAME", "prometheus")
t.Setenv("METRICS_PASSWORD", "right")
for range 2 {
newServer(t).SetupRoutes()
}
}
// TestSentry: with SENTRY_DSN empty there is no Sentry client. With it
// pointing at a local server standing in for Sentry, a panic in a
// handler reaches that server, and the request still gets the 500 from
// the panic recovery.
func TestSentry(t *testing.T) {
const panicMessage = "handler panic for TestSentry"
// sentry.Init sets the client for the whole process; take it away
// again so that no other test reports to Sentry.
t.Cleanup(func() { sentry.CurrentHub().BindClient(nil) })
t.Setenv("SENTRY_DSN", "")
newServer(t)
if sentry.CurrentHub().Client() != nil {
t.Fatal("a Sentry client exists with SENTRY_DSN empty")
}
// The body of the first request the stand-in for Sentry receives.
received := make(chan string, 1)
sentryServer := httptest.NewServer(http.HandlerFunc(
func(_ http.ResponseWriter, r *http.Request) {
body, _ := io.ReadAll(r.Body)
select {
case received <- string(body):
default:
}
},
))
defer sentryServer.Close()
t.Setenv("SENTRY_DSN",
"http://key@"+sentryServer.Listener.Addr().String()+"/1")
srv := newServer(t)
srv.SetupRoutes()
srv.Router().Get("/panic", func(http.ResponseWriter, *http.Request) {
panic(panicMessage)
})
rec := httptest.NewRecorder()
req := httptest.NewRequestWithContext(t.Context(),
http.MethodGet, "/panic", http.NoBody)
srv.ServeHTTP(rec, req)
if rec.Code != http.StatusInternalServerError {
t.Fatalf("status = %d, want %d",
rec.Code, http.StatusInternalServerError)
}
// Sentry sends from a goroutine of its own.
select {
case body := <-received:
if !strings.Contains(body, panicMessage) {
t.Fatalf("the Sentry server received no report of the panic:\n%s",
body)
}
case <-time.After(5 * time.Second):
t.Fatal("nothing reached the Sentry server")
}
}
// TestHealthCheckRejectsOversizeBody sends the health check, which
// never reads its body, a body one byte over the limit. Only the
// router-wide body limit can reject it.
+40 -1
View File
@@ -7,8 +7,10 @@ package server
import (
"context"
"fmt"
"log/slog"
"net/http"
"time"
"sneak.berlin/go/netwatch/internal/config"
"sneak.berlin/go/netwatch/internal/globals"
@@ -16,10 +18,15 @@ import (
"sneak.berlin/go/netwatch/internal/logger"
"sneak.berlin/go/netwatch/internal/middleware"
"github.com/getsentry/sentry-go"
"github.com/go-chi/chi/v5"
"go.uber.org/fx"
)
// sentryFlushTimeout is how long shutdown waits for Sentry to send
// what it still holds.
const sentryFlushTimeout = 2 * time.Second
// Params defines the dependencies for Server.
type Params struct {
fx.In
@@ -56,6 +63,11 @@ func New(
s.log = params.Logger.Get()
s.shutdowner = params.Shutdowner
err := s.enableSentry()
if err != nil {
return nil, err
}
lc.Append(fx.Hook{
OnStart: func(_ context.Context) error {
// Build the router and http.Server synchronously
@@ -87,10 +99,37 @@ func (s *Server) ServeHTTP(
s.router.ServeHTTP(w, r)
}
// enableSentry sets Sentry up when SENTRY_DSN is set, so that
// SetupRoutes can report panics to it. With SENTRY_DSN empty it does
// nothing. A DSN Sentry refuses stops the start.
func (s *Server) enableSentry() error {
if s.params.Config.SentryDSN == "" {
return nil
}
err := sentry.Init(sentry.ClientOptions{
Dsn: s.params.Config.SentryDSN,
Release: s.params.Globals.Appname + "-" + s.params.Globals.Version,
})
if err != nil {
return fmt.Errorf("SENTRY_DSN: %w", err)
}
s.log.Info("sentry error reporting activated")
return nil
}
// shutdown gracefully stops the HTTP server within the
// deadline of the context fx provides for OnStop.
// deadline of the context fx provides for OnStop, then gives
// Sentry, if set up, time to send what it still holds.
func (s *Server) shutdown(ctx context.Context) error {
err := s.httpServer.Shutdown(ctx)
if s.params.Config.SentryDSN != "" {
sentry.Flush(sentryFlushTimeout)
}
if err != nil {
s.log.Error("server clean shutdown failed", "error", err)
+5 -24
View File
@@ -1,32 +1,13 @@
#!/bin/sh
# script/lint: run golangci-lint over the backend. This runs inside the
# lint stage of the root Dockerfile, whose digest-pinned golangci-lint
# image provides the linter; nothing installs golangci-lint on the host.
# From a checkout, run `make lint` at the repo root, which builds that
# stage.
#
# .golangci.yml is standardized org-wide and must never be edited here
# (REPO_POLICIES.md). Its last silent drift replaced the v2 schema with
# v1 keys, which left every threshold in the file inert while the build
# stayed green. So the file is first checked against the canonical
# copy's sha256: a local comparison, no network, nothing unpinned.
# script/lint: lint the whole repo as the root make lint does, by
# building the lint phase of the root Dockerfile. golangci-lint never
# runs on the host (REPO_POLICIES.md).
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
GOLANGCI_CONFIG_SHA256="a79b63a254602a5318db5d0e9a06bc71b84bf0c1d896305229d8bfed1d1b1776"
ROOT="$(cd "$(dirname "$0")/../.." && pwd -P)"
main() {
cd "$ROOT"
actual="$(sha256sum .golangci.yml | cut -d' ' -f1)"
if [ "$actual" != "$GOLANGCI_CONFIG_SHA256" ]; then
echo ".golangci.yml has drifted from the org standard." >&2
echo " expected $GOLANGCI_CONFIG_SHA256" >&2
echo " actual $actual" >&2
echo "Restore it verbatim from sneak/prompts; do not edit it." >&2
exit 1
fi
golangci-lint run ./...
exec "$ROOT/script/lint"
}
main "$@"
+10 -7
View File
@@ -1,18 +1,21 @@
#!/bin/sh
# script/test: run the backend test suite with the race detector and
# coverage. Go's own -timeout bounds the tests and not their compile,
# so a cold build cache cannot fail it. The race detector needs cgo,
# and so a C compiler. If the tests fail, they run again with -v for
# the details, and the script fails even if that run passes.
# script/test: run the backend test suite on the host with the race
# detector and coverage. The root make test runs the same in the test
# phase of the root Dockerfile. Go's own -timeout bounds the tests and
# not their compile, so a cold build cache cannot fail it. -count=1
# keeps Go's test result cache out of both runs, so neither can report
# a stored pass. The race detector needs cgo, and so a C compiler. If
# the tests fail, they run again with -v for the details, and the
# script fails even if that run passes.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
main() {
cd "$ROOT"
go test -timeout 30s -race -cover ./... || {
go test -count=1 -timeout 90s -race -cover ./... || {
echo "--- Rerunning with -v for details ---"
go test -timeout 30s -race -v ./...
go test -count=1 -timeout 90s -race -v ./...
exit 1
}
}
+6 -19
View File
@@ -63,26 +63,13 @@ done > /etc/nginx/trusted-proxies.conf
# netwatch-server keeps its report files in DATA_DIR, on the /data
# volume, which may be a host directory owned by root or by another
# uid. Both are given to the netwatch user here, with the mode the
# server gives a directory it creates, so the host directory needs no
# preparing.
#
# chown and chmod, run as root, change whatever a symbolic link on the
# path points to, anywhere in the container, and the netwatch user can
# put one in /data. So the start stops unless readlink -f, which
# follows every link on a path, gives /data and DATA_DIR back as they
# are. It also writes a path in full, so a DATA_DIR with '.', '..' or
# an extra '/' in it is refused too.
# uid. Here, as root, netwatch-server prepare-data-dir creates DATA_DIR
# and gives /data and everything in it to the netwatch user, so the
# host directory needs no preparing. It stops the start, naming
# DATA_DIR, unless DATA_DIR is /data or a path below it, and it acts on
# nothing outside /data, whatever symbolic links it meets there.
export DATA_DIR="${DATA_DIR:-/data/reports}"
mkdir -p "$DATA_DIR" || exit 1
if [ "$(readlink -f /data)" != /data ] ||
[ "$(readlink -f "$DATA_DIR")" != "$DATA_DIR" ]; then
echo "entrypoint: DATA_DIR must be a full path with no '.', '..'," \
"extra '/' or symbolic link on it or on /data, not '$DATA_DIR'" >&2
exit 1
fi
chown -R netwatch:netwatch /data "$DATA_DIR" || exit 1
chmod 750 /data "$DATA_DIR" || exit 1
netwatch-server prepare-data-dir "$DATA_DIR" || exit 1
# A stop signal is only noted here; the loop below acts on it.
stop_requested=""
+27
View File
@@ -0,0 +1,27 @@
// eslint's recommended rules over every JavaScript file in the repo.
// make lint runs it, in the frontend-lint stage of Dockerfile.
import js from "@eslint/js";
import globals from "globals";
import { defineConfig } from "eslint/config";
export default defineConfig([
js.configs.recommended,
{
// The page, and facts.js, which the viewport test runs inside it.
files: ["src/**/*.js", "test/viewport/facts.js"],
languageOptions: {
globals: {
...globals.browser,
// vite.config.js defines these when it builds the page.
__COMMIT_HASH__: "readonly",
__COMMIT_FULL__: "readonly",
},
},
},
{
// Run by node: the tests, the viewport harness and these configs.
files: ["test/**/*.js", "*.js"],
ignores: ["test/viewport/facts.js"],
languageOptions: { globals: globals.node },
},
]);
+6
View File
@@ -68,4 +68,10 @@ server {
location = /.well-known/healthcheck {
proxy_pass http://127.0.0.1:8081;
}
# The backend's Prometheus metrics, behind its own basic auth. Unless
# METRICS_USERNAME and METRICS_PASSWORD are set, it answers 404.
location = /metrics {
proxy_pass http://127.0.0.1:8081;
}
}
+5 -1
View File
@@ -6,12 +6,16 @@
"scripts": {
"dev": "vite",
"build": "vite build",
"preview": "vite preview"
"preview": "vite preview",
"test": "node --test test/unit/*.test.js"
},
"license": "MIT",
"devDependencies": {
"@eslint/js": "^10.0.1",
"@tailwindcss/vite": "^4.1.18",
"autoprefixer": "^10.4.23",
"eslint": "^10.12.0",
"globals": "^17.13.0",
"postcss": "^8.5.6",
"prettier": "^3.8.1",
"puppeteer-core": "25.5.0",
+30
View File
@@ -0,0 +1,30 @@
#!/bin/sh
# script/add-dependency: add a frontend package, or move one to another
# version, with yarn add, which changes package.json and yarn.lock
# together; then install from yarn.lock with --frozen-lockfile, as
# script/bootstrap does, to show it installs as written. --dev because
# no frontend package is needed when the page runs: it ships as the
# built dist/.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
usage() {
echo "usage: make add-dependency PACKAGE=<name>@<version>" >&2
exit 2
}
main() {
# Exactly one package. A value beginning with - would reach yarn as
# an option; yarn would quietly drop all but the first of several
# packages given in one value.
[ "$#" -eq 1 ] || usage
case "$1" in
"" | -* | *[[:space:]]*) usage ;;
esac
cd "$ROOT"
yarn add --dev "$1"
yarn install --frozen-lockfile
}
main "$@"
+37 -19
View File
@@ -6,27 +6,28 @@
# used directly if it is at least NODE_MIN_VERSION; otherwise it is
# installed at a pinned version via nvm (installing nvm itself first,
# from a hash-verified release archive, never curl | sh). Go, with its
# gofmt, is used directly if it is at least the version backend/go.mod
# asks for; otherwise the pinned Go release is installed from its
# hash-verified archive.
# gofmt, is used directly only if it is exactly GO_VERSION; otherwise
# that release is installed from its hash-verified archive.
#
# What this script installs outside the system package manager lives
# under $HOME and is linked into ~/.local/bin, where make and the git
# hook find it once that directory is on PATH. Nothing in ~/.local/bin
# that this script did not create is ever replaced.
#
# golangci-lint is not installed: make lint runs it in Docker, which
# this script does not install either.
# golangci-lint is not installed: make lint runs it in Docker, as make
# test runs the tests, and this script does not install Docker either.
#
# Unlike the org model: Go and gcc for backend/, a newer node for eslint.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
# Pinned versions, 2026-07-07
# Pinned versions, 2026-07-06
NODE_VERSION="22.17.0"
# The oldest node the frontend's dependencies accept: the "engines"
# field of puppeteer-core 25.5.0, the most demanding of them, asks for
# 22.12.0 or newer, 2026-09-29. An older installed node is not used.
NODE_MIN_VERSION="22.12.0"
# field of eslint 10.12.0, the most demanding of them, asks for 22.13.0
# or newer, 2026-10-03. An older installed node is not used.
NODE_MIN_VERSION="22.13.0"
NVM_VERSION="0.40.3"
# sha256 of https://github.com/nvm-sh/nvm/archive/refs/tags/v0.40.3.tar.gz
NVM_SHA256="5f4d6aaa04a177dc93c985e31dbc411ab6b8c6e1e21d8015dbc1372625fcd1d0"
@@ -184,20 +185,30 @@ ensure_yarn() {
}
# go_ok: the go on PATH has its gofmt beside it (a Go release ships the
# two together) and is at least the version backend/go.mod asks for.
# GOTOOLCHAIN=local makes an older go fail here instead of fetching a
# newer toolchain for itself.
# two together) and go version reports exactly GO_VERSION, compared over
# the whole version, not a prefix of it. A go that fails or prints
# anything else does not pass. GOTOOLCHAIN=local makes go report itself
# rather than a toolchain it would fetch.
go_ok() {
if missing go; then return 1; fi
[ -x "$(dirname "$(command -v go)")/gofmt" ] || return 1
(cd "$ROOT/backend" && GOTOOLCHAIN=local go list -m >/dev/null 2>&1)
version="$(GOTOOLCHAIN=local go version 2>/dev/null)" || return 1
case "$version" in
"go version go$GO_VERSION "*) return 0 ;;
*) return 1 ;;
esac
}
# ensure_go: unless go_ok, install GO_VERSION and link its go and gofmt.
# They are linked on every run that needs them, so a deleted link is put
# back, and the archive is unpacked again if either binary is missing.
# Afterwards go is looked up through PATH again and must pass go_ok, so
# another go that hides the link stops bootstrap.
ensure_go() {
if go_ok; then return 0; fi
if go_ok; then
echo "bootstrap: using $(GOTOOLCHAIN=local go version)"
return 0
fi
go_dir="$TOOLCHAIN/go-$GO_VERSION"
if [ ! -x "$go_dir/bin/go" ] || [ ! -x "$go_dir/bin/gofmt" ]; then
# sha256 of each archive, from https://go.dev/dl/?mode=json
@@ -238,6 +249,13 @@ ensure_go() {
fi
link_bin "$go_dir/bin/go" go
link_bin "$go_dir/bin/gofmt" gofmt
hash -r
if ! go_ok; then
echo "bootstrap: after installing go $GO_VERSION in $go_dir," >&2
echo " the go on PATH, $(command -v go), is not it or has no gofmt" >&2
exit 1
fi
echo "bootstrap: installed $(GOTOOLCHAIN=local go version)"
}
main() {
@@ -254,9 +272,9 @@ main() {
if missing make; then pkg_install gnumake make make make; fi
if missing git; then pkg_install git git git git; fi
# The race detector in make test needs cgo, which Go turns on only
# when it finds its C compiler, gcc on Linux. apt and apk ship the C
# library headers apart from gcc.
# The race detector in backend/'s make test needs cgo, which Go turns
# on only when it finds its C compiler, gcc on Linux. apt and apk
# ship the C library headers apart from gcc.
if missing gcc; then
pkg_install gcc "gcc libc6-dev" gcc "gcc musl-dev"
fi
@@ -269,8 +287,8 @@ main() {
(cd "$ROOT/backend" && go mod download)
if missing docker; then
echo "bootstrap: docker not found; make lint, and so make check" >&2
echo " and the pre-commit hook, need it to run the Go linter" >&2
echo "bootstrap: docker not found; make test and make lint, and so" >&2
echo " make check and the pre-commit hook, need it" >&2
fi
if [ -n "$path_hint" ] && [ -d "$BIN_DIR" ]; then
echo "bootstrap: add $BIN_DIR to the front of your PATH, e.g." >&2
Executable
+13
View File
@@ -0,0 +1,13 @@
#!/bin/sh
# script/build: build the frontend for production into dist/. The Go
# backend is built by backend/script/build.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
main() {
cd "$ROOT"
yarn build
}
main "$@"
+3 -1
View File
@@ -1,6 +1,8 @@
#!/bin/sh
# script/check: run all checks (test, lint, fmt-check). Our own
# extension to scripts-to-rule-them-all. Must not modify any files.
# extension to scripts-to-rule-them-all. test and lint are Docker
# phases; fmt-check is native, because a formatter writes the working
# tree. Must not modify any files.
set -eu
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd -P)"
Executable
+13
View File
@@ -0,0 +1,13 @@
#!/bin/sh
# script/dev: run the frontend's Vite dev server. It proxies /api to a
# netwatch-server running locally (see vite.config.js).
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
main() {
cd "$ROOT"
yarn dev
}
main "$@"
+6 -2
View File
@@ -1,12 +1,16 @@
#!/bin/sh
# script/fmt: format the whole repo (writes): prettier over everything
# it understands, then gofmt over the Go backend.
# script/fmt: format the whole repo (writes): prettier over the
# JavaScript, CSS, HTML and Markdown, then gofmt over the Go backend.
# The org model formats only markdown; this repo also has JS and Go.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
main() {
cd "$ROOT"
# script/bootstrap links the node, yarn and gofmt it installs into
# ~/.local/bin, which the shell that called it may not have on PATH.
PATH="$HOME/.local/bin:$PATH"
"$ROOT/script/frontend-fmt"
"$ROOT/backend/script/fmt"
}
+4
View File
@@ -1,12 +1,16 @@
#!/bin/sh
# script/fmt-check: check formatting across the whole repo (read-only).
# Same scope as script/fmt, but fails instead of writing.
# The org model checks only markdown; this repo also has JS and Go.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
main() {
cd "$ROOT"
# script/bootstrap links the node, yarn and gofmt it installs into
# ~/.local/bin, which the shell that called it may not have on PATH.
PATH="$HOME/.local/bin:$PATH"
"$ROOT/script/frontend-fmt-check"
"$ROOT/backend/script/fmt-check"
}
-18
View File
@@ -1,18 +0,0 @@
#!/bin/sh
# script/frontend-check: run the frontend half of the checks only (test,
# lint, fmt-check). This exists for the frontend stage of Dockerfile, a
# node image with neither Go nor Docker; the Dockerfile's lint and
# backend build stages gate the backend half. Everywhere else, use
# script/check, which covers the whole repo. Must not modify any files.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
main() {
cd "$ROOT"
"$ROOT/script/frontend-test"
"$ROOT/script/frontend-lint"
"$ROOT/script/frontend-fmt-check"
}
main "$@"
+6 -4
View File
@@ -1,14 +1,16 @@
#!/bin/sh
# script/frontend-fmt: format the frontend and every other file prettier
# understands, repo-wide (writes). backend/ is in .prettierignore; Go
# sources are formatted by backend/script/fmt.
# script/frontend-fmt: format the JavaScript, CSS, HTML and Markdown,
# repo-wide (writes), the markdown in backend/ included. Those are the
# languages REPO_POLICIES.md gives prettier; the YAML is left alone, as
# backend/.golangci.yml must stay the org standard byte for byte.
# Prettier does not read Go; backend/script/fmt formats the Go sources.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
main() {
cd "$ROOT"
yarn prettier --write .
yarn prettier --write '**/*.{js,css,html,md}'
}
main "$@"
+1 -1
View File
@@ -7,7 +7,7 @@ ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
main() {
cd "$ROOT"
yarn prettier --check .
yarn prettier --check '**/*.{js,css,html,md}'
}
main "$@"
-12
View File
@@ -1,12 +0,0 @@
#!/bin/sh
# script/frontend-lint: run the frontend linter (prettier in check mode).
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
main() {
cd "$ROOT"
yarn prettier --check .
}
main "$@"
+13 -4
View File
@@ -1,14 +1,23 @@
#!/bin/sh
# script/frontend-test: run the frontend test suite: the unit tests in
# test/unit/ with Node's built-in test runner, then the production
# build, which fails on broken code.
# script/frontend-test: run the frontend test suite on the host: the
# unit tests in test/unit/, through the test script in package.json,
# then the production build, which fails on broken code. make test runs
# the same in the frontend stage of Dockerfile. The tests print a dot
# each; if any fails, they run again with every test listed, and the
# script fails even if that run passes. NODE_OPTIONS chooses the
# reporter because yarn adds its arguments after the test files, where
# node would take a reporter option for one more file.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
main() {
cd "$ROOT"
timeout 30 node --test test/unit/*.test.js
NODE_OPTIONS=--test-reporter=dot timeout 90 yarn --silent run test || {
echo "--- Rerunning with every test listed for details ---"
NODE_OPTIONS=--test-reporter=spec timeout 90 yarn --silent run test
exit 1
}
timeout 30 yarn build
}
+2 -2
View File
@@ -8,8 +8,8 @@ ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
main() {
cd "$ROOT"
hook=".git/hooks/pre-commit"
printf '#!/bin/sh\nset -e\nscript/precommit\n' > "$hook"
chmod +x "$hook"
printf '#!/bin/sh\nset -e\nscript/precommit\n' > .git/hooks/pre-commit
chmod +x .git/hooks/pre-commit
echo "pre-commit hook installed: runs script/precommit"
}
+13 -11
View File
@@ -1,21 +1,23 @@
#!/bin/sh
# script/lint: lint the whole repo: prettier over the frontend, then the
# Go linter over backend/.
# script/lint: run the linter. Linting is a phase of the Dockerfile and
# this builds that phase alone; the linter is never installed or run on
# a developer host, where a shared result cache and a host-global lock
# make its answer untrustworthy.
#
# The Go linter runs only in Docker: this builds the lint stage of
# Dockerfile, the digest-pinned golangci-lint image, which runs the
# backend's fmt-check and lint targets. --no-cache makes the linter
# really run every time rather than reuse an earlier result, and the
# stage is built for its checks alone, so no image is kept.
# The phase is not the last stage in the file, so it is built only when
# --target names it. --no-cache because a cached lint layer is a lint
# that did not run. The tag makes each build replace the previous image
# instead of leaving a dangling one behind.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd -P)"
ROOT="$(cd "$SCRIPT_DIR/.." && pwd -P)"
main() {
cd "$ROOT"
"$ROOT/script/frontend-lint"
timeout 300 docker build --no-cache --target lint \
--output type=cacheonly .
docker build --no-cache \
--target lint \
-t "$("$SCRIPT_DIR/projectname")-lint" .
}
main "$@"
+10 -8
View File
@@ -1,17 +1,19 @@
#!/bin/sh
# script/test: run the test suite for the whole repo: the frontend at
# the repo root, then the Go backend in backend/. Each half has its own
# 30-second limit, and there is none around both: from a cold Go build
# cache, compiling the backend's tests with the race detector can take
# 30 seconds on its own, and Go's -timeout leaves the compile out.
# script/test: run the test suite. Testing is a phase of the Dockerfile
# and this builds that phase alone, on the same terms as script/lint:
# --target because a phase that is not the last stage is built only when
# named, --no-cache because a cached test layer is a test that did not
# run, and a tag so each build replaces the previous image.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd -P)"
ROOT="$(cd "$SCRIPT_DIR/.." && pwd -P)"
main() {
cd "$ROOT"
script/frontend-test
backend/script/test
docker build --no-cache \
--target test \
-t "$("$SCRIPT_DIR/projectname")-test" .
}
main "$@"
Executable
+13
View File
@@ -0,0 +1,13 @@
#!/bin/sh
# script/tidy: run go mod tidy in backend/, which adds the modules the
# Go sources import, drops those they no longer do, and updates go.sum.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
main() {
cd "$ROOT/backend"
go mod tidy
}
main "$@"
+190 -114
View File
@@ -8,6 +8,8 @@
// display their real value in the latency figure. The history buffer holds
// maxHistoryPoints samples (historyDuration / updateInterval).
// reportInterval is how often collected samples are POSTed to the backend.
// The interval menu changes updateInterval while the page runs; the
// getters compute their values from it each time they are read.
export const CONFIG = {
updateInterval: 3000,
maxHistoryPoints: 100,
@@ -28,6 +30,39 @@ export const CONFIG = {
return [0, 1, 2, 3, 4, 5].map((i) => Math.round((d * i) / 5));
},
canvasHeight: 96,
// A latency figure and its sparkline take the color of the first entry
// whose limit, in ms, the latency is below.
latencyColors: [
{ below: 50, hex: "#22c55e", className: "text-green-500" },
{ below: 100, hex: "#84cc16", className: "text-lime-500" },
{ below: 200, hex: "#eab308", className: "text-yellow-500" },
{ below: 500, hex: "#f97316", className: "text-orange-500" },
{ below: Infinity, hex: "#ef4444", className: "text-red-500" },
],
// The health is offline when more than offlineTimeouts WAN hosts timed
// out or were unreachable and at most offlineReachable answered;
// otherwise degraded when more than degradedTimeouts timed out or were
// unreachable; otherwise slow when more than slowHosts answered after
// more than slowLatency ms.
offlineTimeouts: 10,
offlineReachable: 4,
degradedTimeouts: 4,
slowHosts: 3,
slowLatency: 1000,
// The debug log keeps its last maxLogEntries lines.
maxLogEntries: 1000,
// A gateway candidate that has not answered after gatewayTimeout ms is
// passed over.
gatewayTimeout: 1500,
// When no WAN host answers, the recovery probe checks recoveryProbeHosts
// random ones every recoveryProbeInterval ms.
recoveryProbeHosts: 4,
recoveryProbeInterval: 500,
// The rows are sorted after the first round that is not discarded, then
// every roundsPerSort rounds.
roundsPerSort: 10,
// The sparklines are sized and drawn again resizeDelay ms after start.
resizeDelay: 100,
};
// WAN endpoints to monitor. These are used for the aggregate health/stats
@@ -112,7 +147,8 @@ const debugLog = [];
const log = (() => {
function append(level, message) {
debugLog.push({ timestamp: new Date(), level, message });
if (debugLog.length > 1000) debugLog.splice(0, debugLog.length - 1000);
if (debugLog.length > CONFIG.maxLogEntries)
debugLog.splice(0, debugLog.length - CONFIG.maxLogEntries);
const panel = document.getElementById("debug-panel");
if (panel && !panel.classList.contains("hidden")) renderDebugLog();
}
@@ -151,7 +187,7 @@ function formatUTCTimestamp(date) {
// --- Duration Formatting -----------------------------------------------------
function humanDuration(seconds) {
export function humanDuration(seconds) {
const h = Math.floor(seconds / 3600);
const m = Math.floor((seconds % 3600) / 60);
const s = seconds % 60;
@@ -172,7 +208,10 @@ async function detectGateway() {
const result = await Promise.any(
GATEWAY_CANDIDATES.map(async (url) => {
const controller = new AbortController();
const timeoutId = setTimeout(() => controller.abort(), 1500);
const timeoutId = setTimeout(
() => controller.abort(),
CONFIG.gatewayTimeout,
);
try {
await fetch(url, {
method: "GET",
@@ -197,14 +236,40 @@ async function detectGateway() {
// --- App State ---------------------------------------------------------------
class HostState {
// The min, max, median and average of latencies, a list of numbers, or all
// null when it is empty. The median of an even count is the mean of the
// middle two; it and the average are rounded.
function latencyStats(latencies) {
if (latencies.length === 0)
return { min: null, max: null, med: null, avg: null };
const sorted = [...latencies].sort((a, b) => a - b);
const mid = Math.floor(sorted.length / 2);
return {
min: sorted[0],
max: sorted[sorted.length - 1],
med:
sorted.length % 2
? sorted[mid]
: Math.round((sorted[mid - 1] + sorted[mid]) / 2),
avg: Math.round(
latencies.reduce((a, b) => a + b, 0) / latencies.length,
),
};
}
export class HostState {
constructor(host, pinned = false) {
this.name = host.name;
this.url = host.url;
this.history = []; // { timestamp, latency, paused }
// Each entry is either a check's result, { timestamp, latency,
// error }, or a round skipped while paused, { timestamp,
// latency: null, paused: true }.
this.history = [];
this.lastLatency = null;
this.status = "pending"; // 'online' | 'offline' | 'error' | 'pending'
this.pinned = pinned;
// How many recorded checks in a row have failed, up to the last one.
this.consecutiveFailures = 0;
}
pushSample(timestamp, result) {
@@ -218,6 +283,9 @@ class HostState {
if (result.error === "timeout") this.status = "error";
else if (result.error) this.status = "offline";
else this.status = "online";
this.consecutiveFailures = result.error
? this.consecutiveFailures + 1
: 0;
}
pushPaused(timestamp) {
@@ -225,36 +293,14 @@ class HostState {
this._trim();
}
averageLatency() {
const valid = this.history.filter((p) => p.latency !== null);
if (valid.length === 0) return null;
return Math.round(
valid.reduce((s, p) => s + p.latency, 0) / valid.length,
);
}
minLatency() {
const valid = this.history.filter((p) => p.latency !== null);
if (valid.length === 0) return null;
return Math.min(...valid.map((p) => p.latency));
}
maxLatency() {
const valid = this.history.filter((p) => p.latency !== null);
if (valid.length === 0) return null;
return Math.max(...valid.map((p) => p.latency));
}
medianLatency() {
const sorted = this.history
// The min, max, median and average latency of the checks in the history
// that got an answer.
historyStats() {
return latencyStats(
this.history
.filter((p) => p.latency !== null)
.map((p) => p.latency)
.sort((a, b) => a - b);
if (sorted.length === 0) return null;
const mid = Math.floor(sorted.length / 2);
return sorted.length % 2
? sorted[mid]
: Math.round((sorted[mid - 1] + sorted[mid]) / 2);
.map((p) => p.latency),
);
}
_trim() {
@@ -271,6 +317,10 @@ export class AppState {
this.local = localHosts.map((h) => new HostState(h));
this.paused = false;
this.tickCount = 0;
// The recovery probe's timer, null while it is not running, and the
// checks it started last.
this._recoveryProbeId = null;
this._recoveryProbeChecks = null;
}
get allHosts() {
@@ -279,33 +329,13 @@ export class AppState {
/** WAN-only stats from latest sample (excludes local) */
wanStats() {
const reachable = this.wan.filter((h) => h.lastLatency !== null);
const latencies = reachable.map((h) => h.lastLatency);
const total = this.wan.length;
if (latencies.length === 0)
return {
reachable: 0,
total,
min: null,
max: null,
med: null,
avg: null,
};
const sorted = [...latencies].sort((a, b) => a - b);
const mid = Math.floor(sorted.length / 2);
const med =
sorted.length % 2
? sorted[mid]
: Math.round((sorted[mid - 1] + sorted[mid]) / 2);
const latencies = this.wan
.filter((h) => h.lastLatency !== null)
.map((h) => h.lastLatency);
return {
reachable: latencies.length,
total,
min: Math.min(...latencies),
max: Math.max(...latencies),
med,
avg: Math.round(
latencies.reduce((a, b) => a + b, 0) / latencies.length,
),
total: this.wan.length,
...latencyStats(latencies),
};
}
@@ -331,12 +361,16 @@ export class AppState {
const timeouts = this.wan.filter(
(h) => h.status === "error" || h.status === "offline",
).length;
if (timeouts > 10 && reachable <= 4) return "offline";
if (timeouts > 4) return "degraded";
if (
timeouts > CONFIG.offlineTimeouts &&
reachable <= CONFIG.offlineReachable
)
return "offline";
if (timeouts > CONFIG.degradedTimeouts) return "degraded";
const slow = this.wan.filter(
(h) => h.lastLatency !== null && h.lastLatency > 1000,
(h) => h.lastLatency !== null && h.lastLatency > CONFIG.slowLatency,
).length;
if (slow > 3) return "slow";
if (slow > CONFIG.slowHosts) return "slow";
return "healthy";
}
@@ -505,7 +539,13 @@ class Reporter {
// Checks one target. The check times out after CONFIG.requestTimeout; the
// caller can give it up sooner through the optional signal, which also ends
// it as a timeout.
// it as a timeout. A failed check's reason says what went wrong and after
// how long; a check that answered has none.
//
// The URL is fetched as written, with nothing added to it: the Hetzner
// speed-test servers close the connection without an answer when the URL
// has a query string, and cache: "no-store" keeps the browser's cache out
// of the measurement.
export async function measureLatency(url, signal) {
const controller = new AbortController();
const timeoutId = setTimeout(
@@ -514,13 +554,10 @@ export async function measureLatency(url, signal) {
);
signal?.addEventListener("abort", () => controller.abort());
const targetUrl = new URL(url);
targetUrl.searchParams.set("_cb", Date.now().toString());
const start = performance.now();
try {
await fetch(targetUrl.toString(), {
await fetch(url, {
method: "GET",
mode: "no-cors",
cache: "no-store",
@@ -529,40 +566,42 @@ export async function measureLatency(url, signal) {
const latency = Math.round(performance.now() - start);
clearTimeout(timeoutId);
if (latency > CONFIG.maxLatency) {
log.error(`${url} timeout (${latency}ms > ${CONFIG.maxLatency}ms)`);
return { latency: null, error: "timeout" };
return {
latency: null,
error: "timeout",
reason: `answered after ${latency} ms, over the ${CONFIG.maxLatency} ms limit`,
};
}
return { latency, error: null };
return { latency, error: null, reason: null };
} catch (err) {
const took = Math.round(performance.now() - start);
clearTimeout(timeoutId);
if (err.name === "AbortError") {
log.error(`${url} timeout (aborted)`);
return { latency: null, error: "timeout" };
return {
latency: null,
error: "timeout",
reason: `timed out after ${took} ms (limit ${CONFIG.requestTimeout} ms)`,
};
}
log.error(`${url} unreachable`);
return { latency: null, error: "unreachable" };
return {
latency: null,
error: "unreachable",
reason: `network error (${err.name}: ${err.message}) after ${took} ms`,
};
}
}
// --- Color Helpers -----------------------------------------------------------
function latencyHex(latency) {
export function latencyHex(latency) {
if (latency === null) return "#6b7280";
if (latency < 50) return "#22c55e";
if (latency < 100) return "#84cc16";
if (latency < 200) return "#eab308";
if (latency < 500) return "#f97316";
return "#ef4444";
return CONFIG.latencyColors.find((c) => latency < c.below).hex;
}
function latencyClass(latency, status) {
export function latencyClass(latency, status) {
if (status === "offline" || status === "error" || latency === null)
return "text-gray-500";
if (latency < 50) return "text-green-500";
if (latency < 100) return "text-lime-500";
if (latency < 200) return "text-yellow-500";
if (latency < 500) return "text-orange-500";
return "text-red-500";
return CONFIG.latencyColors.find((c) => latency < c.below).className;
}
// --- Sparkline Renderer ------------------------------------------------------
@@ -580,8 +619,8 @@ class SparklineRenderer {
const ch = h - m.top - m.bottom;
ctx.clearRect(0, 0, w, h);
SparklineRenderer._drawYAxis(ctx, w, h, m, ch);
SparklineRenderer._drawXAxis(ctx, w, h, m, cw);
SparklineRenderer._drawYAxis(ctx, w, m, ch);
SparklineRenderer._drawXAxis(ctx, h, m, cw);
const len = history.length;
const pw = cw / (CONFIG.maxHistoryPoints - 1);
@@ -597,7 +636,7 @@ class SparklineRenderer {
SparklineRenderer._drawTip(ctx, history, getX, getY);
}
static _drawYAxis(ctx, w, h, m, ch) {
static _drawYAxis(ctx, w, m, ch) {
ctx.font = "300 12px monospace";
ctx.textAlign = "right";
ctx.textBaseline = "middle";
@@ -614,7 +653,7 @@ class SparklineRenderer {
}
}
static _drawXAxis(ctx, w, h, m, cw) {
static _drawXAxis(ctx, h, m, cw) {
ctx.textAlign = "center";
ctx.textBaseline = "top";
for (const tick of CONFIG.xAxisTicks) {
@@ -702,7 +741,18 @@ class SparklineRenderer {
// horizontally.
const STATUS_TEXT_CLASS = "status-text text-xs text-right col-span-2 mt-5";
function hostRowHTML(host, index, showPin = true) {
// Escapes text for HTML, so it shows as written inside an element or a
// quoted attribute and is never read as markup.
function escapeHTML(text) {
return text
.replaceAll("&", "&amp;")
.replaceAll("<", "&lt;")
.replaceAll(">", "&gt;")
.replaceAll('"', "&quot;")
.replaceAll("'", "&#39;");
}
export function hostRowHTML(host, index, showPin = true) {
const pinColor = host.pinned
? "text-blue-500"
: "text-gray-600 hover:text-gray-400";
@@ -721,12 +771,12 @@ function hostRowHTML(host, index, showPin = true) {
<div class="w-[420px] flex-shrink-0 grid grid-cols-[minmax(0,1fr)_auto] items-center">
<div class="flex items-center gap-2 min-w-[200px]">
<div class="w-3 h-3 rounded-full flex-shrink-0 bg-[#6b7280]"></div>
<span class="font-medium text-white truncate">${host.name}</span>
<span class="font-medium text-white truncate">${escapeHTML(host.name)}</span>
</div>
<div class="latency-value text-4xl font-bold tabular-nums text-right mt-3" data-host="${index}">
<span class="text-gray-500">---</span>
</div>
<a href="${host.url}" target="_blank" rel="noopener" class="text-xs text-gray-500 truncate block col-span-2 -mt-2">${host.url}</a>
<a href="${escapeHTML(host.url)}" target="_blank" rel="noopener" class="text-xs text-gray-500 truncate block col-span-2 -mt-2">${escapeHTML(host.url)}</a>
<div class="${STATUS_TEXT_CLASS} text-gray-500" data-host="${index}">waiting...</div>
</div>
<div class="flex-grow sparkline-container rounded overflow-hidden border border-gray-700/30">
@@ -825,7 +875,7 @@ function buildUI(state) {
</div>
<footer class="mt-8 text-center text-gray-600 text-xs">
<p>Latency measured via GET requests | IPv4 only | CORS restrictions may affect some measurements</p>
<p>Latency measured via GET requests | CORS restrictions may affect some measurements</p>
<p class="mt-2">
<span class="inline-block w-3 h-3 rounded-full bg-green-500 mr-1 align-middle"></span>&lt;50ms
<span class="inline-block w-3 h-3 rounded-full bg-lime-500 mr-1 ml-3 align-middle"></span>&lt;100ms
@@ -892,10 +942,7 @@ function updateHostRow(host, index) {
latencyEl.innerHTML = `<span class="text-gray-500">---</span>`;
}
const avg = host.averageLatency();
const med = host.medianLatency();
const min = host.minLatency();
const max = host.maxLatency();
const { min, med, avg, max } = host.historyStats();
if (host.status === "online" && avg !== null) {
statusEl.innerHTML = statusStatsHTML([
["min", min],
@@ -1042,14 +1089,16 @@ function renderDebugLog() {
info: "text-gray-300",
debug: "text-gray-500",
};
el.innerHTML = debugLog
.map((entry) => {
el.replaceChildren(
...debugLog.map((entry) => {
const ts = formatUTCTimestamp(entry.timestamp);
const cls = levelColors[entry.level] || "text-gray-400";
const lvl = entry.level.toUpperCase().padEnd(7);
return `<div class="${cls}">${ts} ${lvl} ${entry.message}</div>`;
})
.join("");
const line = document.createElement("div");
line.className = levelColors[entry.level] || "text-gray-400";
line.textContent = `${ts} ${lvl} ${entry.message}`;
return line;
}),
);
el.scrollTop = el.scrollHeight;
}
@@ -1105,6 +1154,26 @@ function sortAndRebuildWAN(state) {
// --- Main Loop ---------------------------------------------------------------
// Writes one line to the browser console, and the same line to the debug
// log, for a check the page is about to record: for every check that
// failed, and for a check that answered after checks that failed. Called
// before host.pushSample, while host.consecutiveFailures still counts the
// checks before this one.
function logCheck(host, result) {
const target = `${host.name} ${host.url} at ${new Date().toISOString()}`;
const failed = host.consecutiveFailures;
if (result.error) {
const line = `netwatch: check failed: ${target}: ${result.reason}`;
console.error(line);
log.error(line);
} else if (failed > 0) {
const checks = failed === 1 ? "check" : "checks";
const line = `netwatch: target recovered: ${target}: answered after ${result.latency} ms, following ${failed} failed ${checks} in a row`;
console.info(line);
log.info(line);
}
}
export async function tick(state, signal, onOffline) {
const ts = Date.now();
@@ -1138,6 +1207,7 @@ export async function tick(state, signal, onOffline) {
if (state.paused || signal.aborted || state.tickCount === 0) {
return;
}
logCheck(host, r);
host.pushSample(ts, r);
updateHostRow(host, state.allHosts.indexOf(host));
log.debug(`${host.name}: ${r.error ? r.error : r.latency + "ms"}`);
@@ -1156,8 +1226,13 @@ export async function tick(state, signal, onOffline) {
return;
}
// Sort after the first real check, then every 10 ticks thereafter
if (state.tickCount === 2 || state.tickCount % 10 === 1) {
// Redraw every row: if the user paused and resumed during this round,
// rows whose check ended before the resume still read "paused"
state.allHosts.forEach((host, i) => updateHostRow(host, i));
// Sort after the first real check, then every CONFIG.roundsPerSort
// ticks thereafter
if (state.tickCount === 2 || state.tickCount % CONFIG.roundsPerSort === 1) {
sortAndRebuildWAN(state);
}
@@ -1180,9 +1255,10 @@ export async function tick(state, signal, onOffline) {
// --- Recovery Probe ----------------------------------------------------------
// When offline, check 4 random WAN hosts every 500ms, giving up the checks
// started 500ms before, so at most 4 are ever waiting. As soon as one
// answers, stop probing and start a new round at once.
// When offline, check CONFIG.recoveryProbeHosts random WAN hosts every
// CONFIG.recoveryProbeInterval ms, giving up the checks started one interval
// before, so at most that many are ever waiting. As soon as one answers,
// stop probing and start a new round at once.
function startRecoveryProbe(state, startRounds) {
if (state._recoveryProbeId) return; // already running
const candidates = [...state.wan];
@@ -1190,7 +1266,7 @@ function startRecoveryProbe(state, startRounds) {
const j = Math.floor(Math.random() * (i + 1));
[candidates[i], candidates[j]] = [candidates[j], candidates[i]];
}
const canaries = candidates.slice(0, 4);
const canaries = candidates.slice(0, CONFIG.recoveryProbeHosts);
log.notice(
`Recovery probe started (${canaries.map((h) => h.name).join(", ")})`,
);
@@ -1207,7 +1283,7 @@ function startRecoveryProbe(state, startRounds) {
startRounds();
});
}
}, 500);
}, CONFIG.recoveryProbeInterval);
}
function stopRecoveryProbe(state) {
@@ -1220,7 +1296,7 @@ function stopRecoveryProbe(state) {
// --- Pause / Resume ----------------------------------------------------------
function greyOutUI(state) {
export function greyOutUI(state) {
// Grey out all host rows
state.allHosts.forEach((host, i) => {
const latencyEl = document.querySelector(
@@ -1471,7 +1547,7 @@ async function init() {
});
window.addEventListener("resize", () => handleResize(state));
setTimeout(() => handleResize(state), 100);
setTimeout(() => handleResize(state), CONFIG.resizeDelay);
}
// Bootstrap only when loaded as the page: a real DOM containing the #app
+470 -13
View File
@@ -1,17 +1,38 @@
// Unit tests for src/main.js, run by script/frontend-test with Node's
// built-in test runner. Importing the module does not start the page.
// Unit tests for src/main.js, run with Node's built-in test runner by the
// test script in package.json. Importing the module does not start the
// page.
import { test } from "node:test";
import { beforeEach, test } from "node:test";
import assert from "node:assert/strict";
import { AppState, CONFIG, measureLatency, tick } from "../../src/main.js";
import {
AppState,
CONFIG,
greyOutUI,
hostRowHTML,
HostState,
humanDuration,
latencyClass,
latencyHex,
measureLatency,
tick,
} from "../../src/main.js";
// There is no page here, so the tests stand in for it. The debug log looks
// for its panel by id and finds none. Each element of a host's row that
// tick draws into is a plain object, made the first time it is looked up
// and kept in elements under its selector. Drawing a sparkline does
// nothing; it looks for the pixel ratio on window and finds none.
const elements = {};
// tick or greyOutUI draws into is a plain object, made the first time a
// test looks it up and kept in elements under its selector until the next
// test starts. As on a page, writing its text replaces its markup; the
// status dot greyOutUI looks for in it is not there. Drawing a sparkline
// does nothing; it looks for the pixel ratio on window and finds none. In
// each test, console.error and console.info print nothing and keep what
// they are given.
const doNothing = () => {};
let elements;
beforeEach((t) => {
elements = {};
t.mock.method(console, "error", doNothing);
t.mock.method(console, "info", doNothing);
});
const canvasContext = {
clearRect: doNothing,
beginPath: doNothing,
@@ -27,7 +48,13 @@ globalThis.window = {};
globalThis.document = {
getElementById: () => null,
querySelector: (selector) =>
(elements[selector] ??= { getContext: () => canvasContext }),
(elements[selector] ??= {
getContext: () => canvasContext,
querySelector: () => null,
set textContent(text) {
this.innerHTML = text;
},
}),
};
// What tick last wrote into the latency figure in host's row, or undefined
@@ -37,11 +64,19 @@ function latencyFigure(state, host) {
return elements[`.latency-value[data-host="${index}"]`]?.innerHTML;
}
// What was last written into the status text in host's row.
function statusText(state, host) {
const index = state.allHosts.indexOf(host);
return elements[`.status-text[data-host="${index}"]`]?.innerHTML;
}
// Mocks the clock for test t, so that a check lasting seconds takes no real
// time, and replaces fetch with targets that each answer after
// answerAfter(url) milliseconds of that clock, or never when that is
// Infinity. Both are restored when the test ends.
function mockTargets(t, answerAfter) {
// Infinity. The target at unreachableUrl, if one is given, does not
// answer: after that time its fetch fails with the error a browser gives
// the page for a network error. Both are restored when the test ends.
function mockTargets(t, answerAfter, unreachableUrl) {
t.mock.timers.enable({ apis: ["setTimeout", "Date"] });
t.mock.method(performance, "now", () => Date.now());
t.mock.method(
@@ -49,8 +84,12 @@ function mockTargets(t, answerAfter) {
"fetch",
(url, { signal }) =>
new Promise((resolve, reject) => {
const answer =
url === unreachableUrl
? () => reject(new TypeError("Failed to fetch"))
: resolve;
if (answerAfter(url) !== Infinity) {
setTimeout(resolve, answerAfter(url));
setTimeout(answer, answerAfter(url));
}
signal.addEventListener("abort", () => reject(signal.reason));
}),
@@ -78,6 +117,7 @@ for (const interval of [10000, 30000]) {
assert.deepEqual(await settled(check), {
latency: slowAnswer,
error: null,
reason: null,
});
});
@@ -91,10 +131,37 @@ for (const interval of [10000, 30000]) {
assert.deepEqual(await settled(check), {
latency: null,
error: "timeout",
reason: `timed out after ${timeout} ms (limit ${timeout} ms)`,
});
});
}
// A browser can run a timer late, on a busy page or in a background tab.
// Moving the mocked clock on 30000ms at once runs the 24000ms timeout with
// the clock already at 30000ms.
test("at a 30000ms interval, a check whose timeout runs late gives how long the request took and the time limit", async (t) => {
CONFIG.updateInterval = 30000;
mockTargets(t, () => Infinity);
const check = measureLatency("https://target.test");
t.mock.timers.tick(30000);
assert.deepEqual(await settled(check), {
latency: null,
error: "timeout",
reason: "timed out after 30000 ms (limit 24000 ms)",
});
});
test("a check fetches the target's URL as written, with no query string added", async (t) => {
mockTargets(t, () => 10);
const check = measureLatency("https://fsn1-speed.hetzner.com");
t.mock.timers.tick(10);
assert.notEqual(await settled(check), "still waiting");
assert.equal(
fetch.mock.calls[0].arguments[0],
"https://fsn1-speed.hetzner.com",
);
});
test("at a 30000ms interval, a target answering after 1000ms shows in its row while another target's check is still waiting", async (t) => {
CONFIG.updateInterval = 30000;
const state = new AppState([
@@ -103,7 +170,7 @@ test("at a 30000ms interval, a target answering after 1000ms shows in its row wh
const answering = state.local[0];
const waiting = state.wan[0];
// No target but the answering one ever answers.
mockTargets(t, (url) => (url.startsWith(answering.url) ? 1000 : Infinity));
mockTargets(t, (url) => (url === answering.url ? 1000 : Infinity));
// The third tick: the first is discarded as a whole, and the second ends
// by sorting the rows, which rebuilds a page that is not here.
state.tickCount = 2;
@@ -120,3 +187,393 @@ test("at a 30000ms interval, a target answering after 1000ms shows in its row wh
assert.notEqual(await settled(round), "still waiting");
assert.equal(state.tickCount, 3);
});
// In the next three tests, the answering target's check is still waiting
// when something happens after which its result must not show.
test("at a 30000ms interval, a check still waiting when its round is given up does not show in its row", async (t) => {
CONFIG.updateInterval = 30000;
const state = new AppState([
{ name: "Answering", url: "https://answering.test" },
]);
const answering = state.local[0];
mockTargets(t, (url) => (url === answering.url ? 1000 : Infinity));
state.tickCount = 2;
const roundChecks = new AbortController();
const round = tick(state, roundChecks.signal);
t.mock.timers.tick(500);
// As a round started early does to the last round's checks.
roundChecks.abort();
assert.notEqual(await settled(round), "still waiting");
assert.equal(latencyFigure(state, answering), undefined);
});
test("at a 30000ms interval, a check still waiting when the user pauses does not show in its row", async (t) => {
CONFIG.updateInterval = 30000;
const state = new AppState([
{ name: "Answering", url: "https://answering.test" },
]);
const answering = state.local[0];
mockTargets(t, (url) => (url === answering.url ? 1000 : Infinity));
state.tickCount = 2;
const round = tick(state, new AbortController().signal);
t.mock.timers.tick(500);
state.paused = true;
t.mock.timers.tick(500);
assert.equal(await settled(round), "still waiting");
assert.equal(latencyFigure(state, answering), undefined);
});
test("at a 30000ms interval, a check in the first round does not show in its row", async (t) => {
CONFIG.updateInterval = 30000;
const state = new AppState([
{ name: "Answering", url: "https://answering.test" },
]);
const answering = state.local[0];
mockTargets(t, (url) => (url === answering.url ? 1000 : Infinity));
const round = tick(state, new AbortController().signal);
t.mock.timers.tick(1000);
assert.equal(await settled(round), "still waiting");
assert.equal(latencyFigure(state, answering), undefined);
});
test("at a 30000ms interval, after the user pauses and resumes during a round, no row reads paused once its last check ends", async (t) => {
CONFIG.updateInterval = 30000;
const state = new AppState([
{ name: "Answering", url: "https://answering.test" },
]);
const answering = state.local[0];
mockTargets(t, (url) => (url === answering.url ? 1000 : Infinity));
state.tickCount = 2;
const round = tick(state, new AbortController().signal);
t.mock.timers.tick(1000);
assert.equal(await settled(round), "still waiting");
// The user pauses, which greys out every row, and resumes, which leaves
// the rows as they are, as togglePause does. The answering target's
// check has already ended, so only the redraw of every row at the end
// of the round can take "paused" out of its row.
state.paused = true;
greyOutUI(state);
state.paused = false;
assert.equal(statusText(state, answering), "paused");
// The round ends when the last check times out.
t.mock.timers.tick(CONFIG.requestTimeout - 1000);
assert.notEqual(await settled(round), "still waiting");
for (const host of state.allHosts) {
assert.notEqual(statusText(state, host), "paused", host.name);
}
});
// In the next tests a round checks one target, Target, at a 30000ms
// interval, so a check times out after 24000ms. The mocked clock starts at
// 1970-01-01T00:00:00.000Z.
// An app state with Target and no WAN targets, so its rounds check only
// Target, and whose next round is recorded: it is the third, as the first
// is discarded and the second ends by sorting the rows, which rebuilds a
// page that is not here. Returns it and Target.
function stateWithOneTarget() {
CONFIG.updateInterval = 30000;
const state = new AppState([
{ name: "Target", url: "https://target.test" },
]);
state.wan = [];
state.tickCount = 2;
return { state, target: state.local[0] };
}
// What this test wrote to the browser console with console[method].
function consoleLines(method) {
return console[method].mock.calls.map((call) => call.arguments[0]);
}
test("a check that fails with a network error writes one console line with the target, the time, the error and how long the request took", async (t) => {
const { state, target } = stateWithOneTarget();
mockTargets(t, () => 23, target.url);
const round = tick(state, new AbortController().signal);
t.mock.timers.tick(23);
assert.notEqual(await settled(round), "still waiting");
assert.deepEqual(consoleLines("error"), [
"netwatch: check failed: Target https://target.test at 1970-01-01T00:00:00.023Z: network error (TypeError: Failed to fetch) after 23 ms",
]);
assert.deepEqual(consoleLines("info"), []);
});
test("a check that times out writes one console line with the target, the time, how long the request took and the time limit", async (t) => {
const { state } = stateWithOneTarget();
mockTargets(t, () => Infinity);
const round = tick(state, new AbortController().signal);
t.mock.timers.tick(24000);
assert.notEqual(await settled(round), "still waiting");
assert.deepEqual(consoleLines("error"), [
"netwatch: check failed: Target https://target.test at 1970-01-01T00:00:24.000Z: timed out after 24000 ms (limit 24000 ms)",
]);
assert.deepEqual(consoleLines("info"), []);
});
// An answer that took longer than CONFIG.maxLatency is recorded as a
// timeout. The limit is the check's own timeout, so such an answer is one
// that came in just as the check timed out; here it is lowered to 500ms.
test("a check answered over the time limit writes one console line with the target, the time, how long the answer took and the limit", async (t) => {
const { state } = stateWithOneTarget();
t.mock.getter(CONFIG, "maxLatency", () => 500);
mockTargets(t, () => 1000);
const round = tick(state, new AbortController().signal);
t.mock.timers.tick(1000);
assert.notEqual(await settled(round), "still waiting");
assert.deepEqual(consoleLines("error"), [
"netwatch: check failed: Target https://target.test at 1970-01-01T00:00:01.000Z: answered after 1000 ms, over the 500 ms limit",
]);
assert.deepEqual(consoleLines("info"), []);
});
test("a check that answers writes nothing to the console", async (t) => {
const { state } = stateWithOneTarget();
mockTargets(t, () => 30);
const round = tick(state, new AbortController().signal);
t.mock.timers.tick(30);
assert.notEqual(await settled(round), "still waiting");
assert.deepEqual(consoleLines("error"), []);
assert.deepEqual(consoleLines("info"), []);
});
for (const [failed, checks] of [
[1, "1 failed check"],
[3, "3 failed checks"],
]) {
test(`a check that answers after ${checks} writes one console line saying the target recovered, and the next writes nothing`, async (t) => {
const { state, target } = stateWithOneTarget();
for (let i = 0; i < failed; i++) {
target.pushSample(Date.now(), {
latency: null,
error: "unreachable",
});
}
mockTargets(t, () => 40);
for (let i = 0; i < 2; i++) {
const round = tick(state, new AbortController().signal);
t.mock.timers.tick(40);
assert.notEqual(await settled(round), "still waiting");
}
assert.deepEqual(consoleLines("info"), [
`netwatch: target recovered: Target https://target.test at 1970-01-01T00:00:00.040Z: answered after 40 ms, following ${checks} in a row`,
]);
assert.deepEqual(consoleLines("error"), []);
});
}
test("a check that fails in the first round, which is discarded, writes nothing to the console", async (t) => {
const { state, target } = stateWithOneTarget();
state.tickCount = 0;
mockTargets(t, () => 23, target.url);
const round = tick(state, new AbortController().signal);
t.mock.timers.tick(23);
assert.notEqual(await settled(round), "still waiting");
assert.deepEqual(consoleLines("error"), []);
assert.deepEqual(consoleLines("info"), []);
});
// The page shows &lt; &gt; &quot; &amp; and &#39; in a row's markup as
// < > " & and '.
test(`a target whose name and URL hold < > " & and ' shows those characters in its row`, () => {
const host = new HostState({
name: `<b>"x" & 'y'</b>`,
url: `https://x.test/<b>?a="x"&b='y'`,
});
const row = hostRowHTML(host, 0);
assert.doesNotMatch(row, /<b>/);
assert.ok(
row.includes(
">&lt;b&gt;&quot;x&quot; &amp; &#39;y&#39;&lt;/b&gt;</span>",
),
);
assert.ok(
row.includes(
'href="https://x.test/&lt;b&gt;?a=&quot;x&quot;&amp;b=&#39;y&#39;"',
),
);
assert.ok(
row.includes(
">https://x.test/&lt;b&gt;?a=&quot;x&quot;&amp;b=&#39;y&#39;</a>",
),
);
});
for (const [seconds, text] of [
[0, "0s"],
[1, "1s"],
[59, "59s"],
[60, "1m"],
[61, "1m1s"],
[3599, "59m59s"],
[3600, "1h"],
[3601, "1h1s"],
[3660, "1h1m"],
[3661, "1h1m1s"],
]) {
test(`humanDuration writes ${seconds} seconds as ${text}`, () => {
assert.equal(humanDuration(seconds), text);
});
}
// The colour of a target's latency figure, as a class, and of its sparkline
// for an answer after latency ms, as the colour coding in README.md gives
// them. The rows sit either side of each boundary.
for (const [latency, figure, sparkline] of [
[0, "text-green-500", "#22c55e"],
[49, "text-green-500", "#22c55e"],
[50, "text-lime-500", "#84cc16"],
[99, "text-lime-500", "#84cc16"],
[100, "text-yellow-500", "#eab308"],
[199, "text-yellow-500", "#eab308"],
[200, "text-orange-500", "#f97316"],
[499, "text-orange-500", "#f97316"],
[500, "text-red-500", "#ef4444"],
]) {
test(`an answer after ${latency}ms has a ${figure} figure and a ${sparkline} sparkline`, () => {
assert.equal(latencyClass(latency, "online"), figure);
assert.equal(latencyHex(latency), sparkline);
});
}
test("a check that timed out or found its target unreachable has a grey figure and sparkline", () => {
assert.equal(latencyClass(null, "error"), "text-gray-500");
assert.equal(latencyClass(null, "offline"), "text-gray-500");
assert.equal(latencyHex(null), "#6b7280");
});
// A target whose checks, in turn, answered after each of latencies ms, or,
// for null, found it unreachable.
function hostAfter(latencies) {
const host = new HostState({ name: "Target", url: "https://target.test" });
for (const latency of latencies) {
host.pushSample(
Date.now(),
latency === null
? { latency: null, error: "unreachable" }
: { latency, error: null },
);
}
return host;
}
for (const { history, latencies, statistics } of [
{
history: "no checks",
latencies: [],
statistics: { min: null, max: null, average: null, median: null },
},
{
history: "only unreachable checks",
latencies: [null, null, null],
statistics: { min: null, max: null, average: null, median: null },
},
{
history: "three answers and an unreachable check",
latencies: [30, null, 10, 20],
statistics: { min: 10, max: 30, average: 20, median: 20 },
},
{
// The median of an even number of answers is the mean of the
// middle two. It and the average, 23.75, are rounded.
history: "four answers",
latencies: [10, 40, 20, 25],
statistics: { min: 10, max: 40, average: 24, median: 23 },
},
{
// Sorted as text rather than as numbers, these answers would put
// 100 in the middle. The average, 39.67, is rounded.
history: "answers with different numbers of digits",
latencies: [100, 9, 10],
statistics: { min: 9, max: 100, average: 40, median: 10 },
},
]) {
test(`a target's min, max, average and median latency over ${history}`, () => {
const { min, max, avg, med } = hostAfter(latencies).historyStats();
assert.deepEqual({ min, max, average: avg, median: med }, statistics);
});
}
// The summary's figures come from each WAN target's last check, by the same
// rules as a target's own: here four answered, one was found unreachable
// and the rest have not been checked yet. The median, 22.5, and the
// average, 21.25, are rounded.
test("the summary's min, max, median and average latency over the WAN targets' last checks", () => {
const state = new AppState([]);
[30, 10, null, 25, 20].forEach((latency, i) =>
state.wan[i].pushSample(
Date.now(),
latency === null
? { latency: null, error: "unreachable" }
: { latency, error: null },
),
);
assert.deepEqual(state.wanStats(), {
reachable: 4,
total: state.wan.length,
min: 10,
max: 30,
med: 23,
avg: 21,
});
});
// An app state in which, of the WAN targets, the first timedOut timed out,
// the next unreachable were found unreachable, the next answered answered
// after latency ms, and the rest have not been checked yet.
function stateAfter({ timedOut, unreachable, answered, latency }) {
const state = new AppState([]);
const results = [
...Array(timedOut).fill({ latency: null, error: "timeout" }),
...Array(unreachable).fill({ latency: null, error: "unreachable" }),
...Array(answered).fill({ latency, error: null }),
];
results.forEach((result, i) => state.wan[i].pushSample(Date.now(), result));
return state;
}
// For each number the health is decided by, the rows put it one under, at
// and one over its threshold.
for (const { timedOut, unreachable = 0, answered, latency, health } of [
// Offline: more than 10 timed out and at most 4 answered.
{ timedOut: 9, answered: 4, latency: 30, health: "degraded" },
{ timedOut: 10, answered: 4, latency: 30, health: "degraded" },
{ timedOut: 11, answered: 4, latency: 30, health: "offline" },
{ timedOut: 11, answered: 3, latency: 30, health: "offline" },
{ timedOut: 11, answered: 5, latency: 30, health: "degraded" },
// Otherwise degraded: more than 4 timed out.
{ timedOut: 3, answered: 10, latency: 30, health: "healthy" },
{ timedOut: 4, answered: 10, latency: 30, health: "healthy" },
{ timedOut: 5, answered: 10, latency: 30, health: "degraded" },
// Otherwise slow: more than 3 answered after more than 1000ms.
{ timedOut: 0, answered: 4, latency: 999, health: "healthy" },
{ timedOut: 0, answered: 4, latency: 1000, health: "healthy" },
{ timedOut: 0, answered: 4, latency: 1001, health: "slow" },
{ timedOut: 0, answered: 2, latency: 1001, health: "healthy" },
{ timedOut: 0, answered: 3, latency: 1001, health: "healthy" },
// A target found unreachable counts as one that timed out.
{
timedOut: 5,
unreachable: 6,
answered: 4,
latency: 30,
health: "offline",
},
{
timedOut: 0,
unreachable: 5,
answered: 10,
latency: 30,
health: "degraded",
},
]) {
test(`with ${timedOut} WAN targets timed out, ${unreachable} found unreachable and ${answered} answering after ${latency}ms, the health is ${health}`, () => {
const state = stateAfter({ timedOut, unreachable, answered, latency });
assert.equal(state.healthStatus(), health);
});
}
+15 -13
View File
@@ -13,8 +13,8 @@ width derived from the app's own CSS. Screenshots land in `tmp/viewport/`
alongside a `results.json`; they are artifacts for a human to look at when
something fails, not the evidence. The assertions are the evidence.
The target is deliberately outside `make check`: it needs Docker and takes
minutes, and `make test` has to stay under 20 seconds.
The target is deliberately outside `make check`: it takes minutes, and
`make test` has to stay under 60 seconds.
## How the widths are chosen
@@ -49,10 +49,11 @@ sizes straddling the breakpoint for the rotation case.
excluded: it is a design choice, not breakage.
- **tap-targets-44px** — every interactive control is at least 44x44 CSS px on
touch viewports, _and_ each selector in the control list matched at least the
number of visible elements it declares. The second half is what stops the
check passing vacuously: with size alone, a renamed class would take its
controls out of the measured set and the check would report "all 0 controls
are at least 44x44" and pass. See below.
number of visible elements it declares: one of each single control, and one
pin button per WAN host row. The second half is what stops the check passing
vacuously: with size alone, a renamed class would take its controls out of the
measured set and the check would report "all 0 controls are at least 44x44"
and pass. See below.
- **host-rows-stacked / host-rows-side-by-side** — the rows genuinely reflow.
Computed `flex-direction` _and_ the actual geometry are checked, and in the
narrow layout the info block and the sparkline must each occupy essentially
@@ -76,7 +77,7 @@ the internet, so the app's latency probes cannot reach anything real. The
harness answers them itself from a fixed delay table, with a deterministic
fraction failed outright, so the rows render a realistic spread of one-, two-
and three-digit latencies plus some unreachable rows. That spread is what the
layout has to survive; 24 identical `---` placeholders would not exercise it.
layout has to survive; a `---` placeholder in every row would not exercise it.
## What this cannot verify
@@ -103,10 +104,11 @@ Everything else this issue was actually about — does the layout reflow, does
anything overflow, is content clipped, are the controls big enough — is a
function of viewport width and CSS, and is covered above.
## Relation to the unit test framework (#21)
## Relation to the unit tests
Complementary layers, not two stacks. `vitest` (#21) will exercise module-level
logic in-process with no browser. This harness exercises rendered layout in a
real engine and is the only thing here that can see a media query. Neither
replaces the other; assertions about computed styles and element geometry belong
here, assertions about functions belong in `vitest`.
Complementary layers, not two stacks. The unit tests in `test/unit/`, which
`make test` runs with Node's built-in test runner, exercise the functions
`src/main.js` exports in-process with no browser. This harness exercises
rendered layout in a real engine and is the only thing here that can see a media
query. Neither replaces the other; assertions about computed styles and element
geometry belong here, assertions about functions belong in `test/unit/`.
+14 -13
View File
@@ -15,7 +15,8 @@ export const MIN_TAP_TARGET_PX = 44;
// The controls named in the definition of done, plus the pause button.
// Each carries the smallest number of *visible* instances the page has to
// contain for the tap-target oracle to be measuring anything at all.
// contain, worked out from the facts gathered from that page, for the
// tap-target oracle to be measuring every control it should.
//
// Without those floors the check is inert: `undersized` is empty both when
// every control is large enough and when the selectors have gone stale and
@@ -24,13 +25,12 @@ export const MIN_TAP_TARGET_PX = 44;
// for all three singleton controls vanishing at once — so the floor is per
// selector, and one stale selector out of four fails the check.
export const INTERACTIVE_CONTROLS = [
{ selector: "#pause-btn", minCount: 1 },
{ selector: "#interval-select", minCount: 1 },
// One per pinnable host row. `app-rendered` already requires at least
// 10 host rows, so a count below that means the pin buttons stopped
// being rendered per row rather than that there were fewer hosts.
{ selector: ".pin-btn", minCount: 10 },
{ selector: "#debug-toggle", minCount: 1 },
{ selector: "#pause-btn", minCount: () => 1 },
{ selector: "#interval-select", minCount: () => 1 },
// One per WAN host row, so pin buttons missing from even one row fail
// the check rather than only a drop below some fixed number.
{ selector: ".pin-btn", minCount: (facts) => facts.wanRowCount },
{ selector: "#debug-toggle", minCount: () => 1 },
];
export const INTERACTIVE_SELECTORS = INTERACTIVE_CONTROLS.map(
@@ -105,8 +105,8 @@ export function evaluateChecks(facts, viewport, probes) {
// never rendered. Everything below is only meaningful if this holds.
check(
"app-rendered",
facts.rowCount >= 10 && facts.numericLatencies >= 5,
`${facts.rowCount} host rows, ${facts.numericLatencies} showing a numeric latency`,
facts.wanRowCount >= 10 && facts.numericLatencies >= 5,
`${facts.wanRowCount} WAN host rows, ${facts.numericLatencies} showing a numeric latency`,
);
const viewportWidth = Math.min(facts.innerWidth, facts.documentClientWidth);
@@ -162,7 +162,8 @@ export function evaluateChecks(facts, viewport, probes) {
seen.set(target.selector, (seen.get(target.selector) ?? 0) + 1);
}
const missing = INTERACTIVE_CONTROLS.filter(
(control) => (seen.get(control.selector) ?? 0) < control.minCount,
(control) =>
(seen.get(control.selector) ?? 0) < control.minCount(facts),
);
const undersized = facts.tapTargets.filter(
@@ -184,11 +185,11 @@ export function evaluateChecks(facts, viewport, probes) {
const detail = [];
if (missing.length > 0) {
detail.push(
"oracle is not measuring the page: " +
"oracle is not measuring every control: " +
summarise(
missing,
(c) =>
`${c.selector} matched ${seen.get(c.selector) ?? 0} visible element(s), expected at least ${c.minCount}`,
`${c.selector} matched ${seen.get(c.selector) ?? 0} visible element(s), expected at least ${c.minCount(facts)}`,
4,
),
);
+3 -1
View File
@@ -173,7 +173,9 @@ export function collectLayoutFacts(options) {
clipped,
tapTargets,
rows,
rowCount: document.querySelectorAll(".host-row").length,
// WAN host rows only: each has a pin button, and the tap-target
// check expects one per row. The local host rows have none.
wanRowCount: document.querySelectorAll("#wan-hosts .host-row").length,
numericLatencies: Array.from(
document.querySelectorAll(".latency-value"),
).filter((el) => /\d/.test(el.textContent)).length,
+7 -3
View File
@@ -16,6 +16,10 @@ import { collectLayoutFacts } from "./facts.js";
import { evaluateChecks, INTERACTIVE_SELECTORS } from "./checks.js";
import { deriveViewports } from "./viewports.js";
// page.waitForFunction and page.evaluate run the functions given them in
// the page, where these are defined.
/* global document, requestAnimationFrame */
function required(name) {
const value = process.env[name];
if (!value) {
@@ -62,7 +66,6 @@ function stableHash(text) {
async function connectBrowser() {
const deadline = Date.now() + BROWSER_TIMEOUT_MS;
let lastError;
for (;;) {
try {
const response = await fetch(`${CDP_URL}/json/version`);
@@ -77,9 +80,10 @@ async function connectBrowser() {
});
return { browser, version: info.Browser };
} catch (error) {
lastError = error;
if (Date.now() > deadline) {
throw new Error(`browser never came up: ${lastError}`);
throw new Error(`browser never came up: ${error}`, {
cause: error,
});
}
await sleep(250);
}
+724 -193
View File
File diff suppressed because it is too large Load Diff