Closes #75. `bin/entrypoint.sh`, which already runs as root, now makes the data directory usable before the backend starts: it creates `DATA_DIR` if missing, gives it and `/data` to the `netwatch` user (`chown -R`), and sets mode 750 on both, the mode the backend gives a directory it creates. The backend still runs as `netwatch`. The README "Running under upaas" section loses its first-run step that created and chowned the host directory and names only the path to mount. The Dockerfile's build-time `mkdir` and `chown` of `/data` are gone, since the entrypoint now does this on every start. What the diff does not show: - The host directory mounted at `/data` ends up owned by uid 1000 with mode 750, and everything under `DATA_DIR` is chowned to uid 1000 on every start. - If the directory cannot be created or chowned, the container stops with that tool's error before either process starts. Recorded runs with `--mount type=bind`: an empty directory owned by root (mode 755, and again mode 700), and one holding a `reports` directory and report file owned by uid 1001 with mode 700. Each time the container turned healthy, `netwatch-server` ran as `netwatch`, and a posted report was written to `DATA_DIR`; a second start on the root-owned and the uid 1001 directories did the same. Judgement call: `/data` itself is given to `netwatch` as well as `DATA_DIR`, so the backend can reach `DATA_DIR` inside a host directory with mode 700. Model: opus-5-5 Reviewed-on: #76 Co-authored-by: clawbot <35+clawbot@noreply.example.org>
9.4 KiB
netwatch-server is an MIT-licensed Go HTTP backend by @sneak that receives telemetry reports from the NetWatch SPA and persists them as zstd-compressed JSONL files on disk.
Getting Started
From this directory:
# Build and run locally
make run
From the repo root, whose Dockerfile builds the one image that ships this
backend behind nginx (see Container image):
# Run tests, lint, and format check over the frontend and this backend
make check
# Build the image: nginx, the frontend and this backend
make docker
docker run -p 8080:8080 netwatch
Entrypoints
This directory follows the same
Scripts to Rule Them All
pattern as the repo root: the targets in backend/Makefile are thin shims over
backend/script/. The root Dockerfile runs them, and the root scripts call
test, fmt and fmt-check:
script/build— compile the staticnetwatch-serverbinary with its version and architecture stamped in. The version isVERSIONfrom the environment; when that is unset or empty, it falls back togit describeinside a git checkout, then todevscript/test— run the Go tests under a 30-second timeoutscript/lint— check.golangci.ymlagainst its pinned sha256, then run golangci-lint. It runs inside the golangci-lint image of the lint stage of the rootDockerfile; from a checkout, runmake lintat the repo root, which builds that stagescript/fmt— format the Go sources (writes)script/fmt-check— check Go formatting (read-only)script/run— build and run the server locallyscript/clean— remove build artifacts
There is no check, hooks or docker target here: the root make check
covers this directory, the root make hooks installs the repo's only pre-commit
hook, and the root make docker builds the image that contains this backend.
Rationale
The NetWatch frontend collects latency measurements from the browser but has no
way to persist or aggregate them. This backend provides a minimal
POST /api/v1/reports endpoint that buffers incoming reports in memory and
flushes them to compressed files on disk for later analysis.
Design
The server is structured as an fx-wired Go application under cmd/netwatch-server/.
Internal packages in internal/ follow standard Go project layout:
config: Loads configuration from environment variables and config files via Viper.handlers: HTTP request handlers for the API (health check, report ingestion).reportbuf: In-memory buffer that accumulates JSONL report lines and flushes to zstd-compressed files when the buffer reaches 10 MiB or every 60 seconds.server: Chi-based HTTP server with middleware wiring and route registration.healthcheck,middleware,logger,globals: Supporting infrastructure.
Configuration
| Variable | Default | Description |
|---|---|---|
BIND_ADDRESS |
empty | IP address to listen on; empty listens on every interface |
PORT |
8080 |
HTTP listen port |
DATA_DIR |
./data/reports |
Directory for compressed reports |
DATA_DIR_MAX_BYTES |
1073741824 (1 GiB) |
Largest total size of the report files in DATA_DIR; see Report limits |
DEBUG |
false |
Enable debug logging |
TRUSTED_PROXIES |
loopback + RFC1918 | Comma-separated CIDRs whose X-Forwarded-For / X-Real-IP headers are trusted for client IP resolution |
REPORTS_PER_MINUTE |
60 |
Reports each client address may send a minute; see Report limits |
CORS_ALLOWED_ORIGINS |
empty | Comma-separated origins whose pages may call the API; see CORS |
TRUSTED_PROXIES defaults to 127.0.0.1/32,::1/128,10.0.0.0/8,172.16.0.0/12,192.168.0.0/16.
The loopback entries cover a reverse proxy on the same host. A request whose
direct peer is outside this set has its forwarded headers ignored, and the
direct peer is logged and rate-limited instead. The container image does not use
this default; see Container image.
A variable set to a value the server cannot use, such as PORT=abc,
DEBUG=maybe or a BIND_ADDRESS that is not an IP address, stops it from
starting, with an error naming the variable. An empty variable counts as unset.
Container image
The root Dockerfile builds one image in which nginx listens on the public port
8080, serves the frontend, and proxies /api/ and /.well-known/healthcheck to
this server. The image's entrypoint, bin/entrypoint.sh, starts the server as
user netwatch (uid 1000) with BIND_ADDRESS=127.0.0.1 and PORT=8081, so
only nginx reaches it, and with TRUSTED_PROXIES=127.0.0.1/32, so it takes the
client address nginx passes on and no other. DATA_DIR is /data/reports, on
the /data volume; the entrypoint creates it and gives it and /data to
netwatch before starting the server. nginx replaces the security headers
this server sets with those in the root security-headers.conf, so those are
what clients of the image see.
The container's own TRUSTED_PROXIES goes to nginx instead: IP addresses or
CIDRs, separated by commas, of the reverse proxies in front of the container.
nginx takes the client address from X-Forwarded-For only on a request from one
of them. Unset or empty, nginx trusts no proxy, and the client address is the
one each request comes from, so every client behind a proxy shares one rate
limit. An entry that is not an IP address or CIDR, such as a hostname or
1.2.3, stops the container at start with an error naming TRUSTED_PROXIES:
the entrypoint checks each entry with netwatch-server check-cidr, which parses
it as this server parses its own TRUSTED_PROXIES.
Report storage
Reports are written as reports-<timestamp>-<number>.jsonl.zst files in
DATA_DIR. The timestamp is in UTC to the millisecond, so the names sort by
time. The number starts at 1 when the server starts and goes up by one for each
file the server starts to write, so two files written in the same millisecond
still get different names. A failed write uses up its number, leaving a gap in
the numbers if the file could not be created and otherwise a file under that
number that may be incomplete. Each file contains one JSON object per line,
compressed with zstd. Files are created with O_EXCL to prevent overwrites.
Report limits
POST /api/v1/reports takes reports from anyone who can reach it, without
credentials, so it is bounded instead. Both refusals below answer with the same
{"status":"error"} body as any other error.
- Rate limit. Each client address, resolved through
TRUSTED_PROXIES, may sendREPORTS_PER_MINUTEreports a minute; past that it gets 429 withRetry-After: 60. The minute slides: reports from the minute before still count, fading out over the current one, so an address is sure never to be refused only while it sends at most half ofREPORTS_PER_MINUTEin any 60 seconds. The page sends one report a minute from each open tab, so the default of 60 refuses nothing from up to 30 tabs behind one address, such as a household or an office sharing it, however their reports bunch up. Report responses also carryX-RateLimit-Limit,X-RateLimit-RemainingandX-RateLimit-Resetheaders. - Size cap. The report files in
DATA_DIRmay total at mostDATA_DIR_MAX_BYTES, counting the files already there at start. Reports waiting in memory count at their uncompressed size until they are written, so a report that would take the total past the cap is refused with 507, and nothing of it is stored. Deleting report files frees room only at the next start, when the files are counted again. The default of 1 GiB is small enough for any host; set it to the space you can giveDATA_DIR.
CORS
The page calls the API from the origin it is served from, so by default the
server sends no CORS headers, and browsers let no other origin's pages call it.
To serve the page from elsewhere, list that origin in CORS_ALLOWED_ORIGINS
(for example https://netwatch.example.com); pages from a listed origin may
GET and POST with a Content-Type header. Each entry must be a plain
origin, scheme://host with an optional :port, as browsers send it: no path,
not even a trailing /, and no *. Any other entry stops the server from
starting, with an error naming CORS_ALLOWED_ORIGINS.
TODO
- Add integration test that POSTs a report and verifies the compressed output
- Add report decompression/query endpoint
- Add metrics (Prometheus) for buffer size, flush count, report count
- Add retention policy to prune old report files
License
MIT. See LICENSE.