check / check (push) Successful in 3m32s
Each client address may make 60 requests to /metrics a minute, through the same httprate middleware and TRUSTED_PROXIES resolution the report route uses, with an allowance of its own. The limit runs before the basic auth, so past it the answer is 429 and the password is not checked. backend/README.md says so; a test uses up one client's allowance on wrong passwords, gets 429 with the right one, and checks that another client behind the same nginx still gets in. Model: opus-5-5
233 lines
13 KiB
Markdown
233 lines
13 KiB
Markdown
netwatch-server is an MIT-licensed Go HTTP backend by
|
|
[@sneak](https://sneak.berlin) that receives telemetry reports from the NetWatch
|
|
SPA and persists them as zstd-compressed JSONL files on disk.
|
|
|
|
## Getting Started
|
|
|
|
From this directory:
|
|
|
|
```bash
|
|
# Build and run locally
|
|
make run
|
|
```
|
|
|
|
From the repo root, whose `Dockerfile` builds the one image that ships this
|
|
backend behind nginx (see [Container image](#container-image)):
|
|
|
|
```bash
|
|
# Run tests, lint, and format check over the frontend and this backend
|
|
make check
|
|
|
|
# Build the image: nginx, the frontend and this backend
|
|
make docker
|
|
docker run -p 8080:8080 netwatch
|
|
```
|
|
|
|
## Entrypoints
|
|
|
|
This directory follows the same
|
|
[Scripts to Rule Them All](https://github.com/github/scripts-to-rule-them-all)
|
|
pattern as the repo root: the targets in `backend/Makefile` are thin shims over
|
|
`backend/script/`. The root `Dockerfile` runs them, and the root scripts call
|
|
`test`, `fmt` and `fmt-check`:
|
|
|
|
- `script/build` — compile the static `netwatch-server` binary with its version
|
|
stamped in. The version is `VERSION` from the environment; when that is unset
|
|
or empty, it falls back to `git describe` inside a git checkout, then to `dev`
|
|
- `script/test` — run the Go tests with the race detector and coverage. Go's
|
|
`-timeout 30s` bounds the tests, not their compile. If they fail, they run
|
|
again with `-v` for the details, and the script fails. The race detector needs
|
|
a C compiler
|
|
- `script/lint` — check `.golangci.yml` against its pinned sha256, then run
|
|
golangci-lint. It runs inside the golangci-lint image of the lint stage of the
|
|
root `Dockerfile`; from a checkout, run `make lint` at the repo root, which
|
|
builds that stage
|
|
- `script/fmt` — format the Go sources (writes)
|
|
- `script/fmt-check` — check Go formatting (read-only)
|
|
- `script/run` — build and run the server locally
|
|
- `script/clean` — remove build artifacts
|
|
|
|
There is no `check`, `hooks` or `docker` target here: the root `make check`
|
|
covers this directory, the root `make hooks` installs the repo's only pre-commit
|
|
hook, and the root `make docker` builds the image that contains this backend.
|
|
|
|
## Rationale
|
|
|
|
The NetWatch frontend collects latency measurements from the browser but has no
|
|
way to persist or aggregate them. This backend provides a minimal
|
|
`POST /api/v1/reports` endpoint that buffers incoming reports in memory and
|
|
flushes them to compressed files on disk for later analysis.
|
|
|
|
## Design
|
|
|
|
The server is structured as an `fx`-wired Go application under
|
|
`cmd/netwatch-server/`. Internal packages in `internal/` follow standard Go
|
|
project layout:
|
|
|
|
- **`config`**: Loads configuration from environment variables and config files
|
|
via Viper.
|
|
- **`handlers`**: HTTP request handlers for the API (health check, report
|
|
ingestion).
|
|
- **`reportbuf`**: In-memory buffer that accumulates JSONL report lines and
|
|
flushes to zstd-compressed files when the buffer reaches 10 MiB or every 60
|
|
seconds.
|
|
- **`server`**: Chi-based HTTP server with middleware wiring and route
|
|
registration.
|
|
- **`healthcheck`**, **`middleware`**, **`logger`**, **`globals`**: Supporting
|
|
infrastructure.
|
|
|
|
### Configuration
|
|
|
|
| Variable | Default | Description |
|
|
| ---------------------- | -------------------- | -------------------------------------------------------------------------------------------------------- |
|
|
| `BIND_ADDRESS` | empty | IP address to listen on; empty listens on every interface |
|
|
| `PORT` | `8080` | HTTP listen port |
|
|
| `DATA_DIR` | `./data/reports` | Directory for compressed reports |
|
|
| `DATA_DIR_MAX_BYTES` | `1073741824` (1 GiB) | Most bytes of report files kept in `DATA_DIR`, oldest deleted first; see [Report limits](#report-limits) |
|
|
| `DEBUG` | `false` | Enable debug logging |
|
|
| `TRUSTED_PROXIES` | loopback + RFC1918 | Comma-separated CIDRs whose `X-Forwarded-For` / `X-Real-IP` headers are trusted for client IP resolution |
|
|
| `REPORTS_PER_MINUTE` | `60` | Reports each client address may send a minute; see [Report limits](#report-limits) |
|
|
| `CORS_ALLOWED_ORIGINS` | empty | Comma-separated origins whose pages may call the API; see [CORS](#cors) |
|
|
| `METRICS_USERNAME` | empty | Basic auth user name for `/metrics`; see [Metrics](#metrics) |
|
|
| `METRICS_PASSWORD` | empty | Basic auth password for `/metrics`; see [Metrics](#metrics) |
|
|
| `SENTRY_DSN` | empty | DSN of the Sentry project to send errors to; see [Sentry](#sentry) |
|
|
|
|
`TRUSTED_PROXIES` defaults to
|
|
`127.0.0.1/32,::1/128,10.0.0.0/8,172.16.0.0/12,192.168.0.0/16`. The loopback
|
|
entries cover a reverse proxy on the same host. A request whose direct peer is
|
|
outside this set has its forwarded headers ignored, and the direct peer is
|
|
logged and rate-limited instead. The container image does not use this default;
|
|
see [Container image](#container-image).
|
|
|
|
A variable set to a value the server cannot use, such as `PORT=abc`,
|
|
`DEBUG=maybe` or a `BIND_ADDRESS` that is not an IP address, stops it from
|
|
starting, with an error naming the variable. An empty variable counts as unset.
|
|
|
|
### Container image
|
|
|
|
The root `Dockerfile` builds one image in which nginx listens on the public port
|
|
8080, serves the frontend, and proxies `/api/`, `/.well-known/healthcheck` and
|
|
`/metrics` to this server. The image's entrypoint, `bin/entrypoint.sh`, starts
|
|
the server as user `netwatch` (uid 1000) with `BIND_ADDRESS=127.0.0.1` and
|
|
`PORT=8081`, so only nginx reaches it, and with `TRUSTED_PROXIES=127.0.0.1/32`,
|
|
so it takes the client address nginx passes on and no other. `DATA_DIR` is
|
|
`/data/reports`, on the `/data` volume; before starting the server, the
|
|
entrypoint creates it and gives it and `/data` to `netwatch` with
|
|
`netwatch-server prepare-data-dir`, which acts on nothing outside `/data`. nginx
|
|
replaces the security headers this server sets with those in the root
|
|
`security-headers.conf`, so those are what clients of the image see.
|
|
|
|
The container's own `TRUSTED_PROXIES` goes to nginx instead: IP addresses or
|
|
CIDRs, separated by commas, of the reverse proxies in front of the container.
|
|
nginx takes the client address from `X-Forwarded-For` only on a request from one
|
|
of them. Unset or empty, nginx trusts no proxy, and the client address is the
|
|
one each request comes from, so every client behind a proxy shares one rate
|
|
limit. An entry that is not an IP address or CIDR, such as a hostname or
|
|
`1.2.3`, stops the container at start with an error naming `TRUSTED_PROXIES`:
|
|
the entrypoint checks each entry with `netwatch-server check-cidr`, which parses
|
|
it as this server parses its own `TRUSTED_PROXIES`.
|
|
|
|
### Report storage
|
|
|
|
Reports are written as `reports-<timestamp>-<number>.jsonl.zst` files in
|
|
`DATA_DIR`. The timestamp is in UTC to the millisecond, so the names sort by
|
|
time. The number starts at 1 when the server starts and goes up by one for each
|
|
file the server starts to write, so two files written in the same millisecond
|
|
still get different names. A failed write uses up its number and leaves a gap in
|
|
the numbers: its file, if it was created, is removed. The file stays, counted
|
|
toward `DATA_DIR_MAX_BYTES` from the next start, only if removing it fails too.
|
|
Each file contains one JSON object per line, compressed with zstd. Files are
|
|
created with `O_EXCL` to prevent overwrites.
|
|
|
|
### Report limits
|
|
|
|
`POST /api/v1/reports` takes reports from anyone who can reach it, without
|
|
credentials, so it is bounded instead. Both refusals below answer with the same
|
|
`{"status":"error"}` body as any other error.
|
|
|
|
- **Rate limit.** Each client address, resolved through `TRUSTED_PROXIES`, may
|
|
send `REPORTS_PER_MINUTE` reports a minute; past that it gets 429 with
|
|
`Retry-After: 60`. The minute slides: reports from the minute before still
|
|
count, fading out over the current one, so an address is sure never to be
|
|
refused only while it sends at most half of `REPORTS_PER_MINUTE` in any 60
|
|
seconds. The page sends one report a minute from each open tab, so the default
|
|
of 60 refuses nothing from up to 30 tabs behind one address, such as a
|
|
household or an office sharing it, however their reports bunch up. Report
|
|
responses also carry `X-RateLimit-Limit`, `X-RateLimit-Remaining` and
|
|
`X-RateLimit-Reset` headers.
|
|
- **Size cap.** The report files in `DATA_DIR` may total at most
|
|
`DATA_DIR_MAX_BYTES`, counting the files already there at start. Reports
|
|
waiting in memory count at their uncompressed size until they are written;
|
|
those lost to a failed write stop counting, and the part of its file written
|
|
is removed. When a report would take the total past the cap, the oldest report
|
|
files are deleted to make room, and each deletion is logged with the file's
|
|
name and size; a file still being written is never deleted. A report is
|
|
refused with 507, and nothing of it is stored, only when the reports waiting
|
|
to be written fill the cap on their own, and then no file is deleted. At
|
|
start, report files past the cap, as after lowering it, are deleted the same
|
|
way. So the cap is how much of the newest reports is kept: the default of 1
|
|
GiB is small enough for any host; set it to the space you can give `DATA_DIR`.
|
|
|
|
### CORS
|
|
|
|
The page calls the API from the origin it is served from, so by default the
|
|
server sends no CORS headers, and browsers let no other origin's pages call it.
|
|
To serve the page from elsewhere, list that origin in `CORS_ALLOWED_ORIGINS`
|
|
(for example `https://netwatch.example.com`); pages from a listed origin may
|
|
`GET` and `POST` with a `Content-Type` header. Each entry must be a plain
|
|
origin, `scheme://host` with an optional `:port`, as browsers send it: no path,
|
|
not even a trailing `/`, and no `*`. Any other entry stops the server from
|
|
starting, with an error naming `CORS_ALLOWED_ORIGINS`.
|
|
|
|
### Metrics
|
|
|
|
With both `METRICS_USERNAME` and `METRICS_PASSWORD` set, the server serves
|
|
Prometheus metrics at `GET /metrics` to requests with those as their basic auth
|
|
credentials, and answers any other with 401. For each request that reaches the
|
|
health check or `POST /api/v1/reports`, those the rate limit refuses included,
|
|
the metrics record its duration and response size, labelled with its path,
|
|
method and status; they also count those requests in progress, and include Go's
|
|
runtime and process metrics. No other request is recorded: not those to
|
|
`/metrics` itself, and not those answered before they reach either route, such
|
|
as a CORS preflight, or a request refused with 404 for a path no route has, 405
|
|
for a method its route does not take, or 413 for declaring a body length over
|
|
the 1 MiB limit. A report whose body goes over the limit without declaring its
|
|
length reaches the route, is answered 413 there, and is recorded with that
|
|
status. Clients can make up any number of paths and methods, and each would add
|
|
labels to the metrics for as long as the server runs. With neither set, nothing
|
|
is recorded and `/metrics` answers 404. One without the other stops the server
|
|
from starting, with an error naming both; so does a `METRICS_USERNAME`
|
|
containing `:`, which basic auth cannot carry, with an error naming it.
|
|
|
|
`/metrics` is rate limited, so that its password cannot be guessed quickly: each
|
|
client address, resolved through `TRUSTED_PROXIES`, may make 60 requests to it a
|
|
minute, whatever their credentials. Past that it gets 429 with
|
|
`Retry-After: 60`, and its credentials are not checked. The minute slides as it
|
|
does for reports (see [Report limits](#report-limits)), so a scraper polling
|
|
every 2 seconds or less often is never refused. This allowance is apart from the
|
|
one for reports.
|
|
|
|
### Sentry
|
|
|
|
With `SENTRY_DSN` set, the server sends its errors to that Sentry project: each
|
|
panic in a handler is reported there, under the release `netwatch-server-`
|
|
followed by the server's version, and the request still gets 500 from the
|
|
server's panic recovery. On shutdown the server waits up to 2 seconds for Sentry
|
|
to finish sending. A DSN Sentry refuses stops the server from starting, with an
|
|
error naming `SENTRY_DSN`. With it empty, Sentry is not set up, and nothing is
|
|
sent to it.
|
|
|
|
## TODO
|
|
|
|
- Add integration test that POSTs a report and verifies the compressed output
|
|
- Add report decompression/query endpoint
|
|
- Add metrics (Prometheus) for buffer size, flush count, report count
|
|
|
|
## License
|
|
|
|
MIT. See [LICENSE](LICENSE).
|
|
|
|
## Author
|
|
|
|
[@sneak](https://sneak.berlin)
|