fx wrote its own steps of starting and stopping as plain text to
stderr. It now logs them with its slog event logger through the
backend's logger, so off a terminal every line the backend's own
logger and fx write is JSON. A malformed config file now makes
config.New return the error instead of panicking. A test runs the
server as a child process and checks that all it writes is JSON, on
a normal start and stop and with such a config file. The backend
logs its name, version and architecture once at start.
The health check's uptime keys are now uptime_seconds and
uptime_human; its type and method take the names
GO_HTTP_SERVER_CONVENTIONS.md gives.
Model: opus-5-5
`bin/entrypoint.sh` now runs `netwatch-server prepare-data-dir`, which
refuses a `DATA_DIR` that is not `/data` or a path below it written in
full, then creates `DATA_DIR`, gives `/data` and everything in it to
`netwatch`, and sets mode 750 on `/data` and `DATA_DIR`. Every step goes
through a Go `os.Root` opened on `/data`, and the modes are set on the
opened directories rather than by name, so neither a symbolic link
already there nor one a host process swaps in while the container
starts can make root create or change anything outside `/data`. The
README says which `DATA_DIR` values are accepted.
Model: opus-5-5
When a report would take the report files past DATA_DIR_MAX_BYTES,
reportbuf now deletes the oldest report files until it fits, and does
the same at start when files left by an earlier run are already past
it. A file joins the files that may be deleted, at its place by name,
only once it is completely written, so a file still being written is
never deleted. A report is refused with 507 only when the reports
waiting to be written fill the cap on their own, and then no file is
deleted. The reports of a failed write stop counting, and the part of
its file written is removed. A file whose deletion fails keeps
counting; one already deleted by hand counts as freed.
Model: opus-5-5
backend/script/test runs go test -timeout 30s -race -cover and, if
that fails, runs it again with -v and fails. The root script/test drops
its one 30-second timeout around both halves: from a cold Go build
cache, compiling the tests with -race used it all up. Each half keeps
its own limit. The race detector needs a C compiler: the Dockerfile
builder stage gains gcc and musl-dev, and script/bootstrap installs gcc,
with the C library headers on apt and apk, when gcc is missing; make
build still sets CGO_ENABLED=0. New tests: the health check's answer, a
valid report's answer, a report file's exact lines, and the flush at the
10 MiB threshold. The handlers TestImport stub is gone.
Model: opus-5-5
The architecture is no longer passed in at build time. The Buildarch
variable and field are gone from main and globals, script/build no
longer stamps it in with -X, and the startup and listen log lines
report runtime.GOARCH under the key "arch". The Dockerfile comment and
backend/README.md no longer describe an architecture being stamped in.
Model: opus-5-5
The request log wrote the URL, User-Agent, Referer and other
request-supplied strings with no length limit, and the server accepts
headers up to 1 MiB, so one request could put about 1 MiB per field
into a log line. Every string the request log takes from the request,
including the request ID chi copies from X-Request-Id, is now cut to
the 128-byte bound the report handler already used. That bound and
its helper moved from the handlers package to the logger package so
both use the one copy.
Model: opus-5-5
nginx trusted X-Forwarded-For from every RFC1918 address, so a client
reaching it from one could write a new address on each request and
get a fresh rate-limit allowance. The container's TRUSTED_PROXIES now
names the reverse proxies nginx trusts, none by default.
bin/entrypoint.sh makes each entry a CIDR, checks it with the new
"netwatch-server check-cidr", which runs the server's own
TRUSTED_PROXIES parsing, and writes one set_real_ip_from line per
entry into /etc/nginx/trusted-proxies.conf, which nginx.conf includes.
The backend is started with TRUSTED_PROXIES=127.0.0.1/32, since nginx
is its only client. The viewport test mounts an empty file there.
Model: opus-5-5
Report files were named by a millisecond timestamp and created with
O_EXCL, so two flushes in the same millisecond, such as a flush for
size and the final flush at shutdown, got the same name and the second
failed, losing its reports. Each name now carries a number after the
timestamp that goes up by one for each file the server starts to
write, so names still sort by time and never repeat within a run. A
failed write uses up its number, leaving a gap if the file could not
be created and otherwise a file under that number that may be
incomplete.
Model: opus-5-5
The image's HEALTHCHECK requests /.well-known/healthcheck through
nginx on the port from PORT, so it fails unless both processes answer.
The backend reads PORT and DEBUG with strconv instead of viper, which
turned a bad PORT into 0 and a bad DEBUG into false. Those, and a
BIND_ADDRESS that is not an IP address, now stop the start with an
error naming the variable; the TRUSTED_PROXIES error names it too.
bin/entrypoint.sh also refuses a container PORT outside 1 to 65535,
or 8081, where the backend listens, naming PORT. README.md gains
"Running under upaas". Its first-run steps create the host directory
owned by uid 1000, so the image changes no ownership.
Model: opus-5-5
POST /api/v1/reports stays unauthenticated but is bounded. Each client
address, as the trusted-proxy logic resolves it, may send
REPORTS_PER_MINUTE reports a minute (default 60, counted by
go-chi/httprate over a sliding minute); past that it gets 429 with
Retry-After. reportbuf refuses a report that would take the report
files past DATA_DIR_MAX_BYTES (default 1 GiB), counting the files
already in DATA_DIR and unwritten reports at their uncompressed size;
the handler answers 507. CORS adds nothing unless CORS_ALLOWED_ORIGINS
lists origins. A limit that is not a positive number, or an origin
that is not a plain scheme://host[:port], stops the server from
starting.
Model: opus-5-5
The root Dockerfile builds the only image; Dockerfile.backend is gone.
Its stages: lint, a Go stage that runs the tests and builds
netwatch-server, the node stage, and an nginx runtime. nginx serves
dist/ on 8080 and proxies /api/ and /.well-known/healthcheck to the
backend on 127.0.0.1:8081. bin/entrypoint.sh starts both, turns TERM or
INT into a stop of both, and exits non-zero when either exits on its
own. The backend runs as user netwatch and keeps reports on the /data
volume. New setting BIND_ADDRESS (empty: every interface). STOPSIGNAL is
SIGTERM, since the nginx image's SIGQUIT would miss the entrypoint.
script/docker is the org model verbatim.
Model: opus-5-5
A report the buffer refuses now returns 500 instead of a false `ok`.
Reports reach disk later, so a failed disk write is still answered 200
and shows in the log, and at shutdown as a failed stop with a non-zero
exit. An over-limit body returns 413; malformed JSON stays 400. A
MaxBodyBytes middleware caps every route at 1 MiB; a route group can
only lower that limit. The raw geo blob is no longer logged; client_id,
timestamp and decode error text are cut to 128 bytes before logging.
Panic recovery logs the panic value and stack through slog.
Model: opus-5-5
The old backend/.golangci.yml declared version "2" but used v1 schema
keys, so under v2 it never validated and its thresholds were inert: the
linter ran at defaults. Replace it verbatim with the org-standard file,
repin the Dockerfile.backend lint stage to golangci-lint v2.12.2, and
assert the config's sha256 as the first step of the backend lint target
so it cannot silently drift again -- a local hash check, no network.
The standard config surfaces findings only in the tests: the repeated
IP literals in middleware_test.go become named constants (goconst) and
its request switches to NewRequestWithContext (noctx). reportbuf.go's
gosec suppression gains a plain justification comment. The rest of the
backend, including the fx-based server lifecycle, is already clean.
TODO.md updated.
Model: opus-4-8
The server ran os.Exit at the end of its own goroutine, racing fx's
teardown and sometimes killing the process before reportbuf's OnStop
flushed — silently losing a full flush window of telemetry on every
restart, at exit 0. Shutdown now goes through fx.Shutdowner, so every
OnStop runs in order.
The http.Server is built synchronously in OnStart before the serving
goroutine, so shutdown can no longer race or nil-deref it. A listen
failure exits non-zero via fx.ExitCode(1). reportbuf's OnStop is guarded
by sync.Once. writeTimeout now exceeds the chi per-request budget so that
budget is reachable. Dead startupTime, exitCode, and cancelFunc fields
are gone. A new test asserts a buffered report reaches disk after the
lifecycle stops.
Model: opus-4-8
Add ReadHeaderTimeout and IdleTimeout to the http.Server as named constants beside the existing timeouts. Add a SecurityHeaders middleware (HSTS, a JSON-API CSP of default-src 'none'; frame-ancestors 'none', X-Frame-Options DENY, nosniff, Referrer-Policy, Permissions-Policy), registered before CORS so preflight responses carry it. Resolve the client IP from X-Forwarded-For / X-Real-IP only when the direct peer is in the trusted-proxy allowlist (loopback plus RFC1918 by default, configurable via TRUSTED_PROXIES); an untrusted peer's forwarded headers are ignored. Uses net/netip; no new dependency.
Model: opus-4-8 (implementation and review); claude-fable-5 (merge)
Introduce the Go backend (netwatch-server) with an HTTP API that
accepts telemetry reports and persists them as zstd-compressed JSONL
files. Reports are buffered in memory and flushed to disk when the
buffer reaches 10 MiB or every 60 seconds.