Memory bounded on the live feed (#3): SQLite page cache and heap capped, temporary tables on disk, bounded queues and peering map, two goroutine leaks fixed, GOMEMLIMIT and MALLOC_ARENA_MAX set in the image, a README Memory section.
Batch writes wait for the write lock instead of failing with database is locked.
/api/v1/stats serves its counts from memory and no longer times out on a large database.
Ready for upaas (#31): invalid settings stop the start, the health check follows PORT, a "Running under upaas" README section; the container prepares its own data directory, with no host step (#42).
Clean shutdown: docker stop no longer kills the daemon 2 seconds in (#33) or makes it panic (#34).
Repo policy: an MIT LICENSE (#1), licence and author in the README, a current TODO.md (#38), a .dockerignore (#39).
make check runs offline and inside the container build; base images pinned by digest; a Gitea workflow runs script/cibuild on every push, so merging turns CI on for main.
To deploy:
Set the upaas app up as "Running under upaas" says, including Memory Limit 5g. upaas sets no swap limit.
A 35-hour live-feed run of next at 3898daa peaked at about 1 GiB, with no container memory limit (the build host refuses one). The run under a real 5 GiB limit is yours (#3).
upaas deploys prod, so merging this does not deploy; a deploy is a squash commit from main onto prod. clawbot cannot protect branches; protecting prod is yours.
Model: opus-5-5
Everything on `next` is safe to merge at any time.
What `next` adds to `main`:
- Memory bounded on the live feed (https://git.eeqj.de/sneak/routewatch/issues/3): SQLite page cache and heap capped, temporary tables on disk, bounded queues and peering map, two goroutine leaks fixed, `GOMEMLIMIT` and `MALLOC_ARENA_MAX` set in the image, a README Memory section.
- Batch writes wait for the write lock instead of failing with `database is locked`.
- `/api/v1/stats` serves its counts from memory and no longer times out on a large database.
- Ready for upaas (https://git.eeqj.de/sneak/routewatch/issues/31): invalid settings stop the start, the health check follows `PORT`, a "Running under upaas" README section; the container prepares its own data directory, with no host step (https://git.eeqj.de/sneak/routewatch/issues/42).
- Clean shutdown: `docker stop` no longer kills the daemon 2 seconds in (https://git.eeqj.de/sneak/routewatch/issues/33) or makes it panic (https://git.eeqj.de/sneak/routewatch/issues/34).
- Repo policy: an MIT `LICENSE` (https://git.eeqj.de/sneak/routewatch/issues/1), licence and author in the README, a current `TODO.md` (https://git.eeqj.de/sneak/routewatch/issues/38), a `.dockerignore` (https://git.eeqj.de/sneak/routewatch/issues/39).
- `make check` runs offline and inside the container build; base images pinned by digest; a Gitea workflow runs `script/cibuild` on every push, so merging turns CI on for `main`.
To deploy:
- Set the upaas app up as "Running under upaas" says, including Memory Limit `5g`. upaas sets no swap limit.
- A 35-hour live-feed run of `next` at `3898daa` peaked at about 1 GiB, with no container memory limit (the build host refuses one). The run under a real 5 GiB limit is yours (https://git.eeqj.de/sneak/routewatch/issues/3).
- upaas deploys `prod`, so merging this does not deploy; a deploy is a squash commit from `main` onto `prod`. clawbot cannot protect branches; protecting `prod` is yours.
Model: opus-5-5
clawbot
self-assigned this 2026-09-21 09:05:55 +02:00
TestRouteWatchLiveFeed streams the live RIPE RIS feed for a few seconds, so
the default test run depended on the network and tripped the race detector.
It now skips under testing.Short(), and script/test passes -short in both
go test lines, so make check runs offline and gives the same result every
time. The README says how to run the live test on demand. No non-test code
changed.
The race itself is not fixed: the test reads the peeringHandler field of
RouteWatch while Run is still setting it. Nothing outside the test reads
that field concurrently.
Model: opus-4-8 (implementation, review); fable-5-1 (summary)
REPO_POLICIES.md requires the container build to run the checks; the
Dockerfile only compiled the binary, so script/cibuild checked nothing. A
new lint stage on the golangci-lint v2.7.2 image runs make fmt-check and
make lint, and the build stage waits for it and then runs make test, so
make docker now fails when formatting, lint or a test fails. All three base
images are pinned by digest with their versions unchanged. The final image
and how it starts are unchanged.
make docker is now the gate to use on hosts where the installed linter is
older than the Go toolchain and make lint cannot run directly.
Model: opus-4-8 (implementation, review); fable-5-1 (summary)
The 3 GiB page cache and temp_store=MEMORY were set once in Initialize, so
only one pooled connection carried them and the other nine got no busy_timeout,
which drove many "database is locked" errors.
Move the per-connection settings into the DSN so every pooled connection gets a
64 MiB cache (640 MiB worst case over ten connections), synchronous OFF, a
5 s busy_timeout and WAL. Drop those pragmas from Initialize; DISTINCT temp
B-trees now spill to disk. Add process-wide soft (1 GiB) and hard (1.5 GiB)
heap limits; at the hard limit a statement returns SQLITE_NOMEM and the
existing batch paths log and drop, so nothing panics or exits.
Model: opus-4-8
processPeerings now takes the accumulated AS-path map under the lock and
replaces it with a fresh empty one, so each path is processed once and the
map never grows past a single 30-second interval's traffic. HandleMessage
stops adding new paths once maxTrackedPaths (500000) is reached and counts
the drops, which are logged with each run. The 30-minute prune and its
ticker are removed as dead code under the swap, and the map mutex drops
from RWMutex to Mutex since the read path is gone. RecordPeering already
upserts last_seen, so stored peerings are unchanged.
Model: opus-4-8
Copy .editorconfig verbatim from sneak/dnswatcher (4-space indents,
LF, trailing-whitespace trim, final newline; tabs for the Makefile).
Add .gitea/workflows/check.yml mirroring dnswatcher: it runs
script/cibuild on push, which builds the Dockerfile that runs
make check, so CI gates every push. The checkout action is pinned by
commit SHA as that repo pins it. No LICENSE is added; that part of
the parent issue awaits an owner decision.
Model: opus-4-8
Community and Raw are decoded from every live message and read by no
handler; Raw is the hex of the whole BGP message. Each parsed message
sits in up to four handler queues, so retaining them is a large share of
queue memory. Tagging both json:"-" keeps them out of the decoded
message while the fields handlers use still decode. A decode test over
docs/message-examples.json confirms both stay empty and the used fields
(Path, Announcements) still populate.
Model: opus-4-8
The stats handlers ran the database query in a goroutine that sent on an
unbuffered channel. When the 4s request timeout won, nothing received and
the goroutine blocked forever; the status page polls every 2s, so once the
query exceeds the timeout every poll leaked one goroutine. Give both
channels capacity 1 so the send always completes.
The streamer started two ticker goroutines per connection that exited only
with the streamer's lifetime context, leaking two on every reconnect. Scope
them to a per-connection context cancelled when the stream call returns.
Tests force the stats timeout repeatedly and drive many reconnects, then
assert the goroutine count settles back to its starting value. The streamer
gains an internal endpoint field so a test can point it at a local server.
Model: opus-4-8
Each handler queue held 100,000 message pointers; a message is retained
until the slowest handler drains it, so all four full was a derived worst
case near 800 MiB. Twenty thousand is about four seconds of feed at peak
and caps that at roughly 160 MiB. The streamer already drops rather than
blocks on a full queue, so the smaller bound is safe.
Batch sizes are unchanged. The largest, asnBatchSize, is 30,000 and now
exceeds its queue, but each queued message contributes every ASN in its
path, and every handler also flushes on its own timer regardless of fill,
so batches still flush and no size change is warranted.
Model: opus-4-8
Add ENV GOMEMLIMIT=1536MiB to the runtime stage so the Go runtime keeps a
1.5 GiB soft heap limit. runuser preserves it the way it already does
XDG_DATA_HOME, so the routewatch process inherits it.
Add a Memory section to the README describing the budget (Go 1.5 GiB soft,
SQLite 640 MiB pool cache and 1.5 GiB hard heap, ~0.2 GiB other), the
required 5 GiB container limit, how to override GOMEMLIMIT, what happens at
each limit, and the DEBUG=routewatch System stats line. Every sentence
matches the behaviour already merged to next.
Model: opus-4-8
After the earlier memory work, SQLite's live heap is bounded but process
RSS still climbed about 70 MiB/min, nearly all anonymous and outside the
Go runtime and outside SQLite's own accounting, without any SQLITE_NOMEM.
The cause is glibc: the SQLite C library allocates and frees millions of
small page-cache chunks from many threads, and glibc keeps each arena's
freed chunks resident. With arenas uncapped it creates up to eight per
core, so on a large host RSS grows with the core count.
Set MALLOC_ARENA_MAX=2 in the image to bound the retained memory. Writes
are already serialized, so the two-arena cap costs no throughput. Update
the README Memory section.
Model: opus-4-8
Batch flush paths read before they write, so a deferred transaction starts
as a reader and must upgrade to the write lock on its first INSERT/UPDATE.
When the background maintainer holds the write lock for a WAL checkpoint,
that upgrade fails immediately with "database is locked" and the busy
timeout does not apply, so the batch is dropped. Adding _txlock=immediate
to the DSN makes every transaction take the write lock at BEGIN, so it
waits up to busy_timeout instead of failing.
A regression test drives batch writes against a running checkpoint loop and
fails with "database is locked" without the change.
Model: opus-4-8
Realtime in-memory counters seeded at startup and adjusted on every insert, update and delete; no periodic recompute. Independent review passed: #29 (comment)
model: claude-opus-4-8 (implementation and review); merged by claude-fable-5
A set but invalid PORT (anything but plain digits from 1 to 65535) or a relative XDG_DATA_HOME now stops the start before the database opens; before, a bad PORT left the daemon running without HTTP. entrypoint.sh refuses a MALLOC_ARENA_MAX that is not a positive whole number, since glibc ignores a bad one silently. The HEALTHCHECK probes the port PORT names, 8080 when unset. PORT is now read in internal/config, so server.New takes the config.
The README gives the real Linux state directory, lists XDG_DATA_HOME, and adds "Running under upaas": port, volume, environment, the 5g memory limit and the health check.
Unverified: the 5g memory limit could not be exercised on the build host.
Model: opus-5-5
clawbot
changed title from Next milestone: production memory under 5 GiB to Next milestone: production memory under 5 GiB, ready for upaas2026-09-28 20:09:13 +02:00
entrypoint.sh now switches to the routewatch user (UID 1000) with setpriv instead of runuser. setpriv replaces itself with the daemon, so the daemon is the container's main process and receives docker stop's signal itself. runuser stayed in between, passed the signal on and killed the daemon 2 seconds later, so every stop ended with exit 143. The daemon now gets the whole wait the caller allows, up to its own 60-second limit, and a clean stop exits 0. Taking ownership of the state directory and the MALLOC_ARENA_MAX check still run as root first.
Not fixed here: a stop while the feed is flowing can still panic (#34).
Model: opus-5-5
Stopping the daemon while the RIS Live feed was flowing could panic with "send on closed channel" and skip the rest of the shutdown. Stop cancels the stream and closes the handler queues under the streamer's write lock, but the read loop checked for a stop only before parsing each line. It now checks again under the read lock it already takes just before handing a message to the queues, so it never sends to a closed queue. Stop also clears its cancel function and returns early when there is none, so a second call no longer closes the queues again. A test forces both cases.
Behaviour change: Stop before Start now does nothing.
Model: opus-5-5
Adds LICENSE with the MIT licence sneak chose for routewatch, copied byte for byte from sneak/webhooker's LICENSE: copyright 2026, Jeffrey Paul. The rest of the issue, .editorconfig and the Gitea check workflow, was already on next, and the README already points at LICENSE.
Judgement call: our repos write the copyright holder in several different ways and sneak named no form; webhooker's gives his full name and email address, the most explicit of them.
Model: opus-5-5
Brings README.md and TODO.md in line with REPO_POLICIES.md and with what is on next. The README's first line now names routewatch, what it is, its MIT licence and its author, @sneak; the License section says MIT and links LICENSE; a new Author section names @sneak.
TODO.md's Status, Next Step and Future Steps now describe next as it is: it waits for sneak to merge the milestone PR, after which the upaas deploy and the run under a real 5 GiB limit are his. Completed Steps gains a line for each issue that landed without one.
Docs only. make fmt formats only Go here, so the markdown was wrapped by hand.
Model: opus-5-5
Adds a .dockerignore, which REPO_POLICIES.md lists among the files every repo has. Until now every docker build sent the whole working tree as its build context, and the image's source archive is made from that context. It now keeps out .git, local build and test output, archives, local databases, .env and a local Go workspace. Every tracked file stays in, so the format check, the linter, the tests, the build and the source archive read the same files as before.
Without .git in the context the binary's build info records no git revision; nothing in routewatch reads it.
Model: opus-5-5
sneak's standing rule: the container makes its data directory usable itself, with no step on the host. entrypoint.sh now creates /var/lib/berlin.sneak.app.routewatch if it is missing and stops the start when any step fails (set -euo pipefail); before, a failed cd went on to change the ownership of whatever directory the script was in, and a failed chown still started the daemon. Taking ownership of the directory and switching to the routewatch user through setpriv are unchanged. The README's upaas volume line now says only which path to mount.
The empty-directory and other-uid cases were run by hand on the built image with upaas-style bind mounts, not added as an automated test.
Model: opus-5-5
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Everything on
nextis safe to merge at any time.What
nextadds tomain:GOMEMLIMITandMALLOC_ARENA_MAXset in the image, a README Memory section.database is locked./api/v1/statsserves its counts from memory and no longer times out on a large database.PORT, a "Running under upaas" README section; the container prepares its own data directory, with no host step (#42).docker stopno longer kills the daemon 2 seconds in (#33) or makes it panic (#34).LICENSE(#1), licence and author in the README, a currentTODO.md(#38), a.dockerignore(#39).make checkruns offline and inside the container build; base images pinned by digest; a Gitea workflow runsscript/cibuildon every push, so merging turns CI on formain.To deploy:
5g. upaas sets no swap limit.nextat3898daapeaked at about 1 GiB, with no container memory limit (the build host refuses one). The run under a real 5 GiB limit is yours (#3).prod, so merging this does not deploy; a deploy is a squash commit frommainontoprod. clawbot cannot protect branches; protectingprodis yours.Model: opus-5-5
Next milestone: production memory under 5 GiBto Next milestone: production memory under 5 GiB, ready for upaasView command line instructions
Checkout
From your project repository, check out a new branch and test the changes.