Commit Graph
16 Commits
Author SHA1 Message Date
clawbot d28a59023a Add MIT LICENSE (closes #1)
check / check (push) Successful in 4m5s
Adds LICENSE with the MIT licence sneak chose for routewatch, copied byte for byte from sneak/webhooker's LICENSE: copyright 2026, Jeffrey Paul. The rest of the issue, .editorconfig and the Gitea check workflow, was already on next, and the README already points at LICENSE.

Judgement call: our repos write the copyright holder in several different ways and sneak named no form; webhooker's gives his full name and email address, the most explicit of them.

Model: opus-5-5
2026-09-29 05:05:35 +02:00
clawbot 1ac24669d3 Stop the streamer without sending on closed queues (closes #34)
check / check (push) Successful in 2m38s
Stopping the daemon while the RIS Live feed was flowing could panic with "send on closed channel" and skip the rest of the shutdown. Stop cancels the stream and closes the handler queues under the streamer's write lock, but the read loop checked for a stop only before parsing each line. It now checks again under the read lock it already takes just before handing a message to the queues, so it never sends to a closed queue. Stop also clears its cancel function and returns early when there is none, so a second call no longer closes the queues again. A test forces both cases.

Behaviour change: Stop before Start now does nothing.

Model: opus-5-5
2026-09-28 21:42:33 +02:00
clawbot 6187ac8503 Let the daemon receive docker stop's signal itself (closes #33)
check / check (push) Successful in 2m42s
entrypoint.sh now switches to the routewatch user (UID 1000) with setpriv instead of runuser. setpriv replaces itself with the daemon, so the daemon is the container's main process and receives docker stop's signal itself. runuser stayed in between, passed the signal on and killed the daemon 2 seconds later, so every stop ended with exit 143. The daemon now gets the whole wait the caller allows, up to its own 60-second limit, and a clean stop exits 0. Taking ownership of the state directory and the MALLOC_ARENA_MAX check still run as root first.

Not fixed here: a stop while the feed is flowing can still panic (#34).

Model: opus-5-5
2026-09-28 21:08:33 +02:00
clawbot f2a9e90625 Refuse invalid settings at start, health check follows PORT (closes #31)
check / check (push) Successful in 3m17s
A set but invalid PORT (anything but plain digits from 1 to 65535) or a relative XDG_DATA_HOME now stops the start before the database opens; before, a bad PORT left the daemon running without HTTP. entrypoint.sh refuses a MALLOC_ARENA_MAX that is not a positive whole number, since glibc ignores a bad one silently. The HEALTHCHECK probes the port PORT names, 8080 when unset. PORT is now read in internal/config, so server.New takes the config.

The README gives the real Linux state directory, lists XDG_DATA_HOME, and adds "Running under upaas": port, volume, environment, the 5g memory limit and the health check.

Unverified: the 5g memory limit could not be exercised on the build host.

Model: opus-5-5
2026-09-28 20:08:42 +02:00
clawbot df9e23d503 Serve /api/v1/stats counts from realtime in-memory counters (closes #27)
check / check (push) Successful in 2m36s
Realtime in-memory counters seeded at startup and adjusted on every insert, update and delete; no periodic recompute. Independent review passed: #29 (comment)

model: claude-opus-4-8 (implementation and review); merged by claude-fable-5
2026-09-22 09:41:16 +02:00
clawbot 3898daad4e Add _txlock=immediate so batch writes wait instead of dropping (closes #25)
check / check (push) Failing after 0s
Batch flush paths read before they write, so a deferred transaction starts
as a reader and must upgrade to the write lock on its first INSERT/UPDATE.
When the background maintainer holds the write lock for a WAL checkpoint,
that upgrade fails immediately with "database is locked" and the busy
timeout does not apply, so the batch is dropped. Adding _txlock=immediate
to the DSN makes every transaction take the write lock at BEGIN, so it
waits up to busy_timeout instead of failing.

A regression test drives batch writes against a running checkpoint loop and
fails with "database is locked" without the change.

Model: opus-4-8
2026-09-21 20:12:33 +02:00
clawbot 08f060045b Cap glibc malloc arenas to keep non-Go RSS bounded (closes #23)
check / check (push) Successful in 3m44s
After the earlier memory work, SQLite's live heap is bounded but process
RSS still climbed about 70 MiB/min, nearly all anonymous and outside the
Go runtime and outside SQLite's own accounting, without any SQLITE_NOMEM.
The cause is glibc: the SQLite C library allocates and frees millions of
small page-cache chunks from many threads, and glibc keeps each arena's
freed chunks resident. With arenas uncapped it creates up to eight per
core, so on a large host RSS grows with the core count.

Set MALLOC_ARENA_MAX=2 in the image to bound the retained memory. Writes
are already serialized, so the two-arena cap costs no throughput. Update
the README Memory section.

Model: opus-4-8
2026-09-21 19:29:31 +02:00
clawbot 658aadbb81 Set GOMEMLIMIT in the image and document the memory budget (closes #13)
check / check (push) Failing after 1s
Add ENV GOMEMLIMIT=1536MiB to the runtime stage so the Go runtime keeps a
1.5 GiB soft heap limit. runuser preserves it the way it already does
XDG_DATA_HOME, so the routewatch process inherits it.

Add a Memory section to the README describing the budget (Go 1.5 GiB soft,
SQLite 640 MiB pool cache and 1.5 GiB hard heap, ~0.2 GiB other), the
required 5 GiB container limit, how to override GOMEMLIMIT, what happens at
each limit, and the DEBUG=routewatch System stats line. Every sentence
matches the behaviour already merged to next.

Model: opus-4-8
2026-09-21 17:29:29 +02:00
clawbot efaf79c4e3 Shrink the four handler queues from 100,000 to 20,000 (closes #11)
check / check (push) Successful in 2m52s
Each handler queue held 100,000 message pointers; a message is retained
until the slowest handler drains it, so all four full was a derived worst
case near 800 MiB. Twenty thousand is about four seconds of feed at peak
and caps that at roughly 160 MiB. The streamer already drops rather than
blocks on a full queue, so the smaller bound is safe.

Batch sizes are unchanged. The largest, asnBatchSize, is 30,000 and now
exceeds its queue, but each queued message contributes every ASN in its
path, and every handler also flushes on its own timer regardless of fill,
so batches still flush and no size change is warranted.

Model: opus-4-8
2026-09-21 16:29:39 +02:00
clawbot 63d62b7bc1 Fix two goroutine leaks: stats handlers on timeout, streamer tickers on reconnect (closes #12)
check / check (push) Failing after 1s
The stats handlers ran the database query in a goroutine that sent on an
unbuffered channel. When the 4s request timeout won, nothing received and
the goroutine blocked forever; the status page polls every 2s, so once the
query exceeds the timeout every poll leaked one goroutine. Give both
channels capacity 1 so the send always completes.

The streamer started two ticker goroutines per connection that exited only
with the streamer's lifetime context, leaking two on every reconnect. Scope
them to a per-connection context cancelled when the stream call returns.

Tests force the stats timeout repeatedly and drive many reconnects, then
assert the goroutine count settles back to its starting value. The streamer
gains an internal endpoint field so a test can point it at a local server.

Model: opus-4-8
2026-09-21 16:29:32 +02:00
clawbot cb2374f033 Stop decoding unused Community and Raw fields of RIS messages (closes #9)
check / check (push) Failing after 1s
Community and Raw are decoded from every live message and read by no
handler; Raw is the hex of the whole BGP message. Each parsed message
sits in up to four handler queues, so retaining them is a large share of
queue memory. Tagging both json:"-" keeps them out of the decoded
message while the fields handlers use still decode. A decode test over
docs/message-examples.json confirms both stay empty and the used fields
(Path, Announcements) still populate.

Model: opus-4-8
2026-09-21 16:12:41 +02:00
clawbot 6d8ae7f592 Add .editorconfig and Gitea CI workflow (closes #14)
check / check (push) Failing after 1s
Copy .editorconfig verbatim from sneak/dnswatcher (4-space indents,
LF, trailing-whitespace trim, final newline; tabs for the Makefile).

Add .gitea/workflows/check.yml mirroring dnswatcher: it runs
script/cibuild on push, which builds the Dockerfile that runs
make check, so CI gates every push. The checkout action is pinned by
commit SHA as that repo pins it. No LICENSE is added; that part of
the parent issue awaits an owner decision.

Model: opus-4-8
2026-09-21 16:01:14 +02:00
clawbot 211fdad9c0 Bound peering AS-path map and swap it instead of copying (closes #10)
processPeerings now takes the accumulated AS-path map under the lock and
replaces it with a fresh empty one, so each path is processed once and the
map never grows past a single 30-second interval's traffic. HandleMessage
stops adding new paths once maxTrackedPaths (500000) is reached and counts
the drops, which are logged with each run. The 30-minute prune and its
ticker are removed as dead code under the swap, and the map mutex drops
from RWMutex to Mutex since the read path is gone. RecordPeering already
upserts last_seen, so stored peerings are unchanged.

Model: opus-4-8
2026-09-21 16:01:08 +02:00
clawbot 594e7a504b Bound SQLite memory across the whole connection pool (closes #8)
The 3 GiB page cache and temp_store=MEMORY were set once in Initialize, so
only one pooled connection carried them and the other nine got no busy_timeout,
which drove many "database is locked" errors.

Move the per-connection settings into the DSN so every pooled connection gets a
64 MiB cache (640 MiB worst case over ten connections), synchronous OFF, a
5 s busy_timeout and WAL. Drop those pragmas from Initialize; DISTINCT temp
B-trees now spill to disk. Add process-wide soft (1 GiB) and hard (1.5 GiB)
heap limits; at the hard limit a statement returns SQLITE_NOMEM and the
existing batch paths log and drop, so nothing panics or exits.

Model: opus-4-8
2026-09-21 15:46:34 +02:00
clawbot 54014c88c8 Run make check inside the Docker build with a pinned lint stage (closes #5)
REPO_POLICIES.md requires the container build to run the checks; the
Dockerfile only compiled the binary, so script/cibuild checked nothing. A
new lint stage on the golangci-lint v2.7.2 image runs make fmt-check and
make lint, and the build stage waits for it and then runs make test, so
make docker now fails when formatting, lint or a test fails. All three base
images are pinned by digest with their versions unchanged. The final image
and how it starts are unchanged.

make docker is now the gate to use on hosts where the installed linter is
older than the Go toolchain and make lint cannot run directly.

Model: opus-4-8 (implementation, review); fable-5-1 (summary)
2026-09-21 09:35:16 +02:00
clawbot b0cd884019 Skip live network feed test under -short for deterministic make check (closes #2)
TestRouteWatchLiveFeed streams the live RIPE RIS feed for a few seconds, so
the default test run depended on the network and tripped the race detector.
It now skips under testing.Short(), and script/test passes -short in both
go test lines, so make check runs offline and gives the same result every
time. The README says how to run the live test on demand. No non-test code
changed.

The race itself is not fixed: the test reads the peeringHandler field of
RouteWatch while Run is still setting it. Nothing outside the test reads
that field concurrently.

Model: opus-4-8 (implementation, review); fable-5-1 (summary)
2026-09-21 09:05:33 +02:00