Owner incident + order (sneak, 2026-09-21, chat, verbatim): "i had to kill off routewatch on fsn1app1 because it was using a dozen gigs of memory. figure out how to run it in real internet production while keeping the memory usage under 5GiB"
Definition of done:
A grounded analysis on this issue of where the memory goes on a real internet feed (measured or derived from the code: in-memory route/path state, per-peer buffers, sqlite caches, GC headroom — whatever it actually is), posted for review before implementation.
Implementation of the accepted plan such that steady-state RSS on the real production feed stays under 5 GiB, with headroom for spikes (target well under the cap so the kernel OOM killer and upaas limits never fire).
An enforced ceiling, not just hope: GOMEMLIMIT (or equivalent) set appropriately, plus a documented container memory limit for the upaas deployment; behavior at the limit is graceful (bounded structures / eviction / backpressure), not OOM death.
Verification: a run against the real feed (or a full-table replay) demonstrating the bound, with the measurement method written on this issue.
README/deploy docs updated with the memory requirements and the limit configuration.
model: claude-fable-5
Owner incident + order (sneak, 2026-09-21, chat, verbatim): "i had to kill off routewatch on fsn1app1 because it was using a dozen gigs of memory. figure out how to run it in real internet production while keeping the memory usage under 5GiB"
Definition of done:
- A grounded analysis on this issue of where the memory goes on a real internet feed (measured or derived from the code: in-memory route/path state, per-peer buffers, sqlite caches, GC headroom — whatever it actually is), posted for review before implementation.
- Implementation of the accepted plan such that steady-state RSS on the real production feed stays under 5 GiB, with headroom for spikes (target well under the cap so the kernel OOM killer and upaas limits never fire).
- An enforced ceiling, not just hope: `GOMEMLIMIT` (or equivalent) set appropriately, plus a documented container memory limit for the upaas deployment; behavior at the limit is graceful (bounded structures / eviction / backpressure), not OOM death.
- Verification: a run against the real feed (or a full-table replay) demonstrating the bound, with the measurement method written on this issue.
- README/deploy docs updated with the memory requirements and the limit configuration.
model: claude-fable-5
Manager instruction (top-level sdlc manager, owner directive 2026-09-21): register a recurring session cron now (off-minute, roughly every 15-20 minutes) that wakes you to re-check open issues and PRs, check your spawned workers and reviewers, replace anything wedged, merge passed reviews, and dispatch the next unit; session crons expire after 7 days, re-register before then. Branch workflow addition: if next must ever be frozen, keep spawning workers whose PRs target next2 branched off next; next2 becomes the new next when next merges to main. Both rules are now in the repo-manager agent definition.
model: claude-fable-5
Manager instruction (top-level sdlc manager, owner directive 2026-09-21): register a recurring session cron now (off-minute, roughly every 15-20 minutes) that wakes you to re-check open issues and PRs, check your spawned workers and reviewers, replace anything wedged, merge passed reviews, and dispatch the next unit; session crons expire after 7 days, re-register before then. Branch workflow addition: if `next` must ever be frozen, keep spawning workers whose PRs target `next2` branched off `next`; `next2` becomes the new `next` when `next` merges to `main`. Both rules are now in the repo-manager agent definition.
model: claude-fable-5
Manager instruction (top-level sdlc manager, owner directive 2026-09-21): at most 5 simultaneous worker sessions PER ACCOUNT, all repos combined — issue-to-pr, pr-reviewer, genius-bar and one-off sessions count; managers do not. Before every spawn, count the non-manager sessions on the target account with claudeN agents --json; at 5 or more, do not spawn — wait for your next wake or use the other account if it has a free slot. Excess workers running right now are being stopped by the top-level manager; re-dispatch their units one at a time through free slots, preferring reviews and rework of already-pushed PRs over new implementations.
model: claude-fable-5
Manager instruction (top-level sdlc manager, owner directive 2026-09-21): at most 5 simultaneous worker sessions PER ACCOUNT, all repos combined — issue-to-pr, pr-reviewer, genius-bar and one-off sessions count; managers do not. Before every spawn, count the non-manager sessions on the target account with claudeN agents --json; at 5 or more, do not spawn — wait for your next wake or use the other account if it has a free slot. Excess workers running right now are being stopped by the top-level manager; re-dispatch their units one at a time through free slots, preferring reviews and rework of already-pushed PRs over new implementations.
model: claude-fable-5
Basis: next at ddf0b2f, image built from the repo Dockerfile, run 28 minutes against the live RIS feed on 2026-09-21 from an empty database. Findings were cross-checked by independent verifiers against the code and the vendored go-sqlite3 v1.14.29 / SQLite 3.50.3 source. A 28-minute run shows the ramp, not days of growth; extrapolations are labelled.
1. Measured run
Feed: 4,000-5,100 messages/s, 7.9 M messages, ~723 bytes of JSON per message.
Process VmRSS peak 1.64 GiB, VmHWM 1.86 GiB by minute 30, essentially all anonymous (not file-backed).
Go runtime share (Sys minus HeapReleased): 0.81 GiB peak; heap_sys 777 MiB; heap_alloc swinging 120-580 MiB on a 30-second cycle.
Everything else (RSS minus Go), i.e. SQLite C heap: ~0.82 GiB at 28 minutes and still climbing with the database.
Database file 1.18 GiB at 28 minutes (1.40 GB at minute 30); WAL up to 592 MiB between checkpoints.
Both halves were still rising when the run ended. Nothing plateaued.
2. Where the memory goes
SQLite page cache: up to ~3 GiB of C heap that the Go runtime cannot see (internal/database/database.go:111).
PRAGMA cache_size=-3145728 (3 GiB) and temp_store=MEMORY are run through d.exec -> db.Exec during Initialize, on the single connection db.Ping() opened at :85. cache_size, temp_store, analysis_limit and synchronous are per-connection, so exactly one connection carries the 3 GiB cache, permanently (SetConnMaxLifetime(0) at :94).
The other nine connections of the pool (SetMaxOpenConns(10), :92) use SQLite's default ~1.95 MiB cache: the DSN at :76 is a bare file:PATH with no _cache_size, and go-sqlite3 is not built with SQLITE_DEFAULT_CACHE_SIZE. Total pool ceiling ~3 GiB + ~18 MiB.
3 GiB is a ceiling, not steady state: in WAL mode that connection's cache is discarded whenever another connection has committed since its last read (pager_reset), so the effective bound is min(3 GiB, database size, working set). The freed ~4 KiB chunks return to glibc free lists, not to the kernel, so RSS stays near its high-water mark.
This is malloc memory. GOMEMLIMIT neither sees nor limits it. No mmap_size is set and go-sqlite3 leaves mmap off, so this is not file-backed either (the small -shm wal-index is the only mapping).
Measured: 0.82 GiB of C heap against a 1.18 GiB database. Nothing expires routes (live_routes_v4/v6 unique on (prefix, origin_asn, peer_ip), internal/database/schema.sql:113,127; config.RouteExpirationTimeout at internal/config/config.go:36 is never read), so over days the database grows toward prefixes x peers and this cache fills and stays full.
SQLite temp B-trees for DISTINCT: C heap proportional to the whole IPv6 table, per request, unbounded (internal/database/database.go:1886, :1246).
getIPv6Info and the IPv6 branch of GetASInfoForIPContext select every row of live_routes_v6 joined with asns, DISTINCT, no WHERE.
With temp_store=MEMORY the DISTINCT temp B-tree is held entirely in C heap and is bounded by neither cache_size nor anything else. (Correction to an earlier reading of mine: the sorter does not spill at all under temp_store=MEMORY, and the ORDER BY here can be served by idx_live_routes_v6_mask_length; the temp B-tree, not a sorter PMA, is the consumer.)
Measured at 1.0 M v6 rows: 22.8 s per request; the IPv4 lookup was 18 ms. No RSS rise visible at that database size.
Peering handler AS-path map: largest Go consumer, bounded by time only (internal/routewatch/peeringhandler.go:50).
asPaths keys every distinct AS path of the last 30 minutes by its JSON string; the only removal is the 30-minute prune every 5 minutes.
Measured: 1,118,116 entries at 28 minutes, still growing ~17 k/min; ~184 bytes per entry (derived from heap_alloc lows: 108 MiB at 350 k paths, 246 MiB at 1.1 M). The prune had not yet taken effect; plateau unmeasured, estimated 1.2-2 M entries (220-370 MiB).
Every 30 s processPeerings copies the whole map and calls RecordPeering once per unique peering, each its own transaction under the global Database.mu: 129-150 k transactions per run taking 14-34 s at minute 28, so it runs almost continuously. 110,000 of those calls failed with database is locked during the run. This is what drives the 300 MiB heap_alloc sawtooth, and with default GOGC=100 the runtime holds about twice the live heap.
Handler queues: 4 x 100,000 message pointers, bounded, derived 150-400 MiB when full.
internal/streamer/streamer.go:152; capacities at ashandler.go:15, peerhandler.go:18, prefixhandler.go:19, peeringhandler.go:17.
The same message pointer goes to every interested handler, so a message lives until the slowest handler drains it. Measured high-water marks: PrefixHandler 87,680, ASHandler 80,310, PeerHandler 46,100, PeeringHandler 2,037.
Each parsed message retains two fields no handler reads: Community and Raw (hex of the raw BGP message, present in the live feed), internal/ristypes/ris.go:77,83. Derived retained size 1-2 KiB per message.
The overflow policy is drop, not block: the streamer never blocks on a full queue (non-blocking send plus probabilistic drop above 50% fill, streamer.go:670-716). Measured: PrefixHandler dropped 606,382 of 8.5 M messages, with flushes taking up to 5.3 s on a 1 GB database. That bound works.
Two goroutine leaks, unbounded with uptime.
handleStatusJSON / handleStats (internal/server/handlers.go:183,402) run the stats query in a goroutine that sends on an unbuffered channel; when the 4 s timeout wins, nothing ever receives and the goroutine blocks forever (~4 KiB each). status.html:647 polls every 2 s, so once the stats scans exceed 4 s every poll leaks: derived ~43 k goroutines/day per open status page. Not reached in this run (stats took 1.2 s at 1.4 GB).
Streamer.stream (streamer.go:514,529) starts two ticker goroutines per connection that exit only with the streamer's lifetime context: two leaked per reconnect. Small, permanent.
Smaller and ruled out.
Batch slices are bounded and total ~4 MiB (prefixBatchSize 25000, asnBatchSize 30000, peerBatchSize 10000).
pkg/asinfo (asinfo.go:37-66) decompresses 2.5 MB to 12.4 MB JSON and indexes 130,402 entries: ~15 MiB retained, one time.
JSONResponseMiddleware buffers every response and re-decodes it (internal/server/middleware.go:54); router-wide, skipping only / and /status, so the HTML pages are buffered too. Per-request spike only; responses here are small.
No GOMEMLIMIT or GOGC in Dockerfile or entrypoint.sh, no debug.SetMemoryLimit, no SQLite heap limit, no container memory limit anywhere in the repo.
Memory telemetry only when DEBUG contains routewatch (internal/routewatch/cli.go:26-54), and the image does not set it, so production logs carried no memory figures at all. /api/v1/stats reports Go heap only; nothing reports C-side memory.
What the run does not explain. The bounded consumers sum to roughly 3 GiB C plus 1 GiB Go at steady state. The rest of the ~12 GiB must come from the goroutine leaks, from glibc heap fragmentation (glibc base image, CGO_ENABLED=1, MALLOC_ARENA_MAX unset, SQLite on plain malloc, continuous 4 KiB page churn across threads), or from long periods with all queues full on a slow database. I cannot apportion that from 28 minutes.
3. Plan
Budget inside a 5 GiB container limit: GOMEMLIMIT 1.5 GiB soft, SQLite at most 640 MiB of page cache across the pool plus a 1.5 GiB hard heap limit, other runtime ~0.2 GiB. Hard ceilings sum to ~3.2 GiB, leaving headroom for spikes and the kernel file cache. Expected steady state 1.5-2.5 GiB. Each unit is one commit-sized PR off next; every definition of done includes make check green, which is why U1 is first.
U1. Land #2 first.TestRouteWatchLiveFeed hits the live network and races, so make check is nondeterministic and "green" would mean nothing for the units below. Test-only change, small. Files: internal/routewatch/app_integration_test.go, script/test, README.md. Blocks everything else.
U2. Bound SQLite across the whole pool. Files: internal/database/database.go, internal/database/database_test.go.
Move the per-connection settings into the DSN so every pooled connection gets them: file:PATH?_cache_size=-65536&_synchronous=OFF (64 MiB x 10 connections = 640 MiB worst case).
Do not put the existing -3145728 value in the DSN: applied per connection that is 30 GiB.
Remove cache_size and temp_store=MEMORY from the pragma list in Initialize; temp B-trees then spill to disk instead of growing in C heap.
Add PRAGMA soft_heap_limit=1073741824 and PRAGMA hard_heap_limit=1610612736 (process-wide, so once is enough). At the soft limit the cache recycles instead of allocating; at the hard limit a statement fails with SQLITE_NOMEM, the existing error paths log and drop that batch, the process lives.
Done when: a test holding several connections open asserts PRAGMA cache_size is -65536 on each; a 10-minute live run shows RSS minus Go under 1 GiB with a database over 1 GB. Expect more backpressure drops; that is intended.
U3. Drop dead weight from the parsed message. Tag Community and Rawjson:"-" in internal/ristypes/ris.go; no handler reads either. Done when a decode test over docs/message-examples.json leaves both empty.
U4. Bound the peering map and stop copying it.internal/routewatch/peeringhandler.go.
processPeerings swaps in a fresh map under the lock instead of copying, so each path is processed once; add maxTrackedPaths (500,000) and drop-with-counter when full; the time-based prune goes.
RecordPeering already upserts last_seen, so semantics are unchanged; calls per run fall from the whole 30-minute window to the paths new in the last 30 s, which should also clear most of the database is locked errors.
Done when: tests show the map empty after ProcessPeeringsNow and the cap enforced; a live run shows no 300 MiB heap_alloc sawtooth and paths in the tens of thousands.
U5. Right-size the handler queues. Set all four to 20,000 (about 4 s of feed, twice the largest batch): derived worst case 160 MiB instead of 800 MiB. Touches peeringhandler.go, so after U4.
U6. Fix the two goroutine leaks. Buffer statsChan and errChan (capacity 1) in internal/server/handlers.go; give the two ticker goroutines a per-connection context in internal/streamer/streamer.go:465. Done when goroutines stays flat under a 2-second poll with a forced stats timeout and across a reconnect.
U7. Enforce the ceiling and document it.ENV GOMEMLIMIT=1536MiB in the Dockerfile (inherited through runuser like XDG_DATA_HOME already is); a Memory section in README.md covering the steady-state figure, the required container limit (--memory=5g --memory-swap=5g or the upaas equivalent), how to override GOMEMLIMIT, the SQLite budget, what happens at each limit, and that DEBUG=routewatch emits the System stats line every 60 s. After U1 (both touch README.md).
U8. Verification run per section 4; results posted here. Touches TODO.md only.
Serialisation: U4 before U5 (same file); U1 before U7 (same file). Everything else parallel after U1.
Deliberately out of scope as features, not bounds: route expiry, batching RecordPeering into one transaction, an index-backed IPv6 lookup, cheaper stats queries. After U2 those affect speed and disk, not the memory ceiling.
4. Verification method for U8
Build from next with the repo Dockerfile; run docker run -d --memory=5g --memory-swap=5g -e DEBUG=routewatch with an empty database against the live feed.
Hold one client polling /api/v1/stats every 2 s for the whole run (what the status page does); once an hour request /ip/2001:4860:4860::8888 and /as/3356.
Sample every 60 s: VmHWM, VmRSS, RssAnon, RssFile from /proc/PID/status of the routewatch process; the container's cgroup memory.current; from /api/v1/stats the Go heap figures, goroutines, per-handler queue length, high-water mark and drop count, live_routes, database_size_bytes; sizes of db.sqlite and db.sqlite-wal; and keep the container log, whose System stats line gives sys_mb and heap_released_mb so Go RSS and C RSS can be separated.
Duration at least 24 hours, so the database passes 3 GiB and the peering map and stats path reach steady state. The 28-minute run above is the baseline to compare against.
Pass: VmHWM under 3.5 GiB throughout with no upward trend in the final 12 hours; goroutines flat; docker inspect shows OOMKilled=false and RestartCount=0; RSS minus Go under 1.2 GiB. Post the table of samples and the exact commands here.
5. Disclosures
This measurement ran without --memory: the host's cgroup root does not delegate the memory controller to Docker, so the limit was refused; RSS came from /proc with a 12 GiB watchdog on my own container. U8 needs a host where --memory works.
28 minutes cannot show days-long growth; every extrapolation above says so.
The ~12 GiB is not fully explained by bounded consumers; leaks and glibc fragmentation are the candidates and both are unmeasured.
Per-message retained size (1-2 KiB) and per-path cost (~184 B) are derived from heap figures, not measured directly.
soft_heap_limit throttling of the page cache is read from the SQLite source, not measured; hard_heap_limit is the enforced backstop.
The IPv6 lookup timing (22.8 s at 1.0 M v6 rows) is a single sample.
Whether upaas can set a container memory limit is unknown to me; the README will state the requirement and the owner applies it.
I corrected my own initial reading of the IPv6 query: the consumer is the DISTINCT temp B-tree under temp_store=MEMORY, not a cache_size-derived sorter threshold.
Noticed, out of scope: GetStatsContext logs Failed to get route timestamps on every call (MIN(last_updated) scanned into *time.Time); the Dockerfile does not run make check as policy requires, so nothing gates today except the local pre-commit hook.
On #2: yes, it blocks gating these units and should land first.
model: claude-fable-5
## Memory analysis and plan to run under 5 GiB
Basis: `next` at `ddf0b2f`, image built from the repo `Dockerfile`, run 28 minutes against the live RIS feed on 2026-09-21 from an empty database. Findings were cross-checked by independent verifiers against the code and the vendored `go-sqlite3` v1.14.29 / SQLite 3.50.3 source. A 28-minute run shows the ramp, not days of growth; extrapolations are labelled.
### 1. Measured run
- Feed: 4,000-5,100 messages/s, 7.9 M messages, ~723 bytes of JSON per message.
- Process `VmRSS` peak 1.64 GiB, `VmHWM` 1.86 GiB by minute 30, essentially all anonymous (not file-backed).
- Go runtime share (`Sys` minus `HeapReleased`): 0.81 GiB peak; `heap_sys` 777 MiB; `heap_alloc` swinging 120-580 MiB on a 30-second cycle.
- Everything else (RSS minus Go), i.e. SQLite C heap: ~0.82 GiB at 28 minutes and still climbing with the database.
- Database file 1.18 GiB at 28 minutes (1.40 GB at minute 30); WAL up to 592 MiB between checkpoints.
- Both halves were still rising when the run ended. Nothing plateaued.
### 2. Where the memory goes
- **SQLite page cache: up to ~3 GiB of C heap that the Go runtime cannot see (`internal/database/database.go:111`).**
- `PRAGMA cache_size=-3145728` (3 GiB) and `temp_store=MEMORY` are run through `d.exec` -> `db.Exec` during `Initialize`, on the single connection `db.Ping()` opened at `:85`. `cache_size`, `temp_store`, `analysis_limit` and `synchronous` are per-connection, so exactly one connection carries the 3 GiB cache, permanently (`SetConnMaxLifetime(0)` at `:94`).
- The other nine connections of the pool (`SetMaxOpenConns(10)`, `:92`) use SQLite's default ~1.95 MiB cache: the DSN at `:76` is a bare `file:PATH` with no `_cache_size`, and go-sqlite3 is not built with `SQLITE_DEFAULT_CACHE_SIZE`. Total pool ceiling ~3 GiB + ~18 MiB.
- 3 GiB is a ceiling, not steady state: in WAL mode that connection's cache is discarded whenever another connection has committed since its last read (`pager_reset`), so the effective bound is min(3 GiB, database size, working set). The freed ~4 KiB chunks return to glibc free lists, not to the kernel, so RSS stays near its high-water mark.
- This is `malloc` memory. `GOMEMLIMIT` neither sees nor limits it. No `mmap_size` is set and go-sqlite3 leaves mmap off, so this is not file-backed either (the small `-shm` wal-index is the only mapping).
- Measured: 0.82 GiB of C heap against a 1.18 GiB database. Nothing expires routes (`live_routes_v4/v6` unique on `(prefix, origin_asn, peer_ip)`, `internal/database/schema.sql:113,127`; `config.RouteExpirationTimeout` at `internal/config/config.go:36` is never read), so over days the database grows toward prefixes x peers and this cache fills and stays full.
- **SQLite temp B-trees for `DISTINCT`: C heap proportional to the whole IPv6 table, per request, unbounded (`internal/database/database.go:1886`, `:1246`).**
- `getIPv6Info` and the IPv6 branch of `GetASInfoForIPContext` select every row of `live_routes_v6` joined with `asns`, `DISTINCT`, no `WHERE`.
- With `temp_store=MEMORY` the `DISTINCT` temp B-tree is held entirely in C heap and is bounded by neither `cache_size` nor anything else. (Correction to an earlier reading of mine: the sorter does not spill at all under `temp_store=MEMORY`, and the `ORDER BY` here can be served by `idx_live_routes_v6_mask_length`; the temp B-tree, not a sorter PMA, is the consumer.)
- Measured at 1.0 M v6 rows: 22.8 s per request; the IPv4 lookup was 18 ms. No RSS rise visible at that database size.
- **Peering handler AS-path map: largest Go consumer, bounded by time only (`internal/routewatch/peeringhandler.go:50`).**
- `asPaths` keys every distinct AS path of the last 30 minutes by its JSON string; the only removal is the 30-minute prune every 5 minutes.
- Measured: 1,118,116 entries at 28 minutes, still growing ~17 k/min; ~184 bytes per entry (derived from `heap_alloc` lows: 108 MiB at 350 k paths, 246 MiB at 1.1 M). The prune had not yet taken effect; plateau unmeasured, estimated 1.2-2 M entries (220-370 MiB).
- Every 30 s `processPeerings` copies the whole map and calls `RecordPeering` once per unique peering, each its own transaction under the global `Database.mu`: 129-150 k transactions per run taking 14-34 s at minute 28, so it runs almost continuously. 110,000 of those calls failed with `database is locked` during the run. This is what drives the 300 MiB `heap_alloc` sawtooth, and with default `GOGC=100` the runtime holds about twice the live heap.
- **Handler queues: 4 x 100,000 message pointers, bounded, derived 150-400 MiB when full.**
- `internal/streamer/streamer.go:152`; capacities at `ashandler.go:15`, `peerhandler.go:18`, `prefixhandler.go:19`, `peeringhandler.go:17`.
- The same message pointer goes to every interested handler, so a message lives until the slowest handler drains it. Measured high-water marks: PrefixHandler 87,680, ASHandler 80,310, PeerHandler 46,100, PeeringHandler 2,037.
- Each parsed message retains two fields no handler reads: `Community` and `Raw` (hex of the raw BGP message, present in the live feed), `internal/ristypes/ris.go:77,83`. Derived retained size 1-2 KiB per message.
- The overflow policy is drop, not block: the streamer never blocks on a full queue (non-blocking send plus probabilistic drop above 50% fill, `streamer.go:670-716`). Measured: PrefixHandler dropped 606,382 of 8.5 M messages, with flushes taking up to 5.3 s on a 1 GB database. That bound works.
- **Two goroutine leaks, unbounded with uptime.**
- `handleStatusJSON` / `handleStats` (`internal/server/handlers.go:183,402`) run the stats query in a goroutine that sends on an unbuffered channel; when the 4 s timeout wins, nothing ever receives and the goroutine blocks forever (~4 KiB each). `status.html:647` polls every 2 s, so once the stats scans exceed 4 s every poll leaks: derived ~43 k goroutines/day per open status page. Not reached in this run (stats took 1.2 s at 1.4 GB).
- `Streamer.stream` (`streamer.go:514,529`) starts two ticker goroutines per connection that exit only with the streamer's lifetime context: two leaked per reconnect. Small, permanent.
- **Smaller and ruled out.**
- Batch slices are bounded and total ~4 MiB (`prefixBatchSize` 25000, `asnBatchSize` 30000, `peerBatchSize` 10000).
- `pkg/asinfo` (`asinfo.go:37-66`) decompresses 2.5 MB to 12.4 MB JSON and indexes 130,402 entries: ~15 MiB retained, one time.
- `JSONResponseMiddleware` buffers every response and re-decodes it (`internal/server/middleware.go:54`); router-wide, skipping only `/` and `/status`, so the HTML pages are buffered too. Per-request spike only; responses here are small.
- `Streamer.bgpPeers`, handler metrics, WHOIS state: negligible.
- **No ceiling of any kind exists today.**
- No `GOMEMLIMIT` or `GOGC` in `Dockerfile` or `entrypoint.sh`, no `debug.SetMemoryLimit`, no SQLite heap limit, no container memory limit anywhere in the repo.
- Memory telemetry only when `DEBUG` contains `routewatch` (`internal/routewatch/cli.go:26-54`), and the image does not set it, so production logs carried no memory figures at all. `/api/v1/stats` reports Go heap only; nothing reports C-side memory.
- **What the run does not explain.** The bounded consumers sum to roughly 3 GiB C plus 1 GiB Go at steady state. The rest of the ~12 GiB must come from the goroutine leaks, from glibc heap fragmentation (glibc base image, `CGO_ENABLED=1`, `MALLOC_ARENA_MAX` unset, SQLite on plain `malloc`, continuous 4 KiB page churn across threads), or from long periods with all queues full on a slow database. I cannot apportion that from 28 minutes.
### 3. Plan
Budget inside a 5 GiB container limit: `GOMEMLIMIT` 1.5 GiB soft, SQLite at most 640 MiB of page cache across the pool plus a 1.5 GiB hard heap limit, other runtime ~0.2 GiB. Hard ceilings sum to ~3.2 GiB, leaving headroom for spikes and the kernel file cache. Expected steady state 1.5-2.5 GiB. Each unit is one commit-sized PR off `next`; every definition of done includes `make check` green, which is why U1 is first.
- **U1. Land https://git.eeqj.de/sneak/routewatch/issues/2 first.** `TestRouteWatchLiveFeed` hits the live network and races, so `make check` is nondeterministic and "green" would mean nothing for the units below. Test-only change, small. Files: `internal/routewatch/app_integration_test.go`, `script/test`, `README.md`. Blocks everything else.
- **U2. Bound SQLite across the whole pool.** Files: `internal/database/database.go`, `internal/database/database_test.go`.
- Move the per-connection settings into the DSN so every pooled connection gets them: `file:PATH?_cache_size=-65536&_synchronous=OFF` (64 MiB x 10 connections = 640 MiB worst case).
- Do not put the existing `-3145728` value in the DSN: applied per connection that is 30 GiB.
- Remove `cache_size` and `temp_store=MEMORY` from the pragma list in `Initialize`; temp B-trees then spill to disk instead of growing in C heap.
- Add `PRAGMA soft_heap_limit=1073741824` and `PRAGMA hard_heap_limit=1610612736` (process-wide, so once is enough). At the soft limit the cache recycles instead of allocating; at the hard limit a statement fails with `SQLITE_NOMEM`, the existing error paths log and drop that batch, the process lives.
- Done when: a test holding several connections open asserts `PRAGMA cache_size` is -65536 on each; a 10-minute live run shows RSS minus Go under 1 GiB with a database over 1 GB. Expect more backpressure drops; that is intended.
- **U3. Drop dead weight from the parsed message.** Tag `Community` and `Raw` `json:"-"` in `internal/ristypes/ris.go`; no handler reads either. Done when a decode test over `docs/message-examples.json` leaves both empty.
- **U4. Bound the peering map and stop copying it.** `internal/routewatch/peeringhandler.go`.
- `processPeerings` swaps in a fresh map under the lock instead of copying, so each path is processed once; add `maxTrackedPaths` (500,000) and drop-with-counter when full; the time-based prune goes.
- `RecordPeering` already upserts `last_seen`, so semantics are unchanged; calls per run fall from the whole 30-minute window to the paths new in the last 30 s, which should also clear most of the `database is locked` errors.
- Done when: tests show the map empty after `ProcessPeeringsNow` and the cap enforced; a live run shows no 300 MiB `heap_alloc` sawtooth and `paths` in the tens of thousands.
- **U5. Right-size the handler queues.** Set all four to 20,000 (about 4 s of feed, twice the largest batch): derived worst case 160 MiB instead of 800 MiB. Touches `peeringhandler.go`, so after U4.
- **U6. Fix the two goroutine leaks.** Buffer `statsChan` and `errChan` (capacity 1) in `internal/server/handlers.go`; give the two ticker goroutines a per-connection context in `internal/streamer/streamer.go:465`. Done when `goroutines` stays flat under a 2-second poll with a forced stats timeout and across a reconnect.
- **U7. Enforce the ceiling and document it.** `ENV GOMEMLIMIT=1536MiB` in the `Dockerfile` (inherited through `runuser` like `XDG_DATA_HOME` already is); a Memory section in `README.md` covering the steady-state figure, the required container limit (`--memory=5g --memory-swap=5g` or the upaas equivalent), how to override `GOMEMLIMIT`, the SQLite budget, what happens at each limit, and that `DEBUG=routewatch` emits the `System stats` line every 60 s. After U1 (both touch `README.md`).
- **U8. Verification run** per section 4; results posted here. Touches `TODO.md` only.
- Serialisation: U4 before U5 (same file); U1 before U7 (same file). Everything else parallel after U1.
- Deliberately out of scope as features, not bounds: route expiry, batching `RecordPeering` into one transaction, an index-backed IPv6 lookup, cheaper stats queries. After U2 those affect speed and disk, not the memory ceiling.
### 4. Verification method for U8
- Build from `next` with the repo `Dockerfile`; run `docker run -d --memory=5g --memory-swap=5g -e DEBUG=routewatch` with an empty database against the live feed.
- Hold one client polling `/api/v1/stats` every 2 s for the whole run (what the status page does); once an hour request `/ip/2001:4860:4860::8888` and `/as/3356`.
- Sample every 60 s: `VmHWM`, `VmRSS`, `RssAnon`, `RssFile` from `/proc/PID/status` of the `routewatch` process; the container's cgroup `memory.current`; from `/api/v1/stats` the Go heap figures, `goroutines`, per-handler queue length, high-water mark and drop count, `live_routes`, `database_size_bytes`; sizes of `db.sqlite` and `db.sqlite-wal`; and keep the container log, whose `System stats` line gives `sys_mb` and `heap_released_mb` so Go RSS and C RSS can be separated.
- Duration at least 24 hours, so the database passes 3 GiB and the peering map and stats path reach steady state. The 28-minute run above is the baseline to compare against.
- Pass: `VmHWM` under 3.5 GiB throughout with no upward trend in the final 12 hours; `goroutines` flat; `docker inspect` shows `OOMKilled=false` and `RestartCount=0`; RSS minus Go under 1.2 GiB. Post the table of samples and the exact commands here.
### 5. Disclosures
- This measurement ran without `--memory`: the host's cgroup root does not delegate the memory controller to Docker, so the limit was refused; RSS came from `/proc` with a 12 GiB watchdog on my own container. U8 needs a host where `--memory` works.
- 28 minutes cannot show days-long growth; every extrapolation above says so.
- The ~12 GiB is not fully explained by bounded consumers; leaks and glibc fragmentation are the candidates and both are unmeasured.
- Per-message retained size (1-2 KiB) and per-path cost (~184 B) are derived from heap figures, not measured directly.
- `soft_heap_limit` throttling of the page cache is read from the SQLite source, not measured; `hard_heap_limit` is the enforced backstop.
- The IPv6 lookup timing (22.8 s at 1.0 M v6 rows) is a single sample.
- Whether upaas can set a container memory limit is unknown to me; the README will state the requirement and the owner applies it.
- I corrected my own initial reading of the IPv6 query: the consumer is the `DISTINCT` temp B-tree under `temp_store=MEMORY`, not a `cache_size`-derived sorter threshold.
- Noticed, out of scope: `GetStatsContext` logs `Failed to get route timestamps` on every call (`MIN(last_updated)` scanned into `*time.Time`); the `Dockerfile` does not run `make check` as policy requires, so nothing gates today except the local pre-commit hook.
- On https://git.eeqj.de/sneak/routewatch/issues/2: yes, it blocks gating these units and should land first.
model: claude-fable-5
Plan accepted as posted above, with one addition: busy_timeout is also per-connection and today reaches only one of the ten pooled connections, which likely explains the database is locked errors; it moves into the DSN with the cache setting.
U1 has already merged. The remaining units are filed as implementer's briefs:
U7: #13 (after U2 and U5, so the README describes merged behaviour)
U8, the 24-hour verification run, is reported on this issue. It needs a host where docker run --memory is honoured; the build host refuses it.
Model: fable-5-1
Plan accepted as posted above, with one addition: `busy_timeout` is also per-connection and today reaches only one of the ten pooled connections, which likely explains the `database is locked` errors; it moves into the DSN with the cache setting.
U1 has already merged. The remaining units are filed as implementer's briefs:
- U2: https://git.eeqj.de/sneak/routewatch/issues/8
- U3: https://git.eeqj.de/sneak/routewatch/issues/9
- U4: https://git.eeqj.de/sneak/routewatch/issues/10
- U5: https://git.eeqj.de/sneak/routewatch/issues/11 (after U4, same file)
- U6: https://git.eeqj.de/sneak/routewatch/issues/12
- U7: https://git.eeqj.de/sneak/routewatch/issues/13 (after U2 and U5, so the README describes merged behaviour)
- U8, the 24-hour verification run, is reported on this issue. It needs a host where `docker run --memory` is honoured; the build host refuses it.
Model: fable-5-1
All code units except #13 (GOMEMLIMIT and the README section, in progress) have merged to next.
Needs a decision: where the 24-hour verification run happens. The build host refuses docker run --memory, so the 5 GiB limit cannot be tested there. Options: (a) deploy next to fsn1app1 with a 5 GiB container limit and DEBUG=routewatch, and take the samples there, which is the real target anyway; (b) give clawbot access to another host where Docker honours --memory; (c) run on the build host without the container limit, watching RSS from /proc with a watchdog, which shows the bound but not the behaviour at the limit. Recommendation: (a). Until answered I proceed with (c) once #13 has merged, so a 24-hour measurement exists either way.
Model: fable-5-1
All code units except https://git.eeqj.de/sneak/routewatch/issues/13 (`GOMEMLIMIT` and the README section, in progress) have merged to `next`.
Needs a decision: where the 24-hour verification run happens. The build host refuses `docker run --memory`, so the 5 GiB limit cannot be tested there. Options: (a) deploy `next` to fsn1app1 with a 5 GiB container limit and `DEBUG=routewatch`, and take the samples there, which is the real target anyway; (b) give clawbot access to another host where Docker honours `--memory`; (c) run on the build host without the container limit, watching RSS from `/proc` with a watchdog, which shows the bound but not the behaviour at the limit. Recommendation: (a). Until answered I proceed with (c) once https://git.eeqj.de/sneak/routewatch/issues/13 has merged, so a 24-hour measurement exists either way.
Model: fable-5-1
sneak
was assigned by clawbot2026-09-21 16:30:10 +02:00
All code units have merged; next is at 658aadb. The 24-hour verification run started 2026-09-21 15:32 UTC, using option (c) from my previous comment until you say otherwise.
Method:
Image built from next at 658aadb with the repo Dockerfile; docker run -d -e DEBUG=routewatch -p 127.0.0.1:18643:8080 with an empty database on a named volume, live RIS feed.
Every 60 s: VmHWM, VmRSS, RssAnon, RssFile from /proc/PID/status of the routewatch process; sizes of db.sqlite and db.sqlite-wal; RestartCount and OOMKilled from docker inspect; the full /api/v1/stats response (Go heap, goroutines, queue lengths, high-water marks, drops, route counts). The container log keeps the System stats lines.
One client requests /api/v1/stats every 2 s for the whole run, as an open status page does; once an hour it requests /ip/2001:4860:4860::8888 and /as/3356.
Pass criteria as in section 4 of the plan: VmHWM under 3.5 GiB throughout, no upward trend in the final 12 hours, goroutines flat, no restart.
Disclosures:
No container memory limit: docker run --memory=5g fails on this host with cannot enter cgroupv2 "/sys/fs/cgroup/docker" with domain controllers. A watchdog stops the container if the process passes 6 GiB RSS. Behaviour at the container limit is therefore not tested.
This host appears to be fsn1app1 itself (its storage pool is named pool.fsn1app1). If so, production cannot apply a container memory limit here either until the host's cgroup setup is fixed, and GOMEMLIMIT plus the SQLite heap limit are the only enforced ceilings. That is yours to fix or accept; it changes nothing in the code.
Model: fable-5-1
All code units have merged; `next` is at `658aadb`. The 24-hour verification run started 2026-09-21 15:32 UTC, using option (c) from my previous comment until you say otherwise.
Method:
- Image built from `next` at `658aadb` with the repo `Dockerfile`; `docker run -d -e DEBUG=routewatch -p 127.0.0.1:18643:8080` with an empty database on a named volume, live RIS feed.
- Every 60 s: `VmHWM`, `VmRSS`, `RssAnon`, `RssFile` from `/proc/PID/status` of the `routewatch` process; sizes of `db.sqlite` and `db.sqlite-wal`; `RestartCount` and `OOMKilled` from `docker inspect`; the full `/api/v1/stats` response (Go heap, goroutines, queue lengths, high-water marks, drops, route counts). The container log keeps the `System stats` lines.
- One client requests `/api/v1/stats` every 2 s for the whole run, as an open status page does; once an hour it requests `/ip/2001:4860:4860::8888` and `/as/3356`.
- Pass criteria as in section 4 of the plan: `VmHWM` under 3.5 GiB throughout, no upward trend in the final 12 hours, goroutines flat, no restart.
Disclosures:
- No container memory limit: `docker run --memory=5g` fails on this host with `cannot enter cgroupv2 "/sys/fs/cgroup/docker" with domain controllers`. A watchdog stops the container if the process passes 6 GiB RSS. Behaviour at the container limit is therefore not tested.
- This host appears to be fsn1app1 itself (its storage pool is named `pool.fsn1app1`). If so, production cannot apply a container memory limit here either until the host's cgroup setup is fixed, and `GOMEMLIMIT` plus the SQLite heap limit are the only enforced ceilings. That is yours to fix or accept; it changes nothing in the code.
Model: fable-5-1
For the top-level manager: all fixes are on next (3898daa), but I cannot start the final 24-hour run. My image build for it has been stopped twice by the session host with "system is running low on memory" (about 74 GiB was available each time), as were two earlier container builds today. Worker sessions build the same image without trouble, so this affects builds started from the manager session only. I retry on each wake; if it keeps failing I will have a worker session build the image and start the run.
Baseline for the record (old tree 658aadb, no arena cap, 2 h 40 min): RSS 2.2 GiB at 30 min, 2.9 at 1 h, 3.3 at 1.5 h, 3.5 at 2 h, 3.6 at 2.5 h and still rising; database 3.4 GiB at the end; Go side flat near 150 MiB throughout. That run is stopped and failed the 3.5 GiB criterion, which is what #23 then fixed.
Model: fable-5-1
For the top-level manager: all fixes are on `next` (`3898daa`), but I cannot start the final 24-hour run. My image build for it has been stopped twice by the session host with "system is running low on memory" (about 74 GiB was available each time), as were two earlier container builds today. Worker sessions build the same image without trouble, so this affects builds started from the manager session only. I retry on each wake; if it keeps failing I will have a worker session build the image and start the run.
Baseline for the record (old tree `658aadb`, no arena cap, 2 h 40 min): RSS 2.2 GiB at 30 min, 2.9 at 1 h, 3.3 at 1.5 h, 3.5 at 2 h, 3.6 at 2.5 h and still rising; database 3.4 GiB at the end; Go side flat near 150 MiB throughout. That run is stopped and failed the 3.5 GiB criterion, which is what https://git.eeqj.de/sneak/routewatch/issues/23 then fixed.
Model: fable-5-1
The 24-hour verification run on the final tree (next at 3898daa, all fixes in) started 2026-09-21 18:34 UTC; the build problem from my previous comment is worked around by building from a worker session. Method and pass criteria unchanged from #3 (comment) (container routewatch-verify-3898daa, no container memory limit, 6 GiB watchdog). Result due after 2026-09-22 18:35 UTC.
Model: fable-5-1
The 24-hour verification run on the final tree (`next` at `3898daa`, all fixes in) started 2026-09-21 18:34 UTC; the build problem from my previous comment is worked around by building from a worker session. Method and pass criteria unchanged from https://git.eeqj.de/sneak/routewatch/issues/3#issuecomment-97356 (container `routewatch-verify-3898daa`, no container memory limit, 6 GiB watchdog). Result due after 2026-09-22 18:35 UTC.
Model: fable-5-1
Result of the 24-hour verification run, read by the top-level manager on 2026-09-23 06:04 UTC (the routewatch manager is paused, so nobody had read it). Container routewatch-verify-3898daa (tree 3898daa, option (c): build host, no container memory limit) has run 35.5 hours on the live RIS feed with no restarts and no out-of-memory kill; its health check reports the database and the live feed as ok. The kernel's peak-memory counter for the routewatch process (VmHWM, which covers the whole run) reads 1,031,048 kB, about 0.98 GiB; current VmRSS is 1,005,112 kB, of which 982,704 kB is anonymous memory. That is about a fifth of the 5 GiB target; the old tree reached 3.6 GiB within 2.5 hours on this host and about 12 GiB in production. The per-minute samples the method describes were not posted, and /api/v1/stats now answers HTTP 500 after 4 seconds (filed as #30). Still open for sneak: whether to test under a real 5 GiB container limit on fsn1app1, per #3 (comment) (recommendation (a)). The container is left running.
Model: opus-5-5
Result of the 24-hour verification run, read by the top-level manager on 2026-09-23 06:04 UTC (the routewatch manager is paused, so nobody had read it). Container `routewatch-verify-3898daa` (tree `3898daa`, option (c): build host, no container memory limit) has run 35.5 hours on the live RIS feed with no restarts and no out-of-memory kill; its health check reports the database and the live feed as ok. The kernel's peak-memory counter for the `routewatch` process (`VmHWM`, which covers the whole run) reads 1,031,048 kB, about 0.98 GiB; current `VmRSS` is 1,005,112 kB, of which 982,704 kB is anonymous memory. That is about a fifth of the 5 GiB target; the old tree reached 3.6 GiB within 2.5 hours on this host and about 12 GiB in production. The per-minute samples the method describes were not posted, and `/api/v1/stats` now answers HTTP 500 after 4 seconds (filed as https://git.eeqj.de/sneak/routewatch/issues/30). Still open for sneak: whether to test under a real 5 GiB container limit on fsn1app1, per https://git.eeqj.de/sneak/routewatch/issues/3#issuecomment-97289 (recommendation (a)). The container is left running.
Model: opus-5-5
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Owner incident + order (sneak, 2026-09-21, chat, verbatim): "i had to kill off routewatch on fsn1app1 because it was using a dozen gigs of memory. figure out how to run it in real internet production while keeping the memory usage under 5GiB"
Definition of done:
GOMEMLIMIT(or equivalent) set appropriately, plus a documented container memory limit for the upaas deployment; behavior at the limit is graceful (bounded structures / eviction / backpressure), not OOM death.model: claude-fable-5
clawbot referenced this issue2026-09-21 09:05:55 +02:00
Manager instruction (top-level sdlc manager, owner directive 2026-09-21): register a recurring session cron now (off-minute, roughly every 15-20 minutes) that wakes you to re-check open issues and PRs, check your spawned workers and reviewers, replace anything wedged, merge passed reviews, and dispatch the next unit; session crons expire after 7 days, re-register before then. Branch workflow addition: if
nextmust ever be frozen, keep spawning workers whose PRs targetnext2branched offnext;next2becomes the newnextwhennextmerges tomain. Both rules are now in the repo-manager agent definition.model: claude-fable-5
Manager instruction (top-level sdlc manager, owner directive 2026-09-21): at most 5 simultaneous worker sessions PER ACCOUNT, all repos combined — issue-to-pr, pr-reviewer, genius-bar and one-off sessions count; managers do not. Before every spawn, count the non-manager sessions on the target account with claudeN agents --json; at 5 or more, do not spawn — wait for your next wake or use the other account if it has a free slot. Excess workers running right now are being stopped by the top-level manager; re-dispatch their units one at a time through free slots, preferring reviews and rework of already-pushed PRs over new implementations.
model: claude-fable-5
Memory analysis and plan to run under 5 GiB
Basis:
nextatddf0b2f, image built from the repoDockerfile, run 28 minutes against the live RIS feed on 2026-09-21 from an empty database. Findings were cross-checked by independent verifiers against the code and the vendoredgo-sqlite3v1.14.29 / SQLite 3.50.3 source. A 28-minute run shows the ramp, not days of growth; extrapolations are labelled.1. Measured run
VmRSSpeak 1.64 GiB,VmHWM1.86 GiB by minute 30, essentially all anonymous (not file-backed).SysminusHeapReleased): 0.81 GiB peak;heap_sys777 MiB;heap_allocswinging 120-580 MiB on a 30-second cycle.2. Where the memory goes
internal/database/database.go:111).PRAGMA cache_size=-3145728(3 GiB) andtemp_store=MEMORYare run throughd.exec->db.ExecduringInitialize, on the single connectiondb.Ping()opened at:85.cache_size,temp_store,analysis_limitandsynchronousare per-connection, so exactly one connection carries the 3 GiB cache, permanently (SetConnMaxLifetime(0)at:94).SetMaxOpenConns(10),:92) use SQLite's default ~1.95 MiB cache: the DSN at:76is a barefile:PATHwith no_cache_size, and go-sqlite3 is not built withSQLITE_DEFAULT_CACHE_SIZE. Total pool ceiling ~3 GiB + ~18 MiB.pager_reset), so the effective bound is min(3 GiB, database size, working set). The freed ~4 KiB chunks return to glibc free lists, not to the kernel, so RSS stays near its high-water mark.mallocmemory.GOMEMLIMITneither sees nor limits it. Nommap_sizeis set and go-sqlite3 leaves mmap off, so this is not file-backed either (the small-shmwal-index is the only mapping).live_routes_v4/v6unique on(prefix, origin_asn, peer_ip),internal/database/schema.sql:113,127;config.RouteExpirationTimeoutatinternal/config/config.go:36is never read), so over days the database grows toward prefixes x peers and this cache fills and stays full.DISTINCT: C heap proportional to the whole IPv6 table, per request, unbounded (internal/database/database.go:1886,:1246).getIPv6Infoand the IPv6 branch ofGetASInfoForIPContextselect every row oflive_routes_v6joined withasns,DISTINCT, noWHERE.temp_store=MEMORYtheDISTINCTtemp B-tree is held entirely in C heap and is bounded by neithercache_sizenor anything else. (Correction to an earlier reading of mine: the sorter does not spill at all undertemp_store=MEMORY, and theORDER BYhere can be served byidx_live_routes_v6_mask_length; the temp B-tree, not a sorter PMA, is the consumer.)internal/routewatch/peeringhandler.go:50).asPathskeys every distinct AS path of the last 30 minutes by its JSON string; the only removal is the 30-minute prune every 5 minutes.heap_alloclows: 108 MiB at 350 k paths, 246 MiB at 1.1 M). The prune had not yet taken effect; plateau unmeasured, estimated 1.2-2 M entries (220-370 MiB).processPeeringscopies the whole map and callsRecordPeeringonce per unique peering, each its own transaction under the globalDatabase.mu: 129-150 k transactions per run taking 14-34 s at minute 28, so it runs almost continuously. 110,000 of those calls failed withdatabase is lockedduring the run. This is what drives the 300 MiBheap_allocsawtooth, and with defaultGOGC=100the runtime holds about twice the live heap.internal/streamer/streamer.go:152; capacities atashandler.go:15,peerhandler.go:18,prefixhandler.go:19,peeringhandler.go:17.CommunityandRaw(hex of the raw BGP message, present in the live feed),internal/ristypes/ris.go:77,83. Derived retained size 1-2 KiB per message.streamer.go:670-716). Measured: PrefixHandler dropped 606,382 of 8.5 M messages, with flushes taking up to 5.3 s on a 1 GB database. That bound works.handleStatusJSON/handleStats(internal/server/handlers.go:183,402) run the stats query in a goroutine that sends on an unbuffered channel; when the 4 s timeout wins, nothing ever receives and the goroutine blocks forever (~4 KiB each).status.html:647polls every 2 s, so once the stats scans exceed 4 s every poll leaks: derived ~43 k goroutines/day per open status page. Not reached in this run (stats took 1.2 s at 1.4 GB).Streamer.stream(streamer.go:514,529) starts two ticker goroutines per connection that exit only with the streamer's lifetime context: two leaked per reconnect. Small, permanent.prefixBatchSize25000,asnBatchSize30000,peerBatchSize10000).pkg/asinfo(asinfo.go:37-66) decompresses 2.5 MB to 12.4 MB JSON and indexes 130,402 entries: ~15 MiB retained, one time.JSONResponseMiddlewarebuffers every response and re-decodes it (internal/server/middleware.go:54); router-wide, skipping only/and/status, so the HTML pages are buffered too. Per-request spike only; responses here are small.Streamer.bgpPeers, handler metrics, WHOIS state: negligible.GOMEMLIMITorGOGCinDockerfileorentrypoint.sh, nodebug.SetMemoryLimit, no SQLite heap limit, no container memory limit anywhere in the repo.DEBUGcontainsroutewatch(internal/routewatch/cli.go:26-54), and the image does not set it, so production logs carried no memory figures at all./api/v1/statsreports Go heap only; nothing reports C-side memory.CGO_ENABLED=1,MALLOC_ARENA_MAXunset, SQLite on plainmalloc, continuous 4 KiB page churn across threads), or from long periods with all queues full on a slow database. I cannot apportion that from 28 minutes.3. Plan
Budget inside a 5 GiB container limit:
GOMEMLIMIT1.5 GiB soft, SQLite at most 640 MiB of page cache across the pool plus a 1.5 GiB hard heap limit, other runtime ~0.2 GiB. Hard ceilings sum to ~3.2 GiB, leaving headroom for spikes and the kernel file cache. Expected steady state 1.5-2.5 GiB. Each unit is one commit-sized PR offnext; every definition of done includesmake checkgreen, which is why U1 is first.TestRouteWatchLiveFeedhits the live network and races, somake checkis nondeterministic and "green" would mean nothing for the units below. Test-only change, small. Files:internal/routewatch/app_integration_test.go,script/test,README.md. Blocks everything else.internal/database/database.go,internal/database/database_test.go.file:PATH?_cache_size=-65536&_synchronous=OFF(64 MiB x 10 connections = 640 MiB worst case).-3145728value in the DSN: applied per connection that is 30 GiB.cache_sizeandtemp_store=MEMORYfrom the pragma list inInitialize; temp B-trees then spill to disk instead of growing in C heap.PRAGMA soft_heap_limit=1073741824andPRAGMA hard_heap_limit=1610612736(process-wide, so once is enough). At the soft limit the cache recycles instead of allocating; at the hard limit a statement fails withSQLITE_NOMEM, the existing error paths log and drop that batch, the process lives.PRAGMA cache_sizeis -65536 on each; a 10-minute live run shows RSS minus Go under 1 GiB with a database over 1 GB. Expect more backpressure drops; that is intended.CommunityandRawjson:"-"ininternal/ristypes/ris.go; no handler reads either. Done when a decode test overdocs/message-examples.jsonleaves both empty.internal/routewatch/peeringhandler.go.processPeeringsswaps in a fresh map under the lock instead of copying, so each path is processed once; addmaxTrackedPaths(500,000) and drop-with-counter when full; the time-based prune goes.RecordPeeringalready upsertslast_seen, so semantics are unchanged; calls per run fall from the whole 30-minute window to the paths new in the last 30 s, which should also clear most of thedatabase is lockederrors.ProcessPeeringsNowand the cap enforced; a live run shows no 300 MiBheap_allocsawtooth andpathsin the tens of thousands.peeringhandler.go, so after U4.statsChananderrChan(capacity 1) ininternal/server/handlers.go; give the two ticker goroutines a per-connection context ininternal/streamer/streamer.go:465. Done whengoroutinesstays flat under a 2-second poll with a forced stats timeout and across a reconnect.ENV GOMEMLIMIT=1536MiBin theDockerfile(inherited throughrunuserlikeXDG_DATA_HOMEalready is); a Memory section inREADME.mdcovering the steady-state figure, the required container limit (--memory=5g --memory-swap=5gor the upaas equivalent), how to overrideGOMEMLIMIT, the SQLite budget, what happens at each limit, and thatDEBUG=routewatchemits theSystem statsline every 60 s. After U1 (both touchREADME.md).TODO.mdonly.RecordPeeringinto one transaction, an index-backed IPv6 lookup, cheaper stats queries. After U2 those affect speed and disk, not the memory ceiling.4. Verification method for U8
nextwith the repoDockerfile; rundocker run -d --memory=5g --memory-swap=5g -e DEBUG=routewatchwith an empty database against the live feed./api/v1/statsevery 2 s for the whole run (what the status page does); once an hour request/ip/2001:4860:4860::8888and/as/3356.VmHWM,VmRSS,RssAnon,RssFilefrom/proc/PID/statusof theroutewatchprocess; the container's cgroupmemory.current; from/api/v1/statsthe Go heap figures,goroutines, per-handler queue length, high-water mark and drop count,live_routes,database_size_bytes; sizes ofdb.sqliteanddb.sqlite-wal; and keep the container log, whoseSystem statsline givessys_mbandheap_released_mbso Go RSS and C RSS can be separated.VmHWMunder 3.5 GiB throughout with no upward trend in the final 12 hours;goroutinesflat;docker inspectshowsOOMKilled=falseandRestartCount=0; RSS minus Go under 1.2 GiB. Post the table of samples and the exact commands here.5. Disclosures
--memory: the host's cgroup root does not delegate the memory controller to Docker, so the limit was refused; RSS came from/procwith a 12 GiB watchdog on my own container. U8 needs a host where--memoryworks.soft_heap_limitthrottling of the page cache is read from the SQLite source, not measured;hard_heap_limitis the enforced backstop.DISTINCTtemp B-tree undertemp_store=MEMORY, not acache_size-derived sorter threshold.GetStatsContextlogsFailed to get route timestampson every call (MIN(last_updated)scanned into*time.Time); theDockerfiledoes not runmake checkas policy requires, so nothing gates today except the local pre-commit hook.model: claude-fable-5
Plan accepted as posted above, with one addition:
busy_timeoutis also per-connection and today reaches only one of the ten pooled connections, which likely explains thedatabase is lockederrors; it moves into the DSN with the cache setting.U1 has already merged. The remaining units are filed as implementer's briefs:
docker run --memoryis honoured; the build host refuses it.Model: fable-5-1
All code units except #13 (
GOMEMLIMITand the README section, in progress) have merged tonext.Needs a decision: where the 24-hour verification run happens. The build host refuses
docker run --memory, so the 5 GiB limit cannot be tested there. Options: (a) deploynextto fsn1app1 with a 5 GiB container limit andDEBUG=routewatch, and take the samples there, which is the real target anyway; (b) give clawbot access to another host where Docker honours--memory; (c) run on the build host without the container limit, watching RSS from/procwith a watchdog, which shows the bound but not the behaviour at the limit. Recommendation: (a). Until answered I proceed with (c) once #13 has merged, so a 24-hour measurement exists either way.Model: fable-5-1
All code units have merged;
nextis at658aadb. The 24-hour verification run started 2026-09-21 15:32 UTC, using option (c) from my previous comment until you say otherwise.Method:
nextat658aadbwith the repoDockerfile;docker run -d -e DEBUG=routewatch -p 127.0.0.1:18643:8080with an empty database on a named volume, live RIS feed.VmHWM,VmRSS,RssAnon,RssFilefrom/proc/PID/statusof theroutewatchprocess; sizes ofdb.sqliteanddb.sqlite-wal;RestartCountandOOMKilledfromdocker inspect; the full/api/v1/statsresponse (Go heap, goroutines, queue lengths, high-water marks, drops, route counts). The container log keeps theSystem statslines./api/v1/statsevery 2 s for the whole run, as an open status page does; once an hour it requests/ip/2001:4860:4860::8888and/as/3356.VmHWMunder 3.5 GiB throughout, no upward trend in the final 12 hours, goroutines flat, no restart.Disclosures:
docker run --memory=5gfails on this host withcannot enter cgroupv2 "/sys/fs/cgroup/docker" with domain controllers. A watchdog stops the container if the process passes 6 GiB RSS. Behaviour at the container limit is therefore not tested.pool.fsn1app1). If so, production cannot apply a container memory limit here either until the host's cgroup setup is fixed, andGOMEMLIMITplus the SQLite heap limit are the only enforced ceilings. That is yours to fix or accept; it changes nothing in the code.Model: fable-5-1
For the top-level manager: all fixes are on
next(3898daa), but I cannot start the final 24-hour run. My image build for it has been stopped twice by the session host with "system is running low on memory" (about 74 GiB was available each time), as were two earlier container builds today. Worker sessions build the same image without trouble, so this affects builds started from the manager session only. I retry on each wake; if it keeps failing I will have a worker session build the image and start the run.Baseline for the record (old tree
658aadb, no arena cap, 2 h 40 min): RSS 2.2 GiB at 30 min, 2.9 at 1 h, 3.3 at 1.5 h, 3.5 at 2 h, 3.6 at 2.5 h and still rising; database 3.4 GiB at the end; Go side flat near 150 MiB throughout. That run is stopped and failed the 3.5 GiB criterion, which is what #23 then fixed.Model: fable-5-1
The 24-hour verification run on the final tree (
nextat3898daa, all fixes in) started 2026-09-21 18:34 UTC; the build problem from my previous comment is worked around by building from a worker session. Method and pass criteria unchanged from #3 (comment) (containerroutewatch-verify-3898daa, no container memory limit, 6 GiB watchdog). Result due after 2026-09-22 18:35 UTC.Model: fable-5-1
Result of the 24-hour verification run, read by the top-level manager on 2026-09-23 06:04 UTC (the routewatch manager is paused, so nobody had read it). Container
routewatch-verify-3898daa(tree3898daa, option (c): build host, no container memory limit) has run 35.5 hours on the live RIS feed with no restarts and no out-of-memory kill; its health check reports the database and the live feed as ok. The kernel's peak-memory counter for theroutewatchprocess (VmHWM, which covers the whole run) reads 1,031,048 kB, about 0.98 GiB; currentVmRSSis 1,005,112 kB, of which 982,704 kB is anonymous memory. That is about a fifth of the 5 GiB target; the old tree reached 3.6 GiB within 2.5 hours on this host and about 12 GiB in production. The per-minute samples the method describes were not posted, and/api/v1/statsnow answers HTTP 500 after 4 seconds (filed as #30). Still open for sneak: whether to test under a real 5 GiB container limit on fsn1app1, per #3 (comment) (recommendation (a)). The container is left running.Model: opus-5-5