Found by the top-level manager on 2026-09-23 while reading the 24-hour verification run for #3.
Container routewatch-verify-3898daa (tree 3898daa, started 2026-09-21 18:34 UTC, live RIS feed, empty database at start) is healthy: /.well-known/healthcheck.json reports database and ris_live as ok. GET /api/v1/stats on the same instance returns HTTP 500 with execution_time 4000 ms and the body {"error":{"code":500,"msg":"Internal server error"}}. The 4000 ms suggests a query deadline being hit on a database that has grown over 35 hours.
Definition of done
The cause is found and stated on this issue (which query or step fails, and why).
GET /api/v1/stats answers 200 with its full response within its deadline on a database the size of a 24-hour live-feed run, or the endpoint is changed so it stays within the deadline at that size; a test covers the fixed path.
make check passes.
The repo is paused under the current priority rule; this waits for it to resume.
Model: opus-5-5
Found by the top-level manager on 2026-09-23 while reading the 24-hour verification run for https://git.eeqj.de/sneak/routewatch/issues/3.
Container `routewatch-verify-3898daa` (tree `3898daa`, started 2026-09-21 18:34 UTC, live RIS feed, empty database at start) is healthy: `/.well-known/healthcheck.json` reports `database` and `ris_live` as `ok`. `GET /api/v1/stats` on the same instance returns HTTP 500 with `execution_time` 4000 ms and the body `{"error":{"code":500,"msg":"Internal server error"}}`. The 4000 ms suggests a query deadline being hit on a database that has grown over 35 hours.
## Definition of done
- The cause is found and stated on this issue (which query or step fails, and why).
- `GET /api/v1/stats` answers 200 with its full response within its deadline on a database the size of a 24-hour live-feed run, or the endpoint is changed so it stays within the deadline at that size; a test covers the fixed path.
- `make check` passes.
The repo is paused under the current priority rule; this waits for it to resume.
Model: opus-5-5
Plan. This is very likely the defect of #27: the soak container runs 3898daa, where every stats request ran COUNT(*) over each table and a MIN/MAX scan of both route tables, which pass the handler's 4-second deadline (statsContextTimeout, internal/server/handlers.go) once the database is past about 4.5 GiB. next fixed that in #29: counts come from memory, the timestamps from the ends of an index, with tests. One query still reads a whole index on every request: the prefix distribution (GetPrefixDistributionContext). Nobody has measured it at 24-hour size.
The unit:
Build a database at least the size at which #27 failed (5.3 GiB, 2 M live routes), preferably twice that, without touching the soak container or its volume; run the image built from next on it and request /api/v1/stats repeatedly.
If it answers 200 with the full response within the deadline: state the cause and the result on this issue in one comment, and name the tests that cover the fixed path. An independent reviewer repeats the check; then this issue closes.
If it does not: state the cause (which query, how it grows) on this issue and stop; the fix gets its own plan here.
Model: opus-5-5
Plan. This is very likely the defect of https://git.eeqj.de/sneak/routewatch/issues/27: the soak container runs `3898daa`, where every stats request ran `COUNT(*)` over each table and a `MIN`/`MAX` scan of both route tables, which pass the handler's 4-second deadline (`statsContextTimeout`, `internal/server/handlers.go`) once the database is past about 4.5 GiB. `next` fixed that in https://git.eeqj.de/sneak/routewatch/pulls/29: counts come from memory, the timestamps from the ends of an index, with tests. One query still reads a whole index on every request: the prefix distribution (`GetPrefixDistributionContext`). Nobody has measured it at 24-hour size.
The unit:
- Build a database at least the size at which https://git.eeqj.de/sneak/routewatch/issues/27 failed (5.3 GiB, 2 M live routes), preferably twice that, without touching the soak container or its volume; run the image built from `next` on it and request `/api/v1/stats` repeatedly.
- If it answers 200 with the full response within the deadline: state the cause and the result on this issue in one comment, and name the tests that cover the fixed path. An independent reviewer repeats the check; then this issue closes.
- If it does not: state the cause (which query, how it grows) on this issue and stop; the fix gets its own plan here.
Model: opus-5-5
clawbot
self-assigned this 2026-09-29 07:43:33 +02:00
next still fails at this size. /api/v1/stats answered HTTP 500 at the 4-second deadline on every request, the same symptom as on 3898daa.
Cause: the prefix distribution (GetPrefixDistributionContext in internal/database/database.go), the one query left on the stats path after #29. On every request it reads the whole (mask_length, prefix) index of live_routes_v4, then of live_routes_v6, to count distinct prefixes per mask length. That index has one entry per live route (each prefix once per peer), so the cost grows in step with the number of live routes, not with the number of distinct prefixes or with the file size. At this route count the IPv4 half takes nearly the whole 4 seconds. The IPv6 half is then cut off at the deadline, and the app logs that on every request. The handler's own deadline expires at the same moment, so the whole answer becomes a 500, not a response with only the distribution missing. The counts and the oldest and newest route times are no longer a problem: they come from memory and from the ends of an index.
Tested: 9.7 GiB database created by the app from next (057e0bd), 22.2 M live routes (18.6 M IPv4, 3.5 M IPv6) over 1.0 M IPv4 and 0.24 M IPv6 prefixes from 1,546 peers, mask lengths /8–/32 and /16–/64 in roughly real-table proportions, routes announced, re-announced and withdrawn in feed order. The image built from next ran on it, connected to the live feed, and was polled every 2 seconds.
Not verified: how many live routes a real 24-hour run holds. The soak test's database was not read. This test's route count is far above the count at which #27 appeared, and at that lower count the same query stays well inside the deadline. So the failure comes with route growth, not at a fixed file size.
Deviation: the database is 9.7 GiB, short of the preferred 10.6 GiB. The fill was stopped under the host memory limit.
Stopping here as the plan says; the fix needs its own plan on this issue.
Model: opus-5-5 (test and comment); fable-5-1 (database finishing step)
`next` still fails at this size. `/api/v1/stats` answered HTTP 500 at the 4-second deadline on every request, the same symptom as on `3898daa`.
Cause: the prefix distribution (`GetPrefixDistributionContext` in `internal/database/database.go`), the one query left on the stats path after https://git.eeqj.de/sneak/routewatch/pulls/29. On every request it reads the whole `(mask_length, prefix)` index of `live_routes_v4`, then of `live_routes_v6`, to count distinct prefixes per mask length. That index has one entry per live route (each prefix once per peer), so the cost grows in step with the number of live routes, not with the number of distinct prefixes or with the file size. At this route count the IPv4 half takes nearly the whole 4 seconds. The IPv6 half is then cut off at the deadline, and the app logs that on every request. The handler's own deadline expires at the same moment, so the whole answer becomes a 500, not a response with only the distribution missing. The counts and the oldest and newest route times are no longer a problem: they come from memory and from the ends of an index.
Tested: 9.7 GiB database created by the app from `next` (`057e0bd`), 22.2 M live routes (18.6 M IPv4, 3.5 M IPv6) over 1.0 M IPv4 and 0.24 M IPv6 prefixes from 1,546 peers, mask lengths /8–/32 and /16–/64 in roughly real-table proportions, routes announced, re-announced and withdrawn in feed order. The image built from `next` ran on it, connected to the live feed, and was polled every 2 seconds.
- Not verified: how many live routes a real 24-hour run holds. The soak test's database was not read. This test's route count is far above the count at which https://git.eeqj.de/sneak/routewatch/issues/27 appeared, and at that lower count the same query stays well inside the deadline. So the failure comes with route growth, not at a fixed file size.
- Deviation: the database is 9.7 GiB, short of the preferred 10.6 GiB. The fill was stopped under the host memory limit.
Stopping here as the plan says; the fix needs its own plan on this issue.
Model: opus-5-5 (test and comment); fable-5-1 (database finishing step)
Fix plan, after the finding in #30 (comment): the prefix distribution is the last query on the stats path, and it reads one index entry per live route on every request.
Keep the distribution in memory and current on every write, the way #29 did for the counts (sneak's ruling on #27: realtime, no timer):
Two small arrays of counters, distinct prefixes per mask length for IPv4 and for IPv6, next to the existing counts in internal/database/counts.go, under the same lock.
Seeded once at startup from the existing distribution query, where the counts are seeded now.
When a live route is inserted as new (the write path already tells a new route from a re-announced one), look up by the existing prefix index whether another live route has that prefix; if none, add one for its mask length. When a live route is removed, look up whether any live route with that prefix remains; if none, take one away. One indexed lookup per new or removed route, inside the same transaction.
GetStatsContext fills the distribution from memory; the request path runs no query over the route tables. GetPrefixDistributionContext stays only for the seeding.
No per-prefix table in memory: that would be about 1.2 million entries, the memory this service is tuned to avoid.
Tests: the counters stay exact across inserts, re-announcements, a second peer announcing the same prefix, and withdrawals of the last and of a non-last route for a prefix; and the stats read runs no query on the route tables.
A check on a large database within the host memory limits: /api/v1/stats answers 200 with the distribution, polled every 2 seconds.
Model: opus-5-5
Fix plan, after the finding in https://git.eeqj.de/sneak/routewatch/issues/30#issuecomment-106722: the prefix distribution is the last query on the stats path, and it reads one index entry per live route on every request.
Keep the distribution in memory and current on every write, the way https://git.eeqj.de/sneak/routewatch/pulls/29 did for the counts (sneak's ruling on https://git.eeqj.de/sneak/routewatch/issues/27: realtime, no timer):
- Two small arrays of counters, distinct prefixes per mask length for IPv4 and for IPv6, next to the existing counts in `internal/database/counts.go`, under the same lock.
- Seeded once at startup from the existing distribution query, where the counts are seeded now.
- When a live route is inserted as new (the write path already tells a new route from a re-announced one), look up by the existing `prefix` index whether another live route has that prefix; if none, add one for its mask length. When a live route is removed, look up whether any live route with that prefix remains; if none, take one away. One indexed lookup per new or removed route, inside the same transaction.
- `GetStatsContext` fills the distribution from memory; the request path runs no query over the route tables. `GetPrefixDistributionContext` stays only for the seeding.
- No per-prefix table in memory: that would be about 1.2 million entries, the memory this service is tuned to avoid.
- Tests: the counters stay exact across inserts, re-announcements, a second peer announcing the same prefix, and withdrawals of the last and of a non-last route for a prefix; and the stats read runs no query on the route tables.
- A check on a large database within the host memory limits: `/api/v1/stats` answers 200 with the distribution, polled every 2 seconds.
Model: opus-5-5
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Found by the top-level manager on 2026-09-23 while reading the 24-hour verification run for #3.
Container
routewatch-verify-3898daa(tree3898daa, started 2026-09-21 18:34 UTC, live RIS feed, empty database at start) is healthy:/.well-known/healthcheck.jsonreportsdatabaseandris_liveasok.GET /api/v1/statson the same instance returns HTTP 500 withexecution_time4000 ms and the body{"error":{"code":500,"msg":"Internal server error"}}. The 4000 ms suggests a query deadline being hit on a database that has grown over 35 hours.Definition of done
GET /api/v1/statsanswers 200 with its full response within its deadline on a database the size of a 24-hour live-feed run, or the endpoint is changed so it stays within the deadline at that size; a test covers the fixed path.make checkpasses.The repo is paused under the current priority rule; this waits for it to resume.
Model: opus-5-5
Plan. This is very likely the defect of #27: the soak container runs
3898daa, where every stats request ranCOUNT(*)over each table and aMIN/MAXscan of both route tables, which pass the handler's 4-second deadline (statsContextTimeout,internal/server/handlers.go) once the database is past about 4.5 GiB.nextfixed that in #29: counts come from memory, the timestamps from the ends of an index, with tests. One query still reads a whole index on every request: the prefix distribution (GetPrefixDistributionContext). Nobody has measured it at 24-hour size.The unit:
nexton it and request/api/v1/statsrepeatedly.Model: opus-5-5
nextstill fails at this size./api/v1/statsanswered HTTP 500 at the 4-second deadline on every request, the same symptom as on3898daa.Cause: the prefix distribution (
GetPrefixDistributionContextininternal/database/database.go), the one query left on the stats path after #29. On every request it reads the whole(mask_length, prefix)index oflive_routes_v4, then oflive_routes_v6, to count distinct prefixes per mask length. That index has one entry per live route (each prefix once per peer), so the cost grows in step with the number of live routes, not with the number of distinct prefixes or with the file size. At this route count the IPv4 half takes nearly the whole 4 seconds. The IPv6 half is then cut off at the deadline, and the app logs that on every request. The handler's own deadline expires at the same moment, so the whole answer becomes a 500, not a response with only the distribution missing. The counts and the oldest and newest route times are no longer a problem: they come from memory and from the ends of an index.Tested: 9.7 GiB database created by the app from
next(057e0bd), 22.2 M live routes (18.6 M IPv4, 3.5 M IPv6) over 1.0 M IPv4 and 0.24 M IPv6 prefixes from 1,546 peers, mask lengths /8–/32 and /16–/64 in roughly real-table proportions, routes announced, re-announced and withdrawn in feed order. The image built fromnextran on it, connected to the live feed, and was polled every 2 seconds.Stopping here as the plan says; the fix needs its own plan on this issue.
Model: opus-5-5 (test and comment); fable-5-1 (database finishing step)
Fix plan, after the finding in #30 (comment): the prefix distribution is the last query on the stats path, and it reads one index entry per live route on every request.
Keep the distribution in memory and current on every write, the way #29 did for the counts (sneak's ruling on #27: realtime, no timer):
internal/database/counts.go, under the same lock.prefixindex whether another live route has that prefix; if none, add one for its mask length. When a live route is removed, look up whether any live route with that prefix remains; if none, take one away. One indexed lookup per new or removed route, inside the same transaction.GetStatsContextfills the distribution from memory; the request path runs no query over the route tables.GetPrefixDistributionContextstays only for the seeding./api/v1/statsanswers 200 with the distribution, polled every 2 seconds.Model: opus-5-5