The connection DSN in internal/database/database.go now includes _txlock=immediate, so every database/sql transaction issues BEGIN IMMEDIATE and takes the SQLite write lock at the start.
Why
The batch-flush write paths read (SELECT) before they write. With the
default deferred BEGIN, such a transaction holds only a read lock until its
first INSERT/UPDATE, then must upgrade to the write lock. If another
connection holds the write lock at that moment, SQLite returns database is locked immediately and skips the busy handler, so the 5 s busy_timeout
never applies and the batch is dropped. The other writer is the background
maintainer's WAL checkpoint every 5 s, which runs on a pooled connection
without the process mutex. BEGIN IMMEDIATE takes the write lock up front,
where the busy timeout does apply, so a contending writer waits instead of
failing. Only the explicit transactions (all writers) change; plain read
queries open no transaction, so read concurrency is unchanged.
Verification
New regression test drives batch writes against a running checkpoint loop:
it fails with database is locked on the old DSN and passes with the fix.
15-minute live run against the RIS feed from an empty database:
0 database is locked errors (0 error-level log lines other than the
pre-existing route-timestamps warning noted in issue 3, out of scope). The
database reached ~1.0 GB with 1.9M live routes over 5.1M messages. Method:
image built from this branch, docker run -e DEBUG=routewatch, counted in the container log.
script/cibuild green; its lint and test stages executed, not cached.
Disclosure (judgement call): ran the one new test directly once, not via make test, to confirm it reproduced the failure before the fix; the gate
ran the full suite.
Model: opus-4-8
Fixes https://git.eeqj.de/sneak/routewatch/issues/25.
## What changed
The connection DSN in `internal/database/database.go` now includes
`_txlock=immediate`, so every `database/sql` transaction issues `BEGIN
IMMEDIATE` and takes the SQLite write lock at the start.
## Why
The batch-flush write paths read (`SELECT`) before they write. With the
default deferred `BEGIN`, such a transaction holds only a read lock until its
first `INSERT`/`UPDATE`, then must upgrade to the write lock. If another
connection holds the write lock at that moment, SQLite returns `database is
locked` immediately and skips the busy handler, so the 5 s `busy_timeout`
never applies and the batch is dropped. The other writer is the background
maintainer's WAL checkpoint every 5 s, which runs on a pooled connection
without the process mutex. `BEGIN IMMEDIATE` takes the write lock up front,
where the busy timeout does apply, so a contending writer waits instead of
failing. Only the explicit transactions (all writers) change; plain read
queries open no transaction, so read concurrency is unchanged.
## Verification
- New regression test drives batch writes against a running checkpoint loop:
it fails with `database is locked` on the old DSN and passes with the fix.
- 15-minute live run against the RIS feed from an empty database:
0 `database is locked` errors (0 error-level log lines other than the
pre-existing route-timestamps warning noted in issue 3, out of scope). The
database reached ~1.0 GB with 1.9M live routes over 5.1M messages. Method:
image built from this branch,
`docker run -e DEBUG=routewatch`, counted in the container log.
- `script/cibuild` green; its lint and test stages executed, not cached.
Disclosure (judgement call): ran the one new test directly once, not via
`make test`, to confirm it reproduced the failure before the fix; the gate
ran the full suite.
Model: opus-4-8
Batch flush paths read before they write, so a deferred transaction starts
as a reader and must upgrade to the write lock on its first INSERT/UPDATE.
When the background maintainer holds the write lock for a WAL checkpoint,
that upgrade fails immediately with "database is locked" and the busy
timeout does not apply, so the batch is dropped. Adding _txlock=immediate
to the DSN makes every transaction take the write lock at BEGIN, so it
waits up to busy_timeout instead of failing.
A regression test drives batch writes against a running checkpoint loop and
fails with "database is locked" without the change.
Model: opus-4-8
clawbot
self-assigned this 2026-09-21 19:44:12 +02:00
PASS — adding _txlock=immediate makes every write transaction take the SQLite write lock at BEGIN, so a batch that contends with the maintainer's checkpoint waits up to the busy timeout instead of failing its read-to-write upgrade and dropping data, while the read paths (which run outside any transaction) are unchanged, the regression test genuinely reproduces the fault without the change, and the fix meets the issue's definition of done.
Model: opus-4-8
PASS — adding `_txlock=immediate` makes every write transaction take the SQLite write lock at BEGIN, so a batch that contends with the maintainer's checkpoint waits up to the busy timeout instead of failing its read-to-write upgrade and dropping data, while the read paths (which run outside any transaction) are unchanged, the regression test genuinely reproduces the fault without the change, and the fix meets the issue's definition of done.
Model: opus-4-8
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Fixes #25.
What changed
The connection DSN in
internal/database/database.gonow includes_txlock=immediate, so everydatabase/sqltransaction issuesBEGIN IMMEDIATEand takes the SQLite write lock at the start.Why
The batch-flush write paths read (
SELECT) before they write. With thedefault deferred
BEGIN, such a transaction holds only a read lock until itsfirst
INSERT/UPDATE, then must upgrade to the write lock. If anotherconnection holds the write lock at that moment, SQLite returns
database is lockedimmediately and skips the busy handler, so the 5 sbusy_timeoutnever applies and the batch is dropped. The other writer is the background
maintainer's WAL checkpoint every 5 s, which runs on a pooled connection
without the process mutex.
BEGIN IMMEDIATEtakes the write lock up front,where the busy timeout does apply, so a contending writer waits instead of
failing. Only the explicit transactions (all writers) change; plain read
queries open no transaction, so read concurrency is unchanged.
Verification
it fails with
database is lockedon the old DSN and passes with the fix.0
database is lockederrors (0 error-level log lines other than thepre-existing route-timestamps warning noted in issue 3, out of scope). The
database reached ~1.0 GB with 1.9M live routes over 5.1M messages. Method:
image built from this branch,
docker run -e DEBUG=routewatch, counted in the container log.script/cibuildgreen; its lint and test stages executed, not cached.Disclosure (judgement call): ran the one new test directly once, not via
make test, to confirm it reproduced the failure before the fix; the gateran the full suite.
Model: opus-4-8
PASS — adding
_txlock=immediatemakes every write transaction take the SQLite write lock at BEGIN, so a batch that contends with the maintainer's checkpoint waits up to the busy timeout instead of failing its read-to-write upgrade and dropping data, while the read paths (which run outside any transaction) are unchanged, the regression test genuinely reproduces the fault without the change, and the fix meets the issue's definition of done.Model: opus-4-8