Add _txlock=immediate so batch writes wait instead of dropping (closes #25) #26

Merged
clawbot merged 1 commits from issue-25-txlock-immediate into next 2026-09-21 20:12:34 +02:00
Collaborator

Fixes #25.

What changed

The connection DSN in internal/database/database.go now includes
_txlock=immediate, so every database/sql transaction issues BEGIN IMMEDIATE and takes the SQLite write lock at the start.

Why

The batch-flush write paths read (SELECT) before they write. With the
default deferred BEGIN, such a transaction holds only a read lock until its
first INSERT/UPDATE, then must upgrade to the write lock. If another
connection holds the write lock at that moment, SQLite returns database is locked immediately and skips the busy handler, so the 5 s busy_timeout
never applies and the batch is dropped. The other writer is the background
maintainer's WAL checkpoint every 5 s, which runs on a pooled connection
without the process mutex. BEGIN IMMEDIATE takes the write lock up front,
where the busy timeout does apply, so a contending writer waits instead of
failing. Only the explicit transactions (all writers) change; plain read
queries open no transaction, so read concurrency is unchanged.

Verification

  • New regression test drives batch writes against a running checkpoint loop:
    it fails with database is locked on the old DSN and passes with the fix.
  • 15-minute live run against the RIS feed from an empty database:
    0 database is locked errors (0 error-level log lines other than the
    pre-existing route-timestamps warning noted in issue 3, out of scope). The
    database reached ~1.0 GB with 1.9M live routes over 5.1M messages. Method:
    image built from this branch,
    docker run -e DEBUG=routewatch, counted in the container log.
  • script/cibuild green; its lint and test stages executed, not cached.

Disclosure (judgement call): ran the one new test directly once, not via
make test, to confirm it reproduced the failure before the fix; the gate
ran the full suite.

Model: opus-4-8

Fixes https://git.eeqj.de/sneak/routewatch/issues/25. ## What changed The connection DSN in `internal/database/database.go` now includes `_txlock=immediate`, so every `database/sql` transaction issues `BEGIN IMMEDIATE` and takes the SQLite write lock at the start. ## Why The batch-flush write paths read (`SELECT`) before they write. With the default deferred `BEGIN`, such a transaction holds only a read lock until its first `INSERT`/`UPDATE`, then must upgrade to the write lock. If another connection holds the write lock at that moment, SQLite returns `database is locked` immediately and skips the busy handler, so the 5 s `busy_timeout` never applies and the batch is dropped. The other writer is the background maintainer's WAL checkpoint every 5 s, which runs on a pooled connection without the process mutex. `BEGIN IMMEDIATE` takes the write lock up front, where the busy timeout does apply, so a contending writer waits instead of failing. Only the explicit transactions (all writers) change; plain read queries open no transaction, so read concurrency is unchanged. ## Verification - New regression test drives batch writes against a running checkpoint loop: it fails with `database is locked` on the old DSN and passes with the fix. - 15-minute live run against the RIS feed from an empty database: 0 `database is locked` errors (0 error-level log lines other than the pre-existing route-timestamps warning noted in issue 3, out of scope). The database reached ~1.0 GB with 1.9M live routes over 5.1M messages. Method: image built from this branch, `docker run -e DEBUG=routewatch`, counted in the container log. - `script/cibuild` green; its lint and test stages executed, not cached. Disclosure (judgement call): ran the one new test directly once, not via `make test`, to confirm it reproduced the failure before the fix; the gate ran the full suite. Model: opus-4-8
clawbot added 1 commit 2026-09-21 19:44:12 +02:00
Batch flush paths read before they write, so a deferred transaction starts
as a reader and must upgrade to the write lock on its first INSERT/UPDATE.
When the background maintainer holds the write lock for a WAL checkpoint,
that upgrade fails immediately with "database is locked" and the busy
timeout does not apply, so the batch is dropped. Adding _txlock=immediate
to the DSN makes every transaction take the write lock at BEGIN, so it
waits up to busy_timeout instead of failing.

A regression test drives batch writes against a running checkpoint loop and
fails with "database is locked" without the change.

Model: opus-4-8
clawbot self-assigned this 2026-09-21 19:44:12 +02:00
clawbot added the needs-review label 2026-09-21 19:44:12 +02:00
Author
Collaborator

PASS — adding _txlock=immediate makes every write transaction take the SQLite write lock at BEGIN, so a batch that contends with the maintainer's checkpoint waits up to the busy timeout instead of failing its read-to-write upgrade and dropping data, while the read paths (which run outside any transaction) are unchanged, the regression test genuinely reproduces the fault without the change, and the fix meets the issue's definition of done.

Model: opus-4-8

PASS — adding `_txlock=immediate` makes every write transaction take the SQLite write lock at BEGIN, so a batch that contends with the maintainer's checkpoint waits up to the busy timeout instead of failing its read-to-write upgrade and dropping data, while the read paths (which run outside any transaction) are unchanged, the regression test genuinely reproduces the fault without the change, and the fix meets the issue's definition of done. Model: opus-4-8
clawbot removed the needs-review label 2026-09-21 20:07:39 +02:00
clawbot merged commit 3898daad4e into next 2026-09-21 20:12:34 +02:00
clawbot deleted branch issue-25-txlock-immediate 2026-09-21 20:12:34 +02:00
Sign in to join this conversation.
No Reviewers
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: sneak/routewatch#26