Make SQLite durable under concurrent readers and stop re-delivering stranded webhooks (closes #256)
All checks were successful
check / check (push) Successful in 3m38s
All checks were successful
check / check (push) Successful in 3m38s
An operator running `sqlite3 <db> .dump` against their own per-webhook database wedged it: 60 of 60 inbound webhooks rejected with HTTP 500, 206 delivered webhooks stranded at `pending`, and every one of them POSTed a second time on the next restart while the event log recorded a single attempt. Durability. Every SQLite file — main, per-webhook, and archive — now opens through one path, `internal/database/sqlite_open.go`, in WAL journal mode with a 10-second busy timeout, `BEGIN IMMEDIATE` transactions, and a bounded connection pool. WAL is what stops a reader blocking writers at all. `_txlock=immediate` is what stops a `COMMIT` failing while its transaction stays open on a pooled connection, which is how four `database is locked` errors became 593 `cannot start a transaction within a transaction`: a deferred transaction that upgrades to a write lock mid-flight gets SQLITE_BUSY without the busy handler being consulted. `cache=shared` is gone, because under it an in-process conflict is SQLITE_LOCKED, which the busy handler does not retry. Delivery. `recordResult` and `updateDeliveryStatus` return their errors instead of logging and dropping them, and a caller whose bookkeeping write failed writes nothing at all — the delivery keeps whichever non-terminal status it already held, and both sweeps recover it. Recovery and the sweep now reconcile before re-sending: a pending delivery that already holds a successful `DeliveryResult` is marked delivered rather than sent again, which is the state that did not previously exist. A delivery handed back out is claimed by compare-and-set so successive sweeps cannot send it repeatedly, and it continues its own attempt numbering instead of restarting at 1. The sweep gains a `pending`-with-age-bound arm, so a stranded delivery no longer waits for a restart. Docs. WAL produces `-wal`/`-shm` sidecars, so the backup and restore procedures in README.md are corrected: both documented procedures were re-run against a live instance, and a `-wal` left by a crash carries data the `.db` alone does not. Verified by reproducing the failure on unmodified `next` first — 6 targets, 60 events at 5/s, a concurrent `.dump` reader — which gave 38 HTTP 500s and 112 duplicate POSTs at the sinks across a restart. Both arms of the matched pair now show 0 inbound 500s, 0 engine write errors, and 0 new requests at the sinks after a restart, counted by payload.
This commit is contained in:
147
internal/database/sqlite_open.go
Normal file
147
internal/database/sqlite_open.go
Normal file
@@ -0,0 +1,147 @@
|
||||
package database
|
||||
|
||||
import (
|
||||
"database/sql"
|
||||
"fmt"
|
||||
"net/url"
|
||||
"time"
|
||||
|
||||
_ "modernc.org/sqlite" // Pure Go SQLite driver
|
||||
)
|
||||
|
||||
// Every SQLite file this service opens — the main database, the
|
||||
// per-webhook event databases, and the archive databases — is opened
|
||||
// through OpenSQLite, so the durability settings below are properties
|
||||
// of the service rather than of one call site.
|
||||
//
|
||||
// modernc.org/sqlite installs no busy handler and issues no pragmas of
|
||||
// its own: it executes only the pragmas named in explicit `_pragma=`
|
||||
// DSN parameters, and gorm.io/driver/sqlite adds none when it is
|
||||
// handed an existing *sql.DB. Every setting therefore has to be
|
||||
// spelled out here or it is simply not in effect.
|
||||
// SQLite URI open modes.
|
||||
const (
|
||||
// SQLiteModeCreate creates the database file when it is missing.
|
||||
SQLiteModeCreate = "rwc"
|
||||
|
||||
// SQLiteModeExisting requires the file to exist already.
|
||||
SQLiteModeExisting = "rw"
|
||||
)
|
||||
|
||||
const (
|
||||
// SQLiteBusyTimeout is how long SQLite retries a lock conflict
|
||||
// before returning SQLITE_BUSY.
|
||||
//
|
||||
// Under WAL a reader never blocks a writer, so the only conflict
|
||||
// left is writer against writer: this process's delivery workers
|
||||
// against each other, or against another process holding the write
|
||||
// lock. Those clear in milliseconds. Ten seconds is far above that
|
||||
// and still well inside the receiver's request budget, so an
|
||||
// inbound webhook waits rather than being rejected with a 500.
|
||||
SQLiteBusyTimeout = 10 * time.Second
|
||||
|
||||
// sqliteMaxOpenConns bounds the connection pool for one database
|
||||
// file.
|
||||
//
|
||||
// The pool needs a bound at all because database/sql cannot detect
|
||||
// a connection left mid-transaction: modernc.org/sqlite implements
|
||||
// neither driver.Validator nor driver.SessionResetter, so a
|
||||
// connection whose COMMIT failed is returned to the pool with its
|
||||
// transaction still open and handed out again indefinitely. That is
|
||||
// what turned four `database is locked` errors into 593
|
||||
// `cannot start a transaction within a transaction` in
|
||||
// https://git.eeqj.de/sneak/webhooker/issues/256.
|
||||
//
|
||||
// Four is above the one writer SQLite allows at a time, so reads
|
||||
// still proceed while a write is in flight, and low enough that
|
||||
// contention is resolved by the busy handler rather than by piling
|
||||
// up connections against a lock only one of them can hold.
|
||||
sqliteMaxOpenConns = 4
|
||||
|
||||
// sqliteMaxIdleConns keeps the pool warm without holding every
|
||||
// connection open through an idle period.
|
||||
sqliteMaxIdleConns = 2
|
||||
|
||||
// sqliteConnMaxLifetime and sqliteConnMaxIdleTime retire pooled
|
||||
// connections on a schedule. With _txlock=immediate a failed
|
||||
// COMMIT should no longer be reachable, but these bound the damage
|
||||
// if one happens anyway: a poisoned connection is closed and
|
||||
// replaced within the lifetime instead of wedging the file until
|
||||
// the process restarts.
|
||||
sqliteConnMaxLifetime = 5 * time.Minute
|
||||
sqliteConnMaxIdleTime = time.Minute
|
||||
)
|
||||
|
||||
// SQLiteDSN builds the connection string for one database file.
|
||||
//
|
||||
// mode is the SQLite URI open mode: "rwc" to create the file when it
|
||||
// is missing, "rw" to require that it already exists.
|
||||
//
|
||||
// Three settings carry the fix for
|
||||
// https://git.eeqj.de/sneak/webhooker/issues/256 and none of them is
|
||||
// optional:
|
||||
//
|
||||
// - journal_mode=WAL, so a reader — an operator running
|
||||
// `sqlite3 <db> .dump` over their own data — takes a snapshot
|
||||
// instead of blocking every writer behind it.
|
||||
//
|
||||
// - busy_timeout, so a writer that does meet a lock waits for it.
|
||||
// Without one SQLite gives up immediately; nothing above it
|
||||
// retries.
|
||||
//
|
||||
// - _txlock=immediate, so every transaction takes the write lock at
|
||||
// BEGIN. A deferred transaction acquires it lazily on its first
|
||||
// write, and that upgrade returns SQLITE_BUSY *without* consulting
|
||||
// the busy handler, because SQLite cannot block a transaction that
|
||||
// may already hold a read snapshot. Such a COMMIT then fails while
|
||||
// the transaction stays open on the connection. A busy timeout
|
||||
// alone does not prevent this; BEGIN IMMEDIATE does, by putting
|
||||
// the wait somewhere the handler applies.
|
||||
//
|
||||
// Note what is absent: `cache=shared`. Under a shared cache an
|
||||
// in-process conflict is reported as SQLITE_LOCKED rather than
|
||||
// SQLITE_BUSY, and the busy handler does not retry SQLITE_LOCKED — so
|
||||
// leaving it in would have defeated the busy timeout for exactly the
|
||||
// contention this service generates. Dropping it is part of the fix,
|
||||
// not housekeeping.
|
||||
//
|
||||
// synchronous is deliberately left at SQLite's default of FULL: this
|
||||
// is a webhook receiver whose one promise is that an event it answered
|
||||
// 200 for is durable.
|
||||
func SQLiteDSN(path, mode string) string {
|
||||
q := url.Values{}
|
||||
q.Set("mode", mode)
|
||||
q.Set("_txlock", "immediate")
|
||||
q.Add("_pragma", "journal_mode(WAL)")
|
||||
q.Add(
|
||||
"_pragma",
|
||||
fmt.Sprintf(
|
||||
"busy_timeout(%d)",
|
||||
SQLiteBusyTimeout.Milliseconds(),
|
||||
),
|
||||
)
|
||||
|
||||
return "file:" + path + "?" + q.Encode()
|
||||
}
|
||||
|
||||
// OpenSQLite opens the SQLite file at path with the service's
|
||||
// durability settings and pool bounds applied. mode is the SQLite URI
|
||||
// open mode ("rwc" or "rw").
|
||||
//
|
||||
// The handle is returned rather than a *gorm.DB because the callers
|
||||
// wrap it in gorm themselves with their own logger.
|
||||
func OpenSQLite(path, mode string) (*sql.DB, error) {
|
||||
sqlDB, err := sql.Open("sqlite", SQLiteDSN(path, mode))
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf(
|
||||
"opening sqlite database %s: %w", path, err,
|
||||
)
|
||||
}
|
||||
|
||||
sqlDB.SetMaxOpenConns(sqliteMaxOpenConns)
|
||||
sqlDB.SetMaxIdleConns(sqliteMaxIdleConns)
|
||||
sqlDB.SetConnMaxLifetime(sqliteConnMaxLifetime)
|
||||
sqlDB.SetConnMaxIdleTime(sqliteConnMaxIdleTime)
|
||||
|
||||
return sqlDB, nil
|
||||
}
|
||||
Reference in New Issue
Block a user