All checks were successful
check / check (push) Successful in 3m3s
An operator running `sqlite3 <db> .dump` against their own per-webhook database wedged it: inbound webhooks rejected with HTTP 500, delivered webhooks stranded at `pending`, and every one of them POSTed a second time on the next restart while the event log recorded a single attempt. Durability. Every SQLite file — main, per-webhook, and archive — now opens through one path, `internal/database/sqlite_open.go`, in WAL journal mode with a 10-second busy timeout, `BEGIN IMMEDIATE` transactions, and a bounded connection pool. WAL is what stops a reader blocking writers at all. `_txlock=immediate` is what stops a `COMMIT` failing while its transaction stays open on a pooled connection, which is how four `database is locked` errors became 593 `cannot start a transaction within a transaction`. `cache=shared` is gone, because under it an in-process conflict is SQLITE_LOCKED, which the busy handler does not retry. The busy timeout is applied before journal_mode: the driver runs DSN pragmas in order on every new connection, and `PRAGMA journal_mode` takes a lock, so the reverse order leaves the one pragma that can block uncovered by the handler meant to cover it. Eligibility. `internal/delivery/inflight.go` holds the set of deliveries the engine owns — taken when a task is queued, when a target schedules a retry, and by every recovery path before it re-dispatches; dropped when the worker that ran the task returns. Recovery and both sweep arms re-dispatch only what the set does not hold. Nothing decides that from a row's age: a delivery waiting in a 10000-deep channel is arbitrarily old and perfectly healthy, and reasoning from age re-sends it. `takeForRedispatch` is the single gate every re-dispatch goes through — ownership first, then a conditional update confirming the row is still in the status the batch read. Bookkeeping. `recordResult` and `updateDeliveryStatus` return their errors instead of logging and dropping them, and a caller whose bookkeeping write failed writes nothing at all: the delivery keeps whichever non-terminal status it already held, and the sweeps recover it. Every recovery path — pending and retrying alike — first settles any delivery that already holds a successful `DeliveryResult` rather than sending it again. Recovery continues each delivery's own attempt numbering instead of restarting at 1. The sweep gains a `pending`-with-age-bound arm, so a stranded delivery no longer waits for a restart. Docs. WAL produces `-wal`/`-shm` sidecars, so the backup and restore procedures in README.md are corrected against measurement: both documented procedures were re-run against a live instance, a `-wal` left by a crash carries data the `.db` alone does not, and an archive file normally holds its rows in a `-wal` rather than in the `.db`.
158 lines
6.2 KiB
Go
158 lines
6.2 KiB
Go
package database
|
|
|
|
import (
|
|
"database/sql"
|
|
"fmt"
|
|
"net/url"
|
|
"time"
|
|
|
|
_ "modernc.org/sqlite" // Pure Go SQLite driver
|
|
)
|
|
|
|
// Every SQLite file this service opens — the main database, the
|
|
// per-webhook event databases, and the archive databases — is opened
|
|
// through OpenSQLite, so the durability settings below are properties
|
|
// of the service rather than of one call site.
|
|
//
|
|
// modernc.org/sqlite installs no busy handler and issues no pragmas of
|
|
// its own: it executes only the pragmas named in explicit `_pragma=`
|
|
// DSN parameters, and gorm.io/driver/sqlite adds none when it is
|
|
// handed an existing *sql.DB. Every setting therefore has to be
|
|
// spelled out here or it is simply not in effect.
|
|
// SQLite URI open modes.
|
|
const (
|
|
// SQLiteModeCreate creates the database file when it is missing.
|
|
SQLiteModeCreate = "rwc"
|
|
|
|
// SQLiteModeExisting requires the file to exist already.
|
|
SQLiteModeExisting = "rw"
|
|
)
|
|
|
|
const (
|
|
// SQLiteBusyTimeout is how long SQLite retries a lock conflict
|
|
// before returning SQLITE_BUSY.
|
|
//
|
|
// Under WAL a reader never blocks a writer, so the only conflict
|
|
// left is writer against writer: this process's delivery workers
|
|
// against each other, or against another process holding the write
|
|
// lock. Those clear in milliseconds. Ten seconds is far above that
|
|
// and still well inside the receiver's request budget, so an
|
|
// inbound webhook waits rather than being rejected with a 500.
|
|
SQLiteBusyTimeout = 10 * time.Second
|
|
|
|
// sqliteMaxOpenConns bounds the connection pool for one database
|
|
// file.
|
|
//
|
|
// The pool needs a bound at all because database/sql cannot detect
|
|
// a connection left mid-transaction: modernc.org/sqlite implements
|
|
// neither driver.Validator nor driver.SessionResetter, so a
|
|
// connection whose COMMIT failed is returned to the pool with its
|
|
// transaction still open and handed out again indefinitely. That is
|
|
// what turned four `database is locked` errors into 593
|
|
// `cannot start a transaction within a transaction` in
|
|
// https://git.eeqj.de/sneak/webhooker/issues/256.
|
|
//
|
|
// Four is above the one writer SQLite allows at a time, so reads
|
|
// still proceed while a write is in flight, and low enough that
|
|
// contention is resolved by the busy handler rather than by piling
|
|
// up connections against a lock only one of them can hold.
|
|
sqliteMaxOpenConns = 4
|
|
|
|
// sqliteMaxIdleConns keeps the pool warm without holding every
|
|
// connection open through an idle period.
|
|
sqliteMaxIdleConns = 2
|
|
|
|
// sqliteConnMaxLifetime and sqliteConnMaxIdleTime retire pooled
|
|
// connections on a schedule. With _txlock=immediate a failed
|
|
// COMMIT should no longer be reachable, but these bound the damage
|
|
// if one happens anyway: a poisoned connection is closed and
|
|
// replaced within the lifetime instead of wedging the file until
|
|
// the process restarts.
|
|
sqliteConnMaxLifetime = 5 * time.Minute
|
|
sqliteConnMaxIdleTime = time.Minute
|
|
)
|
|
|
|
// SQLiteDSN builds the connection string for one database file.
|
|
//
|
|
// mode is the SQLite URI open mode: "rwc" to create the file when it
|
|
// is missing, "rw" to require that it already exists.
|
|
//
|
|
// Three settings carry the fix for
|
|
// https://git.eeqj.de/sneak/webhooker/issues/256 and none of them is
|
|
// optional:
|
|
//
|
|
// - journal_mode=WAL, so a reader — an operator running
|
|
// `sqlite3 <db> .dump` over their own data — takes a snapshot
|
|
// instead of blocking every writer behind it.
|
|
//
|
|
// - busy_timeout, so a writer that does meet a lock waits for it.
|
|
// Without one SQLite gives up immediately; nothing above it
|
|
// retries.
|
|
//
|
|
// - _txlock=immediate, so every transaction takes the write lock at
|
|
// BEGIN. A deferred transaction acquires it lazily on its first
|
|
// write, and that upgrade returns SQLITE_BUSY *without* consulting
|
|
// the busy handler, because SQLite cannot block a transaction that
|
|
// may already hold a read snapshot. Such a COMMIT then fails while
|
|
// the transaction stays open on the connection. A busy timeout
|
|
// alone does not prevent this; BEGIN IMMEDIATE does, by putting
|
|
// the wait somewhere the handler applies.
|
|
//
|
|
// Note what is absent: `cache=shared`. Under a shared cache an
|
|
// in-process conflict is reported as SQLITE_LOCKED rather than
|
|
// SQLITE_BUSY, and the busy handler does not retry SQLITE_LOCKED — so
|
|
// leaving it in would have defeated the busy timeout for exactly the
|
|
// contention this service generates. Dropping it is part of the fix,
|
|
// not housekeeping.
|
|
//
|
|
// synchronous is deliberately left at SQLite's default of FULL: this
|
|
// is a webhook receiver whose one promise is that an event it answered
|
|
// 200 for is durable.
|
|
// The order of the _pragma parameters is load-bearing.
|
|
// modernc.org/sqlite executes them in the order they appear, on every
|
|
// new connection, before the connection is handed to the pool. Setting
|
|
// journal_mode first means that pragma itself runs with no busy
|
|
// handler installed: the pool opens connections lazily, so the moment
|
|
// a new one is created is a moment the database is under load, and
|
|
// PRAGMA journal_mode takes a lock. It would fail immediately with
|
|
// SQLITE_BUSY and fail the query that caused the connection to be
|
|
// opened. busy_timeout is therefore set first, so every pragma after
|
|
// it — and the whole life of the connection — is covered.
|
|
func SQLiteDSN(path, mode string) string {
|
|
q := url.Values{}
|
|
q.Set("mode", mode)
|
|
q.Set("_txlock", "immediate")
|
|
q.Add(
|
|
"_pragma",
|
|
fmt.Sprintf(
|
|
"busy_timeout(%d)",
|
|
SQLiteBusyTimeout.Milliseconds(),
|
|
),
|
|
)
|
|
q.Add("_pragma", "journal_mode(WAL)")
|
|
|
|
return "file:" + path + "?" + q.Encode()
|
|
}
|
|
|
|
// OpenSQLite opens the SQLite file at path with the service's
|
|
// durability settings and pool bounds applied. mode is the SQLite URI
|
|
// open mode ("rwc" or "rw").
|
|
//
|
|
// The handle is returned rather than a *gorm.DB because the callers
|
|
// wrap it in gorm themselves with their own logger.
|
|
func OpenSQLite(path, mode string) (*sql.DB, error) {
|
|
sqlDB, err := sql.Open("sqlite", SQLiteDSN(path, mode))
|
|
if err != nil {
|
|
return nil, fmt.Errorf(
|
|
"opening sqlite database %s: %w", path, err,
|
|
)
|
|
}
|
|
|
|
sqlDB.SetMaxOpenConns(sqliteMaxOpenConns)
|
|
sqlDB.SetMaxIdleConns(sqliteMaxIdleConns)
|
|
sqlDB.SetConnMaxLifetime(sqliteConnMaxLifetime)
|
|
sqlDB.SetConnMaxIdleTime(sqliteConnMaxIdleTime)
|
|
|
|
return sqlDB, nil
|
|
}
|