2 Commits
Author SHA1 Message Date
sneak af804be45c Report or refuse each unusable file webhooker reads (closes #290)
check / check (push) Successful in 3m26s
Audit of the files webhooker reads configuration or required state
from. A missing or zero-length database is reported with the "created
a new, empty database" warning and its path: webhooker.db at start,
and a per-webhook database at the latest at the next start, since
restart recovery now opens every webhook's database. The main
database's open errors name webhooker.db
(#459). webhooker resetpw
refuses a zero-length webhooker.db as it refuses a missing one. A
directory in place of a database file or its -wal or -shm is refused,
naming it; beside a -shm directory SQLite opened the database
read-only without a word. The README says how each case is treated.

Model: opus-5-5
2026-10-02 18:12:33 +00:00
clawbot 0945831442 Clamp the HTTP drain by the tail-hook reserve (closes #170)
check / check (push) Successful in 3m17s
The HTTP drain at shutdown waited up to ShutdownTimeout regardless of how much of the stop budget earlier hooks had used, so a slow archive sweeper or retention reaper could eat the reserve the hooks after the server need, and the database close was skipped. The drain now waits at most the shorter of ShutdownTimeout and what is left of the budget less TailHookReserve, as the Sentry flush already does. The reserve is documented as derived from the two timeouts. Tests cover earlier hooks having spent part of the budget, on a clock that host speed cannot move, and pin that a drain on the full budget gets all of ShutdownTimeout.

Model: opus-5-5
2026-10-02 20:11:43 +02:00
18 changed files with 585 additions and 95 deletions
+68 -32
View File
@@ -79,7 +79,8 @@ directory, read once at startup before anything else looks at the
environment.
The file is optional and having none is the normal case for a
deployment. A file that is there but cannot be parsed aborts startup
deployment. An empty file is the same as none: it has nothing in it to
apply. A file that is there but cannot be parsed aborts startup
with a message naming it, because a single malformed line makes none
of the file apply: every variable in it silently reverts to its
default, which is exactly the failure [Invalid values abort
@@ -564,11 +565,13 @@ its Argon2id hash. There is no second account and no forgot-password
flow, so the banner and the reset command below are the only two ways
in.
A start that finds no `webhooker.db` in `DATA_DIR` also logs
A start that finds no `webhooker.db` in `DATA_DIR`, or a zero-length
one (which SQLite opens as an empty database), also logs
`created a new, empty database` at `WARN`, with the file's path,
shortly before the banner. On a deployment that has run before, that
line means `DATA_DIR` was empty, most often because its volume is not
mounted.
line means `webhooker.db` was lost: either `DATA_DIR` was empty, most
often because its volume is not mounted, or the file was zero-length,
as a truncated copy leaves it.
#### Recovering a lost admin password
@@ -612,7 +615,8 @@ What it will not do:
the old password, so a reset underneath it would report a change the
service does not honour.
- **Create anything.** A `DATA_DIR` that does not exist, or that holds
no `webhooker.db`, is an error rather than a new empty deployment —
no `webhooker.db` or a zero-length one, is an error naming the path
rather than a new empty deployment —
a mistyped path must not be built out and then reported as a success.
- **Create an account.** A username that does not exist is an error.
`resetpw` changes an existing account's password and nothing else.
@@ -993,6 +997,16 @@ its sidecars; a killed or crashed instance leaves them, and they must be
carried with the `.db`. An archive the service has not opened since a
crash keeps that crash's sidecars, even across a later clean stop.
A missing sidecar is therefore normal, and SQLite makes new ones, so a
`-wal` lost from a copy cannot be reported: the transactions it held
are simply gone. SQLite reads a `-wal` up to its first damaged frame,
as after a crash, and rebuilds a damaged `-shm`. A sidecar with the
wrong mode is set back to `0600` when its database is opened. A
directory in place of either is refused then, with an error naming it:
for `webhooker.db` the server and `webhooker resetpw` stop, and an event
or archive database fails as a damaged one does (see
[Database Architecture](#database-architecture)).
Configuration is **not** in `DATA_DIR` — it comes from the environment
and from a `.env` file read out of the process working directory. Back
that up with your deployment config, separately.
@@ -1075,13 +1089,17 @@ with any `-wal`/`-shm` beside it, or wait until there are none.
1. Stop the service.
2. Restore the **whole set together**: `webhooker.db` *and* every
`events-*.db` *and* every `archive-*.db`. A partial restore fails
quietly rather than loudly. Every database is opened `mode=rwc`, so a
missing `events-{uuid}.db` is **created empty** on first access
instead of erroring — the webhook comes back with its configuration
intact and its entire event history silently gone. Event databases
restored without `webhooker.db` are simply orphaned; nothing
references their UUIDs.
`events-*.db` *and* every `archive-*.db`. A partial restore is
reported, not refused. Every database is opened `mode=rwc`, so a
missing `events-{uuid}.db` is **created empty**: the webhook comes
back with its configuration intact and its entire event history
gone. The first start after the restore logs
`created a new, empty database` at `WARN` for each such file, with
its path, as it does for a missing `webhooker.db`. A missing
`archive-*.db` is recreated at its target's next delivery without a
warning, since moving one away is a supported workflow. Event
databases restored without `webhooker.db` are simply orphaned;
nothing references their UUIDs.
3. Carry any `*.db-wal` and `*.db-shm` files that are in the backup.
They are part of the database, and dropping a `-wal` silently
@@ -1939,10 +1957,19 @@ encryption key is generated and stored, and an `admin` user is created.
the deliveries per target, kept through retention
Per-webhook databases are created automatically when a webhook is
created (and lazily on first access for webhooks that predate this
feature). They are managed by the `WebhookDBManager` component, which
created. They are managed by the `WebhookDBManager` component, which
handles connection pooling, lazy opening, migrations, and cleanup.
A per-webhook database that is missing or zero-length later means its
webhook's events and pending deliveries are gone. The next time it is
opened, an empty one is created in its place, so the webhook keeps
receiving, and `created a new, empty database` is logged at `WARN` with
the file's path. Every webhook's database is opened when the service
starts, so this appears at the latest at the first start after the
file was lost. A file there that SQLite cannot open fails that
webhook alone, with an `ERROR` naming the webhook on every access and a
500 to its senders, so one damaged file does not stop the others.
This separation provides:
- **Isolation** — a high-volume webhook won't cause lock contention or
@@ -2007,7 +2034,9 @@ After each write the archive handle is closed
and reopened, debounced to at most once per second, so an operator can
move the archive file away for offline archiving without stopping the
service; a moved or removed archive file is recreated automatically on
the next write. An optional `expiry` in the target's config JSON (e.g.
the next write. A zero-length archive file is written to as a new
archive: SQLite opens it as an empty database, so it holds nothing to
lose. An optional `expiry` in the target's config JSON (e.g.
`{"expiry":"720h"}`) is validated when the target is created — the
default (unset or the literal `never`) keeps rows forever — and rows
older than the expiry are pruned each time the archive is (re)opened. An
@@ -3245,9 +3274,9 @@ each hook. The order, read off the fx stop-hook log:
1. `ArchiveSweeper`
2. `RetentionReaper`
3. `server` — the HTTP drain, bounded separately by
`server.ShutdownTimeout` (**3 seconds**), then a Sentry flush if
`SENTRY_DSN` is set
3. `server` — the HTTP drain, bounded by `server.ShutdownTimeout`
(**3 seconds**) and by what the hooks before it left, then a Sentry
flush if `SENTRY_DSN` is set
4. `delivery.Engine` — waits for its workers, then closes the archive
databases
5. `healthcheck`
@@ -3267,23 +3296,30 @@ exhaust the sequence budget at the instant it finished, and every
later hook — the delivery engine, the healthcheck, the webhook DB
manager and the database close — would be skipped in exactly the
case where the drain mattered. 3 seconds leaves 2 seconds
(`server.TailHookReserve`) for the tail, which is far more than the
microseconds it needs.
(`server.TailHookReserve`) for the tail. The reserve is that
remainder, not a figure sized to the tail, which takes about a
millisecond.
That reserve belongs to the tail hooks, not to the server hook, and
the Sentry flush is what could take it: it runs after the drain
**inside the same hook**, and `sentry.Flush` takes a bare duration
and honours no context, so an unreachable Sentry endpoint would add
its own timeout on top of a full-length drain and consume the whole
sequence budget by itself. It is therefore clamped to whatever is
left on the stop context minus the reserve, and skipped when that
leaves too little to be worth attempting — so a full-length drain
means Sentry events are dropped rather than the database close being
skipped.
the server hook could take it in two ways. The hooks before it may
already have spent part of the budget, so a full 3-second drain
would come out of the reserve; the drain is therefore also bounded
by whatever is left on the stop context minus the reserve. And the
Sentry flush runs after the drain **inside the same hook**, and
`sentry.Flush` takes a bare duration and honours no context, so an
unreachable Sentry endpoint would add its own timeout on top of a
full-length drain and consume the whole sequence budget by itself.
It is clamped the same way, and skipped when that leaves too little
to be worth attempting — so a full-length drain means Sentry events
are dropped rather than the database close being skipped.
This does not make the database close unconditional: a wedged
`ArchiveSweeper` or `RetentionReaper` still runs first and can
consume the whole budget on its own.
This does not make the database close unconditional. A slow
`ArchiveSweeper` or `RetentionReaper` is enough to cut the shutdown
short, not only one that consumes the whole budget: what they spend
comes out of the drain first, so after 2 seconds of theirs a request
still in flight gets 1 second to finish, and after 3 it gets none.
Past 3 seconds they spend the reserve itself, and one that takes the
whole budget skips every hook after it, the database close included.
The value is chosen to sit inside the container stop grace period.
Docker's default `docker stop` grace is 10 seconds and the Dockerfile
+9 -7
View File
@@ -38,17 +38,19 @@ import (
// hook that used the whole budget would exhaust it at that instant,
// and fx would skip every hook after the server — the delivery
// engine, the healthcheck, the webhook DB manager and the database
// close. That hook is the 3s HTTP drain plus the Sentry flush that
// follows it in the same hook, so the flush is clamped to the stop
// close. That hook is the HTTP drain plus the Sentry flush that
// follows it in the same hook, and each is clamped to the stop
// context's remaining time less server.TailHookReserve rather than
// running for its own fixed 2s; the reserve is what the tail hooks
// live on, and they are microsecond-scale in normal operation.
// running for its own fixed 3s and 2s; the reserve is what the tail
// hooks live on, and they are microsecond-scale in normal operation.
// TestStopTimeout_LeavesHeadroomForTailHooks pins the arithmetic
// across every drain length.
// across every drain length and every amount of budget the hooks
// before the server may already have spent.
//
// This does not make the database close unconditional: the
// ArchiveSweeper and RetentionReaper hooks run before the server
// and can still consume the whole budget on their own.
// ArchiveSweeper and RetentionReaper hooks run before the server.
// What they spend comes out of the drain first, but past 3s it comes
// out of the reserve, and they can consume the whole budget.
const stopTimeout = 5 * time.Second
// exitUsage is the status for a command line this binary cannot make
+27 -9
View File
@@ -252,22 +252,40 @@ const tailHeadroom = 2 * time.Second
// can produce, since a shorter drain leaves the flush more room and
// the worst case is not necessarily at either extreme.
//
// Shrinking either budget, or unbounding the flush again, must fail
// here rather than silently recreating a hook that swallows the
// whole sequence.
// Nor does the hook start on a full budget: the ArchiveSweeper and
// RetentionReaper hooks run before it, and whatever they spent is
// gone. The outer sweep walks every amount they can spend. Once they
// have eaten into the headroom themselves, the hook must spend
// nothing of what is left. A drain that starts on the full budget
// must still get all of ShutdownTimeout, so a smaller stopTimeout
// cannot silently shorten every drain.
//
// Shrinking either budget, or unbounding the drain or the flush
// again, must fail here rather than silently recreating a hook that
// swallows the whole sequence.
func TestStopTimeout_LeavesHeadroomForTailHooks(t *testing.T) {
t.Parallel()
require.Less(t, server.ShutdownTimeout, stopTimeout)
require.Equal(
t, server.ShutdownTimeout, server.DrainBudget(stopTimeout),
"a drain that starts on the full stop budget is cut short",
)
const step = 10 * time.Millisecond
for drain := time.Duration(0); drain <= server.ShutdownTimeout; drain += step {
hook := drain + server.SentryFlushBudget(stopTimeout-drain)
for spent := time.Duration(0); spent <= stopTimeout; spent += step {
remaining := stopTimeout - spent
longest := max(server.DrainBudget(remaining), 0)
require.LessOrEqual(
t, hook+tailHeadroom, stopTimeout,
"a %s drain leaves the tail hooks short", drain,
)
for drain := time.Duration(0); drain <= longest; drain += step {
hook := drain + server.SentryFlushBudget(remaining-drain)
require.GreaterOrEqual(
t, remaining-hook, min(remaining, tailHeadroom),
"a %s drain after %s of earlier hooks leaves "+
"the tail hooks short", drain, spent,
)
}
}
}
@@ -4,6 +4,7 @@ import (
"bytes"
"context"
"log/slog"
"os"
"path/filepath"
"strings"
"testing"
@@ -119,3 +120,26 @@ func TestNewDatabase_IsLoggedWithItsPath(t *testing.T) {
t, second, created, "an existing database is not new",
)
}
// TestZeroLengthDatabase_IsLoggedAsNew covers what
// https://git.eeqj.de/sneak/webhooker/issues/290 found: SQLite opens a
// zero-length file as an empty database, so a start on one is a first
// start, and it must say so exactly as a start with no file does.
func TestZeroLengthDatabase_IsLoggedAsNew(t *testing.T) {
t.Parallel()
dir := t.TempDir()
path := filepath.Join(dir, database.MainDBFileName)
require.NoError(t, os.WriteFile(path, nil, database.SQLiteFilePerm))
var out bytes.Buffer
db, err := database.Open(dir, slog.New(slog.NewTextHandler(&out, nil)))
require.NoError(t, err)
require.NoError(t, db.Close())
assert.Contains(
t, out.String(),
`level=WARN msg="created a new, empty database" path=`+path,
)
}
+12 -7
View File
@@ -8,7 +8,6 @@ import (
"errors"
"fmt"
"io"
"io/fs"
"log/slog"
"os"
"path/filepath"
@@ -203,8 +202,7 @@ func (d *Database) connectTo(dataDir string) error {
// Checked before opening, which creates the file. A DATA_DIR that
// is unexpectedly empty -- its volume not mounted, say -- looks
// exactly like a first start, so a new database is a warning.
_, statErr := os.Stat(dbPath)
created := errors.Is(statErr, fs.ErrNotExist)
created := missingOrEmpty(dbPath)
// Opened through OpenSQLite so this handle carries the same WAL
// journaling, busy timeout, immediate-transaction locking, and pool
@@ -213,13 +211,15 @@ func (d *Database) connectTo(dataDir string) error {
if err != nil {
d.log.Error(
"failed to open database",
"path", dbPath,
"error", err,
)
return err
}
// Then use it with GORM
// Then use it with GORM. Its errors are SQLite's alone and name no
// file, so the path is added to them here.
db, err := gorm.Open(sqlite.Dialector{
Conn: sqlDB,
}, &gorm.Config{
@@ -229,10 +229,11 @@ func (d *Database) connectTo(dataDir string) error {
if err != nil {
d.log.Error(
"failed to connect to database",
"path", dbPath,
"error", err,
)
return err
return fmt.Errorf("connecting to %s: %w", dbPath, err)
}
d.db = db
@@ -243,8 +244,12 @@ func (d *Database) connectTo(dataDir string) error {
d.log.Info("connected to database", "path", dbPath)
}
// Run migrations
return d.migrate()
err = d.migrate()
if err != nil {
return fmt.Errorf("migrating %s: %w", dbPath, err)
}
return nil
}
func (d *Database) migrate() error {
+25
View File
@@ -1,9 +1,15 @@
package database_test
import (
"bytes"
"context"
"log/slog"
"os"
"path/filepath"
"testing"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
"go.uber.org/fx/fxtest"
"sneak.berlin/go/webhooker/internal/config"
"sneak.berlin/go/webhooker/internal/database"
@@ -100,3 +106,22 @@ func TestDatabaseConnection(t *testing.T) {
)
}
}
// TestOpen_UnreadableDatabaseIsNamed pins
// https://git.eeqj.de/sneak/webhooker/issues/459: when SQLite cannot
// read webhooker.db, the error that stops the server and `webhooker
// resetpw` names the file, not only SQLite's own message.
func TestOpen_UnreadableDatabaseIsNamed(t *testing.T) {
t.Parallel()
dir := t.TempDir()
path := filepath.Join(dir, database.MainDBFileName)
require.NoError(t, os.WriteFile(
path, bytes.Repeat([]byte("junk"), 1024), database.SQLiteFilePerm,
))
_, err := database.Open(dir, slog.New(slog.DiscardHandler))
require.Error(t, err)
assert.Contains(t, err.Error(), path)
assert.Contains(t, err.Error(), "file is not a database")
}
+2 -2
View File
@@ -184,8 +184,8 @@ func (r *RetentionReaper) sweep(ctx context.Context) {
wh := webhooks[i]
// Nothing to reap if the per-webhook database has never
// been created.
// A missing database has nothing to reap. Restart recovery
// reports a lost one (see WebhookDBManager.GetDB).
if !r.dbManager.DBExists(wh.ID) {
continue
}
+23
View File
@@ -182,6 +182,29 @@ func TestOpenSQLiteTightensFilesLeftWorldReadable(t *testing.T) {
requireDatabaseSetOwnerOnly(t, path)
}
// TestOpenSQLiteRefusesADirectorySidecar covers a directory in place
// of -wal or -shm. Beside a -shm directory SQLite opens the database
// read-only without a word, and every write then fails naming no file,
// so the open must stop instead, naming the directory.
func TestOpenSQLiteRefusesADirectorySidecar(t *testing.T) {
t.Parallel()
for _, suffix := range []string{"-wal", "-shm"} {
t.Run(suffix, func(t *testing.T) {
t.Parallel()
path := filepath.Join(t.TempDir(), database.MainDBFileName)
require.NoError(t, os.Mkdir(path+suffix, 0o700))
_, err := database.OpenSQLite(
path, database.SQLiteModeCreate,
)
require.Error(t, err)
assert.Contains(t, err.Error(), path+suffix)
})
}
}
// TestOpenSQLiteExistingModeDoesNotCreateTheFile guards the mechanism
// the fix uses: OpenSQLite now creates the database file itself, and
// must not do so for a caller that asked for an existing database. An
+26 -2
View File
@@ -7,6 +7,7 @@ import (
"io/fs"
"net/url"
"os"
"syscall"
"time"
_ "modernc.org/sqlite" // Pure Go SQLite driver
@@ -93,7 +94,8 @@ const (
const SQLiteFilePerm fs.FileMode = 0o600
// reserveSQLiteFile puts path at SQLiteFilePerm before the driver ever
// touches it, and tightens any sidecar already on disk.
// touches it, and tightens any sidecar already on disk. A directory in
// place of any of them is an error naming it.
//
// The mode has to be settled here rather than by a chmod after opening,
// because SQLite picks it: robust_open substitutes
@@ -143,7 +145,15 @@ func reserveSQLiteFile(path string, create bool) error {
for _, p := range append(
[]string{path}, sqliteSidecarPaths(path)...,
) {
err := os.Chmod(p, SQLiteFilePerm)
// Chmod accepts a directory, and SQLite opens a database whose
// -shm is one read-only, without a word: every write then
// fails naming no file.
info, err := os.Stat(p)
if err == nil && info.IsDir() {
return fmt.Errorf("securing %s: %w", p, syscall.EISDIR)
}
err = os.Chmod(p, SQLiteFilePerm)
if err != nil && !errors.Is(err, fs.ErrNotExist) {
return fmt.Errorf("securing %s: %w", p, err)
}
@@ -152,6 +162,20 @@ func reserveSQLiteFile(path string, create bool) error {
return nil
}
// missingOrEmpty reports whether opening path in SQLiteModeCreate
// would start a new, empty database: the file is not there, or it is
// zero-length, which SQLite opens as an empty database. A file left at
// zero length by an interrupted first start or a truncated copy holds
// as little as a missing one, and must be reported the same way.
func missingOrEmpty(path string) bool {
info, err := os.Stat(path)
if errors.Is(err, fs.ErrNotExist) {
return true
}
return err == nil && info.Size() == 0
}
// sqliteSidecarPaths returns the files SQLite maintains beside a
// database under WAL. They carry the same rows as the database itself,
// so a fix that tightens only the main file has fixed nothing.
+54 -28
View File
@@ -98,34 +98,18 @@ func NewWebhookDBManager(
return m, nil
}
// GetDB returns the database connection for a webhook,
// creating the database file lazily if it doesn't exist.
// GetDB returns the database connection for a webhook, opening it on
// first use.
//
// The file is made by CreateDB when the webhook is created. One that is
// missing or zero-length here means the webhook's events and pending
// deliveries are gone: an empty database is created in its place so
// the webhook keeps receiving, and that is logged as a warning naming
// the file, as a new main database is.
func (m *WebhookDBManager) GetDB(
webhookID string,
) (*gorm.DB, error) {
// Fast path: already open
if val, ok := m.dbs.Load(webhookID); ok {
return asGormDB(val, webhookID)
}
// Slow path: open the database under the lock, looking in the
// cache again first. A caller that raced another one here then
// waits for its handle instead of opening a second one.
m.mu.Lock()
defer m.mu.Unlock()
if val, ok := m.dbs.Load(webhookID); ok {
return asGormDB(val, webhookID)
}
db, err := m.openDB(webhookID)
if err != nil {
return nil, err
}
m.dbs.Store(webhookID, db)
return db, nil
return m.getDB(webhookID, false)
}
// asGormDB returns a value read from the cache as the database
@@ -143,12 +127,12 @@ func asGormDB(val any, webhookID string) (*gorm.DB, error) {
return db, nil
}
// CreateDB explicitly creates a new per-webhook database file
// and runs migrations.
// CreateDB creates a new webhook's database file and runs
// migrations.
func (m *WebhookDBManager) CreateDB(
webhookID string,
) error {
_, err := m.GetDB(webhookID)
_, err := m.getDB(webhookID, true)
return err
}
@@ -266,6 +250,48 @@ func (m *WebhookDBManager) DBPath(
return m.dbPath(webhookID)
}
// getDB is GetDB, and CreateDB when isNew is true: the webhook has just
// been created, so a missing file is expected rather than lost.
func (m *WebhookDBManager) getDB(
webhookID string, isNew bool,
) (*gorm.DB, error) {
// Fast path: already open
if val, ok := m.dbs.Load(webhookID); ok {
return asGormDB(val, webhookID)
}
// Slow path: open the database under the lock, looking in the
// cache again first. A caller that raced another one here then
// waits for its handle instead of opening a second one.
m.mu.Lock()
defer m.mu.Unlock()
if val, ok := m.dbs.Load(webhookID); ok {
return asGormDB(val, webhookID)
}
// Checked before opening, which creates the file. See GetDB.
path := m.dbPath(webhookID)
replaced := !isNew && missingOrEmpty(path)
db, err := m.openDB(webhookID)
if err != nil {
return nil, err
}
if replaced {
m.log.Warn(
"created a new, empty database",
"webhook_id", webhookID,
"path", path,
)
}
m.dbs.Store(webhookID, db)
return db, nil
}
func (m *WebhookDBManager) dbPath(
webhookID string,
) string {
@@ -289,6 +289,75 @@ func TestWebhookDBManager_LazyCreation(t *testing.T) {
assert.True(t, mgr.DBExists(webhookID))
}
// A webhook's database is made by CreateDB along with the webhook. One
// that GetDB finds missing or zero-length has lost the webhook's events
// and pending deliveries, so the empty database made in its place is
// logged as a warning naming the file
// (https://git.eeqj.de/sneak/webhooker/issues/290). CreateDB, and
// reopening a database that is there, log no such warning.
func TestWebhookDBManager_LostDatabaseIsLogged(t *testing.T) {
t.Parallel()
const created = `level=WARN msg="created a new, empty database"`
open := func(
t *testing.T, prepare func(*database.WebhookDBManager, string),
) (string, string) {
t.Helper()
var logs bytes.Buffer
mgr := database.NewTestWebhookDBManagerWithLogger(
t.TempDir(),
slog.New(slog.NewTextHandler(&logs, nil)),
)
webhookID := uuid.New().String()
prepare(mgr, webhookID)
_, err := mgr.GetDB(webhookID)
require.NoError(t, err)
require.NoError(t, mgr.CloseAll())
return logs.String(),
" webhook_id=" + webhookID + " path=" + mgr.DBPath(webhookID)
}
t.Run("missing", func(t *testing.T) {
t.Parallel()
logs, fields := open(
t, func(*database.WebhookDBManager, string) {},
)
assert.Contains(t, logs, created+fields)
})
t.Run("zero-length", func(t *testing.T) {
t.Parallel()
logs, fields := open(
t, func(mgr *database.WebhookDBManager, webhookID string) {
require.NoError(t, os.WriteFile(
mgr.DBPath(webhookID), nil, database.SQLiteFilePerm,
))
},
)
assert.Contains(t, logs, created+fields)
})
t.Run("created with the webhook, then reopened", func(t *testing.T) {
t.Parallel()
logs, _ := open(
t, func(mgr *database.WebhookDBManager, webhookID string) {
require.NoError(t, mgr.CreateDB(webhookID))
require.NoError(t, mgr.CloseAll())
},
)
assert.NotContains(t, logs, created)
})
}
func TestWebhookDBManager_DeliveryWorkflow(t *testing.T) {
t.Parallel()
+3 -4
View File
@@ -699,10 +699,9 @@ func (e *Engine) recoverInFlight(ctx context.Context) {
default:
}
if !e.dbManager.DBExists(webhookID) {
continue
}
// Opened even when its file is missing, so that GetDB reports
// a lost database at start, not when the webhook next receives
// an event, which for a quiet webhook may be never.
e.recoverWebhookDeliveries(ctx, webhookID)
}
}
@@ -1,6 +1,7 @@
package delivery_test
import (
"bytes"
"context"
"encoding/json"
"fmt"
@@ -1134,6 +1135,40 @@ func TestRecoverInFlight_WithPendingDeliveries(
}
}
// TestRecoverInFlight_ReportsAMissingWebhookDatabase covers a webhook
// whose database file is gone, after a partial restore say. Restart
// recovery opens every webhook's database, so the empty one made in its
// place is reported at start, naming the file
// (https://git.eeqj.de/sneak/webhooker/issues/290).
func TestRecoverInFlight_ReportsAMissingWebhookDatabase(t *testing.T) {
t.Parallel()
mainDB := iMainDB(t)
webhookID := uuid.New().String()
iCreateWebhook(t, mainDB, webhookID, "lost-database")
var logs bytes.Buffer
dbMgr := database.NewTestWebhookDBManagerWithLogger(
t.TempDir(), slog.New(slog.NewTextHandler(&logs, nil)),
)
t.Cleanup(func() { _ = dbMgr.CloseAll() })
engine := delivery.NewTestEngineWithDB(
database.NewTestDatabase(mainDB), dbMgr,
slog.New(slog.DiscardHandler),
&http.Client{Timeout: 5 * time.Second}, 1,
)
engine.ExportRecoverInFlight(context.Background())
assert.Contains(
t, logs.String(),
`level=WARN msg="created a new, empty database" webhook_id=`+
webhookID+" path="+dbMgr.DBPath(webhookID),
)
}
// --- HTTP Config with custom headers ---
func TestDeliverHTTP_CustomTargetHeaders(t *testing.T) {
+10 -1
View File
@@ -308,7 +308,7 @@ func checkDataDir(dir string) error {
dbPath := filepath.Join(dir, database.MainDBFileName)
_, err = os.Stat(dbPath)
dbInfo, err := os.Stat(dbPath)
switch {
case errors.Is(err, fs.ErrNotExist):
@@ -319,6 +319,15 @@ func checkDataDir(dir string) error {
)
case err != nil:
return fmt.Errorf("checking %s: %w", dbPath, err)
case dbInfo.Size() == 0:
// SQLite opens a zero-length file as an empty database, so
// it holds no deployment either, and opening it would write
// an empty schema into it.
return fmt.Errorf(
"%w: %s is zero-length. The admin account is created by "+
"the first server start",
ErrNoDatabase, dbPath,
)
}
return nil
+27
View File
@@ -377,6 +377,33 @@ func TestMissingDatabaseCreatesNothing(t *testing.T) {
)
}
// TestZeroLengthDatabaseCreatesNothing covers a webhooker.db left at
// zero length, as a truncated copy leaves it. SQLite would open it as
// an empty database, so it is refused like a missing one and left as
// it is.
func TestZeroLengthDatabaseCreatesNothing(t *testing.T) {
dir := t.TempDir()
t.Setenv("DATA_DIR", dir)
dbPath := filepath.Join(dir, database.MainDBFileName)
require.NoError(
t, os.WriteFile(dbPath, nil, database.SQLiteFilePerm),
)
code, _, stderr := run(t, newPassword+"\n", operatorUser)
require.Equal(t, exitFailure, code)
assert.Contains(t, stderr, dbPath)
entries, err := os.ReadDir(dir)
require.NoError(t, err)
assert.Len(t, entries, 1, "nothing may be created beside it")
info, err := os.Stat(dbPath)
require.NoError(t, err)
assert.Zero(t, info.Size(), "nothing may be written into it")
}
// TestUnknownUserFails states the decision: resetpw changes an
// existing account's password and never creates an account. A typo in
// the username must say so rather than quietly adding a second user.
+10
View File
@@ -1,6 +1,8 @@
package server
import (
"context"
"log/slog"
"net/http"
"testing"
@@ -37,6 +39,14 @@ func SentryClientOptionsForTest(
return sentryClientOptions(dsn, release)
}
// CleanShutdownForTest runs the server's stop hook, cleanShutdown,
// against hs: a server the test started itself, so it can hold a
// request open across the drain. Sentry is off.
func CleanShutdownForTest(ctx context.Context, hs *http.Server) {
s := &Server{log: slog.New(slog.DiscardHandler), httpServer: hs}
s.cleanShutdown(ctx)
}
// newServerForTest builds a Server through New, as the application
// does, on a lifecycle that is never started: the hooks New adds to
// it never run, so nothing listens.
+26 -3
View File
@@ -39,6 +39,12 @@ const (
// refuses to spend, leaving it for the hooks that run after the
// server: the delivery engine, the healthcheck, the webhook DB
// manager and the database close.
//
// Its value is not tuned to those hooks, which take about a
// millisecond between them. It is what the 5s fx stop timeout in
// cmd/webhooker leaves after a full ShutdownTimeout drain, so a
// drain that starts on a full budget still gets all of
// ShutdownTimeout.
TailHookReserve = 2 * time.Second
// sentryFlushTimeout is the longest wait for Sentry to flush
@@ -59,6 +65,16 @@ const (
// key off it, and a zero exit would read as a deliberate stop.
const StartupFailureExitCode = 1
// DrainBudget reports how long the HTTP drain may wait for in-flight
// requests when remaining is the time left on the fx stop context as
// the server's stop hook starts. The hooks before the server can
// already have spent part of the budget, so the drain takes its time
// out of what they left, never out of TailHookReserve. Zero or less
// means no wait at all.
func DrainBudget(remaining time.Duration) time.Duration {
return min(ShutdownTimeout, remaining-TailHookReserve)
}
// SentryFlushBudget reports how long the Sentry flush may run when
// remaining is the time left on the fx stop context after the HTTP
// drain. sentry.Flush takes a bare duration and honours no context,
@@ -261,10 +277,17 @@ func (s *Server) cleanupForExit() {
s.log.Info("cleaning up")
}
// cleanShutdown drains the HTTP server and flushes Sentry inside what
// is left of the fx stop budget. A context carrying no deadline — a
// caller outside the fx lifecycle — gets the full ShutdownTimeout.
func (s *Server) cleanShutdown(ctx context.Context) {
ctxShutdown, shutdownCancel := context.WithTimeout(
ctx, ShutdownTimeout,
)
drain := ShutdownTimeout
if deadline, ok := ctx.Deadline(); ok {
drain = DrainBudget(time.Until(deadline))
}
ctxShutdown, shutdownCancel := context.WithTimeout(ctx, drain)
defer shutdownCancel()
err := s.httpServer.Shutdown(ctxShutdown)
+135
View File
@@ -1,13 +1,148 @@
package server_test
import (
"context"
"net"
"net/http"
"testing"
"testing/synctest"
"time"
"github.com/stretchr/testify/require"
"sneak.berlin/go/webhooker/internal/server"
)
// TestDrainBudget covers the clamp that keeps the HTTP drain from
// spending the tail hooks' share of the fx stop budget when the hooks
// before the server have already used part of it.
func TestDrainBudget(t *testing.T) {
t.Parallel()
tests := []struct {
name string
remaining time.Duration
want time.Duration
}{
{
name: "only the reserve is left",
remaining: server.TailHookReserve,
want: 0,
},
{
name: "earlier hooks spent part of the budget",
remaining: server.TailHookReserve + time.Second,
want: time.Second,
},
{
name: "capped at the nominal timeout",
remaining: time.Hour,
want: server.ShutdownTimeout,
},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
t.Parallel()
require.Equal(t, tt.want, server.DrainBudget(tt.remaining))
})
}
}
// TestCleanShutdown_LeavesTailHookReserve stops the server with a
// request still in flight, after the hooks before it have spent all
// of the stop budget but TailHookReserve. The drain must give up at
// once rather than wait for the request: what is left belongs to the
// hooks after the server, the database close among them. A drain
// bounded only by ShutdownTimeout waits until the stop context
// expires, and fx then skips those hooks.
//
// The test runs in a synctest bubble, whose clock moves only while
// every goroutine in it is blocked, so a drain that gives up at once
// leaves the stop context unexpired however slow the host is. The
// request travels over net.Pipe because a goroutine waiting on a
// real socket would stop that clock from moving at all.
func TestCleanShutdown_LeavesTailHookReserve(t *testing.T) {
t.Parallel()
synctest.Test(t, func(t *testing.T) {
entered := make(chan struct{})
release := make(chan struct{})
hs := &http.Server{
Handler: http.HandlerFunc(
func(http.ResponseWriter, *http.Request) {
close(entered)
<-release
},
),
ReadHeaderTimeout: time.Second,
}
srvConn, cliConn := net.Pipe()
listener := pipeListener{
conns: make(chan net.Conn, 1),
closed: make(chan struct{}),
}
listener.conns <- srvConn
go func() { _ = hs.Serve(listener) }()
// Cleanups run last first: the handler returns, then closing
// the client end ends the server's write of the response.
t.Cleanup(func() { _ = cliConn.Close() })
t.Cleanup(func() { close(release) })
_, err := cliConn.Write(
[]byte("GET / HTTP/1.1\r\nHost: webhooker.test\r\n\r\n"),
)
require.NoError(t, err)
<-entered
stopCtx, cancel := context.WithTimeout(
t.Context(), server.TailHookReserve,
)
defer cancel()
server.CleanShutdownForTest(stopCtx, hs)
require.NoError(
t, stopCtx.Err(), "the drain spent the tail hooks' reserve",
)
})
}
// pipeListener is the net.Listener http.Server.Serve needs to serve
// the server end of a net.Pipe: Accept returns that one connection,
// then waits until Close, as a real listener with no more clients
// does.
type pipeListener struct {
conns chan net.Conn
closed chan struct{}
}
func (l pipeListener) Accept() (net.Conn, error) {
select {
case conn := <-l.conns:
return conn, nil
case <-l.closed:
return nil, net.ErrClosed
}
}
func (l pipeListener) Close() error {
close(l.closed)
return nil
}
// Addr is never called by http.Server.Serve.
func (pipeListener) Addr() net.Addr {
return nil
}
// TestSentryFlushBudget covers the clamp that keeps the Sentry flush
// from spending the tail hooks' share of the fx stop budget.
// sentry.Flush ignores the stop context, so without the clamp a