Close archive writers when the delivery engine stops (closes #280)
check / check (push) Successful in 3m45s
check / check (push) Successful in 3m45s
The engine cached archive writers and never closed them at shutdown, so after a clean stop an archive's rows could sit in its -wal while the .db held no table. The engine's stop hook now evicts every cached writer once its workers have returned, the same way deleting a webhook does, so a clean stop leaves each archive as one file and a late write is refused. If the workers do not return within the stop budget, the writers are left open as a kill would leave them: closing would wait on a write in progress, and a still-running worker would open new ones. The README no longer says archives keep their sidecars across a clean stop. Model: opus-5-5
This commit was merged in pull request #332.
This commit is contained in:
@@ -968,15 +968,10 @@ scratch file**: it holds committed transactions that are not yet in the
|
||||
have no readable schema at all. `-shm` is regenerable, but there is no
|
||||
reason to separate the two — copy the directory and you have them.
|
||||
|
||||
A clean shutdown closes `webhooker.db` and every `events-*.db`, which
|
||||
checkpoints and removes their sidecars; a killed or crashed instance
|
||||
leaves them, and they must be carried with the `.db`. **Archive
|
||||
databases are different**: their handle is not closed at shutdown, so
|
||||
`archive-*.db-wal` and `-shm` normally survive a clean stop and the
|
||||
`-wal` can hold every row the archive has. Measured on a stopped
|
||||
instance: `archive-….db` 4096 bytes with no table, its `-wal` 157 KB
|
||||
holding all 8 archived events. Copying `DATA_DIR` in full is what makes
|
||||
this a non-issue; copying `.db` files out of it by name is not.
|
||||
A clean shutdown closes every database, which checkpoints and removes
|
||||
its sidecars; a killed or crashed instance leaves them, and they must be
|
||||
carried with the `.db`. An archive the service has not opened since a
|
||||
crash keeps that crash's sidecars, even across a later clean stop.
|
||||
|
||||
Configuration is **not** in `DATA_DIR` — it comes from the environment
|
||||
and from a `.env` file read out of the process working directory. Back
|
||||
@@ -1051,10 +1046,9 @@ The file becomes self-contained again when the handle closes, which
|
||||
happens on the next write past the debounce window, when the connection
|
||||
pool retires the idle connection (about a minute after the last write),
|
||||
or at the idle archive sweep — measured, the same file was a complete
|
||||
20 KB `.db` with no sidecars about a minute after its last write.
|
||||
Shutdown is **not** on that list: the archive handle is not closed when
|
||||
the service stops. So either move `archive-{uuid}.db` together with any
|
||||
`-wal`/`-shm` beside it, or wait until there are none.
|
||||
20 KB `.db` with no sidecars about a minute after its last write. A
|
||||
clean stop closes it too. So either move `archive-{uuid}.db` together
|
||||
with any `-wal`/`-shm` beside it, or wait until there are none.
|
||||
|
||||
### Restore
|
||||
|
||||
@@ -1073,12 +1067,10 @@ the service stops. So either move `archive-{uuid}.db` together with any
|
||||
They are part of the database, and dropping a `-wal` silently
|
||||
discards every transaction it still holds. An `.backup` set will not
|
||||
contain any: it writes a single consolidated file per database. A
|
||||
stop-and-copy set has none for `webhooker.db` or the `events-*.db`,
|
||||
because a clean stop closes those and checkpoints their sidecars
|
||||
away — but it will normally have them for `archive-*.db`, whose
|
||||
handle stays open across shutdown, and those carry the archive's
|
||||
rows. A copy salvaged from a crashed instance has them for
|
||||
everything, and needs all of them.
|
||||
stop-and-copy set normally has none, because a clean stop closes
|
||||
every database and checkpoints its sidecars away; the exception is an
|
||||
archive not opened since a crash. A copy salvaged from a crashed
|
||||
instance has them for everything, and needs all of them.
|
||||
|
||||
4. **Fix ownership.** The container runs as the non-root `webhooker`
|
||||
user, UID 1000 / GID 1000. Restored files must be owned by (or
|
||||
@@ -3088,7 +3080,8 @@ each hook. The order, read off the fx stop-hook log:
|
||||
3. `server` — the HTTP drain, bounded separately by
|
||||
`server.ShutdownTimeout` (**3 seconds**), then a Sentry flush if
|
||||
`SENTRY_DSN` is set
|
||||
4. `delivery.Engine`
|
||||
4. `delivery.Engine` — waits for its workers, then closes the archive
|
||||
databases
|
||||
5. `healthcheck`
|
||||
6. `WebhookDBManager`
|
||||
7. the database close
|
||||
|
||||
Reference in New Issue
Block a user