Clamp the HTTP drain by the tail-hook reserve (closes #170)
check / check (push) Waiting to run

The HTTP drain at shutdown waited up to ShutdownTimeout regardless of how much of the stop budget earlier hooks had used, so a slow archive sweeper or retention reaper could eat the reserve the hooks after the server need, and the database close was skipped. The drain now waits at most the shorter of ShutdownTimeout and what is left of the budget less TailHookReserve, as the Sentry flush already does. The reserve is documented as derived from the two timeouts. Tests cover earlier hooks having spent part of the budget, on a clock that host speed cannot move, and pin that a drain on the full budget gets all of ShutdownTimeout.

Model: opus-5-5
This commit was merged in pull request #454.
This commit is contained in:
2026-10-02 20:11:43 +02:00
parent f82b730c31
commit 0945831442
6 changed files with 231 additions and 36 deletions
+24 -17
View File
@@ -3245,9 +3245,9 @@ each hook. The order, read off the fx stop-hook log:
1. `ArchiveSweeper`
2. `RetentionReaper`
3. `server` — the HTTP drain, bounded separately by
`server.ShutdownTimeout` (**3 seconds**), then a Sentry flush if
`SENTRY_DSN` is set
3. `server` — the HTTP drain, bounded by `server.ShutdownTimeout`
(**3 seconds**) and by what the hooks before it left, then a Sentry
flush if `SENTRY_DSN` is set
4. `delivery.Engine` — waits for its workers, then closes the archive
databases
5. `healthcheck`
@@ -3267,23 +3267,30 @@ exhaust the sequence budget at the instant it finished, and every
later hook — the delivery engine, the healthcheck, the webhook DB
manager and the database close — would be skipped in exactly the
case where the drain mattered. 3 seconds leaves 2 seconds
(`server.TailHookReserve`) for the tail, which is far more than the
microseconds it needs.
(`server.TailHookReserve`) for the tail. The reserve is that
remainder, not a figure sized to the tail, which takes about a
millisecond.
That reserve belongs to the tail hooks, not to the server hook, and
the Sentry flush is what could take it: it runs after the drain
**inside the same hook**, and `sentry.Flush` takes a bare duration
and honours no context, so an unreachable Sentry endpoint would add
its own timeout on top of a full-length drain and consume the whole
sequence budget by itself. It is therefore clamped to whatever is
left on the stop context minus the reserve, and skipped when that
leaves too little to be worth attempting — so a full-length drain
means Sentry events are dropped rather than the database close being
skipped.
the server hook could take it in two ways. The hooks before it may
already have spent part of the budget, so a full 3-second drain
would come out of the reserve; the drain is therefore also bounded
by whatever is left on the stop context minus the reserve. And the
Sentry flush runs after the drain **inside the same hook**, and
`sentry.Flush` takes a bare duration and honours no context, so an
unreachable Sentry endpoint would add its own timeout on top of a
full-length drain and consume the whole sequence budget by itself.
It is clamped the same way, and skipped when that leaves too little
to be worth attempting — so a full-length drain means Sentry events
are dropped rather than the database close being skipped.
This does not make the database close unconditional: a wedged
`ArchiveSweeper` or `RetentionReaper` still runs first and can
consume the whole budget on its own.
This does not make the database close unconditional. A slow
`ArchiveSweeper` or `RetentionReaper` is enough to cut the shutdown
short, not only one that consumes the whole budget: what they spend
comes out of the drain first, so after 2 seconds of theirs a request
still in flight gets 1 second to finish, and after 3 it gets none.
Past 3 seconds they spend the reserve itself, and one that takes the
whole budget skips every hook after it, the database close included.
The value is chosen to sit inside the container stop grace period.
Docker's default `docker stop` grace is 10 seconds and the Dockerfile