Commit Graph
2 Commits
Author SHA1 Message Date
sneak 9b44189a91 Clamp the HTTP drain by the tail-hook reserve (closes #170)
check / check (push) Successful in 3m23s
The server's stop hook bounded the drain by ShutdownTimeout alone, so
once the archive sweeper or retention reaper had spent part of the fx
stop budget, a request held open could use up the reserve and fx
skipped every hook after the server, the database close included.
The drain now gets the shorter of ShutdownTimeout and what is left
less TailHookReserve, the clamp the Sentry flush already has.

The headroom test also sweeps the time earlier hooks spent and checks
that a drain on the full budget gets all of ShutdownTimeout. A new
test, run on synctest's clock, holds a request open over net.Pipe
against a stop context with only the reserve left. The README and the
reserve's comment say how the reserve is derived and what a slow
sweeper now costs.

Model: opus-5-5
2026-10-02 17:38:15 +00:00
clawbot bef9986542 Set fx.StopTimeout inside the container stop grace (closes #134)
check / check (push) Successful in 3m3s
fx defaults to a 15s stop timeout and the Dockerfile sets no grace
override, so Docker SIGKILLed at 10s and the bounded shutdown #130 built
— including the log line that tells an operator a component is wedged —
was unreachable in the image this repo produces.

Sets fx.StopTimeout to 5s, and lowers the HTTP drain to 3s so a
full-length drain no longer exhausts the whole sequence budget and skip
every later hook, database close included. The Sentry flush, which runs
in the same hook and honours no context, is clamped to the remaining
stop budget less a 2s tail reserve, so a stalled flush drops Sentry
events rather than the database close.

Also fixes a latent coin flip in the shared stop-hook waiter, which
reported "shutdown timed out" about half the time for a component that
drained cleanly against an already-expired context.

Independently reviewed three times. The final reviewer derived a
stronger invariant than the implementation claims — the server hook's
absolute end is bounded at stopTimeout minus the reserve regardless of
drain length or of time consumed by preceding hooks — and confirmed the
guard's 10ms sweep cannot step over the maximum, since both breakpoints
land on its grid. Both Sentry probe arms, the docker stop demo and every
mutation were reproduced independently.

Known residual, filed separately: the HTTP drain itself is not clamped
by the reserve, so slow preceding hooks can still jointly exhaust the
budget. Demonstrated with a 2.2s sweeper delay.
2026-08-18 00:12:51 +02:00