Clamp the HTTP drain by the tail-hook reserve (closes #170)
check / check (push) Successful in 3m21s

The server's stop hook bounded the drain by ShutdownTimeout alone, so
once the archive sweeper or retention reaper had spent part of the fx
stop budget, a request held open could use up the reserve and fx
skipped every hook after the server, the database close included.
The drain now gets the shorter of ShutdownTimeout and what is left
less TailHookReserve, the clamp the Sentry flush already has.

The headroom test also sweeps the time earlier hooks spent; a new
test holds a request open against a stop context with only the
reserve left. The README and the reserve's comment say how the
reserve is derived and what a slow sweeper now costs.

Model: opus-5-5
This commit is contained in:
2026-10-02 16:05:09 +00:00
parent bf3df0312b
commit 6aa9907c8e
6 changed files with 180 additions and 36 deletions
+26 -3
View File
@@ -39,6 +39,12 @@ const (
// refuses to spend, leaving it for the hooks that run after the
// server: the delivery engine, the healthcheck, the webhook DB
// manager and the database close.
//
// Its value is not tuned to those hooks, which take about a
// millisecond between them. It is what the 5s fx stop timeout in
// cmd/webhooker leaves after a full ShutdownTimeout drain, so a
// drain that starts on a full budget still gets all of
// ShutdownTimeout.
TailHookReserve = 2 * time.Second
// sentryFlushTimeout is the longest wait for Sentry to flush
@@ -59,6 +65,16 @@ const (
// key off it, and a zero exit would read as a deliberate stop.
const StartupFailureExitCode = 1
// DrainBudget reports how long the HTTP drain may wait for in-flight
// requests when remaining is the time left on the fx stop context as
// the server's stop hook starts. The hooks before the server can
// already have spent part of the budget, so the drain takes its time
// out of what they left, never out of TailHookReserve. Zero or less
// means no wait at all.
func DrainBudget(remaining time.Duration) time.Duration {
return min(ShutdownTimeout, remaining-TailHookReserve)
}
// SentryFlushBudget reports how long the Sentry flush may run when
// remaining is the time left on the fx stop context after the HTTP
// drain. sentry.Flush takes a bare duration and honours no context,
@@ -261,10 +277,17 @@ func (s *Server) cleanupForExit() {
s.log.Info("cleaning up")
}
// cleanShutdown drains the HTTP server and flushes Sentry inside what
// is left of the fx stop budget. A context carrying no deadline — a
// caller outside the fx lifecycle — gets the full ShutdownTimeout.
func (s *Server) cleanShutdown(ctx context.Context) {
ctxShutdown, shutdownCancel := context.WithTimeout(
ctx, ShutdownTimeout,
)
drain := ShutdownTimeout
if deadline, ok := ctx.Deadline(); ok {
drain = DrainBudget(time.Until(deadline))
}
ctxShutdown, shutdownCancel := context.WithTimeout(ctx, drain)
defer shutdownCancel()
err := s.httpServer.Shutdown(ctxShutdown)