All checks were successful
check / check (push) Successful in 3m3s
fx defaults to a 15s stop timeout and the Dockerfile sets no grace override, so Docker SIGKILLed at 10s and the bounded shutdown #130 built — including the log line that tells an operator a component is wedged — was unreachable in the image this repo produces. Sets fx.StopTimeout to 5s, and lowers the HTTP drain to 3s so a full-length drain no longer exhausts the whole sequence budget and skip every later hook, database close included. The Sentry flush, which runs in the same hook and honours no context, is clamped to the remaining stop budget less a 2s tail reserve, so a stalled flush drops Sentry events rather than the database close. Also fixes a latent coin flip in the shared stop-hook waiter, which reported "shutdown timed out" about half the time for a component that drained cleanly against an already-expired context. Independently reviewed three times. The final reviewer derived a stronger invariant than the implementation claims — the server hook's absolute end is bounded at stopTimeout minus the reserve regardless of drain length or of time consumed by preceding hooks — and confirmed the guard's 10ms sweep cannot step over the maximum, since both breakpoints land on its grid. Both Sentry probe arms, the docker stop demo and every mutation were reproduced independently. Known residual, filed separately: the HTTP drain itself is not clamped by the reserve, so slow preceding hooks can still jointly exhaust the budget. Demonstrated with a 2.2s sweeper delay.
81 lines
2.0 KiB
Go
81 lines
2.0 KiB
Go
// Package lifecycle holds helpers shared by the components that
|
|
// register fx start and stop hooks.
|
|
package lifecycle
|
|
|
|
import (
|
|
"context"
|
|
"fmt"
|
|
"log/slog"
|
|
"sync"
|
|
)
|
|
|
|
// WaitForShutdown waits for wg to drain, bounded by ctx.
|
|
//
|
|
// fx hands OnStop a context carrying the application's stop
|
|
// timeout. A bare wg.Wait() discards that deadline, so a single
|
|
// goroutine that never observes cancellation — a delivery target
|
|
// that never returns, a SQLite operation blocked on a lock —
|
|
// hangs the process forever instead of letting it exit when the
|
|
// timeout expires, which is exactly when a clean shutdown matters
|
|
// most.
|
|
//
|
|
// On timeout it logs at error naming component and returns an
|
|
// error: the goroutines are still running, and reporting success
|
|
// would hide an unclean shutdown from the operator. The waiting
|
|
// goroutine outlives this call and exits when (if) wg drains; it
|
|
// holds nothing but the channel it closes.
|
|
func WaitForShutdown(
|
|
ctx context.Context,
|
|
log *slog.Logger,
|
|
component string,
|
|
wg *sync.WaitGroup,
|
|
) error {
|
|
done := make(chan struct{})
|
|
|
|
go func() {
|
|
defer close(done)
|
|
|
|
wg.Wait()
|
|
}()
|
|
|
|
return waitDone(ctx, log, component, done)
|
|
}
|
|
|
|
// waitDone waits for done to close, bounded by ctx.
|
|
//
|
|
// The non-blocking preamble is load-bearing. When the component has
|
|
// already drained and ctx has already expired, both cases of the
|
|
// bounded select are ready and Go picks between them uniformly at
|
|
// random, so a clean shutdown would be reported as a timeout about
|
|
// half the time. Draining wins: the goroutines are gone, and there
|
|
// is nothing left for the operator to act on.
|
|
func waitDone(
|
|
ctx context.Context,
|
|
log *slog.Logger,
|
|
component string,
|
|
done <-chan struct{},
|
|
) error {
|
|
select {
|
|
case <-done:
|
|
return nil
|
|
default:
|
|
}
|
|
|
|
select {
|
|
case <-done:
|
|
return nil
|
|
case <-ctx.Done():
|
|
log.Error(
|
|
"shutdown timed out, goroutines still running",
|
|
"component", component,
|
|
"error", ctx.Err(),
|
|
)
|
|
|
|
return fmt.Errorf(
|
|
"%s: shutdown timed out, "+
|
|
"goroutines still running: %w",
|
|
component, ctx.Err(),
|
|
)
|
|
}
|
|
}
|