Fail loudly on an unparseable SENTRY_DSN and a malformed .env (closes #283)
All checks were successful
check / check (push) Successful in 3m32s
All checks were successful
check / check (push) Successful in 3m32s
Two configuration paths still failed silently, against the rule every other variable follows: a value that is set but cannot be parsed must abort startup rather than substitute a default. SENTRY_DSN is now parsed in loadFromEnv, with sentry.NewDsn — the same call sentry.Init makes on the DSN it is handed, so configuration and initialisation cannot disagree about what a valid DSN is. That costs internal/config an import of the Sentry SDK, which is already a module dependency already linked into the binary, and buys a single definition of validity rather than a hand-rolled second one free to drift. A typo in a DSN used to log one error line and leave the process serving with error reporting off forever, which nothing downstream can notice: the variable is still set, so every later signal reports it as on. hasSentryDSN is replaced by Config.SentryEnabled(), following MetricsAuthEnabled(): one method read by the startup log field, by the SDK initialisation and by the sentryhttp middleware, so the log cannot report reporting as on while nothing is sending. enableSentry's error branch is now fatal, and Run gives up before it listens rather than binding a port it is about to release. Fatal there means what a listen failure already meant — Shutdowner.Shutdown(fx.ExitCode(1)), through fx's normal stop sequence — so shutdownOnListenFailure is now shutdownWithFailure and ListenFailureExitCode is StartupFailureExitCode. The godotenv/autoload blank import is replaced by config.LoadDotEnv, called at the top of dispatch. autoload discarded Load's error, and godotenv applies nothing at all when a file will not parse, so one mistyped line reverted every variable in the file to its default and started the server with no log line naming the file. A missing file stays fine — it is optional and most deployments have none. The call sits in dispatch rather than in loadFromEnv because autoload ran in an init(), ahead of config.DataDir(), which both the DATA_DIR lock and resetpw call outside the fx graph; loading any later would let a .env that sets DATA_DIR lock one directory while the config opened databases in another. Both defects were reproduced against the previous build first: an unparseable DSN served traffic while logging "hasSentryDSN":true, and a malformed .env started on the default port with the file unmentioned.
This commit is contained in:
@@ -51,12 +51,13 @@ const (
|
||||
minSentryFlush = 250 * time.Millisecond
|
||||
)
|
||||
|
||||
// ListenFailureExitCode is the status the process exits with when the
|
||||
// HTTP listener cannot be established, or dies for a reason other
|
||||
// than a requested shutdown. It must stay non-zero: systemd
|
||||
// `Restart=on-failure` and Docker's restart policies key off it, and a
|
||||
// zero exit would read as a deliberate stop.
|
||||
const ListenFailureExitCode = 1
|
||||
// StartupFailureExitCode is the status the process exits with when
|
||||
// the serving goroutine gives up: the HTTP listener cannot be
|
||||
// established or dies for a reason other than a requested shutdown, or
|
||||
// error reporting is configured and cannot be started. It must stay
|
||||
// non-zero: systemd `Restart=on-failure` and Docker's restart policies
|
||||
// key off it, and a zero exit would read as a deliberate stop.
|
||||
const StartupFailureExitCode = 1
|
||||
|
||||
// SentryFlushBudget reports how long the Sentry flush may run when
|
||||
// remaining is the time left on the fx stop context after the HTTP
|
||||
@@ -135,11 +136,25 @@ func New(lc fx.Lifecycle, params ServerParams) (*Server, error) {
|
||||
}
|
||||
|
||||
// Run configures Sentry and starts serving HTTP requests.
|
||||
//
|
||||
// A Sentry failure ends the application instead of listening. It runs
|
||||
// before the listener rather than after it so that the process never
|
||||
// binds a port it is about to give up.
|
||||
func (s *Server) Run() {
|
||||
s.configure()
|
||||
|
||||
// logging before sentry, because sentry logs
|
||||
s.enableSentry()
|
||||
err := s.enableSentry()
|
||||
if err != nil {
|
||||
s.log.Error(
|
||||
"SENTRY_DSN is set but error reporting could not be "+
|
||||
"started; refusing to serve with it off",
|
||||
"error", err,
|
||||
)
|
||||
s.shutdownWithFailure()
|
||||
|
||||
return
|
||||
}
|
||||
|
||||
s.serve()
|
||||
}
|
||||
@@ -150,11 +165,23 @@ func (s *Server) MaintenanceMode() bool {
|
||||
return s.params.Config.MaintenanceMode
|
||||
}
|
||||
|
||||
func (s *Server) enableSentry() {
|
||||
// enableSentry initialises the Sentry SDK when error reporting is
|
||||
// configured, and reports the failure when it is configured and cannot
|
||||
// be initialised. A DSN that is not set is not a failure: reporting
|
||||
// stays off and the server starts normally.
|
||||
//
|
||||
// There is no fallback to running with reporting off. An operator who
|
||||
// set SENTRY_DSN asked for failures to be visible, and serving traffic
|
||||
// with reporting quietly off is the one state nothing can ever tell
|
||||
// them about — the DSN is still set, so every later signal says it is
|
||||
// on. Config already refused a DSN the SDK cannot parse, which is what
|
||||
// a typo produces, so reaching this branch means the SDK refused
|
||||
// something that parsed: not a condition to guess at either.
|
||||
func (s *Server) enableSentry() error {
|
||||
s.sentryEnabled.Store(false)
|
||||
|
||||
if s.params.Config.SentryDSN == "" {
|
||||
return
|
||||
if !s.params.Config.SentryEnabled() {
|
||||
return nil
|
||||
}
|
||||
|
||||
err := sentry.Init(sentryClientOptions(
|
||||
@@ -166,19 +193,19 @@ func (s *Server) enableSentry() {
|
||||
),
|
||||
))
|
||||
if err != nil {
|
||||
s.log.Error("sentry init failure", "error", err)
|
||||
// Don't use fatal since we still want the service to run
|
||||
return
|
||||
return fmt.Errorf("initialising sentry: %w", err)
|
||||
}
|
||||
|
||||
s.log.Info("sentry error reporting activated")
|
||||
s.sentryEnabled.Store(true)
|
||||
|
||||
return nil
|
||||
}
|
||||
|
||||
// serve installs the signal watcher, starts the listener and blocks
|
||||
// until the server's context is cancelled. The process exit status is
|
||||
// fx's to decide — from a signal, or from the code
|
||||
// shutdownOnListenFailure hands the Shutdowner — so this reports
|
||||
// shutdownWithFailure hands the Shutdowner — so this reports
|
||||
// nothing back to its caller.
|
||||
func (s *Server) serve() {
|
||||
ctx, cancelFunc := context.WithCancel(context.Background())
|
||||
@@ -208,20 +235,24 @@ func (s *Server) serve() {
|
||||
// Do not call cleanShutdown() here to avoid double invocation.
|
||||
}
|
||||
|
||||
// shutdownOnListenFailure ends the application after the HTTP
|
||||
// listener failed. The fx OnStart hook returns as soon as the serving
|
||||
// goroutine is spawned, so nothing downstream of it ever learns that
|
||||
// the listen failed: fx reports RUNNING and the process sits alive
|
||||
// with nothing bound, which is invisible to systemd and Docker
|
||||
// restart policies. Asking the Shutdowner to stop the app with a
|
||||
// non-zero code is what turns that into a visible failure.
|
||||
// shutdownWithFailure ends the application non-zero from the serving
|
||||
// goroutine. It is how anything on that goroutine fails fatally: the
|
||||
// fx OnStart hook returns as soon as the goroutine is spawned, so
|
||||
// nothing downstream of it ever learns that the goroutine gave up. fx
|
||||
// reports RUNNING and the process sits alive having done neither what
|
||||
// it was asked nor anything visible instead, which systemd and
|
||||
// Docker restart policies cannot see. Asking the Shutdowner to stop
|
||||
// the app with a non-zero code is what turns that into a visible
|
||||
// failure, and it is the whole of "fatal" here — no panic, no
|
||||
// os.Exit, and every stop hook still runs.
|
||||
//
|
||||
// The context cancel that follows only unwinds serve()'s own wait.
|
||||
// The shutdown itself runs through fx's normal stop sequence, so the
|
||||
// clean-shutdown drain in cleanShutdown is reached unchanged.
|
||||
func (s *Server) shutdownOnListenFailure() {
|
||||
// The context cancel that follows only unwinds serve()'s own wait,
|
||||
// and is skipped before serve has installed one. The shutdown itself
|
||||
// runs through fx's normal stop sequence, so the clean-shutdown drain
|
||||
// in cleanShutdown is reached unchanged.
|
||||
func (s *Server) shutdownWithFailure() {
|
||||
err := s.params.Shutdowner.Shutdown(
|
||||
fx.ExitCode(ListenFailureExitCode),
|
||||
fx.ExitCode(StartupFailureExitCode),
|
||||
)
|
||||
if err != nil {
|
||||
s.log.Error("shutdown request failed", "error", err)
|
||||
|
||||
Reference in New Issue
Block a user