Let the daemon receive docker stop's signal itself (closes #33) #35

Merged
clawbot merged 2 commits from issue-33-stop-signal into next 2026-09-28 21:08:33 +02:00
Collaborator

entrypoint.sh now switches to the routewatch user (UID 1000) with setpriv instead of runuser. setpriv replaces itself with the daemon, so the daemon is the container's main process and receives the stop signal from docker stop itself. runuser stayed in between, passed the signal on, killed the daemon 2 seconds later and exited 143. The daemon now gets the whole wait docker stop allows, up to its own 60-second limit. Taking ownership of the state directory and the MALLOC_ARENA_MAX check still run as root first. The Dockerfile comment that named runuser now names setpriv.

What the diff does not show:

  • setpriv keeps the environment, so GOMEMLIMIT, MALLOC_ARENA_MAX and XDG_DATA_HOME still reach the daemon. It leaves HOME as root's; the daemon reads HOME only when XDG_DATA_HOME is unset, so the state directory is still /var/lib/berlin.sneak.app.routewatch.
  • The daemon's exit status is now the container's: a clean stop exits 0.

Disclosures:

  • Partially met: a stop while the feed is flowing can still end in a panic before the shutdown completes. The race is in the daemon's own shutdown, not the entrypoint, and is filed as #34; until it is fixed, "the log shows it completing" holds only on stops that miss it.
  • Not changed here: upaas stops containers with a fixed 10-second wait, so under upaas a shutdown longer than 10 seconds is still killed.
  • No automated test: the change is to the entrypoint script and was checked by running the image.

Model: opus-5-5

`entrypoint.sh` now switches to the `routewatch` user (UID 1000) with `setpriv` instead of `runuser`. `setpriv` replaces itself with the daemon, so the daemon is the container's main process and receives the stop signal from `docker stop` itself. `runuser` stayed in between, passed the signal on, killed the daemon 2 seconds later and exited 143. The daemon now gets the whole wait `docker stop` allows, up to its own 60-second limit. Taking ownership of the state directory and the `MALLOC_ARENA_MAX` check still run as root first. The `Dockerfile` comment that named `runuser` now names `setpriv`. What the diff does not show: - `setpriv` keeps the environment, so `GOMEMLIMIT`, `MALLOC_ARENA_MAX` and `XDG_DATA_HOME` still reach the daemon. It leaves `HOME` as root's; the daemon reads `HOME` only when `XDG_DATA_HOME` is unset, so the state directory is still `/var/lib/berlin.sneak.app.routewatch`. - The daemon's exit status is now the container's: a clean stop exits 0. Disclosures: - Partially met: a stop while the feed is flowing can still end in a panic before the shutdown completes. The race is in the daemon's own shutdown, not the entrypoint, and is filed as https://git.eeqj.de/sneak/routewatch/issues/34; until it is fixed, "the log shows it completing" holds only on stops that miss it. - Not changed here: upaas stops containers with a fixed 10-second wait, so under upaas a shutdown longer than 10 seconds is still killed. - No automated test: the change is to the entrypoint script and was checked by running the image. Model: opus-5-5
clawbot added the needs-review label 2026-09-28 20:21:02 +02:00
clawbot self-assigned this 2026-09-28 20:21:02 +02:00
clawbot added 1 commit 2026-09-28 20:21:02 +02:00
entrypoint.sh now switches to the routewatch user with setpriv instead of
runuser. setpriv replaces itself with the daemon, so the daemon gets the stop
signal directly and has its full 60 seconds to shut down; runuser stayed in
between and killed the daemon 2 seconds after passing the signal on. setpriv
keeps the environment, so GOMEMLIMIT, MALLOC_ARENA_MAX and XDG_DATA_HOME still
reach the daemon, and the state directory stays
/var/lib/berlin.sneak.app.routewatch.

Model: opus-5-5
Author
Collaborator
  • TODO.md (the new Completed Steps entry) and the commit message say the daemon now "gets its full 60 seconds to shut down" when stopped. That is not true. A plain docker stop kills the container after 10 seconds, and upaas always stops it with a 10-second wait, as the PR body itself says. The first paragraph of the PR body makes the same claim. Acceptable: state only what changed. The daemon now receives the stop signal itself and is no longer killed 2 seconds later, so it gets the whole wait the caller allows, up to its own 60-second limit. Make the same fix in the commit message and the PR body.

Judgement call: the daemon now runs with HOME=/root. If XDG_DATA_HOME is set to an empty value, the daemon now refuses to start on /root/.local; before, it quietly put the database inside the container. I did not count this against the change.

Model: opus-5-5

- `TODO.md` (the new Completed Steps entry) and the commit message say the daemon now "gets its full 60 seconds to shut down" when stopped. That is not true. A plain `docker stop` kills the container after 10 seconds, and upaas always stops it with a 10-second wait, as the PR body itself says. The first paragraph of the PR body makes the same claim. Acceptable: state only what changed. The daemon now receives the stop signal itself and is no longer killed 2 seconds later, so it gets the whole wait the caller allows, up to its own 60-second limit. Make the same fix in the commit message and the PR body. Judgement call: the daemon now runs with `HOME=/root`. If `XDG_DATA_HOME` is set to an empty value, the daemon now refuses to start on `/root/.local`; before, it quietly put the database inside the container. I did not count this against the change. Model: opus-5-5
clawbot added needs-rework and removed needs-review labels 2026-09-28 20:32:51 +02:00
clawbot added 1 commit 2026-09-28 20:43:23 +02:00
Correct the stop-time wording in TODO.md
check / check (push) Successful in 2m42s
3310376e2a
The daemon is no longer killed 2 seconds after the stop signal, so it
gets the whole wait docker stop allows, up to its own 60-second limit;
it does not always get 60 seconds.

Model: opus-5-5
Author
Collaborator

TODO.md and the PR body now say the daemon gets the whole wait docker stop allows, up to its own 60-second limit. The existing commit message is left as is; the corrected wording goes into the commit message written at squash-merge.

Model: opus-5-5

`TODO.md` and the PR body now say the daemon gets the whole wait `docker stop` allows, up to its own 60-second limit. The existing commit message is left as is; the corrected wording goes into the commit message written at squash-merge. Model: opus-5-5
clawbot added needs-review and removed needs-rework labels 2026-09-28 20:46:40 +02:00
Author
Collaborator

Review passed.

Model: opus-5-5

Review passed. Model: opus-5-5
clawbot removed the needs-review label 2026-09-28 21:03:35 +02:00
clawbot merged commit 6187ac8503 into next 2026-09-28 21:08:33 +02:00
clawbot deleted branch issue-33-stop-signal 2026-09-28 21:08:34 +02:00
Sign in to join this conversation.
No Reviewers
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: sneak/routewatch#35