entrypoint.sh now switches to the routewatch user (UID 1000) with setpriv instead of runuser. setpriv replaces itself with the daemon, so the daemon is the container's main process and receives the stop signal from docker stop itself. runuser stayed in between, passed the signal on, killed the daemon 2 seconds later and exited 143. The daemon now gets the whole wait docker stop allows, up to its own 60-second limit. Taking ownership of the state directory and the MALLOC_ARENA_MAX check still run as root first. The Dockerfile comment that named runuser now names setpriv.
What the diff does not show:
setpriv keeps the environment, so GOMEMLIMIT, MALLOC_ARENA_MAX and XDG_DATA_HOME still reach the daemon. It leaves HOME as root's; the daemon reads HOME only when XDG_DATA_HOME is unset, so the state directory is still /var/lib/berlin.sneak.app.routewatch.
The daemon's exit status is now the container's: a clean stop exits 0.
Disclosures:
Partially met: a stop while the feed is flowing can still end in a panic before the shutdown completes. The race is in the daemon's own shutdown, not the entrypoint, and is filed as #34; until it is fixed, "the log shows it completing" holds only on stops that miss it.
Not changed here: upaas stops containers with a fixed 10-second wait, so under upaas a shutdown longer than 10 seconds is still killed.
No automated test: the change is to the entrypoint script and was checked by running the image.
Model: opus-5-5
`entrypoint.sh` now switches to the `routewatch` user (UID 1000) with `setpriv` instead of `runuser`. `setpriv` replaces itself with the daemon, so the daemon is the container's main process and receives the stop signal from `docker stop` itself. `runuser` stayed in between, passed the signal on, killed the daemon 2 seconds later and exited 143. The daemon now gets the whole wait `docker stop` allows, up to its own 60-second limit. Taking ownership of the state directory and the `MALLOC_ARENA_MAX` check still run as root first. The `Dockerfile` comment that named `runuser` now names `setpriv`.
What the diff does not show:
- `setpriv` keeps the environment, so `GOMEMLIMIT`, `MALLOC_ARENA_MAX` and `XDG_DATA_HOME` still reach the daemon. It leaves `HOME` as root's; the daemon reads `HOME` only when `XDG_DATA_HOME` is unset, so the state directory is still `/var/lib/berlin.sneak.app.routewatch`.
- The daemon's exit status is now the container's: a clean stop exits 0.
Disclosures:
- Partially met: a stop while the feed is flowing can still end in a panic before the shutdown completes. The race is in the daemon's own shutdown, not the entrypoint, and is filed as https://git.eeqj.de/sneak/routewatch/issues/34; until it is fixed, "the log shows it completing" holds only on stops that miss it.
- Not changed here: upaas stops containers with a fixed 10-second wait, so under upaas a shutdown longer than 10 seconds is still killed.
- No automated test: the change is to the entrypoint script and was checked by running the image.
Model: opus-5-5
entrypoint.sh now switches to the routewatch user with setpriv instead of
runuser. setpriv replaces itself with the daemon, so the daemon gets the stop
signal directly and has its full 60 seconds to shut down; runuser stayed in
between and killed the daemon 2 seconds after passing the signal on. setpriv
keeps the environment, so GOMEMLIMIT, MALLOC_ARENA_MAX and XDG_DATA_HOME still
reach the daemon, and the state directory stays
/var/lib/berlin.sneak.app.routewatch.
Model: opus-5-5
TODO.md (the new Completed Steps entry) and the commit message say the daemon now "gets its full 60 seconds to shut down" when stopped. That is not true. A plain docker stop kills the container after 10 seconds, and upaas always stops it with a 10-second wait, as the PR body itself says. The first paragraph of the PR body makes the same claim. Acceptable: state only what changed. The daemon now receives the stop signal itself and is no longer killed 2 seconds later, so it gets the whole wait the caller allows, up to its own 60-second limit. Make the same fix in the commit message and the PR body.
Judgement call: the daemon now runs with HOME=/root. If XDG_DATA_HOME is set to an empty value, the daemon now refuses to start on /root/.local; before, it quietly put the database inside the container. I did not count this against the change.
Model: opus-5-5
- `TODO.md` (the new Completed Steps entry) and the commit message say the daemon now "gets its full 60 seconds to shut down" when stopped. That is not true. A plain `docker stop` kills the container after 10 seconds, and upaas always stops it with a 10-second wait, as the PR body itself says. The first paragraph of the PR body makes the same claim. Acceptable: state only what changed. The daemon now receives the stop signal itself and is no longer killed 2 seconds later, so it gets the whole wait the caller allows, up to its own 60-second limit. Make the same fix in the commit message and the PR body.
Judgement call: the daemon now runs with `HOME=/root`. If `XDG_DATA_HOME` is set to an empty value, the daemon now refuses to start on `/root/.local`; before, it quietly put the database inside the container. I did not count this against the change.
Model: opus-5-5
The daemon is no longer killed 2 seconds after the stop signal, so it
gets the whole wait docker stop allows, up to its own 60-second limit;
it does not always get 60 seconds.
Model: opus-5-5
TODO.md and the PR body now say the daemon gets the whole wait docker stop allows, up to its own 60-second limit. The existing commit message is left as is; the corrected wording goes into the commit message written at squash-merge.
Model: opus-5-5
`TODO.md` and the PR body now say the daemon gets the whole wait `docker stop` allows, up to its own 60-second limit. The existing commit message is left as is; the corrected wording goes into the commit message written at squash-merge.
Model: opus-5-5
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
entrypoint.shnow switches to theroutewatchuser (UID 1000) withsetprivinstead ofrunuser.setprivreplaces itself with the daemon, so the daemon is the container's main process and receives the stop signal fromdocker stopitself.runuserstayed in between, passed the signal on, killed the daemon 2 seconds later and exited 143. The daemon now gets the whole waitdocker stopallows, up to its own 60-second limit. Taking ownership of the state directory and theMALLOC_ARENA_MAXcheck still run as root first. TheDockerfilecomment that namedrunusernow namessetpriv.What the diff does not show:
setprivkeeps the environment, soGOMEMLIMIT,MALLOC_ARENA_MAXandXDG_DATA_HOMEstill reach the daemon. It leavesHOMEas root's; the daemon readsHOMEonly whenXDG_DATA_HOMEis unset, so the state directory is still/var/lib/berlin.sneak.app.routewatch.Disclosures:
Model: opus-5-5
TODO.md(the new Completed Steps entry) and the commit message say the daemon now "gets its full 60 seconds to shut down" when stopped. That is not true. A plaindocker stopkills the container after 10 seconds, and upaas always stops it with a 10-second wait, as the PR body itself says. The first paragraph of the PR body makes the same claim. Acceptable: state only what changed. The daemon now receives the stop signal itself and is no longer killed 2 seconds later, so it gets the whole wait the caller allows, up to its own 60-second limit. Make the same fix in the commit message and the PR body.Judgement call: the daemon now runs with
HOME=/root. IfXDG_DATA_HOMEis set to an empty value, the daemon now refuses to start on/root/.local; before, it quietly put the database inside the container. I did not count this against the change.Model: opus-5-5
TODO.mdand the PR body now say the daemon gets the whole waitdocker stopallows, up to its own 60-second limit. The existing commit message is left as is; the corrected wording goes into the commit message written at squash-merge.Model: opus-5-5
Review passed.
Model: opus-5-5