Correct four inaccurate comments in cancel_test.go and rename a constant (closes #33)
check / check (push) Failing after 2s

Fix the comments the re-review of
#6 found misdescribing their
tests; no test's behaviour changes.

State the property the walkClock tests rely on, that the index load's
Done cost does not grow with the record count, instead of a wrong fixed
figure. Describe poolUnwind so it is true of every use, and mark the
tests that have no bound. Record that hashWorker's results-send exit is
reachable through stop but has no test that fails without it. Rename
walkCancelInFlightDirs to walkCancelInFlightFiles, since it counts
files. Note at the top why the file departs from the
one-test-file-per-source-file convention.

Model: opus-4-8 (implementation); opus-5-5 (rework)
This commit was merged in pull request #59.
This commit is contained in:
2026-10-04 13:01:31 +02:00
parent 7ff0dbb6e1
commit 948fb03630
2 changed files with 64 additions and 19 deletions
+60 -19
View File
@@ -18,11 +18,21 @@ import (
"time"
)
// poolUnwind bounds how long a goroutine is given to leave a pool
// after its context is cancelled. Only a failing run ever waits this
// long: a pool that ignored its cancellation parks forever, and this
// is what turns that into a failed assertion instead of a suite that
// hangs until the test binary's own timeout.
// This file gathers the tests for scan cancellation and worker-pool
// unwinding. Everything it exercises lives in scan.go, so by the repo's
// convention of one test file per source file it would belong in
// scan_test.go. It is kept separate on purpose: cancellation behaviour
// cuts across both the walk pool and the hash pool as a single concern,
// and scan_test.go is already over 1,600 lines. That is the deliberate
// exception the convention otherwise expects to be stated.
// poolUnwind bounds how long a test waits for a cancellation to take
// effect: for a goroutine to return or a channel to close once its
// context is cancelled, or for a signal to cancel the scan's context.
// Only a failing run waits this long, and the bound is what makes that
// failure an assertion instead of a hang. A call made without it, as
// most of this file's scans are, has no bound: a regression that parks
// it is caught only as the test binary's own timeout.
const poolUnwind = 2 * time.Second
// walkClock is a context whose cancellation is driven by the scan's
@@ -33,9 +43,14 @@ const poolUnwind = 2 * time.Second
//
// The accounting behind the n chosen by each test: every blocking
// channel operation in the walk selects on Done, so the walk spends
// one consultation per file event plus a couple per directory, while
// the index load that runs ahead of it spends a small fixed number
// (three) whatever the record count.
// one consultation per file event plus a couple per directory. The
// index load that runs ahead of it also consults Done, but a bounded
// number of times that does not grow with the record count. The tests
// depend on that property, not on the bound's exact value: each test
// sets n from the consultations of the walk, plus those of the hash
// phase when it cancels mid-hash, far from both ends of the phase it
// interrupts, so the cancellation lands inside that phase whatever the
// record count.
type walkClock struct {
n int64
seen atomic.Int64
@@ -91,11 +106,14 @@ func (c *walkClock) Value(_ any) any {
// directory still queued and only the handful already in flight can
// emit anything more.
const (
walkCancelDirs = 100
walkCancelFilesPerDir = 20
walkCancelFiles = walkCancelDirs * walkCancelFilesPerDir
walkCancelWorkers = 4
walkCancelInFlightDirs = walkCancelWorkers * walkCancelFilesPerDir
walkCancelDirs = 100
walkCancelFilesPerDir = 20
walkCancelFiles = walkCancelDirs * walkCancelFilesPerDir
walkCancelWorkers = 4
// The most files the walkCancelWorkers directories already in
// flight when the scan is cancelled can still emit, at
// walkCancelFilesPerDir each. A file count, not a directory count.
walkCancelInFlightFiles = walkCancelWorkers * walkCancelFilesPerDir
)
// walkCancelAtDone is the consultation on which the fixture's context
@@ -158,6 +176,11 @@ func assertRecordsIntact(t *testing.T, db *sql.DB, before []string) {
// unreachable, makes the scan carry its truncated view into the update
// phase, which counts every record the walk never reached for removal.
//
// The syncScan call here is not bounded by poolUnwind: a regression
// that left a worker pool parked would hang it, and that regression is
// caught only by the test binary's own timeout, not by a quick
// assertion.
//
//nolint:paralleltest // counts goroutines: must not run beside others
func TestSyncScanCancelledMidWalkKeepsRecords(t *testing.T) {
dir := buildWalkCancelTree(t)
@@ -209,10 +232,11 @@ func assertWalkGuardAborted(t *testing.T, st scanStats, err error) {
}
// The workers drop every directory still queued once the scan is
// cancelled, so only the directories already in flight can add to
// the census after the fact. A census beyond that bound would mean
// the cancellation was not observed where it should have been.
limit := walkCancelAtDone + walkCancelInFlightDirs
// cancelled, so only the files in the directories already in flight
// can add to the census after the fact. A census beyond that bound
// would mean the cancellation was not observed where it should have
// been.
limit := walkCancelAtDone + walkCancelInFlightFiles
if st.unchanged > limit {
t.Errorf("census covers %d files, want at most %d: the walk kept "+
"taking directories off the queue after cancellation",
@@ -646,7 +670,11 @@ func TestDispatchDirsClosesJobsWhenCancelled(t *testing.T) {
// TestFeedHashJobsClosesJobsWhenCancelled checks that the hash feeder
// abandons the runs it has not queued yet and still closes the job
// channel, which is what lets the workers' range terminate.
// channel, which is what lets the workers' range terminate. The
// receive on jobs below is not bounded: a feeder that returned without
// closing jobs would leave that receive with no sender and no close, so
// this regression is caught by the test binary's timeout rather than by
// a bounded assertion.
func TestFeedHashJobsClosesJobsWhenCancelled(t *testing.T) {
t.Parallel()
@@ -674,6 +702,16 @@ func TestFeedHashJobsClosesJobsWhenCancelled(t *testing.T) {
// nobody wants the hashes of — while still letting the range run out
// so the pool tears down. The queued run names a file that does not
// exist, so a worker that hashed it anyway would produce a result.
//
// hashWorker's other cancellation exit, abandoning the send of a
// result, is reachable from the scan: stop cancels the pool before it
// drains results, so a worker waiting on that send can leave through
// it. The tests that stop a scan mid-hash, among them
// TestScanHashWriteFailureUnwindsPool, reach it in some runs only,
// depending on timing, and no test fails without it, since stop's
// drain frees a waiting worker anyway. This test, for its part, catches
// a removed drop check in some runs only: a worker that hashes the run
// anyway then picks at random between sending the result and leaving.
func TestHashWorkerDropsQueuedRuns(t *testing.T) {
t.Parallel()
@@ -706,7 +744,10 @@ func TestHashWorkerDropsQueuedRuns(t *testing.T) {
// TestHashPhaseCancelledReturnsContextError checks the result loop's
// own exit: with the pool cancelled, no result will ever arrive, and
// the loop must leave through the cancellation rather than wait for a
// receive that cannot happen.
// receive that cannot happen. This call is not bounded by poolUnwind: a
// loop that dropped its cancellation case would block on that receive,
// so the regression surfaces as the test binary's timeout rather than
// as a bounded assertion.
func TestHashPhaseCancelledReturnsContextError(t *testing.T) {
t.Parallel()