Give golangci-lint per-checkout cache and lock state (closes #30)
All checks were successful
check / check (push) Successful in 7s

A golangci-lint result on a host running many concurrent workers does not
reliably belong to the tree that asked for it. Two independent mechanisms,
which have repeatedly been mistaken for one:

The result cache is keyed on file content, not location, so two checkouts
of the same commit hold byte-identical files, share cache entries, and one
tree's findings are served for the other under the other tree's path. This
produced a confirmed false green as well as the loud false reds. Moving
workers from shared worktrees to their own clones does not address it —
two clones collide exactly as two worktrees did — and removes only the
foreign-path artefact that made the defect noticeable.

The concurrency lock is $TMPDIR/golangci-lint.lock (pkg/commands/run.go,
acquireFileLock), host-global and independent of GOLANGCI_LINT_CACHE, with
a five-second acquire timeout, so it fails when the host is busiest. A
private cache directory does not isolate it. Setting only the cache closes
the contamination half and leaves runs failing red on a condition that is
not a result at all.

REPO_POLICIES.md now carries the canonical Go script/lint: both variables
scoped into a .lint-cache/ directory inside the checkout, above any
container-versus-host branch so every path reaching the linter gets them;
--allow-serial-runners, which keeps the mutual-exclusion guard and queues
rather than aborting, for the same-checkout overlap TMPDIR scoping cannot
cover, with --allow-parallel-runners rejected because it deletes the
guard; and a bounded retry that treats the lock error as VOID rather than
as findings, exiting 75 on exhaustion so it is neither a pass nor a
failure. Detection is on the stderr stream and never on exit status:
findings go to stdout, so a finding quoting the lock message in source
cannot be retried away, and the exit status is not a stable discriminator
anyway. The interim void rule is recorded with the ../ clause that the
original filter missed, and with its limit stated — it catches
contamination that names foreign files, not contamination that suppresses
findings. Both checklists gained the corresponding items, since a half-fix
that sets only the cache reads as complete.

GOCACHE was measured rather than assumed and does not need isolating: with
the two variables scoped per checkout and GOCACHE shared at the host
default, each checkout reported its own paths.

Verified with the snippet extracted from the committed document and
executed as a consuming repo would adopt it, each control paired against
the pre-fix form: contamination reproduced on the pre-fix script and
absent on the adopted one; a stub linter colliding twice then clearing,
with the retry engaging and succeeding; exhaustion exiting 75 with a VOID
message; a genuine finding whose text quotes the lock message reported as
findings with no retry; and a real held lock failing the pre-fix script
with exit 3 while the adopted script, inheriting the same environment,
completed in one second.
This commit is contained in:
clawbot
2026-08-09 17:35:48 +00:00
parent 3a218497b8
commit 33fb5dde98
4 changed files with 389 additions and 0 deletions

29
TODO.md
View File

@@ -21,6 +21,35 @@ fmt-check, and commit.
# Completed Steps
- 2026-08-09: Made a golangci-lint result belong to the tree that asked for it.
REPO_POLICIES.md now carries the canonical Go `script/lint`, which gives the
linter per-checkout `GOLANGCI_LINT_CACHE` and per-checkout `TMPDIR`. The two
are separate defects and the second is the one that gets dropped: the result
cache is keyed on file content rather than location, so checkouts holding
identical files serve each other's findings under the other's path, while the
concurrency lock is `$TMPDIR/golangci-lint.lock` — host-global, independent of
the cache, and unaffected by isolating it. Moving workers from worktrees to
their own clones does not help either half; it only removes the foreign-path
artefact that made the defect visible. The lock error is retried rather than
surfaced, because it is not a result: it exits non-zero exactly as findings
do, and reporting it as findings sends a correct branch back for rework.
Detection is on the stderr stream and never on exit status, so a finding
quoting the lock message in source cannot be retried away, and exhaustion
exits 75 with a VOID message rather than passing or failing quietly.
`--allow-serial-runners` (which keeps the guard and queues) covers the
same-checkout overlap that `TMPDIR` scoping cannot; `--allow-parallel-runners`
is rejected outright. The stdout and stderr capture files are per invocation
rather than per checkout, because serialising the linter does not serialise
the shell's redirections: two runs in one checkout — the overlap the flag
exists to support — would otherwise truncate and read each other's output,
which is the same defect one layer above where it was fixed. Both checklists
gained the corresponding items, since a half-fix that sets only the cache
reads as complete. `GOCACHE` was measured and does not need isolating.
Verified with the snippet extracted from the committed document and executed
as a consuming repo would adopt it, against paired controls: contamination
reproduced on the pre-fix form and absent on the adopted one, retry engaged,
exhaustion loud, a genuine finding still reported, and a held host lock
failing the pre-fix script while leaving the adopted one untouched.
- 2026-08-09: Kept in-repo agent scratch out of the Docker build context and out
of version control. `.claude/` holds one worktree — an entire additional
checkout of the repo — per in-flight agent, and under `COPY . .` all of it was