Fold the August fleet findings into the policies, or drop them #62

Open
opened 2026-09-25 12:25:17 +02:00 by clawbot · 0 comments
Collaborator

Moved from the project-management repo's TODO.md (sneak/project-management#9), where it sat since 2026-08-09 as "cross-cutting findings worth remembering". Three of them became #25, #26 and #27, all closed. The rest were never folded into the policies. Git worktrees have since been banned outright, so the worktree items only explain history.

Definition of done

  • Each finding below is either written into the policies or dropped here with the reason.

The findings, as recorded on 2026-08-09

Cross-cutting findings worth remembering:

  • The pre-migration .golangci.yml (sha256 33ba2bf7...) declared
    version: "2" with v1 layout, so golangci-lint silently fell back to defaults
    and every threshold was ignored — any repo still on it has been reporting
    meaningless 0 issues.. Detect with golangci-lint config verify, but NEVER
    wire that into make lint or CI: it fetches its schema over an unpinned live
    HTTPS call.

  • Pinning golangci-lint by Docker image digest (@sha256:...) satisfies the
    hash-pinning policy. Demanding the git commit c0d3ddc9... from a repo with a
    separate pinned lint stage is a phantom violation — the top-sdlc-manager's
    launch prompts got this wrong and several managers correctly refused it.

  • The gomodguard deprecation under v2.12 belongs upstream: filed as prompts
    #25, since the canonical config is untouchable in consuming repos.

  • ORDERING HAZARD, fleet-wide, and the rule is BROADER than "repos with no
    .dockerignore": ANY directory in the Docker build context that churns
    accidentally protects a repo, because it invalidates COPY . . and forces the
    check layers to re-run. Tightening .dockerignore removes that protection.
    Two known churning directories:

    • .git — churns on nearly every git operation. Protects repos that have no
      .dockerignore at all; the canonical file excludes it.
    • .claude/ — holds worktrees/, created and destroyed constantly wherever
      agents run. This one hits repos that ALREADY exclude .git and therefore
      look unaffected. prompts #27 adds the exclusion.

    CONFIRMED by controlled experiment on rfscan 2026-08-09 (isolated copy,
    one variable per pair, only files inside .git touched): with no
    .dockerignore a .git-only change forces a real ~10s run; with .git
    excluded the identical change is a 0.37s cache hit, and the toggle reverses
    cleanly three times. ABSENCE OF A .dockerignore IS NOT A SCREENING TEST —
    a repo that has always excluded .git never had the protection and is
    exposed RIGHT NOW. EVERY repo needs the cache-bust regardless of its
    .dockerignore state. A cache hit is SUB-SECOND: a "warm" run of a minute
    or more is a real run, and the question becomes what else in the context is
    mutating (.claude/worktrees being the usual answer). Measure in an
    isolated copy with zero git activity between paired runs. Where a repo has
    NO .dockerignore, the canonical file and the cache-bust must land TOGETHER
    or cache-bust-first — never .dockerignore-first, and they are sometimes
    atomic in one PR. Order per repo: land the ARG CHECK_EPOCH fix (prompts
    #26) → verify two consecutive runs execute the checks AND that the bootstrap
    layer still caches → land the .dockerignore tightening (prompts #27) →
    RE-VERIFY, because the context changed underneath the earlier verification.
    That last step is the one most likely to be skipped, and skipping it is the
    same mistake as the lora.vegas outage: validating against one environment
    and assuming the result transfers to a changed one.

  • NESTED AGENT WORKTREES POLLUTE TOOLS THAT WALK THE FILESYSTEM — but which
    tools, exactly, depends on their DISCOVERY semantics, not on being a "test
    runner". Corrected by bsfirehose after an over-broad first statement:

    • SAFE: go test ./.... The go tool ignores directories beginning with .
      or _ (go help packages), so it never descends into
      .claude/worktrees/. No N+1 inflation; a 30s timeout measures what it
      claims to.
    • UNSAFE: gofmt -l ., and anything else doing a plain filesystem walk. It
      descends into nested worktrees and can fail make check on files that are
      not part of the repo at all. Tracked as bsfirehose #24.
    • UNSAFE: JS test runners. quak measured 1050 tests against a real 210 (18
      files x 5) — vitest's default exclude does not cover .claude/ and it
      does not read .gitignore for discovery. There the wall-clock inflation
      is real and does distort a timed gate. Reviewers inside their own
      worktrees are unaffected either way. This only shows in the SHARED clone —
      exactly where a manager does post-merge verification and is most likely to
      trust the number.
  • NEVER docker builder prune ON THIS HOST. It is shared mutable state across
    ~18 concurrent sessions. One subagent destroyed ~41 GB of shared BuildKit
    cache on 2026-08-09 while trying to prove a build was uncached. To prove that,
    scope the invalidation: docker build --no-cache on the one image, or
    --no-cache-filter=<stage> on the check stage. Note the prohibition must
    travel WITH the "a green can lie to you" caution, not after it — the moment an
    agent is told the cache is untrustworthy, destroying it is the obvious next
    thought, which is exactly how this happened.

    • Fallout to expect after any prune: cold-cache timeout flakes that look
      like new bugs (netwatch saw timeout 30 go test ./... die on an empty Go
      build cache, then pass in 11s on retry), and cache measurements taken
      across the boundary being invalid pairs.
  • DISPATCH RULE: idle in ListAgents does NOT mean stalled. A repo-manager is
    idle between turns — having just reported, or waiting on a respawned
    reviewer. Nudging on that signal alone produces messages asserting things that
    are not true (it happened twice on 2026-08-09, to eight sessions and then to
    two more, and every one pushed back with evidence). Before nudging, re-read
    the repo's tracker; and phrase the nudge as a question about state rather than
    a claim about it.

  • DISPATCH RULE for the top-sdlc-manager: resume-or-dispatch is an EITHER/OR.
    Sending a new requirement to an already-running agent AND spawning a rework
    agent for the same task puts two agents on one branch. It happened on
    cattbox 2026-08-07; no conflicting history resulted only because the
    no-force-push rule held. Pick one before acting.

  • When fixing the cache hole, verify the BOOTSTRAP layer still caches, not just
    that the check re-runs. A blanket --no-cache "fixes" the hole while turning
    a ~10s check into a full toolchain reinstall on every run.

  • A check that passes for an accidental reason fails silently the moment the
    accident is tidied up. Three instances so far: .git churn forcing real runs;
    netwatch's schema-invalid .golangci.yml making 0 issues. meaningless;
    netwatch's root make check genuinely passing while never touching the
    backend at all.

  • Evidence standard, in descending strength: a NEGATIVE CONTROL (a cached layer
    cannot produce a specifically predicted failure, and cannot fail then pass
    after a fix) beats layer inspection, which beats wall-clock. A green with
    neither duration nor a negative control is UNVERIFIED, not passed. Related
    discipline: "corroborated by a subagent" is not "verified" — say which.

  • "Pin by hash" is not "bump to the latest major". lora.vegas PR #17 passed
    two independent reviews and still broke production, because an
    upload-artifact v3->v4 bump assumed a protocol this Gitea does not serve. A
    CI path that cannot be executed pre-merge cannot be validated by review at all
    — use a temporary branch trigger plus if: github.ref_name == 'main' on
    publishing jobs, and read verdicts from the commit-status API.

Model: opus-5-5

Moved from the project-management repo's `TODO.md` (https://git.eeqj.de/sneak/project-management/issues/9), where it sat since 2026-08-09 as "cross-cutting findings worth remembering". Three of them became https://git.eeqj.de/sneak/prompts/issues/25, https://git.eeqj.de/sneak/prompts/issues/26 and https://git.eeqj.de/sneak/prompts/issues/27, all closed. The rest were never folded into the policies. Git worktrees have since been banned outright, so the worktree items only explain history. ## Definition of done - Each finding below is either written into the policies or dropped here with the reason. ## The findings, as recorded on 2026-08-09 Cross-cutting findings worth remembering: - The pre-migration `.golangci.yml` (sha256 `33ba2bf7...`) declared `version: "2"` with v1 layout, so golangci-lint silently fell back to defaults and every threshold was ignored — any repo still on it has been reporting meaningless `0 issues.`. Detect with `golangci-lint config verify`, but NEVER wire that into `make lint` or CI: it fetches its schema over an unpinned live HTTPS call. - Pinning golangci-lint by Docker image digest (`@sha256:...`) satisfies the hash-pinning policy. Demanding the git commit `c0d3ddc9...` from a repo with a separate pinned lint stage is a phantom violation — the top-sdlc-manager's launch prompts got this wrong and several managers correctly refused it. - The `gomodguard` deprecation under v2.12 belongs upstream: filed as `prompts` #25, since the canonical config is untouchable in consuming repos. - ORDERING HAZARD, fleet-wide, and the rule is BROADER than "repos with no `.dockerignore`": ANY directory in the Docker build context that churns accidentally protects a repo, because it invalidates `COPY . .` and forces the check layers to re-run. Tightening `.dockerignore` removes that protection. Two known churning directories: - `.git` — churns on nearly every git operation. Protects repos that have no `.dockerignore` at all; the canonical file excludes it. - `.claude/` — holds `worktrees/`, created and destroyed constantly wherever agents run. This one hits repos that ALREADY exclude `.git` and therefore look unaffected. `prompts` #27 adds the exclusion. CONFIRMED by controlled experiment on `rfscan` 2026-08-09 (isolated copy, one variable per pair, only files inside `.git` touched): with no `.dockerignore` a `.git`-only change forces a real ~10s run; with `.git` excluded the identical change is a 0.37s cache hit, and the toggle reverses cleanly three times. ABSENCE OF A `.dockerignore` IS NOT A SCREENING TEST — a repo that has always excluded `.git` never had the protection and is exposed RIGHT NOW. EVERY repo needs the cache-bust regardless of its `.dockerignore` state. A cache hit is SUB-SECOND: a "warm" run of a minute or more is a real run, and the question becomes what else in the context is mutating (`.claude/worktrees` being the usual answer). Measure in an isolated copy with zero git activity between paired runs. Where a repo has NO `.dockerignore`, the canonical file and the cache-bust must land TOGETHER or cache-bust-first — never `.dockerignore`-first, and they are sometimes atomic in one PR. Order per repo: land the `ARG CHECK_EPOCH` fix (`prompts` #26) → verify two consecutive runs execute the checks AND that the bootstrap layer still caches → land the `.dockerignore` tightening (`prompts` #27) → RE-VERIFY, because the context changed underneath the earlier verification. That last step is the one most likely to be skipped, and skipping it is the same mistake as the lora.vegas outage: validating against one environment and assuming the result transfers to a changed one. - NESTED AGENT WORKTREES POLLUTE TOOLS THAT WALK THE FILESYSTEM — but which tools, exactly, depends on their DISCOVERY semantics, not on being a "test runner". Corrected by `bsfirehose` after an over-broad first statement: - SAFE: `go test ./...`. The go tool ignores directories beginning with `.` or `_` (`go help packages`), so it never descends into `.claude/worktrees/`. No N+1 inflation; a 30s timeout measures what it claims to. - UNSAFE: `gofmt -l .`, and anything else doing a plain filesystem walk. It descends into nested worktrees and can fail `make check` on files that are not part of the repo at all. Tracked as `bsfirehose` #24. - UNSAFE: JS test runners. `quak` measured 1050 tests against a real 210 (18 files x 5) — vitest's default exclude does not cover `.claude/` and it does not read `.gitignore` for discovery. There the wall-clock inflation is real and does distort a timed gate. Reviewers inside their own worktrees are unaffected either way. This only shows in the SHARED clone — exactly where a manager does post-merge verification and is most likely to trust the number. - NEVER `docker builder prune` ON THIS HOST. It is shared mutable state across ~18 concurrent sessions. One subagent destroyed ~41 GB of shared BuildKit cache on 2026-08-09 while trying to prove a build was uncached. To prove that, scope the invalidation: `docker build --no-cache` on the one image, or `--no-cache-filter=<stage>` on the check stage. Note the prohibition must travel WITH the "a green can lie to you" caution, not after it — the moment an agent is told the cache is untrustworthy, destroying it is the obvious next thought, which is exactly how this happened. - Fallout to expect after any prune: cold-cache timeout flakes that look like new bugs (netwatch saw `timeout 30 go test ./...` die on an empty Go build cache, then pass in 11s on retry), and cache measurements taken across the boundary being invalid pairs. - DISPATCH RULE: `idle` in `ListAgents` does NOT mean stalled. A repo-manager is `idle` between turns — having just reported, or waiting on a respawned reviewer. Nudging on that signal alone produces messages asserting things that are not true (it happened twice on 2026-08-09, to eight sessions and then to two more, and every one pushed back with evidence). Before nudging, re-read the repo's tracker; and phrase the nudge as a question about state rather than a claim about it. - DISPATCH RULE for the top-sdlc-manager: resume-or-dispatch is an EITHER/OR. Sending a new requirement to an already-running agent AND spawning a rework agent for the same task puts two agents on one branch. It happened on `cattbox` 2026-08-07; no conflicting history resulted only because the no-force-push rule held. Pick one before acting. - When fixing the cache hole, verify the BOOTSTRAP layer still caches, not just that the check re-runs. A blanket `--no-cache` "fixes" the hole while turning a ~10s check into a full toolchain reinstall on every run. - A check that passes for an accidental reason fails silently the moment the accident is tidied up. Three instances so far: `.git` churn forcing real runs; netwatch's schema-invalid `.golangci.yml` making `0 issues.` meaningless; netwatch's root `make check` genuinely passing while never touching the backend at all. - Evidence standard, in descending strength: a NEGATIVE CONTROL (a cached layer cannot produce a specifically predicted failure, and cannot fail then pass after a fix) beats layer inspection, which beats wall-clock. A green with neither duration nor a negative control is UNVERIFIED, not passed. Related discipline: "corroborated by a subagent" is not "verified" — say which. - "Pin by hash" is not "bump to the latest major". `lora.vegas` PR #17 passed two independent reviews and still broke production, because an `upload-artifact` v3->v4 bump assumed a protocol this Gitea does not serve. A CI path that cannot be executed pre-merge cannot be validated by review at all — use a temporary branch trigger plus `if: github.ref_name == 'main'` on publishing jobs, and read verdicts from the commit-status API. Model: opus-5-5
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: sneak/prompts#62