Hash-pin every external reference in deploy.yml, verified on a real runner (closes #7) #22

Merged
clawbot merged 5 commits from pin-deploy-refs-observable into main 2026-08-09 07:03:41 +02:00

5 Commits

Author SHA1 Message Date
54ed6376af Hash-pin every external reference in deploy.yml (closes #7)
All checks were successful
check / check (push) Successful in 6s
deploy.yml was the last file in the repo carrying mutable external references.
Both job container images are now pinned by digest, all three `uses:` by a full
40-hex commit SHA, and the wrangler install by exact version, each with a
version/date comment above the reference.

- build container: klakegg/hugo:ext-alpine (abandoned since 2021, mutable tag)
  replaced by the exact alpine 3.21 digest the Dockerfile already pins, with a
  pre-checkout `apk add --no-cache nodejs git tar` step, `shell: sh` as the job
  default, then script/bootstrap and script/test. One pinned base and one
  dependency list now serve both the check build and the deploy build.
- deploy container: node:20 -> node@sha256:8f693eaa... (node 20.20.2 bookworm).
- actions/checkout: v4 -> 11bd7190... (v4.2.2), the same SHA check.yml pins.
- actions/upload-artifact: -> ff15f030... (v3.2.1).
- actions/download-artifact: -> 9bc31d5c... (v3.0.2).
- wrangler: `npm install -g wrangler` -> `wrangler@4.86.0`.

Also drops the dead feat/initial-site push trigger, reindents to 4-space YAML
to match check.yml, and adds `if: github.ref_name == 'main'` to the deploy job
so it can never publish from a branch.

This is the second attempt. The first passed two adversarial reviews, merged,
and broke the deploy, because deploy.yml triggers only on push to main and so
nobody could execute what they were reviewing. This time the workflow was
temporarily triggered on the branch, with the deploy job guarded off, and
iterated against the commit-status API until the build job ran green for real.
Doing that found two independent breaks that review had not:

1. actions/upload-artifact v4 fails on this Gitea Actions instance -- artifacts
   v4 is a different wire protocol and it is not served here. Two otherwise
   identical branch jobs, one with the v4 upload step and one without, failed
   and passed respectively. The issue asked for the v3 -> v4 bump; the
   artifact actions instead stay on the v3 line, pinned by SHA, at the exact
   commits the mutable @v3 references were already resolving to. Tracked
   separately in issue 20.
2. wrangler 4.120.0 requires node >= 22 and refuses to start on the pinned
   node 20 container. `npm install` only warns about engines, so the install
   step would have passed and the deploy step would have failed. The unpinned
   command this replaces was never installing `latest` either: npm resolves a
   bare name to the newest version whose engines the running node satisfies,
   which on node 20 is 4.86.0. So 4.86.0 is what has actually been deploying
   this site, and that is what is pinned. Tracked separately in issue 21.

The temporary branch trigger and the temporary probe workflow used to bisect
this are removed in this commit; the deploy guard is deliberately kept.

Verified: make check and script/cibuild green; the build job observed green on
the branch under act_runner (commit 73f912c, "Successful in 7s"); a probe job
pair rehearsed the deploy job end to end -- same pinned node image, same pinned
download action, same pinned wrangler, real site tarball extracted -- stopping
at `wrangler pages deploy --help` instead of publishing. The real deploy job
remains unexercised: it needs CLOUDFLARE_API_TOKEN and would publish, so it can
only run on main. The main run must still be watched and the live site
confirmed.
2026-08-09 03:06:21 +00:00
73f912c7ed Pin wrangler to the version that actually runs on the pinned node image
All checks were successful
check / check (push) Successful in 6s
Build and Deploy to Cloudflare Pages / build (push) Successful in 7s
probe / s1-build (push) Successful in 13s
Build and Deploy to Cloudflare Pages / deploy (push) Has been skipped
probe / s2-deploy-dryrun (push) Successful in 10s
Round 3 (07af755) cleared the artifact path and left one failure:

    check / check                        success   8s
    Build and Deploy .../ build          success   8s   <- green
    Build and Deploy .../ deploy         skipped        <- if: guard
    probe / r1-wrangler-only             failure   7s
    probe / r2a-upload-proven            success  12s
    probe / r2b-download-proven          success   2s
    probe / r3a-upload-node20            success   8s
    probe / r3b-download-node20          success   2s

r2a/r2b and r3a/r3b upload and download the real site tarball across the two
job containers, so the artifact round trip is sound. r1 does nothing but
install wrangler and invoke it, and it fails.

Reproduced locally in the pinned node image, which is faster than another CI
round:

    $ docker run --rm node@sha256:8f693eaa... sh -c \
        'npm install -g wrangler@4.120.0; wrangler --version'
    install exit=0            (with EBADENGINE warnings)
    Wrangler requires at least Node.js v22.0.0. You are using v20.20.2.
    version exit=1

npm treats engines as a warning on an explicit version, so the install step
would have passed and the deploy step would have failed -- a second break,
independent of the artifact one, in the same job nobody could run.

The instructive part is what the unpinned command it replaced was doing:

    $ docker run --rm node@sha256:8f693eaa... sh -c \
        'npm install -g wrangler; wrangler --version'
    `-- wrangler@4.86.0
    4.86.0

npm resolves a bare name to the newest version whose engines the running node
satisfies, so `npm install -g wrangler` on node 20 has been installing 4.86.0,
not the 4.120.0 that `latest` points at. Pinning 4.120.0 was therefore not
"pin the version we are already getting", it was an unnoticed major-ish bump
onto a node the container does not have.

So this pins wrangler 4.86.0 (engines: node >= 20.3.0, published 2026-04-28),
which is exactly the version that has been deploying this site, verified to
install and run on the pinned node 20 digest. The node image digest is left
alone. Bumping the container to node 22 to keep 4.120.0 is the alternative,
but that changes the deploy runtime for no benefit this issue asks for.

Round 4 replaces the probe jobs with a single end-to-end rehearsal: the build
job as written, then the deploy job as written with `wrangler pages deploy
--help` in place of the publish call.
2026-08-09 02:59:50 +00:00
07af755d1e Move the artifact pair to the exact commits @v3 was resolving to
Some checks failed
check / check (push) Successful in 8s
Build and Deploy to Cloudflare Pages / build (push) Successful in 8s
probe / r1-wrangler-only (push) Failing after 7s
probe / r2a-upload-proven (push) Successful in 12s
probe / r3a-upload-node20 (push) Successful in 8s
Build and Deploy to Cloudflare Pages / deploy (push) Has been skipped
probe / r2b-download-proven (push) Successful in 2s
probe / r3b-download-node20 (push) Successful in 2s
Round 2 (602fd60) put the build job green:

    check / check                        success   6s
    Build and Deploy .../ build          success  20s   <- green
    Build and Deploy .../ deploy         skipped        <- if: guard
    probe / q1-upload-v3-node16          success   7s
    probe / q2-upload-v3-node20          success  22s
    probe / q3-build-for-roundtrip       success  11s
    probe / q4-deploy-dryrun             failure  43s

Every v3 upload works and the build job is fixed. But q4 -- the deploy-side
rehearsal, which downloads the artifact in the pinned node container and
installs the pinned wrangler, stopping short of the publish call -- failed.
That is a break the deploy job would have hit on main, in a job nobody has
ever been able to run.

q4 bundled two things together, so round 3 splits them:

- r1 runs only the wrangler install and invocation. Worth measuring rather
  than assuming: wrangler 4.120.0 declares engines.node >= 22 and the deploy
  container is node 20, though the pre-issue deploy did run an unpinned
  wrangler on node:20 successfully.
- r2a/r2b run the artifact round trip with no wrangler at all.
- r3a/r3b do the same for the newer node20 artifact builds, so the choice
  between the two pairs is made on measurement.

deploy.yml meanwhile moves to the artifact commits that the mutable `@v3`
references were actually resolving to while this site was deploying, rather
than to the newest thing on the v3 line:

- upload-artifact   -> ff15f030 (v3.2.1)
- download-artifact -> 9bc31d5c (v3.0.2)

That is the conservative reading of what this issue is for: pin what is known
to work, do not take a version bump for free on the way past.
2026-08-09 02:55:51 +00:00
602fd609e7 Pin the artifact actions on v3: v4 does not work on this instance
Some checks failed
check / check (push) Successful in 6s
Build and Deploy to Cloudflare Pages / build (push) Successful in 20s
probe / q1-upload-v3-node16 (push) Successful in 7s
probe / q2-upload-v3-node20 (push) Successful in 22s
probe / q3-build-for-roundtrip (push) Successful in 11s
Build and Deploy to Cloudflare Pages / deploy (push) Has been skipped
probe / q4-deploy-dryrun (push) Failing after 43s
Round 1 of the branch probes reproduced the main failure and localised it.
Observed commit-status output for 2d328e7:

    check / check                        success  10s
    Build and Deploy .../ build          failure  15s   <- reproduced
    Build and Deploy .../ deploy         skipped        <- if: guard working
    probe / p1-bare-alpine-checkout      failure   3s
    probe / p2-alpine-apk-checkout       success   5s
    probe / p3-alpine-apk-build          success  15s
    probe / p4-alpine-apk-upload         failure  11s
    probe / p5-node20alpine-checkout     success   8s
    probe / p6-node20slim-checkout       success  11s

Reading that:

- p1 vs p2: act_runner does not supply node for JavaScript actions, so the
  `apk add --no-cache nodejs git tar` prerequisite step is genuinely required
  and genuinely sufficient. checkout then runs on musl.
- p3: script/bootstrap and script/test complete inside the Actions container
  on the pinned alpine digest. The mandated image replacement was never the
  problem.
- p2 vs p4: the only difference is a trailing upload-artifact v4 step, and it
  is the difference between success and failure.
- p5/p6: musl is not the issue -- checkout runs on both musl and glibc images.

So what broke the deploy was not the image swap that everyone reviewed, it was
the v3 -> v4 artifact bump that nobody questioned. Gitea 1.25.4's artifact
backend and this runner do not serve the v4 protocol; the workflow used v3
before this issue and that is what worked.

The artifact actions therefore move back to the v3 line, still pinned by full
commit SHA, which satisfies the hash-pinning requirement this issue is actually
about. Both are the node20 builds rather than the node16 defaults, so nothing
depends on a node16 runtime:

- upload-artifact  -> c6a3b2bd (v3.2.2-node20)
- download-artifact -> ad191675 (v3.1.0-node20)

Round 2 probes: the two fallback v3 builds in case the node20 ones do not
resolve, plus a producer/consumer pair that rehearses the deploy job -- same
pinned node image, same pinned download action, same pinned wrangler version,
stopping short of `wrangler pages deploy` so it touches nothing external.
2026-08-09 02:50:36 +00:00
2d328e759b Re-apply deploy.yml pinning behind a deploy guard, and probe the failure
Some checks failed
check / check (push) Successful in 10s
Build and Deploy to Cloudflare Pages / build (push) Failing after 15s
probe / p1-bare-alpine-checkout (push) Failing after 3s
probe / p2-alpine-apk-checkout (push) Successful in 5s
probe / p3-alpine-apk-build (push) Successful in 15s
probe / p4-alpine-apk-upload (push) Failing after 11s
probe / p5-node20alpine-checkout (push) Successful in 8s
probe / p6-node20slim-checkout (push) Successful in 11s
Build and Deploy to Cloudflare Pages / deploy (push) Has been skipped
Restores the hash-pinning work reverted in 3d17e22 (originally 3f91a7c and
b157bfd) verbatim -- all six pinned values were independently re-resolved and
confirmed correct twice, so they are reused, not re-derived.

What is different this time is that the path is observable before it reaches
main. The previous attempt broke the deploy because deploy.yml triggers only on
push to main, so every pre-merge check simulated the runner instead of being
it, and two adversarial reviews could not catch what neither could execute.

Three changes on top of the restored work:

- A temporary development-only branch trigger on on.push.branches, so the
  build job actually executes under act_runner. Removed before merge.
- if: github.ref_name == 'main' on the deploy job. Without it, a branch push
  would run wrangler pages deploy against the real Cloudflare project with the
  real token on every iteration. This guard is permanent: it is one line and it
  makes any future branch trigger, deliberate or accidental, unable to reach
  Cloudflare.
- A temporary .gitea/workflows/probe.yml, also deleted before merge. The
  Actions jobs and logs API is not readable by this account; the commit-status
  API is, and it reports one entry per job. So the diagnosis is encoded as job
  topology rather than log output: six jobs, each isolating one hypothesis
  about the 22s failure (bare alpine vs apk prerequisites, checkout vs site
  build vs artifact upload, musl node vs glibc node), each surfacing as its own
  status context so a single push tests them all in parallel.

make check is green. No pinned value is touched.
2026-08-09 02:46:29 +00:00