Round 2 (602fd60) put the build job green:
check / check success 6s
Build and Deploy .../ build success 20s <- green
Build and Deploy .../ deploy skipped <- if: guard
probe / q1-upload-v3-node16 success 7s
probe / q2-upload-v3-node20 success 22s
probe / q3-build-for-roundtrip success 11s
probe / q4-deploy-dryrun failure 43s
Every v3 upload works and the build job is fixed. But q4 -- the deploy-side
rehearsal, which downloads the artifact in the pinned node container and
installs the pinned wrangler, stopping short of the publish call -- failed.
That is a break the deploy job would have hit on main, in a job nobody has
ever been able to run.
q4 bundled two things together, so round 3 splits them:
- r1 runs only the wrangler install and invocation. Worth measuring rather
than assuming: wrangler 4.120.0 declares engines.node >= 22 and the deploy
container is node 20, though the pre-issue deploy did run an unpinned
wrangler on node:20 successfully.
- r2a/r2b run the artifact round trip with no wrangler at all.
- r3a/r3b do the same for the newer node20 artifact builds, so the choice
between the two pairs is made on measurement.
deploy.yml meanwhile moves to the artifact commits that the mutable `@v3`
references were actually resolving to while this site was deploying, rather
than to the newest thing on the v3 line:
- upload-artifact -> ff15f030 (v3.2.1)
- download-artifact -> 9bc31d5c (v3.0.2)
That is the conservative reading of what this issue is for: pin what is known
to work, do not take a version bump for free on the way past.
Round 1 of the branch probes reproduced the main failure and localised it.
Observed commit-status output for 2d328e7:
check / check success 10s
Build and Deploy .../ build failure 15s <- reproduced
Build and Deploy .../ deploy skipped <- if: guard working
probe / p1-bare-alpine-checkout failure 3s
probe / p2-alpine-apk-checkout success 5s
probe / p3-alpine-apk-build success 15s
probe / p4-alpine-apk-upload failure 11s
probe / p5-node20alpine-checkout success 8s
probe / p6-node20slim-checkout success 11s
Reading that:
- p1 vs p2: act_runner does not supply node for JavaScript actions, so the
`apk add --no-cache nodejs git tar` prerequisite step is genuinely required
and genuinely sufficient. checkout then runs on musl.
- p3: script/bootstrap and script/test complete inside the Actions container
on the pinned alpine digest. The mandated image replacement was never the
problem.
- p2 vs p4: the only difference is a trailing upload-artifact v4 step, and it
is the difference between success and failure.
- p5/p6: musl is not the issue -- checkout runs on both musl and glibc images.
So what broke the deploy was not the image swap that everyone reviewed, it was
the v3 -> v4 artifact bump that nobody questioned. Gitea 1.25.4's artifact
backend and this runner do not serve the v4 protocol; the workflow used v3
before this issue and that is what worked.
The artifact actions therefore move back to the v3 line, still pinned by full
commit SHA, which satisfies the hash-pinning requirement this issue is actually
about. Both are the node20 builds rather than the node16 defaults, so nothing
depends on a node16 runtime:
- upload-artifact -> c6a3b2bd (v3.2.2-node20)
- download-artifact -> ad191675 (v3.1.0-node20)
Round 2 probes: the two fallback v3 builds in case the node20 ones do not
resolve, plus a producer/consumer pair that rehearses the deploy job -- same
pinned node image, same pinned download action, same pinned wrangler version,
stopping short of `wrangler pages deploy` so it touches nothing external.
Restores the hash-pinning work reverted in 3d17e22 (originally 3f91a7c and
b157bfd) verbatim -- all six pinned values were independently re-resolved and
confirmed correct twice, so they are reused, not re-derived.
What is different this time is that the path is observable before it reaches
main. The previous attempt broke the deploy because deploy.yml triggers only on
push to main, so every pre-merge check simulated the runner instead of being
it, and two adversarial reviews could not catch what neither could execute.
Three changes on top of the restored work:
- A temporary development-only branch trigger on on.push.branches, so the
build job actually executes under act_runner. Removed before merge.
- if: github.ref_name == 'main' on the deploy job. Without it, a branch push
would run wrangler pages deploy against the real Cloudflare project with the
real token on every iteration. This guard is permanent: it is one line and it
makes any future branch trigger, deliberate or accidental, unable to reach
Cloudflare.
- A temporary .gitea/workflows/probe.yml, also deleted before merge. The
Actions jobs and logs API is not readable by this account; the commit-status
API is, and it reports one entry per job. So the diagnosis is encoded as job
topology rather than log output: six jobs, each isolating one hypothesis
about the 22s failure (bare alpine vs apk prerequisites, checkout vs site
build vs artifact upload, musl node vs glibc node), each surfacing as its own
status context so a single push tests them all in parallel.
make check is green. No pinned value is touched.