Raise org-wide make test cap to 60s, backstop timeout to 90s
#42
Reference in New Issue
Block a user
Delete Branch "org-wide-60s-test-cap"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
The question this answers
Is the 60-second test-time cap the new org-wide ceiling, or an approved
dnswatcher-only divergence?
This PR implements org-wide.
The ruling
On sneak/dnswatcher#93 (comment) (2026-08-09),
verbatim:
> make the cap 60s in both and never use mocking, always use live resolvers and
> assume the build and run environments have full unmodified unrestricted
> internet access. it is ok if they fail due to a bad build environment that
> alters dns packets.
And on #41 (comment), disambiguating
the scope:
> org wide. the hard cap is 60 for ci/green, but over 20s should be filed as an
> improvement bug.
That second comment landed after this work was started, and it confirms the
option implemented here. The two-tier shape it describes (60s hard, 20s target,
overage filed as a bug) is encoded in the policy text.
What changed, and the number chosen
The backstop moves from
30sto90s.The old pairing was incoherent under the new cap: a 60-second ceiling with a
30-second
-timeoutmeans the timeout kills the suite long before the ceilingis reached, so the ceiling would never be the thing that fails. The backstop has
to sit above the cap, where it does its actual job of catching a hung test
rather than a merely slow one.
90spreserves the 1.5x backstop-to-cap ratiothe old
20s/30spair already had, so the relationship between the twonumbers is unchanged and only the scale moves.
Every place a number changed:
prompts/REPO_POLICIES.md20 secondsto60 seconds, plus the new 20s target / improvement-bug tierprompts/REPO_POLICIES.md30-second timeoutto90-second timeout, with the rationale for why it exceeds the capprompts/REPO_POLICIES.mdgo test -timeout 30sto-timeout 90sprompts/REPO_POLICIES.mdgo test -timeout 30sto-timeout 90sprompts/EXISTING_REPO_CHECKLIST.mdmake testhas a30-secondtimeout, to90-secondtimeout plus the 60s hard cap and the 20s filing ruleprompts/NEW_REPO_CHECKLIST.mdscript/test/make testentrypoint line:30-second timeoutto90-second timeout, 60-second hard cap on wall timeI swept the whole repo for
20 second,30-second,20s,30s,timeout 20,timeout 30, andunder 20/under 30. After this change the only remainingoccurrences of
20in a test-timing context are the two intentional referencesto the new 20-second target. The other
timeouthits in the repo are unrelated(
.golangci.ymllint timeout, HTTP serverReadTimeout/WriteTimeoutexamples,middleware.Timeout) and were left alone.One deliberate non-change: the Python example Makefile snippet in
prompts/REPO_POLICIES.mdcarries no timeout flag today and still carries none.pytesthas no built-in timeout, so adding one would mean mandating thepytest-timeoutplugin org-wide, which is a new dependency requirement ratherthan a renumbering, and outside what was ruled on. It is a pre-existing gap
between the prose and that snippet, not one this PR introduces. Happy to file it
separately or add
--timeout=90here if you want it in scope.The alternative that was not implemented
Keep canonical at 20s and let
sneak/dnswatchercarry a documented per-repodivergence. That was the recommendation originally written up in
#41, on the reasoning that the 20s
ceiling is doing real work in repos with fast deterministic suites and only
dnswatcher needs the headroom.
Org-wide was chosen instead for two reasons. First, the pressure is not specific
to DNS: any repo whose tests exercise real infrastructure over the network
inherits the same variance, and there is no principled line that admits
dnswatcher and excludes the next such repo. Second, and more decisively,
REPO_POLICIES.mdis a vendored file. A sanctioned per-repo divergence in avendored file is indistinguishable, on inspection, from a stale vendored copy:
the next re-vendoring silently reverts the divergence, and nobody reading a
consuming repo can tell whether the number they are looking at is an intentional
exception or drift. That is the bidirectional-drift problem already tracked in
#31. The two-tier cap gets the same
outcome without the drift, since a fast repo that regresses from 4s to 45s still
generates an improvement bug.
Status
This PR was opened speculatively, ahead of a decision, on the standing "open it
rather than wait" instruction. The scope question has since been answered
org-wide in the issue, but the specific backstop value of
90sand the two-tierwording are still my proposals rather than anything ruled on, so closing this or
sending it back for a different number is a perfectly fine outcome.
make checkpasses;make fmtwas run and the result is included (it was ano-op, the edits were already prettier-conformant).
sneak/dnswatcheris landing the matching 60s edit to its vendored copy inparallel, and will match whichever way this is decided.