Long-lived integration branch. One commit per work unit lands here; this PR accumulates them until it is merged to `main`. ## Landed units - **Run all linting in Docker via `Dockerfile.lint` + `script/lint`** — #134 golangci-lint is no longer installed or run on the host. New root `Dockerfile.lint` COPYs the repo into the digest-pinned `golangci/golangci-lint:v2.12.2` image and lints as a build step, so a successful build IS a clean lint; `script/lint` is reduced to a thin wrapper that builds it. This also works where the docker daemon is remote and bind mounts are impossible. **Pinned digest and how it was verified.** `golangci/golangci-lint:v2.12.2@sha256:5cceeef04e53efe1470638d4b4b4f5ceefd574955ab3941b2d9a68a8c9ad5240`, exactly as quoted in the issue. It resolves, and it is genuinely v2.12.2: ``` $ docker buildx imagetools inspect golangci/golangci-lint:v2.12.2 Name: docker.io/golangci/golangci-lint:v2.12.2 MediaType: application/vnd.oci.image.index.v1+json Digest: sha256:5cceeef04e53efe1470638d4b4b4f5ceefd574955ab3941b2d9a68a8c9ad5240 $ docker run --rm golangci/golangci-lint@sha256:5cceeef04e...ad5240 golangci-lint --version golangci-lint has version 2.12.2 built with go1.26.2 from c0d3ddc9 on 2026-05-06T11:07:58Z ``` The tag's index digest is the quoted digest, and the binary inside reports commit `c0d3ddc9`, matching the org's canonical pin `c0d3ddc9cf3faa61a4e378e879ece580256d76e5`. **Forcing the linter to actually run.** `Dockerfile.lint` is split into a `deps` stage (base image + `go mod download`) and a `lint` stage (source copy + linter run). `script/lint` runs: ``` docker build --progress=plain --no-cache-filter=lint --target lint -f Dockerfile.lint . ``` Caching is explicitly waived for linting, and a cached build lints nothing, so the `lint` stage is invalidated on every invocation. The invalidation is scoped: the `deps` stage stays cached and no global cache wipe is performed. `--progress=plain` keeps the linter's own output visible. **`golangci-lint config verify`: deliberately NOT included.** It fetches its JSON schema over a live, unpinned HTTPS call, which would make linting network-dependent and defeat hash-pinning. Omitted for that reason, and the reason is recorded in a comment at the top of `Dockerfile.lint`. **`script/bootstrap`** no longer installs golangci-lint (and its pinned ref is gone); it warns non-fatally when `docker` is absent instead. The `goimports` install stays, because `script/fmt` and `script/fmt-check` still run on the host. Header comment updated accordingly. **Root `Dockerfile`** — required consequence, not scope creep. Its builder stage ran `make check`, which now calls `script/lint`, which shells out to `docker build`; there is no docker daemon inside a docker build, so `script/cibuild` and `script/docker` would have broken. It gains its own lint stage on the same pinned image (linter invoked directly, with a comment explaining why not `make lint`), with the builder stage depending on it via `COPY --from=lint /src/go.sum /dev/null` and running `make fmt-check`, `make test`, `make build`. The now-unneeded golangci-lint install is gone from the builder stage. **README** `Entrypoints` and `Building` sections now describe linting as a docker-only operation. `TODO.md` updated in the same commit. ### Verification All runs via `make` / `script/` entrypoints only. Two consecutive `make lint` runs on an unchanged tree, both executing the linter: ``` # run 1 #10 [lint 2/2] RUN golangci-lint run --config .golangci.yml ./... #10 10.98 0 issues. #10 DONE 12.0s # run 2, tree untouched #10 [lint 2/2] RUN golangci-lint run --config .golangci.yml ./... #10 14.63 0 issues. ``` A third run shows the cache scoping is working as intended — `deps` served from cache, `lint` re-executed: ``` #6 [deps 2/4] WORKDIR /src #6 CACHED #7 [deps 3/4] COPY go.mod go.sum ./ #7 CACHED #8 [deps 4/4] RUN go mod download #8 CACHED #9 [lint 1/2] COPY . . #10 [lint 2/2] RUN golangci-lint run --config .golangci.yml ./... ``` **Negative control.** A deliberate violation (an unused function containing an ineffectual assignment) was added to `internal/config/config.go`: ``` #10 11.26 internal/config/config.go:29:2: ineffectual assignment to x (ineffassign) #10 11.26 internal/config/config.go:28:6: func negativeControlUnused is unused (unused) #10 11.26 2 issues: #10 ERROR: process "/bin/sh -c golangci-lint run --config .golangci.yml ./..." did not complete successfully: exit code: 1 ERROR: failed to build: failed to solve: process "/bin/sh -c golangci-lint run --config .golangci.yml ./..." did not complete successfully: exit code: 1 make: *** [Makefile:26: lint] Error 1 ``` `make lint` exited non-zero naming both findings and their exact lines. After reverting the file, `make lint` was clean again (`0 issues.`). **`make check`** green end to end (test, lint, fmt-check), exit 0. **`script/cibuild`** green, confirming the `Dockerfile` restructure does not recurse: the lint stage ran (`#16 12.56 0 issues.`), then `#22 [builder 8/9] RUN make test` with `PASS` lines, then `#23 [builder 9/9] RUN make build`. ### Notes for the owner This supersedes two PRs you still have queued for merge, both of which tune host linting that no longer exists after this change: #128 (isolates the host golangci-lint cache and lock) and #131 (always installs the pinned lint tools in `script/bootstrap`). Neither was merged or incorporated here. Made moot by this change: #121 and #130. ### Review outcome Reviewed at #136 (comment) — **PASS**. Three non-blocking comment-accuracy findings were recorded there for the next touch of those files; the reviewer independently confirmed the lint gate is live by negative control from a warm cache. - **Live DNS tests made robust instead of gated; test caps moved to the org-wide 60s/20s/90s values** — #93 The resolver's live-DNS tests failed nondeterministically, a different subset each run. Fixed by engineering the nondeterminism out, not by routing around the network. **Nothing is mocked, faked, stubbed, recorded or replayed; there is no `-short` flag, no build tag, no skip, and no environment-tolerance for restricted egress.** Production resolver behaviour is unchanged. ### Root causes, all test-side 1. **Burst fan-out at one root server.** Every test in `internal/resolver` calls `t.Parallel()` and the build hosts have many cores (48 here), so all ~35 iterative resolutions started within milliseconds of each other, and because `queryServers` walks `rootServerList()` in fixed order they all aimed their first query at `198.41.0.4`. Root servers rate-limit that, which fits the reported symptom of a different arbitrary subset failing each run. 2. **No retry anywhere.** One dropped UDP packet in a delegation chain failed a test outright. 3. **Unanimity assertions.** `TestQueryAllNameservers_AllReturnOK` and `_NXDomainFromAllNS` required *every* nameserver of a domain to answer — four independent chances to fail per run, with no tolerance for one being slow. ### What was built New `internal/resolver/livedns_test.go` holds all the live-DNS machinery, so `resolver_test.go` itself takes only call-site edits: - **Bounded live concurrency.** A package-wide semaphore (`liveConcurrency = 6`) caps how many live resolutions are in flight at once. Tests keep `t.Parallel()`; only their network work is throttled. This is the direct fix for cause 1, and the 60s budget is what makes it affordable. - **Retry with exponential backoff.** Three attempts per live operation, 8s deadline each, 500ms base backoff doubling. The retry predicate is deliberately **transport-level** — "did a nameserver answer at all" — and never the assertion the test is making, so a resolver that answers *incorrectly* still fails on the first attempt rather than being retried into a false green. - **Quorum instead of unanimity, tolerating SILENCE ONLY.** A strict majority of the discovered nameservers must answer as expected, and every individual result must additionally fall inside a closed **allowlist** of statuses the test explicitly sanctions: `ok`/`timeout`/`error` for the all-OK test, `nxdomain`/`timeout`/`error` for the NXDOMAIN test. A nameserver that stays silent is tolerated; one that answers **wrongly** is not, at any count. The allowlist is the load-bearing part — see the rework note below for why a blocklist was not enough. New `internal/resolver/livedns_harness_test.go` tests that machinery directly — quorum arithmetic, status counting, the allowlist, the gate's concurrency bound, per-attempt deadlines, and recovery from a transient failure. It performs no DNS resolution of any kind, so it neither mocks DNS nor depends on it. ### Rework after review — the quorum could not fail on a wrong answer The review at #136 (comment) returned **FAIL** on `9cb2c2b`, correctly. Fixed in `87bce43`. **The defect.** The claim above was, as first written, false for `resolver.StatusNoData`. Each test banned exactly one wrong status — `_AllReturnOK` banned only `nxdomain`, `_NXDomainFromAllNS` banned only `ok` — and `nodata` is neither. It is a **wrong answer, not silence**: `answeredCount` counted it as answered, so it did not even trigger a retry, and with a quorum of 3-of-4 a single wrong nameserver slid through undetected. The pre-change unanimity assertions would have caught it. That is robustness work quietly becoming assertion-loosening, which is exactly what this repo cannot afford. **The fix.** Tolerance is now an allowlist, not a blocklist of one status. New `unsanctionedStatuses()` returns every per-nameserver result whose status the caller did not explicitly sanction, and each test asserts that list is empty in addition to its quorum. A blocklist bans the one wrong answer its author thought of and silently admits everything else, including any status added to the resolver later; an allowlist fails on anything nobody sanctioned. `answeredCount` was reframed the same way — it now counts the closed set `ok`/`nxdomain`/`nodata`, so an unfamiliar status is treated as silence and can only ever cause a retry and then a loud failure, never a quiet pass. **Evidence — the reviewer's exact probe, re-run.** `queryEachNS` in `internal/resolver/iterative.go` was patched to force one of `google.com`'s four nameservers to return `StatusNoData` with empty records. `make test` now goes **red**, naming the offending nameserver and status: ``` exit=2 --- FAIL: TestQueryAllNameservers_AllReturnOK (1.12s) Error: Should be empty, but was [ns1.google.com.=nodata] Messages: every nameserver must answer OK or not answer at all: ns1.google.com.=nodata ns2.google.com.=ok ns3.google.com.=ok ns4.google.com.=ok --- FAIL: TestQueryAllNameservers_NXDomainFromAllNS (1.34s) Error: Should be empty, but was [ns1.google.com.=nodata] Messages: every nameserver must report NXDOMAIN or not answer at all: ns1.google.com.=nodata ns2.google.com.=nxdomain ns3.google.com.=nxdomain ns4.google.com.=nxdomain FAIL sneak.berlin/go/dnswatcher/internal/resolver 1.704s ``` That is the same input that returned `exit=0` with both tests **passing** under the old assertions. Probe reverted, tree clean, suite green again: ``` exit=0 ok sneak.berlin/go/dnswatcher/internal/resolver 4.185s coverage: 77.1% of statements (zero `(cached)` lines) ``` Two harness tests lock the regression in without any probe: three OK plus one `nodata` (quorum satisfied, no NXDOMAIN present — the exact input that used to pass) is reported as unsanctioned, and an unknown status is neither counted as answered nor tolerated. Also fixed from the review: the per-attempt deadline assertion in `livedns_harness_test.go` had no lower bound, so it passed for a deadline far shorter than intended. It now asserts the remaining time exceeds `liveAttemptTimeout/2` as well. Deliberately **not** done in this rework, per the review and the owner: no `-count=1` in `script/test` (the test-cache issue is real but pre-existing and repo-wide, filed separately); the remaining non-DNS mocks stay for #97; `queryServers` root-ordering stays untouched under #138. ### The `-timeout` backstop value: 90s Per the ruling at sneak/prompts#41 (comment) the cap is org-wide with two tiers: **60s hard cap for CI green, 20s target, and anything between the two must be filed as an improvement bug.** The backstop is **`90s`**, matching sneak/prompts#42 and preserving the 1.5x backstop-to-cap ratio the old 20s/30s pair had. It must strictly exceed the 60s cap or the cap is unreachable — the old `-timeout 30s` would have killed a 60s-capped suite at half its allowance. Applied to `script/test`; nothing else in the repo carried the old `30s`. Worst case for one live operation is 3 attempts x 8s plus ~1.5s of backoff, about 26s — comfortably inside the 90s backstop even if several operations exhaust their attempts at once. ### `REPO_POLICIES.md` is re-vendored, not hand-edited The file is org-canonical, so it was **copied byte-for-byte** from `prompts/REPO_POLICIES.md` on `sneak/prompts` branch `org-wide-60s-test-cap` (commit `52b5192`) rather than reworded to approximately the same thing. Verified: ``` $ cmp prompts/REPO_POLICIES.md dnswatcher/REPO_POLICIES.md && echo identical identical $ sha256sum REPO_POLICIES.md bcf11c312a1bee18a0e937eb412b51914411c1ab23308b8362409f3f88379ff7 ``` **What that byte-identity does and does not certify.** The source branch `org-wide-60s-test-cap` is an **unmerged proposal** — sneak/prompts#42 — not `prompts` `main`. So, precisely: - The **60s hard cap and 20s improvement-bug tier ARE the owner's ruling** (sneak/prompts#41 (comment)). - The **`90s` backstop is our own proposed number and is NOT ratified** (sneak/prompts#41 (comment)). - The vendored text is therefore **the proposed canonical text, pending** sneak/prompts#42. If that PR lands with different numbers, this file must be re-vendored to match; it should not be hand-edited here either way. **Known mismatch with this repo's actual state, recorded not papered over.** Re-vendoring also picked up the paragraph at `REPO_POLICIES.md:266-271` mandating that canonical golangci-lint be installed commit-pinned via `go install ...@c0d3ddc9...`. This repo does **not** comply with that mechanism: `cc86473` in this same PR made linting Docker-only, and `script/bootstrap` now installs golangci-lint nowhere. The **version and commit match** (`v2.12.2` / `c0d3ddc9`); the **installation mechanism does not**. The vendored file is org-canonical and must not be edited downstream, so this is being raised upstream for the org text to accommodate Docker-only linting rather than patched here. `TESTING.md`'s stale "within the 30-second target" follows to 60. That edit is deliberately a single line so it merges cleanly when #97 lands. ### Verification All runs through `make` / `script/` entrypoints only; lint runs in Docker. **Ten consecutive `make check` runs, all green, none served from cache.** Go's test cache will happily report `ok pkg (cached)` without executing anything, which proves nothing about nondeterminism, so every run was forced to actually execute and each log was checked for zero `(cached)` lines: ``` check#1 exit=0 wall=32s resolver=2.895s fails=0 cached=0 check#2 exit=0 wall=27s resolver=2.872s fails=0 cached=0 check#3 exit=0 wall=47s resolver=2.729s fails=0 cached=0 check#4 exit=0 wall=36s resolver=2.885s fails=0 cached=0 check#5 exit=0 wall=32s resolver=2.899s fails=0 cached=0 check#6 exit=0 (harness tests added) fails=0 cached=0 check#7 exit=0 wall=41s resolver=2.891s fails=0 cached=0 check#8 exit=0 wall=36s resolver=2.871s fails=0 cached=0 check#9 exit=0 wall=26s resolver=2.846s fails=0 cached=0 check#10 exit=0 wall=46s resolver=2.833s fails=0 cached=0 ``` After the rework commit `87bce43`, `make check` green again end to end, zero `(cached)` test lines, Docker lint stage demonstrably executed rather than served from cache: ``` exit=0 cached=0 #8 [deps 4/4] RUN go mod download #8 CACHED #10 [lint 2/2] RUN golangci-lint run --config .golangci.yml ./... #10 30.72 0 issues. ``` The Docker lint stage was confirmed to execute rather than cache on each run: ``` #8 [deps 4/4] RUN go mod download #8 CACHED #9 [lint 1/2] COPY . . #9 DONE 0.6s #10 [lint 2/2] RUN golangci-lint run --config .golangci.yml ./... #10 DONE 39.7s ``` **`make test` wall time: 3.6-4.1s** across three timed uncached runs (`4093ms`, `3649ms`, `3729ms`). `internal/resolver` went from 2.0s to ~2.9s — the concurrency gate's cost. That is inside the 20s target, so no improvement bug is owed under the new two-tier rule. **Honest note on what these runs do and do not prove.** Live DNS was healthy throughout: **no live-DNS retry fired even once**, and no flake was observed either before or after the change (six pre-change baseline runs were also clean). So these runs demonstrate the change is not itself flaky and does not slow the suite; they do **not** demonstrate recovery from a real DNS failure, because no real DNS failure occurred. The original flakiness is *not reproduced* rather than *shown fixed*. The retry path is instead proven by `TestRetryLiveRecoversFromTransientFailure`, the only source of the single `retrying in 500ms` line in each log: ``` livedns_harness_test.go:90: transient: attempt 1 of 3 failed (no answer from live DNS), retrying in 500ms ``` ### Observation, not acted on The single most effective remaining lever against root-server rate limiting would be to stop `queryServers` always trying `rootServerList()` in the same order, so that load spreads across all thirteen roots instead of concentrating on `a.root-servers.net`. That is **production** code and this issue scopes the work as test-side, so it was left alone rather than changed quietly. It is now tracked for the owner's decision at #138. ### Interaction with #97 `TESTING.md` and `internal/resolver/resolver_test.go` auto-merge — that PR touches `resolver_test.go` only at the import block and the final timeout-test section, while this change touches the body of the file and adds two new files, and it leaves the mock-`DNSClient` timeout test at the tail of `resolver_test.go` entirely alone since removing it is that PR's job. `TODO.md` does conflict; that PR is already labelled `needs-rebase`, so this adds nothing material to its rebase. - **Go's test cache disabled, so every `make test` actually queries live DNS** — #139 `script/test` did not pass `-count=1`, so on an unchanged tree Go served the whole suite from cache: exit 0 in ~0.2s, every package marked `(cached)`, and not one DNS query made. This repo's suite exists to exercise live resolution on every run (`TESTING.md`), so that green asserted nothing — and it is exactly the green used as evidence that a flakiness fix works, since "run it a few times" stops being runs after the first. It had already misled two agents, each of whom forced uncached runs by hand. `-count=1` now disables caching on every invocation. **The conditional verbose rerun was missing and is added here.** `REPO_POLICIES.md` mandates it; the primary run had been unconditionally `-v`, which is the failure mode the policy exists to prevent (unreadable CI and `docker build` logs on success). Tests now run quiet, and only a failure triggers the `-v` rerun. Two properties matter and both are covered: the rerun carries `-count=1` too, so it cannot replay a cached copy of the failure it is meant to diagnose; and its exit status is discarded in favour of a forced `1`, so a flake that passes the second time cannot turn the build green — the first failure already proved the suite broken. **`-timeout 90s` untouched.** It is a deliberate backstop that must strictly exceed the 60s hard cap. **No special-casing for the Docker build**, which reaches the same script via `RUN make test`: a fresh container's test cache is empty, so `-count=1` is a no-op there, and carving out an exception would only create a second code path that could drift. ### Verification **Proven from a warm cache, not a cold one.** The suite was run first to populate the cache, and the pre-change state confirmed: ``` ok sneak.berlin/go/dnswatcher/internal/config (cached) coverage: 92.6% of statements ok sneak.berlin/go/dnswatcher/internal/resolver (cached) coverage: 77.1% of statements ...8 of 8 packages (cached)... real 0m0.203s ``` With the change applied to that same warm cache, three back-to-back runs on an unchanged tree, **zero `(cached)` markers** in all three: ``` # run 1 # run 2 ok .../internal/config 1.045s ok .../internal/config 1.040s ok .../internal/handlers 1.029s ok .../internal/handlers 1.021s ok .../internal/notify 1.148s ok .../internal/notify 1.252s ok .../internal/portcheck 1.026s ok .../internal/portcheck 1.025s ok .../internal/resolver 3.005s ok .../internal/resolver 2.846s ok .../internal/state 1.078s ok .../internal/state 1.097s ok .../internal/tlscheck 1.067s ok .../internal/tlscheck 1.079s ok .../internal/watcher 1.566s ok .../internal/watcher 1.544s real 0m4.174s real 0m4.015s # run 3: grep -c '(cached)' => 0 real 0m4.519s ``` **Measured uncached wall time: 4.0-4.5s** (was ~0.2s served from cache). Inside the 20s target, so no improvement bug is owed under the two-tier rule at sneak/prompts#41 (comment). **`-count=1` composes with `-race` and `-cover`**: both still present in the primary run, and the per-package coverage percentages above are identical to the pre-change values. **The failure path was exercised, not assumed.** A purpose-built flaky test that fails on its first run and passes every run after (marker file kept outside the module, so the tree stays byte-identical and a cached result would be served if caching were on) was run through the script: ``` exit code: 1 --- FAIL: TestFlaky (0.00s) FAIL flakeproof 0.013s --- Rerunning with -v for details --- --- PASS: TestFlaky (0.00s) ok flakeproof 1.014s ``` Quiet failure, verbose rerun that genuinely re-executed (it passed, so it did not replay the cached `FAIL`), and exit `1` regardless of the rerun passing. Scratch module removed afterwards. **`make check` green**, exit 0, with the Docker lint stage demonstrably executed rather than served from cache: ``` #8 [deps 4/4] RUN go mod download #8 CACHED #10 [lint 2/2] RUN golangci-lint run --config .golangci.yml ./... #10 21.36 0 issues. #10 DONE 26.2s ``` `README.md` and `TESTING.md` record why the cache is waived, and `TODO.md` is updated in the same commit. ### Question for the owner, not filed as a defect The Docker lint run emits `The linter 'gomodguard' is deprecated (since v2.12.0) due to: new major version. Replaced by gomodguard_v2.` It is pre-existing and out of this issue's scope. It is not filed as an issue here because `.golangci.yml` tracks the org-canonical config, so switching to `gomodguard_v2` looks like an upstream `sneak/prompts` decision rather than a per-repo fix. Say the word and it gets filed in whichever place you consider canonical. - **MIT `LICENSE` added; README states the licence** — #102 The repo had no licence file at all, so publicly readable code was all-rights-reserved by default and nobody could legally use it. `LICENSE` was also the last file missing from `REPO_POLICIES.md`'s required minimum. MIT, by standing org policy rather than a per-repo call: any public repo lacking a licence gets MIT, and a private one with no licence is already all-rights-reserved. `sneak/dnswatcher` is public (`private: false` on the Gitea repo record). `README.md`'s first line now names the licence, per the Description requirement, and the License section states MIT and points at the file instead of recording the decision as pending. `TODO.md` updated in the same commit. ### Verification `LICENSE` is the canonical MIT text byte-for-byte with only the copyright line filled in (`Copyright (c) 2026 sneak`) — no clauses added, removed, reworded, or reflowed. It was not typed from memory: the file was copied from an existing verbatim MIT template on disk and only the copyright line edited (`diff` against that template shows that one line and nothing else), then the result was word-diffed against SPDX `MIT.txt` fetched from `spdx/license-list-data`, ignoring only line wrapping and the placeholder — identical. `make fmt` did **not** touch `LICENSE`, and cannot: `script/fmt` runs `gofmt -s -w .` and `goimports -w .` only, with no prettier or markdown step in the repo, so no exclusion was needed. `make check` green, exit 0. Tests executed rather than replayed (zero `(cached)` lines, `internal/resolver 3.098s`), and the Docker lint stage ran rather than cached: ``` #10 [deps 4/4] RUN go mod download #10 CACHED #12 [lint 2/2] RUN golangci-lint run --config .golangci.yml ./... #12 36.44 0 issues. #12 DONE 36.7s ``` - **Comment-only corrections in `script/bootstrap`, `script/cibuild`, and `Dockerfile.lint`** — #137 Follow-up to this PR's own review. Nothing executable changed: the diff touches comment lines, one warning string, and `TODO.md`. - `script/bootstrap`'s header justified the pinned `goimports` install by claiming `script/fmt-check` runs it on the host. Verified against the script: `script/fmt-check` runs `gofmt -l .` and nothing else. The header now credits `script/fmt` alone. That `fmt-check` does not verify goimports at all is #119 and was deliberately left alone. - `script/cibuild`'s header still said the `Dockerfile` runs `make check`. It now describes the current file: lint stage runs `make fmt-check` and `golangci-lint`, builder stage runs `make test` and `make build`. - The `docker`-missing warning was three fragments, each re-prefixed with `bootstrap:` mid-clause. Now one sentence: `bootstrap: WARNING: docker not found; install it to run make lint and make docker.` - `Dockerfile.lint`'s comment explained why `golangci-lint config verify` is omitted but read as though the omission were free. It now states the residual risk: unknown top-level keys in `.golangci.yml` are silently ignored, so a mistyped or wrong-schema key lints clean while applying nothing. `config verify` was **not** added — the network-dependence reasoning stands. ### Verification `make check` green, exit 0. Tests executed rather than replayed (zero `(cached)` lines, `internal/resolver 2.820s`), Docker lint stage executed rather than cached: ``` #8 [deps 4/4] RUN go mod download #8 CACHED #10 [lint 2/2] RUN golangci-lint run --config .golangci.yml ./... #10 13.30 0 issues. ``` Comment-only confirmed by reading the whole diff: no statement, flag, or command changed anywhere. --- ## Issues closed by this merge The commits on `next` each carry a bare `(closes #N)` in their subject, but the references in the prose above are full URLs, which Gitea's auto-close parser does not act on. Listing them here in bare form so the merge to `main` definitively closes them rather than leaving them open to be re-picked up as idle work: Closes #93 Closes #102 Closes #134 Closes #137 Closes #139 Co-authored-by: sneak <sneak@sneak.berlin> Reviewed-on: #136 Co-authored-by: clawbot <clawbot@noreply.example.org> Co-committed-by: clawbot <clawbot@noreply.example.org>
This commit was merged in pull request #136.
This commit is contained in:
+26
-12
@@ -1,13 +1,9 @@
|
|||||||
# Build stage
|
# Lint stage - fast feedback on lint issues, before the build starts.
|
||||||
# golang 1.25-alpine, 2026-02-28
|
# The linter is invoked directly rather than through `make lint`: that
|
||||||
FROM golang@sha256:f6751d823c26342f9506c03797d2527668d095b0a15f1862cddb4d927a7a4ced AS builder
|
# target shells out to `docker build -f Dockerfile.lint`, and there is
|
||||||
|
# no docker daemon inside a docker build.
|
||||||
RUN apk add --no-cache git make gcc musl-dev binutils-gold
|
# golangci/golangci-lint:v2.12.2 (Debian-based), 2026-08-10
|
||||||
|
FROM golangci/golangci-lint:v2.12.2@sha256:5cceeef04e53efe1470638d4b4b4f5ceefd574955ab3941b2d9a68a8c9ad5240 AS lint
|
||||||
# golangci-lint v2.12.2, 2026-08-07
|
|
||||||
RUN go install github.com/golangci/golangci-lint/v2/cmd/golangci-lint@c0d3ddc9cf3faa61a4e378e879ece580256d76e5
|
|
||||||
# goimports v0.42.0
|
|
||||||
RUN go install golang.org/x/tools/cmd/goimports@009367f5c17a8d4c45a961a3a509277190a9a6f0
|
|
||||||
|
|
||||||
WORKDIR /src
|
WORKDIR /src
|
||||||
COPY go.mod go.sum ./
|
COPY go.mod go.sum ./
|
||||||
@@ -15,8 +11,26 @@ RUN go mod download
|
|||||||
|
|
||||||
COPY . .
|
COPY . .
|
||||||
|
|
||||||
# Run all checks - build fails if any check fails
|
RUN make fmt-check
|
||||||
RUN make check
|
RUN golangci-lint run --config .golangci.yml ./...
|
||||||
|
|
||||||
|
# Build stage
|
||||||
|
# golang 1.25-alpine, 2026-02-28
|
||||||
|
FROM golang@sha256:f6751d823c26342f9506c03797d2527668d095b0a15f1862cddb4d927a7a4ced AS builder
|
||||||
|
|
||||||
|
RUN apk add --no-cache git make gcc musl-dev binutils-gold
|
||||||
|
|
||||||
|
# Force BuildKit to run the lint stage before proceeding
|
||||||
|
COPY --from=lint /src/go.sum /dev/null
|
||||||
|
|
||||||
|
WORKDIR /src
|
||||||
|
COPY go.mod go.sum ./
|
||||||
|
RUN go mod download
|
||||||
|
|
||||||
|
COPY . .
|
||||||
|
|
||||||
|
# Run the tests - build fails if any test fails
|
||||||
|
RUN make test
|
||||||
|
|
||||||
# Build the binary
|
# Build the binary
|
||||||
RUN make build
|
RUN make build
|
||||||
|
|||||||
@@ -0,0 +1,29 @@
|
|||||||
|
# Lint-only image: used by script/lint. golangci-lint is never run on
|
||||||
|
# the host — the repo is COPYed into the build context and the linter
|
||||||
|
# runs as a build step, so a successful build IS a clean lint. This
|
||||||
|
# also works where the docker daemon is remote and bind mounts are
|
||||||
|
# impossible.
|
||||||
|
#
|
||||||
|
# `golangci-lint config verify` is deliberately NOT run here: it
|
||||||
|
# fetches its JSON schema over a live, unpinned HTTPS call, which would
|
||||||
|
# make linting network-dependent and defeat hash-pinning. The cost of
|
||||||
|
# that: unknown top-level keys in .golangci.yml are silently ignored,
|
||||||
|
# so a mistyped or wrong-schema key lints clean while applying nothing.
|
||||||
|
#
|
||||||
|
# golangci/golangci-lint:v2.12.2 (Debian-based), 2026-08-10
|
||||||
|
FROM golangci/golangci-lint:v2.12.2@sha256:5cceeef04e53efe1470638d4b4b4f5ceefd574955ab3941b2d9a68a8c9ad5240 AS deps
|
||||||
|
|
||||||
|
WORKDIR /src
|
||||||
|
|
||||||
|
# Dependencies first, so this stage stays cached across lint runs.
|
||||||
|
COPY go.mod go.sum ./
|
||||||
|
RUN go mod download
|
||||||
|
|
||||||
|
# Everything below is invalidated on every run by the
|
||||||
|
# --no-cache-filter=lint that script/lint passes: caching is explicitly
|
||||||
|
# waived for linting, and a cached build lints nothing.
|
||||||
|
FROM deps AS lint
|
||||||
|
|
||||||
|
COPY . .
|
||||||
|
|
||||||
|
RUN golangci-lint run --config .golangci.yml ./...
|
||||||
@@ -0,0 +1,21 @@
|
|||||||
|
MIT License
|
||||||
|
|
||||||
|
Copyright (c) 2026 sneak
|
||||||
|
|
||||||
|
Permission is hereby granted, free of charge, to any person obtaining a copy
|
||||||
|
of this software and associated documentation files (the "Software"), to deal
|
||||||
|
in the Software without restriction, including without limitation the rights
|
||||||
|
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
||||||
|
copies of the Software, and to permit persons to whom the Software is
|
||||||
|
furnished to do so, subject to the following conditions:
|
||||||
|
|
||||||
|
The above copyright notice and this permission notice shall be included in all
|
||||||
|
copies or substantial portions of the Software.
|
||||||
|
|
||||||
|
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
||||||
|
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
||||||
|
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
||||||
|
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
||||||
|
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
||||||
|
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
||||||
|
SOFTWARE.
|
||||||
@@ -1,6 +1,6 @@
|
|||||||
# dnswatcher
|
# dnswatcher
|
||||||
|
|
||||||
dnswatcher is a pre-1.0 Go daemon by [@sneak](https://sneak.berlin) that monitors DNS records, TCP port availability, and TLS certificates, delivering real-time change notifications via Slack, Mattermost, and ntfy webhooks.
|
dnswatcher is an MIT-licensed, pre-1.0 Go daemon by [@sneak](https://sneak.berlin) that monitors DNS records, TCP port availability, and TLS certificates, delivering real-time change notifications via Slack, Mattermost, and ntfy webhooks.
|
||||||
|
|
||||||
> ⚠️ Pre-1.0 software. APIs, configuration, and behavior may change without notice.
|
> ⚠️ Pre-1.0 software. APIs, configuration, and behavior may change without notice.
|
||||||
|
|
||||||
@@ -218,7 +218,8 @@ internal/
|
|||||||
- **Structured logging**: All logs use `log/slog` with JSON output in
|
- **Structured logging**: All logs use `log/slog` with JSON output in
|
||||||
production (TTY detection for development).
|
production (TTY detection for development).
|
||||||
- **Graceful shutdown**: All background goroutines respect context
|
- **Graceful shutdown**: All background goroutines respect context
|
||||||
cancellation and the fx lifecycle.
|
cancellation and the fx lifecycle. In-flight notification deliveries
|
||||||
|
are drained on shutdown, bounded by the shutdown timeout.
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
@@ -380,14 +381,25 @@ standard: normalized scripts in `script/` are the entrypoints for the
|
|||||||
development workflow, and the Makefile targets are thin shims that call
|
development workflow, and the Makefile targets are thin shims that call
|
||||||
them. We provide:
|
them. We provide:
|
||||||
|
|
||||||
- `script/bootstrap` — install all dependencies (go, pinned
|
- `script/bootstrap` — install all dependencies (go, pinned goimports,
|
||||||
golangci-lint and goimports, `go mod download`)
|
`go mod download`). It does not install golangci-lint: see
|
||||||
|
`script/lint` below.
|
||||||
- `script/setup` — make a fresh clone ready for development: bootstrap
|
- `script/setup` — make a fresh clone ready for development: bootstrap
|
||||||
plus the git pre-commit hook
|
plus the git pre-commit hook
|
||||||
- `script/projectname` — print the project name (used for the Docker
|
- `script/projectname` — print the project name (used for the Docker
|
||||||
image tag)
|
image tag)
|
||||||
- `script/test` — run the test suite (race detector, coverage)
|
- `script/test` — run the test suite (race detector, coverage). Caching
|
||||||
- `script/lint` — run golangci-lint
|
is waived for testing, exactly as it is for linting: `-count=1`
|
||||||
|
forces every invocation to execute, because the suite queries live
|
||||||
|
DNS and a cached pass queries nothing. Failures are rerun with `-v`
|
||||||
|
automatically, and the build fails even if that rerun passes.
|
||||||
|
- `script/lint` — run golangci-lint, always inside Docker: it builds
|
||||||
|
`Dockerfile.lint`, which COPYs the repo into the digest-pinned
|
||||||
|
`golangci-lint` image and lints as a build step, so a successful
|
||||||
|
build is a clean lint. The linter is never installed or run on the
|
||||||
|
host, and Docker is the only prerequisite. Caching is waived for
|
||||||
|
linting: the lint stage is forced to execute on every run with
|
||||||
|
`--no-cache-filter`, because a cached build lints nothing.
|
||||||
- `script/fmt` — format all code (gofmt -s, goimports)
|
- `script/fmt` — format all code (gofmt -s, goimports)
|
||||||
- `script/fmt-check` — check formatting (read-only)
|
- `script/fmt-check` — check formatting (read-only)
|
||||||
- `script/check` — run test, lint, and fmt-check
|
- `script/check` — run test, lint, and fmt-check
|
||||||
@@ -403,7 +415,7 @@ them. We provide:
|
|||||||
```sh
|
```sh
|
||||||
make build # Build binary to bin/dnswatcher
|
make build # Build binary to bin/dnswatcher
|
||||||
make test # Run tests with race detector
|
make test # Run tests with race detector
|
||||||
make lint # Run golangci-lint
|
make lint # Run golangci-lint in Docker (requires docker)
|
||||||
make fmt # Format code
|
make fmt # Format code
|
||||||
make check # Run all checks (test, lint, fmt-check)
|
make check # Run all checks (test, lint, fmt-check)
|
||||||
make clean # Remove build artifacts
|
make clean # Remove build artifacts
|
||||||
@@ -452,8 +464,14 @@ docker run -d \
|
|||||||
from a previous cycle.
|
from a previous cycle.
|
||||||
4. **On change detection**: Send notifications to all configured
|
4. **On change detection**: Send notifications to all configured
|
||||||
endpoints, update in-memory state, persist to disk.
|
endpoints, update in-memory state, persist to disk.
|
||||||
5. **Shutdown**: Persist final state to disk, complete in-flight
|
5. **Shutdown**: Persist final state to disk, wait for in-flight
|
||||||
notifications, stop gracefully.
|
notification deliveries to complete, stop gracefully. The wait is
|
||||||
|
bounded by the fx shutdown timeout (15s by default): deliveries still
|
||||||
|
retrying against an unreachable endpoint when that expires are
|
||||||
|
abandoned, and the number abandoned is logged at warn level rather
|
||||||
|
than dropped silently. Notifications generated after shutdown has
|
||||||
|
begun are refused and logged, so a late burst cannot extend the
|
||||||
|
shutdown.
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
@@ -475,8 +493,9 @@ Viper for configuration.
|
|||||||
|
|
||||||
## License
|
## License
|
||||||
|
|
||||||
License has not yet been chosen for this project. Pending decision by the
|
dnswatcher is released under the MIT License, Copyright (c) 2026
|
||||||
author (MIT, GPL, or WTFPL).
|
[@sneak](https://sneak.berlin). See the [`LICENSE`](./LICENSE) file in the
|
||||||
|
repository root for the full text.
|
||||||
|
|
||||||
## Author
|
## Author
|
||||||
|
|
||||||
|
|||||||
+14
-6
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: Repository Policies
|
title: Repository Policies
|
||||||
last_modified: 2026-07-06
|
last_modified: 2026-08-07
|
||||||
---
|
---
|
||||||
|
|
||||||
This document covers repository structure, tooling, and workflow standards. Code
|
This document covers repository structure, tooling, and workflow standards. Code
|
||||||
@@ -189,8 +189,13 @@ style conventions are in separate documents:
|
|||||||
module under test to verify it compiles/parses. There is no excuse for
|
module under test to verify it compiles/parses. There is no excuse for
|
||||||
`make test` to be a no-op.
|
`make test` to be a no-op.
|
||||||
|
|
||||||
- `make test` must complete in under 20 seconds. Add a 30-second timeout in the
|
- `make test` must complete in under 60 seconds. That is the hard cap, and a
|
||||||
Makefile.
|
suite that exceeds it fails. Under 20 seconds is the target. A suite between
|
||||||
|
20 and 60 seconds is still green, but the overage must be filed as an
|
||||||
|
improvement bug against that repo. Add a 90-second timeout to the test
|
||||||
|
invocation in the Makefile (`go test -timeout 90s`). The backstop deliberately
|
||||||
|
sits above the hard cap so that it catches a genuinely hung test rather than a
|
||||||
|
merely slow one.
|
||||||
|
|
||||||
- **`make test` should use the conditional verbose rerun pattern.** Run tests
|
- **`make test` should use the conditional verbose rerun pattern.** Run tests
|
||||||
without `-v` (verbose) first. If tests fail, automatically rerun with `-v` to
|
without `-v` (verbose) first. If tests fail, automatically rerun with `-v` to
|
||||||
@@ -209,9 +214,9 @@ style conventions are in separate documents:
|
|||||||
|
|
||||||
```makefile
|
```makefile
|
||||||
test:
|
test:
|
||||||
@go test -timeout 30s -race -cover ./... || \
|
@go test -timeout 90s -race -cover ./... || \
|
||||||
{ echo "--- Rerunning with -v for details ---"; \
|
{ echo "--- Rerunning with -v for details ---"; \
|
||||||
go test -timeout 30s -race -v ./...; exit 1; }
|
go test -timeout 90s -race -v ./...; exit 1; }
|
||||||
```
|
```
|
||||||
|
|
||||||
Python example:
|
Python example:
|
||||||
@@ -260,7 +265,10 @@ style conventions are in separate documents:
|
|||||||
|
|
||||||
- `.golangci.yml` is standardized and must _NEVER_ be modified by an agent, only
|
- `.golangci.yml` is standardized and must _NEVER_ be modified by an agent, only
|
||||||
manually by the user. Fetch from
|
manually by the user. Fetch from
|
||||||
`https://git.eeqj.de/sneak/prompts/raw/branch/main/.golangci.yml`.
|
`https://git.eeqj.de/sneak/prompts/raw/branch/main/.golangci.yml`. The
|
||||||
|
canonical golangci-lint version is v2.12.2 (released 2026-05-06), installed
|
||||||
|
commit-pinned via
|
||||||
|
`go install github.com/golangci/golangci-lint/v2/cmd/golangci-lint@c0d3ddc9cf3faa61a4e378e879ece580256d76e5`.
|
||||||
|
|
||||||
- When pinning images or packages by hash, add a comment above the reference
|
- When pinning images or packages by hash, add a comment above the reference
|
||||||
with the version and date (YYYY-MM-DD).
|
with the version and date (YYYY-MM-DD).
|
||||||
|
|||||||
+4
-1
@@ -17,7 +17,7 @@ real servers ensures the resolver works correctly in production.
|
|||||||
|
|
||||||
- Tests hit real DNS infrastructure and require network access
|
- Tests hit real DNS infrastructure and require network access
|
||||||
- Test duration depends on network conditions; timeout tuning keeps
|
- Test duration depends on network conditions; timeout tuning keeps
|
||||||
the suite within the 30-second target
|
the suite within the 60-second target
|
||||||
- Query timeout is calibrated to 3× maximum antipodal RTT (~300ms)
|
- Query timeout is calibrated to 3× maximum antipodal RTT (~300ms)
|
||||||
plus processing margin
|
plus processing margin
|
||||||
- Root server fan-out is limited to reduce parallel query load
|
- Root server fan-out is limited to reduce parallel query load
|
||||||
@@ -31,4 +31,7 @@ real servers ensures the resolver works correctly in production.
|
|||||||
exists for unit-testing other packages that consume the resolver)
|
exists for unit-testing other packages that consume the resolver)
|
||||||
- **Do not add `-short` flags** to skip slow tests
|
- **Do not add `-short` flags** to skip slow tests
|
||||||
- **Do not increase `-timeout`** to hide hanging queries
|
- **Do not increase `-timeout`** to hide hanging queries
|
||||||
|
- **Do not remove `-count=1` from `script/test`** — Go's test cache
|
||||||
|
replays a previous run's output without querying anything, so a
|
||||||
|
cached pass is not evidence that live resolution works
|
||||||
- **Do not modify linter configuration** to suppress findings
|
- **Do not modify linter configuration** to suppress findings
|
||||||
|
|||||||
@@ -18,13 +18,89 @@ iterative resolver implementation with hermetic mocked tests.
|
|||||||
|
|
||||||
# Next Step
|
# Next Step
|
||||||
|
|
||||||
Policy scaffold commit: add LICENSE, REPO_POLICIES.md, .editorconfig,
|
Add the README sections required by policy (Description, Getting Started,
|
||||||
.dockerignore, and .gitea/workflows/check.yml, and add the missing
|
Rationale, Design, TODO, License, Author) if any are still missing.
|
||||||
fmt-check, docker, and hooks targets to the Makefile. One commit, then
|
|
||||||
confirm make check still passes.
|
|
||||||
|
|
||||||
# Completed Steps
|
# Completed Steps
|
||||||
|
|
||||||
|
- 2026-08-10: comment-only corrections to `script/bootstrap`,
|
||||||
|
`script/cibuild`, and `Dockerfile.lint`. The `goimports` pin in
|
||||||
|
`script/bootstrap` was justified by a claim that `script/fmt-check`
|
||||||
|
runs it on the host; it does not (it runs `gofmt -l .` only), so the
|
||||||
|
header now credits `script/fmt` alone. `script/cibuild` still claimed
|
||||||
|
the `Dockerfile` runs `make check`, which stopped being true when
|
||||||
|
linting moved to its own stage; it now describes the lint stage
|
||||||
|
(`make fmt-check` plus `golangci-lint`) and the builder stage
|
||||||
|
(`make test`, `make build`). The `docker`-missing warning in
|
||||||
|
`script/bootstrap` reads as one sentence instead of three fragments
|
||||||
|
each re-prefixed with `bootstrap:`. `Dockerfile.lint` now records the
|
||||||
|
residual risk of omitting `golangci-lint config verify`: unknown
|
||||||
|
top-level keys in `.golangci.yml` are silently ignored, so a mistyped
|
||||||
|
key lints clean while applying nothing. No behaviour changed
|
||||||
|
- 2026-08-10: MIT `LICENSE` added at the repository root, closing the
|
||||||
|
last gap in `REPO_POLICIES.md`'s required-minimum file list and
|
||||||
|
removing the all-rights-reserved default that would otherwise have
|
||||||
|
shipped with a 1.0 tag. The licence choice is the standing org policy
|
||||||
|
(any public repo lacking a licence gets MIT; a private repo with no
|
||||||
|
licence is already all-rights-reserved), and this repo is public. The
|
||||||
|
file holds the canonical MIT text byte-for-byte with only the
|
||||||
|
copyright line filled in (`Copyright (c) 2026 sneak`); no clauses were
|
||||||
|
added, removed, or reflowed. `README.md`'s first line now names the
|
||||||
|
licence, as the Description requirement demands, and the License
|
||||||
|
section states MIT and points at the file instead of saying the choice
|
||||||
|
is pending. `make fmt` covers only Go sources (`gofmt -s`,
|
||||||
|
`goimports`), so it cannot reflow `LICENSE`
|
||||||
|
- 2026-08-10: the policy scaffold (`REPO_POLICIES.md`, `.editorconfig`,
|
||||||
|
`.dockerignore`, `.gitea/workflows/check.yml`, and the `fmt-check`,
|
||||||
|
`docker`, and hooks Makefile targets) is present; it landed piecemeal
|
||||||
|
across the scripts-to-rule-them-all and policy commits rather than as
|
||||||
|
the single commit this file once planned
|
||||||
|
- 2026-08-10: Go's test cache disabled for `script/test` via `-count=1`,
|
||||||
|
so every invocation actually executes. A cached pass replays an
|
||||||
|
earlier run's output without querying DNS at all, which in this repo
|
||||||
|
means the suite's entire premise goes unexercised while the run
|
||||||
|
reports green in under a second. The conditional verbose rerun that
|
||||||
|
`REPO_POLICIES.md` mandates was added at the same time (the primary
|
||||||
|
run had been unconditionally `-v`): quiet first, `-v` only on
|
||||||
|
failure, `-count=1` on both, and exit 1 forced regardless of the
|
||||||
|
rerun's result so a flake passing the second time cannot turn the
|
||||||
|
build green. `-timeout 90s` left alone as the deliberate backstop
|
||||||
|
above the 60s hard cap. Uncached suite runs ~4s, well inside the 20s
|
||||||
|
target
|
||||||
|
- 2026-08-10: live-DNS test flakiness addressed by robustness rather
|
||||||
|
than gating, per the owner's ruling on #93: new
|
||||||
|
`internal/resolver/livedns_test.go` adds a package-wide concurrency
|
||||||
|
gate (so parallel tests stop bursting at the first root server),
|
||||||
|
retry with exponential backoff on transport failures only, and
|
||||||
|
quorum instead of unanimity for multi-nameserver assertions. Quorum
|
||||||
|
tolerates silence only: every per-nameserver status must be in a
|
||||||
|
closed allowlist (`ok`/`timeout`/`error`, or
|
||||||
|
`nxdomain`/`timeout`/`error`), so a wrong answer from a minority —
|
||||||
|
`nodata` today, any status added later — fails the test instead of
|
||||||
|
sliding through under the majority. The
|
||||||
|
`make test` cap moved to the new org-wide 60s hard cap / 20s target
|
||||||
|
with a 90s `-timeout` backstop; `REPO_POLICIES.md` re-vendored
|
||||||
|
byte-identical from `sneak/prompts`. No mocks, no `-short`, no build
|
||||||
|
tags, no skips, and no change to production resolver behaviour
|
||||||
|
- 2026-08-10: all linting moved into Docker: new root `Dockerfile.lint`
|
||||||
|
on the digest-pinned `golangci/golangci-lint:v2.12.2` image,
|
||||||
|
`script/lint` reduced to a thin wrapper that builds it with
|
||||||
|
`--no-cache-filter=lint` so the linter actually executes every run,
|
||||||
|
golangci-lint install dropped from `script/bootstrap` (goimports
|
||||||
|
stays, `script/fmt` needs it on the host), and the root `Dockerfile`
|
||||||
|
given its own lint stage so its build no longer recurses through
|
||||||
|
`make check` into `script/lint`. `golangci-lint config verify` is
|
||||||
|
deliberately omitted: it fetches its schema over an unpinned live
|
||||||
|
HTTPS call
|
||||||
|
- 2026-08-09: in-flight notification deliveries are now drained at
|
||||||
|
shutdown (#106): `notify.New` registers an fx `OnStop` hook that waits
|
||||||
|
on a `sync.WaitGroup` of tracked delivery goroutines, bounded by the
|
||||||
|
`OnStop` context; on expiry the outstanding count is logged at warn
|
||||||
|
level and parked retry backoffs are released instead of being dropped
|
||||||
|
silently, and deliveries submitted after the drain begins are refused
|
||||||
|
so shutdown cannot be extended indefinitely; an `OnStop` context that
|
||||||
|
is already expired on entry with nothing outstanding drains quietly
|
||||||
|
rather than warning about deliveries that were never abandoned
|
||||||
- 2026-08-07: golangci-lint bumped to v2.12.2 (commit-pinned installs
|
- 2026-08-07: golangci-lint bumped to v2.12.2 (commit-pinned installs
|
||||||
in `Dockerfile` and `script/bootstrap`); `.golangci.yml` set to the
|
in `Dockerfile` and `script/bootstrap`); `.golangci.yml` set to the
|
||||||
org-standard v2-schema config used across the org's repos
|
org-standard v2-schema config used across the org's repos
|
||||||
@@ -53,8 +129,6 @@ confirm make check still passes.
|
|||||||
|
|
||||||
Compliance:
|
Compliance:
|
||||||
|
|
||||||
- Add README sections required by policy (Description, Getting Started,
|
|
||||||
Rationale, Design, TODO, License, Author) if any are missing
|
|
||||||
- Pin Dockerfile base images by sha256 and ensure the Docker build runs
|
- Pin Dockerfile base images by sha256 and ensure the Docker build runs
|
||||||
make check
|
make check
|
||||||
|
|
||||||
|
|||||||
@@ -32,11 +32,27 @@ func NewRequestForTest(
|
|||||||
// NewTestService creates a Service suitable for unit testing.
|
// NewTestService creates a Service suitable for unit testing.
|
||||||
// It discards log output and uses the given transport.
|
// It discards log output and uses the given transport.
|
||||||
func NewTestService(transport http.RoundTripper) *Service {
|
func NewTestService(transport http.RoundTripper) *Service {
|
||||||
return &Service{
|
return newService(slog.New(slog.DiscardHandler), transport)
|
||||||
log: slog.New(slog.DiscardHandler),
|
}
|
||||||
transport: transport,
|
|
||||||
history: NewAlertHistory(),
|
// NewTestServiceWithLogger creates a Service that writes to the
|
||||||
}
|
// given handler, so tests can assert on emitted log records.
|
||||||
|
func NewTestServiceWithLogger(
|
||||||
|
transport http.RoundTripper,
|
||||||
|
handler slog.Handler,
|
||||||
|
) *Service {
|
||||||
|
return newService(slog.New(handler), transport)
|
||||||
|
}
|
||||||
|
|
||||||
|
// Drain exports drain for testing.
|
||||||
|
func (svc *Service) Drain(ctx context.Context) {
|
||||||
|
svc.drain(ctx)
|
||||||
|
}
|
||||||
|
|
||||||
|
// OutstandingDeliveries reports how many delivery goroutines
|
||||||
|
// are currently tracked as in flight.
|
||||||
|
func (svc *Service) OutstandingDeliveries() int64 {
|
||||||
|
return svc.outstanding.Load()
|
||||||
}
|
}
|
||||||
|
|
||||||
// SetNtfyURL sets the ntfy URL on a Service for testing.
|
// SetNtfyURL sets the ntfy URL on a Service for testing.
|
||||||
|
|||||||
+81
-64
@@ -12,6 +12,8 @@ import (
|
|||||||
"log/slog"
|
"log/slog"
|
||||||
"net/http"
|
"net/http"
|
||||||
"net/url"
|
"net/url"
|
||||||
|
"sync"
|
||||||
|
"sync/atomic"
|
||||||
"time"
|
"time"
|
||||||
|
|
||||||
"go.uber.org/fx"
|
"go.uber.org/fx"
|
||||||
@@ -115,19 +117,41 @@ type Service struct {
|
|||||||
history *AlertHistory
|
history *AlertHistory
|
||||||
retryConfig RetryConfig
|
retryConfig RetryConfig
|
||||||
sleepFn func(time.Duration) <-chan time.Time
|
sleepFn func(time.Duration) <-chan time.Time
|
||||||
|
|
||||||
|
// Shutdown draining state. drainMu guards draining and
|
||||||
|
// serialises it against the counter increment in
|
||||||
|
// startDelivery; inFlight tracks the delivery goroutines
|
||||||
|
// themselves and outstanding mirrors its count so a timed
|
||||||
|
// out drain can report how many were abandoned.
|
||||||
|
drainMu sync.Mutex
|
||||||
|
draining bool
|
||||||
|
inFlight sync.WaitGroup
|
||||||
|
outstanding atomic.Int64
|
||||||
|
abandon chan struct{}
|
||||||
|
abandonOnce sync.Once
|
||||||
|
}
|
||||||
|
|
||||||
|
// newService builds a Service with the fields every Service
|
||||||
|
// needs regardless of how it was constructed.
|
||||||
|
func newService(
|
||||||
|
log *slog.Logger,
|
||||||
|
transport http.RoundTripper,
|
||||||
|
) *Service {
|
||||||
|
return &Service{
|
||||||
|
log: log,
|
||||||
|
transport: transport,
|
||||||
|
history: NewAlertHistory(),
|
||||||
|
abandon: make(chan struct{}),
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
// New creates a new notify Service.
|
// New creates a new notify Service.
|
||||||
func New(
|
func New(
|
||||||
_ fx.Lifecycle,
|
lifecycle fx.Lifecycle,
|
||||||
params Params,
|
params Params,
|
||||||
) (*Service, error) {
|
) (*Service, error) {
|
||||||
svc := &Service{
|
svc := newService(params.Logger.Get(), http.DefaultTransport)
|
||||||
log: params.Logger.Get(),
|
svc.config = params.Config
|
||||||
transport: http.DefaultTransport,
|
|
||||||
config: params.Config,
|
|
||||||
history: NewAlertHistory(),
|
|
||||||
}
|
|
||||||
|
|
||||||
if params.Config.NtfyTopic != "" {
|
if params.Config.NtfyTopic != "" {
|
||||||
u, err := ValidateWebhookURL(
|
u, err := ValidateWebhookURL(
|
||||||
@@ -168,6 +192,14 @@ func New(
|
|||||||
svc.mattermostWebhookURL = u
|
svc.mattermostWebhookURL = u
|
||||||
}
|
}
|
||||||
|
|
||||||
|
lifecycle.Append(fx.Hook{
|
||||||
|
OnStop: func(ctx context.Context) error {
|
||||||
|
svc.drain(ctx)
|
||||||
|
|
||||||
|
return nil
|
||||||
|
},
|
||||||
|
})
|
||||||
|
|
||||||
return svc, nil
|
return svc, nil
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -194,6 +226,32 @@ func (svc *Service) SendNotification(
|
|||||||
svc.dispatchMattermost(ctx, title, message, priority)
|
svc.dispatchMattermost(ctx, title, message, priority)
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// dispatch delivers a notification to one endpoint on a
|
||||||
|
// tracked background goroutine.
|
||||||
|
//
|
||||||
|
// The delivery context is detached from ctx with
|
||||||
|
// context.WithoutCancel so that a cancelled caller does not
|
||||||
|
// kill a delivery already under way; the shutdown drain, not
|
||||||
|
// the caller, decides how long deliveries may keep running.
|
||||||
|
func (svc *Service) dispatch(
|
||||||
|
ctx context.Context,
|
||||||
|
endpoint string,
|
||||||
|
send func(context.Context) error,
|
||||||
|
) {
|
||||||
|
notifyCtx := context.WithoutCancel(ctx)
|
||||||
|
|
||||||
|
svc.startDelivery(endpoint, func() {
|
||||||
|
err := svc.deliverWithRetry(notifyCtx, endpoint, send)
|
||||||
|
if err != nil {
|
||||||
|
svc.log.Error(
|
||||||
|
"failed to send notification after retries",
|
||||||
|
"endpoint", endpoint,
|
||||||
|
"error", err,
|
||||||
|
)
|
||||||
|
}
|
||||||
|
})
|
||||||
|
}
|
||||||
|
|
||||||
func (svc *Service) dispatchNtfy(
|
func (svc *Service) dispatchNtfy(
|
||||||
ctx context.Context,
|
ctx context.Context,
|
||||||
title, message, priority string,
|
title, message, priority string,
|
||||||
@@ -202,26 +260,11 @@ func (svc *Service) dispatchNtfy(
|
|||||||
return
|
return
|
||||||
}
|
}
|
||||||
|
|
||||||
go func() {
|
svc.dispatch(ctx, "ntfy", func(c context.Context) error {
|
||||||
notifyCtx := context.WithoutCancel(ctx)
|
return svc.sendNtfy(
|
||||||
|
c, svc.ntfyURL, title, message, priority,
|
||||||
err := svc.deliverWithRetry(
|
|
||||||
notifyCtx, "ntfy",
|
|
||||||
func(c context.Context) error {
|
|
||||||
return svc.sendNtfy(
|
|
||||||
c, svc.ntfyURL,
|
|
||||||
title, message, priority,
|
|
||||||
)
|
|
||||||
},
|
|
||||||
)
|
)
|
||||||
if err != nil {
|
})
|
||||||
svc.log.Error(
|
|
||||||
"failed to send ntfy notification "+
|
|
||||||
"after retries",
|
|
||||||
"error", err,
|
|
||||||
)
|
|
||||||
}
|
|
||||||
}()
|
|
||||||
}
|
}
|
||||||
|
|
||||||
func (svc *Service) dispatchSlack(
|
func (svc *Service) dispatchSlack(
|
||||||
@@ -232,26 +275,11 @@ func (svc *Service) dispatchSlack(
|
|||||||
return
|
return
|
||||||
}
|
}
|
||||||
|
|
||||||
go func() {
|
svc.dispatch(ctx, "slack", func(c context.Context) error {
|
||||||
notifyCtx := context.WithoutCancel(ctx)
|
return svc.sendSlack(
|
||||||
|
c, svc.slackWebhookURL, title, message, priority,
|
||||||
err := svc.deliverWithRetry(
|
|
||||||
notifyCtx, "slack",
|
|
||||||
func(c context.Context) error {
|
|
||||||
return svc.sendSlack(
|
|
||||||
c, svc.slackWebhookURL,
|
|
||||||
title, message, priority,
|
|
||||||
)
|
|
||||||
},
|
|
||||||
)
|
)
|
||||||
if err != nil {
|
})
|
||||||
svc.log.Error(
|
|
||||||
"failed to send slack notification "+
|
|
||||||
"after retries",
|
|
||||||
"error", err,
|
|
||||||
)
|
|
||||||
}
|
|
||||||
}()
|
|
||||||
}
|
}
|
||||||
|
|
||||||
func (svc *Service) dispatchMattermost(
|
func (svc *Service) dispatchMattermost(
|
||||||
@@ -262,26 +290,15 @@ func (svc *Service) dispatchMattermost(
|
|||||||
return
|
return
|
||||||
}
|
}
|
||||||
|
|
||||||
go func() {
|
svc.dispatch(
|
||||||
notifyCtx := context.WithoutCancel(ctx)
|
ctx, "mattermost",
|
||||||
|
func(c context.Context) error {
|
||||||
err := svc.deliverWithRetry(
|
return svc.sendSlack(
|
||||||
notifyCtx, "mattermost",
|
c, svc.mattermostWebhookURL,
|
||||||
func(c context.Context) error {
|
title, message, priority,
|
||||||
return svc.sendSlack(
|
|
||||||
c, svc.mattermostWebhookURL,
|
|
||||||
title, message, priority,
|
|
||||||
)
|
|
||||||
},
|
|
||||||
)
|
|
||||||
if err != nil {
|
|
||||||
svc.log.Error(
|
|
||||||
"failed to send mattermost notification "+
|
|
||||||
"after retries",
|
|
||||||
"error", err,
|
|
||||||
)
|
)
|
||||||
}
|
},
|
||||||
}()
|
)
|
||||||
}
|
}
|
||||||
|
|
||||||
func (svc *Service) sendNtfy(
|
func (svc *Service) sendNtfy(
|
||||||
|
|||||||
@@ -2,6 +2,7 @@ package notify
|
|||||||
|
|
||||||
import (
|
import (
|
||||||
"context"
|
"context"
|
||||||
|
"fmt"
|
||||||
"math"
|
"math"
|
||||||
"math/rand/v2"
|
"math/rand/v2"
|
||||||
"time"
|
"time"
|
||||||
@@ -121,6 +122,14 @@ func (svc *Service) deliverWithRetry(
|
|||||||
select {
|
select {
|
||||||
case <-ctx.Done():
|
case <-ctx.Done():
|
||||||
return ctx.Err()
|
return ctx.Err()
|
||||||
|
case <-svc.abandon:
|
||||||
|
// Shutdown drained past its deadline; stop
|
||||||
|
// sleeping rather than outlive the process.
|
||||||
|
// A nil channel (Service built without a
|
||||||
|
// constructor) simply never fires.
|
||||||
|
return fmt.Errorf(
|
||||||
|
"%w: %s", ErrDeliveryAbandoned, endpoint,
|
||||||
|
)
|
||||||
case <-svc.sleepFunc(delay):
|
case <-svc.sleepFunc(delay):
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -0,0 +1,119 @@
|
|||||||
|
package notify
|
||||||
|
|
||||||
|
import (
|
||||||
|
"context"
|
||||||
|
"errors"
|
||||||
|
)
|
||||||
|
|
||||||
|
// ErrDeliveryAbandoned is returned by a retry loop that was
|
||||||
|
// cut short because shutdown drained past its deadline.
|
||||||
|
var ErrDeliveryAbandoned = errors.New(
|
||||||
|
"notification delivery abandoned at shutdown",
|
||||||
|
)
|
||||||
|
|
||||||
|
// startDelivery runs fn on its own goroutine while tracking it,
|
||||||
|
// so that drain can wait for it during shutdown.
|
||||||
|
//
|
||||||
|
// The WaitGroup counter is incremented here, on the caller's
|
||||||
|
// goroutine, before the worker exists: incrementing it inside
|
||||||
|
// the worker would race with drain's Wait and could let
|
||||||
|
// shutdown sail past a delivery that had not started yet.
|
||||||
|
//
|
||||||
|
// Once draining has begun the delivery is refused outright
|
||||||
|
// rather than queued, so a steady stream of newly submitted
|
||||||
|
// notifications cannot keep extending the drain.
|
||||||
|
func (svc *Service) startDelivery(endpoint string, fn func()) {
|
||||||
|
svc.drainMu.Lock()
|
||||||
|
|
||||||
|
if svc.draining {
|
||||||
|
svc.drainMu.Unlock()
|
||||||
|
|
||||||
|
svc.log.Warn(
|
||||||
|
"notification not dispatched: shutdown in progress",
|
||||||
|
"endpoint", endpoint,
|
||||||
|
)
|
||||||
|
|
||||||
|
return
|
||||||
|
}
|
||||||
|
|
||||||
|
svc.outstanding.Add(1)
|
||||||
|
|
||||||
|
// WaitGroup.Go increments the counter synchronously, here,
|
||||||
|
// and only then starts the goroutine.
|
||||||
|
svc.inFlight.Go(func() {
|
||||||
|
// Runs before the WaitGroup counter is decremented, so
|
||||||
|
// a drain that times out reports an accurate count.
|
||||||
|
defer svc.outstanding.Add(-1)
|
||||||
|
|
||||||
|
fn()
|
||||||
|
})
|
||||||
|
|
||||||
|
svc.drainMu.Unlock()
|
||||||
|
}
|
||||||
|
|
||||||
|
// drain waits for in-flight notification deliveries to finish.
|
||||||
|
//
|
||||||
|
// It first stops accepting new deliveries, then waits until
|
||||||
|
// either every outstanding delivery has completed or ctx
|
||||||
|
// expires — whichever comes first. ctx is the context fx
|
||||||
|
// passes to the OnStop hook, so a permanently dead webhook
|
||||||
|
// cannot hang shutdown indefinitely.
|
||||||
|
//
|
||||||
|
// When the deadline arrives with deliveries still outstanding,
|
||||||
|
// the count is logged at warn level and the abandon channel is
|
||||||
|
// closed, which releases any retry loop sleeping in backoff.
|
||||||
|
// Deliveries already inside an HTTP round trip are bounded by
|
||||||
|
// the existing httpClientTimeout instead.
|
||||||
|
//
|
||||||
|
// A ctx that is already expired on entry is not by itself cause
|
||||||
|
// for alarm: if nothing is outstanding there is nothing to
|
||||||
|
// abandon, and the drain says so at debug level rather than
|
||||||
|
// warning about deliveries that do not exist.
|
||||||
|
func (svc *Service) drain(ctx context.Context) {
|
||||||
|
svc.drainMu.Lock()
|
||||||
|
svc.draining = true
|
||||||
|
svc.drainMu.Unlock()
|
||||||
|
|
||||||
|
done := make(chan struct{})
|
||||||
|
|
||||||
|
go func() {
|
||||||
|
svc.inFlight.Wait()
|
||||||
|
close(done)
|
||||||
|
}()
|
||||||
|
|
||||||
|
select {
|
||||||
|
case <-done:
|
||||||
|
svc.log.Debug(
|
||||||
|
"all in-flight notifications completed",
|
||||||
|
)
|
||||||
|
case <-ctx.Done():
|
||||||
|
// outstanding is decremented before the WaitGroup
|
||||||
|
// counter, and startDelivery can no longer add to it
|
||||||
|
// now that draining is set, so a zero here means every
|
||||||
|
// delivery really did finish. ctx expiring in that
|
||||||
|
// state (an OnStop context that was already cancelled
|
||||||
|
// on entry is the usual way) abandons nothing, so it
|
||||||
|
// must not close abandon or warn about it.
|
||||||
|
abandoned := svc.outstanding.Load()
|
||||||
|
if abandoned == 0 {
|
||||||
|
svc.log.Debug(
|
||||||
|
"all in-flight notifications completed",
|
||||||
|
)
|
||||||
|
|
||||||
|
return
|
||||||
|
}
|
||||||
|
|
||||||
|
svc.abandonOnce.Do(func() {
|
||||||
|
if svc.abandon != nil {
|
||||||
|
close(svc.abandon)
|
||||||
|
}
|
||||||
|
})
|
||||||
|
|
||||||
|
svc.log.Warn(
|
||||||
|
"shutdown deadline reached with notifications "+
|
||||||
|
"still in flight; abandoning them",
|
||||||
|
"abandoned", abandoned,
|
||||||
|
"error", ctx.Err(),
|
||||||
|
)
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,531 @@
|
|||||||
|
package notify_test
|
||||||
|
|
||||||
|
import (
|
||||||
|
"bytes"
|
||||||
|
"context"
|
||||||
|
"log/slog"
|
||||||
|
"net/http"
|
||||||
|
"net/http/httptest"
|
||||||
|
"net/url"
|
||||||
|
"strings"
|
||||||
|
"sync"
|
||||||
|
"sync/atomic"
|
||||||
|
"testing"
|
||||||
|
"time"
|
||||||
|
|
||||||
|
"go.uber.org/fx"
|
||||||
|
|
||||||
|
"sneak.berlin/go/dnswatcher/internal/config"
|
||||||
|
"sneak.berlin/go/dnswatcher/internal/globals"
|
||||||
|
"sneak.berlin/go/dnswatcher/internal/logger"
|
||||||
|
"sneak.berlin/go/dnswatcher/internal/notify"
|
||||||
|
)
|
||||||
|
|
||||||
|
// Timings used by the drain tests. They stay in the same
|
||||||
|
// 10-100ms band as the retry tests so the suite never waits on
|
||||||
|
// a real backoff delay.
|
||||||
|
const (
|
||||||
|
// inFlightHold is how long a delivery is kept mid-request
|
||||||
|
// before the handler is released.
|
||||||
|
inFlightHold = 30 * time.Millisecond
|
||||||
|
|
||||||
|
// drainDeadline bounds a drain that is expected to time
|
||||||
|
// out.
|
||||||
|
drainDeadline = 50 * time.Millisecond
|
||||||
|
|
||||||
|
// drainSlack is the upper bound on how long a bounded
|
||||||
|
// drain may take; generous enough for a loaded CI box,
|
||||||
|
// still far below the 20s test ceiling.
|
||||||
|
drainSlack = 2 * time.Second
|
||||||
|
|
||||||
|
// settleDelay is how long to wait before asserting that
|
||||||
|
// something did *not* happen.
|
||||||
|
settleDelay = 50 * time.Millisecond
|
||||||
|
|
||||||
|
// idleDrainBound is the upper bound on a drain that has
|
||||||
|
// nothing in flight. It is deliberately far above the cost
|
||||||
|
// of the goroutine hop through inFlight.Wait() — which
|
||||||
|
// reached 57ms on a loaded box under -race with the package's
|
||||||
|
// parallel tests — and far below drainSlack, the deadline
|
||||||
|
// such a drain is given. A drain that blocked until its
|
||||||
|
// deadline instead of returning on the WaitGroup therefore
|
||||||
|
// still fails this bound, but scheduling delay alone cannot.
|
||||||
|
idleDrainBound = 500 * time.Millisecond
|
||||||
|
)
|
||||||
|
|
||||||
|
// syncBuffer is an io.Writer safe for concurrent use, so log
|
||||||
|
// output written from delivery goroutines can be inspected.
|
||||||
|
type syncBuffer struct {
|
||||||
|
mu sync.Mutex
|
||||||
|
buf bytes.Buffer
|
||||||
|
}
|
||||||
|
|
||||||
|
func (sb *syncBuffer) Write(p []byte) (int, error) {
|
||||||
|
sb.mu.Lock()
|
||||||
|
defer sb.mu.Unlock()
|
||||||
|
|
||||||
|
return sb.buf.Write(p) //nolint:wrapcheck // test helper
|
||||||
|
}
|
||||||
|
|
||||||
|
func (sb *syncBuffer) String() string {
|
||||||
|
sb.mu.Lock()
|
||||||
|
defer sb.mu.Unlock()
|
||||||
|
|
||||||
|
return sb.buf.String()
|
||||||
|
}
|
||||||
|
|
||||||
|
// newLoggingService returns a Service writing JSON logs into
|
||||||
|
// the returned buffer.
|
||||||
|
func newLoggingService(
|
||||||
|
transport http.RoundTripper,
|
||||||
|
) (*notify.Service, *syncBuffer) {
|
||||||
|
logs := &syncBuffer{}
|
||||||
|
handler := slog.NewJSONHandler(logs, nil)
|
||||||
|
|
||||||
|
return notify.NewTestServiceWithLogger(transport, handler),
|
||||||
|
logs
|
||||||
|
}
|
||||||
|
|
||||||
|
// blockingNtfyServer returns a server whose handler signals on
|
||||||
|
// entered, waits for release, and then responds 200.
|
||||||
|
func blockingNtfyServer(
|
||||||
|
entered chan<- struct{},
|
||||||
|
release <-chan struct{},
|
||||||
|
served *atomic.Bool,
|
||||||
|
) *httptest.Server {
|
||||||
|
var once sync.Once
|
||||||
|
|
||||||
|
return httptest.NewServer(
|
||||||
|
http.HandlerFunc(
|
||||||
|
func(w http.ResponseWriter, _ *http.Request) {
|
||||||
|
once.Do(func() { close(entered) })
|
||||||
|
<-release
|
||||||
|
|
||||||
|
served.Store(true)
|
||||||
|
|
||||||
|
w.WriteHeader(http.StatusOK)
|
||||||
|
}),
|
||||||
|
)
|
||||||
|
}
|
||||||
|
|
||||||
|
// TestDrainWaitsForInFlightDelivery verifies that a delivery
|
||||||
|
// already under way when shutdown starts is allowed to finish.
|
||||||
|
func TestDrainWaitsForInFlightDelivery(t *testing.T) {
|
||||||
|
t.Parallel()
|
||||||
|
|
||||||
|
var served atomic.Bool
|
||||||
|
|
||||||
|
entered := make(chan struct{})
|
||||||
|
release := make(chan struct{})
|
||||||
|
|
||||||
|
srv := blockingNtfyServer(entered, release, &served)
|
||||||
|
defer srv.Close()
|
||||||
|
|
||||||
|
topicURL, _ := url.Parse(srv.URL)
|
||||||
|
|
||||||
|
svc := notify.NewTestService(http.DefaultTransport)
|
||||||
|
svc.SetNtfyURL(topicURL)
|
||||||
|
|
||||||
|
svc.SendNotification(
|
||||||
|
context.Background(), "t", "m", prioInfo,
|
||||||
|
)
|
||||||
|
|
||||||
|
// Make sure the delivery really is mid-request before the
|
||||||
|
// drain begins.
|
||||||
|
select {
|
||||||
|
case <-entered:
|
||||||
|
case <-time.After(drainSlack):
|
||||||
|
t.Fatal("delivery never reached the endpoint")
|
||||||
|
}
|
||||||
|
|
||||||
|
// As in TestDrainBoundedByContextDeadline: start is captured
|
||||||
|
// before the clock it is compared against, here the timer
|
||||||
|
// holding the delivery open, so elapsed covers the whole hold
|
||||||
|
// and the lower bound cannot come out short from scheduling
|
||||||
|
// delay alone.
|
||||||
|
start := time.Now()
|
||||||
|
|
||||||
|
timer := time.AfterFunc(inFlightHold, func() {
|
||||||
|
close(release)
|
||||||
|
})
|
||||||
|
defer timer.Stop()
|
||||||
|
|
||||||
|
ctx, cancel := context.WithTimeout(
|
||||||
|
context.Background(), drainSlack,
|
||||||
|
)
|
||||||
|
defer cancel()
|
||||||
|
|
||||||
|
svc.Drain(ctx)
|
||||||
|
|
||||||
|
elapsed := time.Since(start)
|
||||||
|
|
||||||
|
if !served.Load() {
|
||||||
|
t.Error(
|
||||||
|
"drain returned before the in-flight delivery " +
|
||||||
|
"completed",
|
||||||
|
)
|
||||||
|
}
|
||||||
|
|
||||||
|
if elapsed < inFlightHold {
|
||||||
|
t.Errorf(
|
||||||
|
"drain took %v, want at least %v",
|
||||||
|
elapsed, inFlightHold,
|
||||||
|
)
|
||||||
|
}
|
||||||
|
|
||||||
|
if got := svc.OutstandingDeliveries(); got != 0 {
|
||||||
|
t.Errorf("outstanding deliveries = %d, want 0", got)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// neverFires returns a channel that never delivers, standing in
|
||||||
|
// for a long backoff sleep without actually sleeping.
|
||||||
|
func neverFires(_ time.Duration) <-chan time.Time {
|
||||||
|
return make(chan time.Time)
|
||||||
|
}
|
||||||
|
|
||||||
|
// TestDrainBoundedByContextDeadline verifies that a delivery
|
||||||
|
// stuck retrying against a dead endpoint does not hold shutdown
|
||||||
|
// past the OnStop context deadline, and that the abandoned
|
||||||
|
// deliveries are logged at warn level rather than dropped
|
||||||
|
// silently.
|
||||||
|
func TestDrainBoundedByContextDeadline(t *testing.T) {
|
||||||
|
t.Parallel()
|
||||||
|
|
||||||
|
var requests atomic.Int64
|
||||||
|
|
||||||
|
srv := httptest.NewServer(
|
||||||
|
http.HandlerFunc(
|
||||||
|
func(w http.ResponseWriter, _ *http.Request) {
|
||||||
|
requests.Add(1)
|
||||||
|
|
||||||
|
w.WriteHeader(http.StatusInternalServerError)
|
||||||
|
}),
|
||||||
|
)
|
||||||
|
defer srv.Close()
|
||||||
|
|
||||||
|
topicURL, _ := url.Parse(srv.URL)
|
||||||
|
|
||||||
|
svc, logs := newLoggingService(http.DefaultTransport)
|
||||||
|
svc.SetNtfyURL(topicURL)
|
||||||
|
// Never let the backoff sleep complete: the delivery is
|
||||||
|
// parked in its retry wait until shutdown releases it.
|
||||||
|
svc.SetSleepFunc(neverFires)
|
||||||
|
svc.SetRetryConfig(notify.RetryConfig{
|
||||||
|
MaxRetries: 5,
|
||||||
|
BaseDelay: time.Hour,
|
||||||
|
MaxDelay: time.Hour,
|
||||||
|
})
|
||||||
|
|
||||||
|
svc.SendNotification(
|
||||||
|
context.Background(), "t", "m", prioError,
|
||||||
|
)
|
||||||
|
|
||||||
|
waitForCondition(t, func() bool {
|
||||||
|
return requests.Load() >= 1 &&
|
||||||
|
svc.OutstandingDeliveries() == 1
|
||||||
|
})
|
||||||
|
|
||||||
|
// start must be captured *before* the deadline clock starts,
|
||||||
|
// so that the measured interval is a superset of the deadline
|
||||||
|
// interval. Capturing it after context.WithTimeout would
|
||||||
|
// make elapsed structurally smaller than drainDeadline and
|
||||||
|
// the lower bound below unfalsifiable-by-luck: it would fail
|
||||||
|
// whenever the two statements were separated by any
|
||||||
|
// scheduling delay, and pass otherwise, regardless of what
|
||||||
|
// the drain did.
|
||||||
|
start := time.Now()
|
||||||
|
|
||||||
|
ctx, cancel := context.WithTimeout(
|
||||||
|
context.Background(), drainDeadline,
|
||||||
|
)
|
||||||
|
defer cancel()
|
||||||
|
|
||||||
|
// The upper bound is enforced by a watchdog rather than by
|
||||||
|
// measuring after the fact: a drain that is not bounded at
|
||||||
|
// all never returns here (the delivery is parked in a backoff
|
||||||
|
// that never fires), so an unbounded drain must fail this
|
||||||
|
// test promptly instead of hanging the package until the test
|
||||||
|
// binary's 30s timeout.
|
||||||
|
returned := make(chan struct{})
|
||||||
|
|
||||||
|
go func() {
|
||||||
|
defer close(returned)
|
||||||
|
|
||||||
|
svc.Drain(ctx)
|
||||||
|
}()
|
||||||
|
|
||||||
|
select {
|
||||||
|
case <-returned:
|
||||||
|
case <-time.After(drainSlack):
|
||||||
|
t.Fatalf(
|
||||||
|
"drain did not return within %v; its %v deadline "+
|
||||||
|
"did not bound it",
|
||||||
|
drainSlack, drainDeadline,
|
||||||
|
)
|
||||||
|
}
|
||||||
|
|
||||||
|
// The lower bound is the real assertion: the drain must have
|
||||||
|
// waited for its whole deadline rather than giving up on the
|
||||||
|
// outstanding delivery early. With start captured above, an
|
||||||
|
// early return is the only thing that can make it fail.
|
||||||
|
if elapsed := time.Since(start); elapsed < drainDeadline {
|
||||||
|
t.Errorf(
|
||||||
|
"drain returned after %v, before its %v deadline",
|
||||||
|
elapsed, drainDeadline,
|
||||||
|
)
|
||||||
|
}
|
||||||
|
|
||||||
|
assertAbandonLogged(t, logs.String())
|
||||||
|
|
||||||
|
// The abandoned delivery must stop retrying rather than
|
||||||
|
// outlive the drain.
|
||||||
|
waitForCondition(t, func() bool {
|
||||||
|
return svc.OutstandingDeliveries() == 0
|
||||||
|
})
|
||||||
|
}
|
||||||
|
|
||||||
|
// assertAbandonLogged checks that the drain logged the
|
||||||
|
// abandoned deliveries at warn level with a count.
|
||||||
|
func assertAbandonLogged(t *testing.T, output string) {
|
||||||
|
t.Helper()
|
||||||
|
|
||||||
|
if !strings.Contains(output, `"level":"WARN"`) {
|
||||||
|
t.Errorf(
|
||||||
|
"abandoned deliveries not logged at warn level; "+
|
||||||
|
"log output: %s",
|
||||||
|
output,
|
||||||
|
)
|
||||||
|
}
|
||||||
|
|
||||||
|
if !strings.Contains(output, `"abandoned":1`) {
|
||||||
|
t.Errorf(
|
||||||
|
"abandoned delivery count not logged; "+
|
||||||
|
"log output: %s",
|
||||||
|
output,
|
||||||
|
)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// TestDrainRefusesNewDeliveries verifies that notifications
|
||||||
|
// submitted after the drain has begun are refused and logged,
|
||||||
|
// so a stream of new work cannot extend shutdown indefinitely.
|
||||||
|
func TestDrainRefusesNewDeliveries(t *testing.T) {
|
||||||
|
t.Parallel()
|
||||||
|
|
||||||
|
var requests atomic.Int64
|
||||||
|
|
||||||
|
srv := httptest.NewServer(
|
||||||
|
http.HandlerFunc(
|
||||||
|
func(w http.ResponseWriter, _ *http.Request) {
|
||||||
|
requests.Add(1)
|
||||||
|
|
||||||
|
w.WriteHeader(http.StatusOK)
|
||||||
|
}),
|
||||||
|
)
|
||||||
|
defer srv.Close()
|
||||||
|
|
||||||
|
target, _ := url.Parse(srv.URL)
|
||||||
|
|
||||||
|
svc, logs := newLoggingService(http.DefaultTransport)
|
||||||
|
svc.SetNtfyURL(target)
|
||||||
|
svc.SetSlackWebhookURL(target)
|
||||||
|
svc.SetMattermostWebhookURL(target)
|
||||||
|
|
||||||
|
ctx, cancel := context.WithTimeout(
|
||||||
|
context.Background(), drainSlack,
|
||||||
|
)
|
||||||
|
defer cancel()
|
||||||
|
|
||||||
|
// Nothing is in flight, so this returns immediately and
|
||||||
|
// leaves the service refusing further deliveries.
|
||||||
|
svc.Drain(ctx)
|
||||||
|
|
||||||
|
for range 3 {
|
||||||
|
svc.SendNotification(
|
||||||
|
context.Background(), "t", "m", prioInfo,
|
||||||
|
)
|
||||||
|
}
|
||||||
|
|
||||||
|
time.Sleep(settleDelay)
|
||||||
|
|
||||||
|
if got := requests.Load(); got != 0 {
|
||||||
|
t.Errorf(
|
||||||
|
"%d requests reached the endpoint after drain, "+
|
||||||
|
"want 0",
|
||||||
|
got,
|
||||||
|
)
|
||||||
|
}
|
||||||
|
|
||||||
|
if got := svc.OutstandingDeliveries(); got != 0 {
|
||||||
|
t.Errorf("outstanding deliveries = %d, want 0", got)
|
||||||
|
}
|
||||||
|
|
||||||
|
output := logs.String()
|
||||||
|
if !strings.Contains(output, "shutdown in progress") {
|
||||||
|
t.Errorf(
|
||||||
|
"refused deliveries not logged; log output: %s",
|
||||||
|
output,
|
||||||
|
)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// recordingLifecycle is a minimal fx.Lifecycle that records the
|
||||||
|
// hooks appended to it, so the wiring done by notify.New can be
|
||||||
|
// inspected without standing up a whole fx application.
|
||||||
|
type recordingLifecycle struct {
|
||||||
|
hooks []fx.Hook
|
||||||
|
}
|
||||||
|
|
||||||
|
func (l *recordingLifecycle) Append(hook fx.Hook) {
|
||||||
|
l.hooks = append(l.hooks, hook)
|
||||||
|
}
|
||||||
|
|
||||||
|
// newNotifyService builds a Service through the real
|
||||||
|
// constructor, wired to the given lifecycle.
|
||||||
|
func newNotifyService(
|
||||||
|
t *testing.T,
|
||||||
|
lifecycle fx.Lifecycle,
|
||||||
|
ntfyTopic string,
|
||||||
|
) *notify.Service {
|
||||||
|
t.Helper()
|
||||||
|
|
||||||
|
g, err := globals.New(nil)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("globals.New: %v", err)
|
||||||
|
}
|
||||||
|
|
||||||
|
log, err := logger.New(nil, logger.Params{Globals: g})
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("logger.New: %v", err)
|
||||||
|
}
|
||||||
|
|
||||||
|
svc, err := notify.New(lifecycle, notify.Params{
|
||||||
|
Logger: log,
|
||||||
|
Config: &config.Config{NtfyTopic: ntfyTopic},
|
||||||
|
})
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("notify.New: %v", err)
|
||||||
|
}
|
||||||
|
|
||||||
|
return svc
|
||||||
|
}
|
||||||
|
|
||||||
|
// TestNewRegistersDrainingStopHook verifies that notify.New
|
||||||
|
// wires an OnStop hook into the fx lifecycle and that the hook
|
||||||
|
// waits for in-flight deliveries.
|
||||||
|
func TestNewRegistersDrainingStopHook(t *testing.T) {
|
||||||
|
t.Parallel()
|
||||||
|
|
||||||
|
var served atomic.Bool
|
||||||
|
|
||||||
|
entered := make(chan struct{})
|
||||||
|
release := make(chan struct{})
|
||||||
|
|
||||||
|
srv := blockingNtfyServer(entered, release, &served)
|
||||||
|
defer srv.Close()
|
||||||
|
|
||||||
|
lifecycle := &recordingLifecycle{}
|
||||||
|
svc := newNotifyService(t, lifecycle, srv.URL)
|
||||||
|
|
||||||
|
if len(lifecycle.hooks) != 1 {
|
||||||
|
t.Fatalf(
|
||||||
|
"appended %d lifecycle hooks, want 1",
|
||||||
|
len(lifecycle.hooks),
|
||||||
|
)
|
||||||
|
}
|
||||||
|
|
||||||
|
stop := lifecycle.hooks[0].OnStop
|
||||||
|
if stop == nil {
|
||||||
|
t.Fatal("lifecycle hook has no OnStop function")
|
||||||
|
}
|
||||||
|
|
||||||
|
svc.SendNotification(
|
||||||
|
context.Background(), "t", "m", prioInfo,
|
||||||
|
)
|
||||||
|
|
||||||
|
select {
|
||||||
|
case <-entered:
|
||||||
|
case <-time.After(drainSlack):
|
||||||
|
t.Fatal("delivery never reached the endpoint")
|
||||||
|
}
|
||||||
|
|
||||||
|
timer := time.AfterFunc(inFlightHold, func() {
|
||||||
|
close(release)
|
||||||
|
})
|
||||||
|
defer timer.Stop()
|
||||||
|
|
||||||
|
ctx, cancel := context.WithTimeout(
|
||||||
|
context.Background(), drainSlack,
|
||||||
|
)
|
||||||
|
defer cancel()
|
||||||
|
|
||||||
|
err := stop(ctx)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("OnStop returned error: %v", err)
|
||||||
|
}
|
||||||
|
|
||||||
|
if !served.Load() {
|
||||||
|
t.Error(
|
||||||
|
"OnStop returned before the in-flight delivery " +
|
||||||
|
"completed",
|
||||||
|
)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// TestDrainWithoutDeliveriesReturnsImmediately verifies the
|
||||||
|
// common case: nothing in flight, shutdown is not delayed.
|
||||||
|
func TestDrainWithoutDeliveriesReturnsImmediately(t *testing.T) {
|
||||||
|
t.Parallel()
|
||||||
|
|
||||||
|
svc := notify.NewTestService(http.DefaultTransport)
|
||||||
|
|
||||||
|
// Captured before the deadline clock, as elsewhere in this
|
||||||
|
// file; for an upper bound that is the conservative
|
||||||
|
// direction, since the measured interval can then only be
|
||||||
|
// longer than the drain itself.
|
||||||
|
start := time.Now()
|
||||||
|
|
||||||
|
ctx, cancel := context.WithTimeout(
|
||||||
|
context.Background(), drainSlack,
|
||||||
|
)
|
||||||
|
defer cancel()
|
||||||
|
|
||||||
|
svc.Drain(ctx)
|
||||||
|
|
||||||
|
if elapsed := time.Since(start); elapsed > idleDrainBound {
|
||||||
|
t.Errorf(
|
||||||
|
"drain of an idle service took %v, want well "+
|
||||||
|
"under its %v deadline",
|
||||||
|
elapsed, drainSlack,
|
||||||
|
)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// TestDrainWithCancelledContextDoesNotWarn verifies that an
|
||||||
|
// OnStop context that is already dead on entry does not produce
|
||||||
|
// an "abandoning them" warning when there was nothing in flight
|
||||||
|
// to abandon. The expired context wins the select immediately,
|
||||||
|
// so only the outstanding count can tell the difference between
|
||||||
|
// a genuine timeout and a shutdown that had simply already run
|
||||||
|
// out of time with no work left.
|
||||||
|
func TestDrainWithCancelledContextDoesNotWarn(t *testing.T) {
|
||||||
|
t.Parallel()
|
||||||
|
|
||||||
|
svc, logs := newLoggingService(http.DefaultTransport)
|
||||||
|
|
||||||
|
ctx, cancel := context.WithCancel(context.Background())
|
||||||
|
cancel()
|
||||||
|
|
||||||
|
svc.Drain(ctx)
|
||||||
|
|
||||||
|
if output := logs.String(); strings.Contains(
|
||||||
|
output, `"level":"WARN"`,
|
||||||
|
) {
|
||||||
|
t.Errorf(
|
||||||
|
"drain with nothing in flight warned about "+
|
||||||
|
"abandoned deliveries; log output: %s",
|
||||||
|
output,
|
||||||
|
)
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,284 @@
|
|||||||
|
package resolver_test
|
||||||
|
|
||||||
|
import (
|
||||||
|
"context"
|
||||||
|
"sync"
|
||||||
|
"testing"
|
||||||
|
"time"
|
||||||
|
|
||||||
|
"github.com/stretchr/testify/assert"
|
||||||
|
|
||||||
|
"sneak.berlin/go/dnswatcher/internal/resolver"
|
||||||
|
)
|
||||||
|
|
||||||
|
// Tests for the live-DNS harness in livedns_test.go itself. These
|
||||||
|
// exercise pure logic and the retry/concurrency plumbing; they
|
||||||
|
// perform no DNS resolution of any kind, so they neither mock DNS
|
||||||
|
// nor depend on it.
|
||||||
|
|
||||||
|
// Names for the synthetic status maps below. Nothing is ever queried
|
||||||
|
// at them: they are map keys handed to the package's pure counting
|
||||||
|
// helpers, not a stand-in for a nameserver.
|
||||||
|
const (
|
||||||
|
nsExample1 = "ns1.example."
|
||||||
|
nsExample2 = "ns2.example."
|
||||||
|
nsExample3 = "ns3.example."
|
||||||
|
nsExample4 = "ns4.example."
|
||||||
|
)
|
||||||
|
|
||||||
|
func TestLiveQuorumIsStrictMajority(t *testing.T) {
|
||||||
|
t.Parallel()
|
||||||
|
|
||||||
|
cases := map[int]int{
|
||||||
|
0: 1,
|
||||||
|
1: 1,
|
||||||
|
2: 2,
|
||||||
|
3: 2,
|
||||||
|
4: 3,
|
||||||
|
5: 3,
|
||||||
|
13: 7,
|
||||||
|
}
|
||||||
|
|
||||||
|
for total, want := range cases {
|
||||||
|
assert.Equal(
|
||||||
|
t, want, liveQuorum(total),
|
||||||
|
"liveQuorum(%d)", total,
|
||||||
|
)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestStatusCountingIgnoresSilentNameservers(t *testing.T) {
|
||||||
|
t.Parallel()
|
||||||
|
|
||||||
|
results := map[string]*resolver.NameserverResponse{
|
||||||
|
nsExample1: {
|
||||||
|
Nameserver: nsExample1,
|
||||||
|
Status: resolver.StatusOK,
|
||||||
|
},
|
||||||
|
nsExample2: {
|
||||||
|
Nameserver: nsExample2,
|
||||||
|
Status: resolver.StatusOK,
|
||||||
|
},
|
||||||
|
nsExample3: {
|
||||||
|
Nameserver: nsExample3,
|
||||||
|
Status: resolver.StatusTimeout,
|
||||||
|
},
|
||||||
|
nsExample4: {
|
||||||
|
Nameserver: nsExample4,
|
||||||
|
Status: resolver.StatusError,
|
||||||
|
},
|
||||||
|
}
|
||||||
|
|
||||||
|
assert.Equal(
|
||||||
|
t, 2, countStatus(results, resolver.StatusOK),
|
||||||
|
)
|
||||||
|
assert.Equal(
|
||||||
|
t, 0, countStatus(results, resolver.StatusNXDomain),
|
||||||
|
)
|
||||||
|
|
||||||
|
// Two of four answered, which is short of the quorum of
|
||||||
|
// three: this is the state that triggers a retry rather
|
||||||
|
// than an assertion failure.
|
||||||
|
assert.Equal(t, 2, answeredCount(results))
|
||||||
|
assert.Less(t, answeredCount(results), liveQuorum(len(results)))
|
||||||
|
|
||||||
|
assert.Equal(
|
||||||
|
t,
|
||||||
|
"ns1.example.=ok ns2.example.=ok "+
|
||||||
|
"ns3.example.=timeout ns4.example.=error",
|
||||||
|
describeStatuses(results),
|
||||||
|
)
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestRetryLiveRecoversFromTransientFailure(t *testing.T) {
|
||||||
|
t.Parallel()
|
||||||
|
|
||||||
|
const wantAttempts = 2
|
||||||
|
|
||||||
|
attempts := 0
|
||||||
|
|
||||||
|
retryLive(t, "transient", func(_ context.Context) error {
|
||||||
|
attempts++
|
||||||
|
|
||||||
|
if attempts < wantAttempts {
|
||||||
|
return errLiveNoAnswer
|
||||||
|
}
|
||||||
|
|
||||||
|
return nil
|
||||||
|
})
|
||||||
|
|
||||||
|
assert.Equal(t, wantAttempts, attempts)
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestRetryLiveGivesEachAttemptADeadline(t *testing.T) {
|
||||||
|
t.Parallel()
|
||||||
|
|
||||||
|
retryLive(t, "deadline", func(ctx context.Context) error {
|
||||||
|
deadline, ok := ctx.Deadline()
|
||||||
|
assert.True(t, ok, "attempt should carry a deadline")
|
||||||
|
|
||||||
|
remaining := time.Until(deadline)
|
||||||
|
|
||||||
|
assert.LessOrEqual(t, remaining, liveAttemptTimeout)
|
||||||
|
|
||||||
|
// Lower bound too: without one this passes for a
|
||||||
|
// deadline far shorter than intended, which would
|
||||||
|
// silently turn every live attempt into an instant
|
||||||
|
// timeout.
|
||||||
|
assert.Greater(t, remaining, liveAttemptTimeout/2)
|
||||||
|
|
||||||
|
return nil
|
||||||
|
})
|
||||||
|
}
|
||||||
|
|
||||||
|
// TestUnsanctionedStatusesRejectsWrongAnswers is the regression test
|
||||||
|
// for the defect this allowlist exists to prevent: a minority of
|
||||||
|
// nameservers answering WRONGLY while quorum keeps the suite green.
|
||||||
|
// nodata is the case that motivated it — it is a wrong answer, not
|
||||||
|
// silence, and it was previously banned by neither test.
|
||||||
|
func TestUnsanctionedStatusesRejectsWrongAnswers(t *testing.T) {
|
||||||
|
t.Parallel()
|
||||||
|
|
||||||
|
// Four nameservers, three OK and one answering nodata: a
|
||||||
|
// quorum of three is satisfied and no NXDOMAIN is present, so
|
||||||
|
// the old blocklist assertions both passed on this input.
|
||||||
|
results := map[string]*resolver.NameserverResponse{
|
||||||
|
nsExample1: {
|
||||||
|
Nameserver: nsExample1,
|
||||||
|
Status: resolver.StatusOK,
|
||||||
|
},
|
||||||
|
nsExample2: {
|
||||||
|
Nameserver: nsExample2,
|
||||||
|
Status: resolver.StatusOK,
|
||||||
|
},
|
||||||
|
nsExample3: {
|
||||||
|
Nameserver: nsExample3,
|
||||||
|
Status: resolver.StatusOK,
|
||||||
|
},
|
||||||
|
nsExample4: {
|
||||||
|
Nameserver: nsExample4,
|
||||||
|
Status: resolver.StatusNoData,
|
||||||
|
},
|
||||||
|
}
|
||||||
|
|
||||||
|
assert.GreaterOrEqual(
|
||||||
|
t,
|
||||||
|
countStatus(results, resolver.StatusOK),
|
||||||
|
liveQuorum(len(results)),
|
||||||
|
)
|
||||||
|
assert.Zero(t, countStatus(results, resolver.StatusNXDomain))
|
||||||
|
|
||||||
|
// nodata is an ANSWER, so it never triggers a retry: nothing
|
||||||
|
// but the allowlist stands between it and a false green.
|
||||||
|
assert.Equal(t, len(results), answeredCount(results))
|
||||||
|
|
||||||
|
assert.Equal(
|
||||||
|
t,
|
||||||
|
[]string{nsExample4 + "=nodata"},
|
||||||
|
unsanctionedStatuses(
|
||||||
|
results,
|
||||||
|
resolver.StatusOK,
|
||||||
|
resolver.StatusTimeout,
|
||||||
|
resolver.StatusError,
|
||||||
|
),
|
||||||
|
"nodata must be reported as an unsanctioned status",
|
||||||
|
)
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestUnsanctionedStatusesToleratesSilenceOnly(t *testing.T) {
|
||||||
|
t.Parallel()
|
||||||
|
|
||||||
|
results := map[string]*resolver.NameserverResponse{
|
||||||
|
nsExample1: {
|
||||||
|
Nameserver: nsExample1,
|
||||||
|
Status: resolver.StatusNXDomain,
|
||||||
|
},
|
||||||
|
nsExample2: {
|
||||||
|
Nameserver: nsExample2,
|
||||||
|
Status: resolver.StatusTimeout,
|
||||||
|
},
|
||||||
|
nsExample3: {
|
||||||
|
Nameserver: nsExample3,
|
||||||
|
Status: resolver.StatusError,
|
||||||
|
},
|
||||||
|
}
|
||||||
|
|
||||||
|
allowed := []string{
|
||||||
|
resolver.StatusNXDomain,
|
||||||
|
resolver.StatusTimeout,
|
||||||
|
resolver.StatusError,
|
||||||
|
}
|
||||||
|
|
||||||
|
assert.Empty(
|
||||||
|
t,
|
||||||
|
unsanctionedStatuses(results, allowed...),
|
||||||
|
"timeout and error are non-answers and are tolerated",
|
||||||
|
)
|
||||||
|
|
||||||
|
// The same silent nameservers do not count towards a quorum.
|
||||||
|
assert.Equal(t, 1, answeredCount(results))
|
||||||
|
|
||||||
|
// An unknown status is treated as silence by answeredCount —
|
||||||
|
// so it retries and fails loudly — and is unsanctioned by the
|
||||||
|
// allowlist rather than quietly permitted.
|
||||||
|
const laterStatus = "some-status-added-later"
|
||||||
|
|
||||||
|
results[nsExample4] = &resolver.NameserverResponse{
|
||||||
|
Nameserver: nsExample4,
|
||||||
|
Status: laterStatus,
|
||||||
|
}
|
||||||
|
|
||||||
|
assert.Equal(t, 1, answeredCount(results))
|
||||||
|
assert.Equal(
|
||||||
|
t,
|
||||||
|
[]string{nsExample4 + "=" + laterStatus},
|
||||||
|
unsanctionedStatuses(results, allowed...),
|
||||||
|
)
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestRunLiveBoundsConcurrency(t *testing.T) {
|
||||||
|
t.Parallel()
|
||||||
|
|
||||||
|
const workers = 24
|
||||||
|
|
||||||
|
var (
|
||||||
|
mu sync.Mutex
|
||||||
|
wg sync.WaitGroup
|
||||||
|
inFlight int
|
||||||
|
maxSeen int
|
||||||
|
)
|
||||||
|
|
||||||
|
wg.Add(workers)
|
||||||
|
|
||||||
|
for range workers {
|
||||||
|
go func() {
|
||||||
|
defer wg.Done()
|
||||||
|
|
||||||
|
_ = runLive(func(_ context.Context) error {
|
||||||
|
mu.Lock()
|
||||||
|
inFlight++
|
||||||
|
|
||||||
|
if inFlight > maxSeen {
|
||||||
|
maxSeen = inFlight
|
||||||
|
}
|
||||||
|
mu.Unlock()
|
||||||
|
|
||||||
|
time.Sleep(time.Millisecond)
|
||||||
|
|
||||||
|
mu.Lock()
|
||||||
|
inFlight--
|
||||||
|
mu.Unlock()
|
||||||
|
|
||||||
|
return nil
|
||||||
|
})
|
||||||
|
}()
|
||||||
|
}
|
||||||
|
|
||||||
|
wg.Wait()
|
||||||
|
|
||||||
|
assert.Positive(t, maxSeen)
|
||||||
|
assert.LessOrEqual(
|
||||||
|
t, maxSeen, liveConcurrency,
|
||||||
|
"live queries must stay under the package-wide gate",
|
||||||
|
)
|
||||||
|
}
|
||||||
@@ -0,0 +1,495 @@
|
|||||||
|
package resolver_test
|
||||||
|
|
||||||
|
import (
|
||||||
|
"context"
|
||||||
|
"errors"
|
||||||
|
"fmt"
|
||||||
|
"slices"
|
||||||
|
"sort"
|
||||||
|
"strings"
|
||||||
|
"testing"
|
||||||
|
"time"
|
||||||
|
|
||||||
|
"sneak.berlin/go/dnswatcher/internal/resolver"
|
||||||
|
)
|
||||||
|
|
||||||
|
// ----------------------------------------------------------------
|
||||||
|
// Live DNS test support
|
||||||
|
// ----------------------------------------------------------------
|
||||||
|
//
|
||||||
|
// Every test in this package resolves against the real, live DNS —
|
||||||
|
// see TESTING.md. Nothing here mocks, fakes, stubs, records or
|
||||||
|
// replays DNS, and nothing here skips or gates a test: the helpers
|
||||||
|
// below only change *how* the live queries are issued, so that a
|
||||||
|
// single dropped UDP packet or one slow authoritative server does
|
||||||
|
// not turn a correct resolver into a red build.
|
||||||
|
//
|
||||||
|
// Three mechanisms, all test-side:
|
||||||
|
//
|
||||||
|
// 1. Bounded concurrency. The package's tests are parallel and the
|
||||||
|
// build hosts have many cores, so without a limit every test
|
||||||
|
// starts its own iterative resolution at the same instant and
|
||||||
|
// they all hit the first root server in rootServerList() within
|
||||||
|
// a few milliseconds of each other. Root servers rate-limit
|
||||||
|
// that, which shows up as a different arbitrary subset of tests
|
||||||
|
// failing on each run. liveGate caps how many resolutions are
|
||||||
|
// in flight at once.
|
||||||
|
//
|
||||||
|
// 2. Retry with exponential backoff. Each live operation gets
|
||||||
|
// several attempts with its own timeout. The retry predicate is
|
||||||
|
// strictly transport-level — "did a nameserver answer at all" —
|
||||||
|
// never the assertion the test is making. A resolver that
|
||||||
|
// answers incorrectly still fails on the first attempt.
|
||||||
|
//
|
||||||
|
// 3. Quorum. Where an assertion spans several independent
|
||||||
|
// nameservers, a strict majority answering as expected is
|
||||||
|
// enough; a server that fails to answer is tolerated, while a
|
||||||
|
// server that answers *wrongly* still fails the test.
|
||||||
|
//
|
||||||
|
// The tolerance in (3) is expressed as an ALLOWLIST of sanctioned
|
||||||
|
// statuses, never as a blocklist of known-bad ones. A blocklist bans
|
||||||
|
// the one wrong answer its author thought of and silently admits
|
||||||
|
// every other status, including any added to the resolver later; an
|
||||||
|
// allowlist fails on anything nobody explicitly sanctioned. Silence
|
||||||
|
// (timeout, error) is the only thing quorum exists to tolerate. A
|
||||||
|
// *wrong answer* — nxdomain for a name that exists, ok for one that
|
||||||
|
// does not, nodata for either — is never tolerated at any count.
|
||||||
|
|
||||||
|
const (
|
||||||
|
// liveAttempts is how many times a live DNS operation is
|
||||||
|
// attempted before the test fails.
|
||||||
|
liveAttempts = 3
|
||||||
|
|
||||||
|
// liveAttemptTimeout bounds one attempt. Worst case for an
|
||||||
|
// operation is liveAttempts * liveAttemptTimeout plus the
|
||||||
|
// backoff — about 26 seconds, well inside the 90-second
|
||||||
|
// `go test -timeout` backstop even when several operations
|
||||||
|
// exhaust their attempts.
|
||||||
|
liveAttemptTimeout = 8 * time.Second
|
||||||
|
|
||||||
|
// liveBackoffBase is the delay after the first failed
|
||||||
|
// attempt; it is multiplied by liveBackoffFactor each time.
|
||||||
|
liveBackoffBase = 500 * time.Millisecond
|
||||||
|
|
||||||
|
// liveBackoffFactor is the exponential backoff multiplier.
|
||||||
|
liveBackoffFactor = 2
|
||||||
|
|
||||||
|
// liveConcurrency caps how many live resolutions may be in
|
||||||
|
// flight across the whole package at once.
|
||||||
|
liveConcurrency = 6
|
||||||
|
|
||||||
|
// minNameservers is the smallest nameserver count a
|
||||||
|
// well-run zone is expected to publish.
|
||||||
|
minNameservers = 2
|
||||||
|
)
|
||||||
|
|
||||||
|
// liveGate bounds concurrent live resolutions package-wide. It has
|
||||||
|
// to be package scoped: the whole point is that it is shared by
|
||||||
|
// every parallel test in the package.
|
||||||
|
//
|
||||||
|
//nolint:gochecknoglobals // package-wide live query rate limit
|
||||||
|
var liveGate = make(chan struct{}, liveConcurrency)
|
||||||
|
|
||||||
|
var (
|
||||||
|
// errLiveNoAnswer reports that a live operation produced no
|
||||||
|
// usable answer, which is retried rather than asserted on.
|
||||||
|
errLiveNoAnswer = errors.New("no answer from live DNS")
|
||||||
|
|
||||||
|
// errLiveNoQuorum reports that too few of a domain's
|
||||||
|
// nameservers answered for a quorum assertion to be made.
|
||||||
|
errLiveNoQuorum = errors.New("no nameserver quorum")
|
||||||
|
)
|
||||||
|
|
||||||
|
// runLive executes one attempt of a live operation, holding a slot
|
||||||
|
// in liveGate for its duration and bounding it with its own
|
||||||
|
// timeout.
|
||||||
|
func runLive(op func(ctx context.Context) error) error {
|
||||||
|
liveGate <- struct{}{}
|
||||||
|
defer func() { <-liveGate }()
|
||||||
|
|
||||||
|
ctx, cancel := context.WithTimeout(
|
||||||
|
context.Background(), liveAttemptTimeout,
|
||||||
|
)
|
||||||
|
defer cancel()
|
||||||
|
|
||||||
|
return op(ctx)
|
||||||
|
}
|
||||||
|
|
||||||
|
// retryLive runs op until it reports success, retrying transport
|
||||||
|
// failures with exponential backoff, and fails the test if every
|
||||||
|
// attempt fails. op returns an error only for a failure to obtain
|
||||||
|
// an answer — never for an answer the test disagrees with, which
|
||||||
|
// belongs in an assertion so that it fails immediately. op stores
|
||||||
|
// whatever it obtained where its caller can find it.
|
||||||
|
func retryLive(
|
||||||
|
t *testing.T,
|
||||||
|
what string,
|
||||||
|
op func(ctx context.Context) error,
|
||||||
|
) {
|
||||||
|
t.Helper()
|
||||||
|
|
||||||
|
var last error
|
||||||
|
|
||||||
|
backoff := liveBackoffBase
|
||||||
|
|
||||||
|
for attempt := range liveAttempts {
|
||||||
|
if attempt > 0 {
|
||||||
|
t.Logf(
|
||||||
|
"%s: attempt %d of %d failed (%v), "+
|
||||||
|
"retrying in %s",
|
||||||
|
what, attempt, liveAttempts, last, backoff,
|
||||||
|
)
|
||||||
|
time.Sleep(backoff)
|
||||||
|
|
||||||
|
backoff *= liveBackoffFactor
|
||||||
|
}
|
||||||
|
|
||||||
|
last = runLive(op)
|
||||||
|
if last == nil {
|
||||||
|
return
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
t.Fatalf(
|
||||||
|
"%s: no answer after %d live attempts: %v",
|
||||||
|
what, liveAttempts, last,
|
||||||
|
)
|
||||||
|
}
|
||||||
|
|
||||||
|
// liveQuorum is how many of total nameservers must agree for a
|
||||||
|
// multi-nameserver assertion to hold: a strict majority.
|
||||||
|
func liveQuorum(total int) int {
|
||||||
|
if total < 1 {
|
||||||
|
return 1
|
||||||
|
}
|
||||||
|
|
||||||
|
return total/2 + 1
|
||||||
|
}
|
||||||
|
|
||||||
|
// countStatus counts the responses carrying the given status.
|
||||||
|
func countStatus(
|
||||||
|
results map[string]*resolver.NameserverResponse,
|
||||||
|
status string,
|
||||||
|
) int {
|
||||||
|
n := 0
|
||||||
|
|
||||||
|
for _, resp := range results {
|
||||||
|
if resp.Status == status {
|
||||||
|
n++
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
return n
|
||||||
|
}
|
||||||
|
|
||||||
|
// liveAnswerStatuses is the closed set of statuses that count as a
|
||||||
|
// nameserver having ANSWERED at all, whether or not the test agrees
|
||||||
|
// with the answer. It is deliberately an allowlist: a status added
|
||||||
|
// to the resolver later is treated as silence, so it can only ever
|
||||||
|
// cause a retry and then a loud failure, never a quiet pass.
|
||||||
|
func liveAnswerStatuses() []string {
|
||||||
|
return []string{
|
||||||
|
resolver.StatusOK,
|
||||||
|
resolver.StatusNXDomain,
|
||||||
|
resolver.StatusNoData,
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// answeredCount counts the nameservers that produced an answer of
|
||||||
|
// any kind, as opposed to failing or timing out.
|
||||||
|
func answeredCount(
|
||||||
|
results map[string]*resolver.NameserverResponse,
|
||||||
|
) int {
|
||||||
|
answers := liveAnswerStatuses()
|
||||||
|
|
||||||
|
n := 0
|
||||||
|
|
||||||
|
for _, resp := range results {
|
||||||
|
if slices.Contains(answers, resp.Status) {
|
||||||
|
n++
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
return n
|
||||||
|
}
|
||||||
|
|
||||||
|
// unsanctionedStatuses returns "nameserver=status" for every result
|
||||||
|
// whose status the caller did not explicitly sanction, sorted for a
|
||||||
|
// stable failure message. Callers pass the full closed set they will
|
||||||
|
// accept — the expected answer plus whichever non-answers (timeout,
|
||||||
|
// error) quorum is allowed to tolerate — so that any status outside
|
||||||
|
// it fails the test by name.
|
||||||
|
func unsanctionedStatuses(
|
||||||
|
results map[string]*resolver.NameserverResponse,
|
||||||
|
allowed ...string,
|
||||||
|
) []string {
|
||||||
|
offenders := make([]string, 0, len(results))
|
||||||
|
|
||||||
|
for ns, resp := range results {
|
||||||
|
if slices.Contains(allowed, resp.Status) {
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
|
||||||
|
offenders = append(
|
||||||
|
offenders, fmt.Sprintf("%s=%s", ns, resp.Status),
|
||||||
|
)
|
||||||
|
}
|
||||||
|
|
||||||
|
sort.Strings(offenders)
|
||||||
|
|
||||||
|
return offenders
|
||||||
|
}
|
||||||
|
|
||||||
|
// describeStatuses renders per-nameserver statuses for use in
|
||||||
|
// assertion failure messages.
|
||||||
|
func describeStatuses(
|
||||||
|
results map[string]*resolver.NameserverResponse,
|
||||||
|
) string {
|
||||||
|
parts := make([]string, 0, len(results))
|
||||||
|
for ns, resp := range results {
|
||||||
|
parts = append(
|
||||||
|
parts, fmt.Sprintf("%s=%s", ns, resp.Status),
|
||||||
|
)
|
||||||
|
}
|
||||||
|
|
||||||
|
sort.Strings(parts)
|
||||||
|
|
||||||
|
return strings.Join(parts, " ")
|
||||||
|
}
|
||||||
|
|
||||||
|
// ----------------------------------------------------------------
|
||||||
|
// Live operation wrappers
|
||||||
|
// ----------------------------------------------------------------
|
||||||
|
|
||||||
|
// liveFindAuthoritative resolves a domain's authoritative
|
||||||
|
// nameservers, retrying until the delegation chain can be walked.
|
||||||
|
func liveFindAuthoritative(
|
||||||
|
t *testing.T,
|
||||||
|
r *resolver.Resolver,
|
||||||
|
domain string,
|
||||||
|
) []string {
|
||||||
|
t.Helper()
|
||||||
|
|
||||||
|
var out []string
|
||||||
|
|
||||||
|
retryLive(
|
||||||
|
t,
|
||||||
|
"FindAuthoritativeNameservers("+domain+")",
|
||||||
|
func(ctx context.Context) error {
|
||||||
|
ns, err := r.FindAuthoritativeNameservers(ctx, domain)
|
||||||
|
if err != nil {
|
||||||
|
return err
|
||||||
|
}
|
||||||
|
|
||||||
|
if len(ns) == 0 {
|
||||||
|
return fmt.Errorf(
|
||||||
|
"%w: %s has no nameservers",
|
||||||
|
errLiveNoAnswer, domain,
|
||||||
|
)
|
||||||
|
}
|
||||||
|
|
||||||
|
out = ns
|
||||||
|
|
||||||
|
return nil
|
||||||
|
},
|
||||||
|
)
|
||||||
|
|
||||||
|
return out
|
||||||
|
}
|
||||||
|
|
||||||
|
// liveLookupNS is liveFindAuthoritative through the LookupNS entry
|
||||||
|
// point, so that both entry points stay independently exercised.
|
||||||
|
func liveLookupNS(
|
||||||
|
t *testing.T,
|
||||||
|
r *resolver.Resolver,
|
||||||
|
domain string,
|
||||||
|
) []string {
|
||||||
|
t.Helper()
|
||||||
|
|
||||||
|
var out []string
|
||||||
|
|
||||||
|
retryLive(
|
||||||
|
t,
|
||||||
|
"LookupNS("+domain+")",
|
||||||
|
func(ctx context.Context) error {
|
||||||
|
ns, err := r.LookupNS(ctx, domain)
|
||||||
|
if err != nil {
|
||||||
|
return err
|
||||||
|
}
|
||||||
|
|
||||||
|
if len(ns) == 0 {
|
||||||
|
return fmt.Errorf(
|
||||||
|
"%w: %s has no nameservers",
|
||||||
|
errLiveNoAnswer, domain,
|
||||||
|
)
|
||||||
|
}
|
||||||
|
|
||||||
|
out = ns
|
||||||
|
|
||||||
|
return nil
|
||||||
|
},
|
||||||
|
)
|
||||||
|
|
||||||
|
return out
|
||||||
|
}
|
||||||
|
|
||||||
|
// liveQueryNameserver queries one nameserver, retrying while that
|
||||||
|
// nameserver fails to answer. NXDOMAIN and NODATA are answers and
|
||||||
|
// are returned to the caller to assert on.
|
||||||
|
func liveQueryNameserver(
|
||||||
|
t *testing.T,
|
||||||
|
r *resolver.Resolver,
|
||||||
|
nameserver string,
|
||||||
|
hostname string,
|
||||||
|
) *resolver.NameserverResponse {
|
||||||
|
t.Helper()
|
||||||
|
|
||||||
|
what := fmt.Sprintf(
|
||||||
|
"QueryNameserver(%s, %s)", nameserver, hostname,
|
||||||
|
)
|
||||||
|
|
||||||
|
var out *resolver.NameserverResponse
|
||||||
|
|
||||||
|
retryLive(
|
||||||
|
t,
|
||||||
|
what,
|
||||||
|
func(ctx context.Context) error {
|
||||||
|
resp, err := r.QueryNameserver(
|
||||||
|
ctx, nameserver, hostname,
|
||||||
|
)
|
||||||
|
if err != nil {
|
||||||
|
return err
|
||||||
|
}
|
||||||
|
|
||||||
|
if resp.Status == resolver.StatusTimeout ||
|
||||||
|
resp.Status == resolver.StatusError {
|
||||||
|
return fmt.Errorf(
|
||||||
|
"%w: %s returned %s: %s",
|
||||||
|
errLiveNoAnswer, nameserver,
|
||||||
|
resp.Status, resp.Error,
|
||||||
|
)
|
||||||
|
}
|
||||||
|
|
||||||
|
out = resp
|
||||||
|
|
||||||
|
return nil
|
||||||
|
},
|
||||||
|
)
|
||||||
|
|
||||||
|
return out
|
||||||
|
}
|
||||||
|
|
||||||
|
// liveQueryAllNameservers queries every authoritative nameserver
|
||||||
|
// for a hostname, retrying until a quorum of them has answered.
|
||||||
|
// Individual nameservers that stay silent are left in the result
|
||||||
|
// for the caller to account for.
|
||||||
|
func liveQueryAllNameservers(
|
||||||
|
t *testing.T,
|
||||||
|
r *resolver.Resolver,
|
||||||
|
hostname string,
|
||||||
|
) map[string]*resolver.NameserverResponse {
|
||||||
|
t.Helper()
|
||||||
|
|
||||||
|
var out map[string]*resolver.NameserverResponse
|
||||||
|
|
||||||
|
retryLive(
|
||||||
|
t,
|
||||||
|
"QueryAllNameservers("+hostname+")",
|
||||||
|
func(ctx context.Context) error {
|
||||||
|
results, err := r.QueryAllNameservers(ctx, hostname)
|
||||||
|
if err != nil {
|
||||||
|
return err
|
||||||
|
}
|
||||||
|
|
||||||
|
if len(results) == 0 {
|
||||||
|
return fmt.Errorf(
|
||||||
|
"%w: no nameservers queried for %s",
|
||||||
|
errLiveNoAnswer, hostname,
|
||||||
|
)
|
||||||
|
}
|
||||||
|
|
||||||
|
answered := answeredCount(results)
|
||||||
|
if answered < liveQuorum(len(results)) {
|
||||||
|
return fmt.Errorf(
|
||||||
|
"%w: %d of %d answered: %s",
|
||||||
|
errLiveNoQuorum, answered,
|
||||||
|
len(results), describeStatuses(results),
|
||||||
|
)
|
||||||
|
}
|
||||||
|
|
||||||
|
out = results
|
||||||
|
|
||||||
|
return nil
|
||||||
|
},
|
||||||
|
)
|
||||||
|
|
||||||
|
return out
|
||||||
|
}
|
||||||
|
|
||||||
|
// liveResolveIPs resolves a hostname that is expected to have
|
||||||
|
// addresses, retrying until at least one is returned.
|
||||||
|
func liveResolveIPs(
|
||||||
|
t *testing.T,
|
||||||
|
r *resolver.Resolver,
|
||||||
|
hostname string,
|
||||||
|
) []string {
|
||||||
|
t.Helper()
|
||||||
|
|
||||||
|
var out []string
|
||||||
|
|
||||||
|
retryLive(
|
||||||
|
t,
|
||||||
|
"ResolveIPAddresses("+hostname+")",
|
||||||
|
func(ctx context.Context) error {
|
||||||
|
ips, err := r.ResolveIPAddresses(ctx, hostname)
|
||||||
|
if err != nil {
|
||||||
|
return err
|
||||||
|
}
|
||||||
|
|
||||||
|
if len(ips) == 0 {
|
||||||
|
return fmt.Errorf(
|
||||||
|
"%w: no addresses for %s",
|
||||||
|
errLiveNoAnswer, hostname,
|
||||||
|
)
|
||||||
|
}
|
||||||
|
|
||||||
|
out = ips
|
||||||
|
|
||||||
|
return nil
|
||||||
|
},
|
||||||
|
)
|
||||||
|
|
||||||
|
return out
|
||||||
|
}
|
||||||
|
|
||||||
|
// liveResolveIPsAllowingEmpty resolves a hostname that may legitimately
|
||||||
|
// have no addresses, so the empty result is returned rather than
|
||||||
|
// retried. Used for names that must not exist; the corresponding
|
||||||
|
// QueryAllNameservers test is what proves the nameservers actively
|
||||||
|
// said NXDOMAIN rather than merely staying silent.
|
||||||
|
func liveResolveIPsAllowingEmpty(
|
||||||
|
t *testing.T,
|
||||||
|
r *resolver.Resolver,
|
||||||
|
hostname string,
|
||||||
|
) []string {
|
||||||
|
t.Helper()
|
||||||
|
|
||||||
|
var out []string
|
||||||
|
|
||||||
|
retryLive(
|
||||||
|
t,
|
||||||
|
"ResolveIPAddresses("+hostname+")",
|
||||||
|
func(ctx context.Context) error {
|
||||||
|
ips, err := r.ResolveIPAddresses(ctx, hostname)
|
||||||
|
if err != nil {
|
||||||
|
return err
|
||||||
|
}
|
||||||
|
|
||||||
|
out = ips
|
||||||
|
|
||||||
|
return nil
|
||||||
|
},
|
||||||
|
)
|
||||||
|
|
||||||
|
return out
|
||||||
|
}
|
||||||
+101
-205
@@ -32,32 +32,17 @@ func newTestResolver(t *testing.T) *resolver.Resolver {
|
|||||||
return resolver.NewFromLogger(log)
|
return resolver.NewFromLogger(log)
|
||||||
}
|
}
|
||||||
|
|
||||||
func testContext(t *testing.T) context.Context {
|
// findOneNSForDomain picks one authoritative nameserver to aim a
|
||||||
t.Helper()
|
// test at. Live-DNS retry, concurrency and quorum handling live in
|
||||||
|
// livedns_test.go.
|
||||||
ctx, cancel := context.WithTimeout(
|
|
||||||
context.Background(), 60*time.Second,
|
|
||||||
)
|
|
||||||
t.Cleanup(cancel)
|
|
||||||
|
|
||||||
return ctx
|
|
||||||
}
|
|
||||||
|
|
||||||
func findOneNSForDomain(
|
func findOneNSForDomain(
|
||||||
t *testing.T,
|
t *testing.T,
|
||||||
r *resolver.Resolver,
|
r *resolver.Resolver,
|
||||||
ctx context.Context, //nolint:revive // test helper
|
|
||||||
domain string,
|
domain string,
|
||||||
) string {
|
) string {
|
||||||
t.Helper()
|
t.Helper()
|
||||||
|
|
||||||
nameservers, err := r.FindAuthoritativeNameservers(
|
return liveFindAuthoritative(t, r, domain)[0]
|
||||||
ctx, domain,
|
|
||||||
)
|
|
||||||
require.NoError(t, err)
|
|
||||||
require.NotEmpty(t, nameservers)
|
|
||||||
|
|
||||||
return nameservers[0]
|
|
||||||
}
|
}
|
||||||
|
|
||||||
// ----------------------------------------------------------------
|
// ----------------------------------------------------------------
|
||||||
@@ -70,13 +55,7 @@ func TestFindAuthoritativeNameservers_ValidDomain(
|
|||||||
t.Parallel()
|
t.Parallel()
|
||||||
|
|
||||||
r := newTestResolver(t)
|
r := newTestResolver(t)
|
||||||
ctx := testContext(t)
|
nameservers := liveFindAuthoritative(t, r, "google.com")
|
||||||
|
|
||||||
nameservers, err := r.FindAuthoritativeNameservers(
|
|
||||||
ctx, "google.com",
|
|
||||||
)
|
|
||||||
require.NoError(t, err)
|
|
||||||
require.NotEmpty(t, nameservers)
|
|
||||||
|
|
||||||
hasGoogleNS := false
|
hasGoogleNS := false
|
||||||
|
|
||||||
@@ -99,13 +78,9 @@ func TestFindAuthoritativeNameservers_Subdomain(
|
|||||||
t.Parallel()
|
t.Parallel()
|
||||||
|
|
||||||
r := newTestResolver(t)
|
r := newTestResolver(t)
|
||||||
ctx := testContext(t)
|
nameservers := liveFindAuthoritative(t, r, "www.google.com")
|
||||||
|
|
||||||
nameservers, err := r.FindAuthoritativeNameservers(
|
assert.NotEmpty(t, nameservers)
|
||||||
ctx, "www.google.com",
|
|
||||||
)
|
|
||||||
require.NoError(t, err)
|
|
||||||
require.NotEmpty(t, nameservers)
|
|
||||||
}
|
}
|
||||||
|
|
||||||
func TestFindAuthoritativeNameservers_ReturnsSorted(
|
func TestFindAuthoritativeNameservers_ReturnsSorted(
|
||||||
@@ -114,12 +89,7 @@ func TestFindAuthoritativeNameservers_ReturnsSorted(
|
|||||||
t.Parallel()
|
t.Parallel()
|
||||||
|
|
||||||
r := newTestResolver(t)
|
r := newTestResolver(t)
|
||||||
ctx := testContext(t)
|
nameservers := liveFindAuthoritative(t, r, "google.com")
|
||||||
|
|
||||||
nameservers, err := r.FindAuthoritativeNameservers(
|
|
||||||
ctx, "google.com",
|
|
||||||
)
|
|
||||||
require.NoError(t, err)
|
|
||||||
|
|
||||||
assert.True(
|
assert.True(
|
||||||
t,
|
t,
|
||||||
@@ -134,17 +104,8 @@ func TestFindAuthoritativeNameservers_Deterministic(
|
|||||||
t.Parallel()
|
t.Parallel()
|
||||||
|
|
||||||
r := newTestResolver(t)
|
r := newTestResolver(t)
|
||||||
ctx := testContext(t)
|
first := liveFindAuthoritative(t, r, "google.com")
|
||||||
|
second := liveFindAuthoritative(t, r, "google.com")
|
||||||
first, err := r.FindAuthoritativeNameservers(
|
|
||||||
ctx, "google.com",
|
|
||||||
)
|
|
||||||
require.NoError(t, err)
|
|
||||||
|
|
||||||
second, err := r.FindAuthoritativeNameservers(
|
|
||||||
ctx, "google.com",
|
|
||||||
)
|
|
||||||
require.NoError(t, err)
|
|
||||||
|
|
||||||
assert.Equal(t, first, second)
|
assert.Equal(t, first, second)
|
||||||
}
|
}
|
||||||
@@ -155,17 +116,8 @@ func TestFindAuthoritativeNameservers_TrailingDot(
|
|||||||
t.Parallel()
|
t.Parallel()
|
||||||
|
|
||||||
r := newTestResolver(t)
|
r := newTestResolver(t)
|
||||||
ctx := testContext(t)
|
ns1 := liveFindAuthoritative(t, r, "google.com")
|
||||||
|
ns2 := liveFindAuthoritative(t, r, "google.com.")
|
||||||
ns1, err := r.FindAuthoritativeNameservers(
|
|
||||||
ctx, "google.com",
|
|
||||||
)
|
|
||||||
require.NoError(t, err)
|
|
||||||
|
|
||||||
ns2, err := r.FindAuthoritativeNameservers(
|
|
||||||
ctx, "google.com.",
|
|
||||||
)
|
|
||||||
require.NoError(t, err)
|
|
||||||
|
|
||||||
assert.Equal(t, ns1, ns2)
|
assert.Equal(t, ns1, ns2)
|
||||||
}
|
}
|
||||||
@@ -176,13 +128,7 @@ func TestFindAuthoritativeNameservers_CloudflareDomain(
|
|||||||
t.Parallel()
|
t.Parallel()
|
||||||
|
|
||||||
r := newTestResolver(t)
|
r := newTestResolver(t)
|
||||||
ctx := testContext(t)
|
nameservers := liveFindAuthoritative(t, r, "cloudflare.com")
|
||||||
|
|
||||||
nameservers, err := r.FindAuthoritativeNameservers(
|
|
||||||
ctx, "cloudflare.com",
|
|
||||||
)
|
|
||||||
require.NoError(t, err)
|
|
||||||
require.NotEmpty(t, nameservers)
|
|
||||||
|
|
||||||
for _, ns := range nameservers {
|
for _, ns := range nameservers {
|
||||||
assert.True(t, strings.HasSuffix(ns, "."),
|
assert.True(t, strings.HasSuffix(ns, "."),
|
||||||
@@ -199,13 +145,9 @@ func TestQueryNameserver_BasicA(t *testing.T) {
|
|||||||
t.Parallel()
|
t.Parallel()
|
||||||
|
|
||||||
r := newTestResolver(t)
|
r := newTestResolver(t)
|
||||||
ctx := testContext(t)
|
ns := findOneNSForDomain(t, r, "google.com")
|
||||||
ns := findOneNSForDomain(t, r, ctx, "google.com")
|
resp := liveQueryNameserver(t, r, ns, "www.google.com")
|
||||||
|
|
||||||
resp, err := r.QueryNameserver(
|
|
||||||
ctx, ns, "www.google.com",
|
|
||||||
)
|
|
||||||
require.NoError(t, err)
|
|
||||||
require.NotNil(t, resp)
|
require.NotNil(t, resp)
|
||||||
|
|
||||||
assert.Equal(t, resolver.StatusOK, resp.Status)
|
assert.Equal(t, resolver.StatusOK, resp.Status)
|
||||||
@@ -222,13 +164,8 @@ func TestQueryNameserver_AAAA(t *testing.T) {
|
|||||||
t.Parallel()
|
t.Parallel()
|
||||||
|
|
||||||
r := newTestResolver(t)
|
r := newTestResolver(t)
|
||||||
ctx := testContext(t)
|
ns := findOneNSForDomain(t, r, "cloudflare.com")
|
||||||
ns := findOneNSForDomain(t, r, ctx, "cloudflare.com")
|
resp := liveQueryNameserver(t, r, ns, "cloudflare.com")
|
||||||
|
|
||||||
resp, err := r.QueryNameserver(
|
|
||||||
ctx, ns, "cloudflare.com",
|
|
||||||
)
|
|
||||||
require.NoError(t, err)
|
|
||||||
|
|
||||||
aaaaRecords := resp.Records["AAAA"]
|
aaaaRecords := resp.Records["AAAA"]
|
||||||
require.NotEmpty(t, aaaaRecords,
|
require.NotEmpty(t, aaaaRecords,
|
||||||
@@ -247,13 +184,8 @@ func TestQueryNameserver_MX(t *testing.T) {
|
|||||||
t.Parallel()
|
t.Parallel()
|
||||||
|
|
||||||
r := newTestResolver(t)
|
r := newTestResolver(t)
|
||||||
ctx := testContext(t)
|
ns := findOneNSForDomain(t, r, "google.com")
|
||||||
ns := findOneNSForDomain(t, r, ctx, "google.com")
|
resp := liveQueryNameserver(t, r, ns, "google.com")
|
||||||
|
|
||||||
resp, err := r.QueryNameserver(
|
|
||||||
ctx, ns, "google.com",
|
|
||||||
)
|
|
||||||
require.NoError(t, err)
|
|
||||||
|
|
||||||
mxRecords := resp.Records["MX"]
|
mxRecords := resp.Records["MX"]
|
||||||
require.NotEmpty(t, mxRecords,
|
require.NotEmpty(t, mxRecords,
|
||||||
@@ -265,13 +197,8 @@ func TestQueryNameserver_TXT(t *testing.T) {
|
|||||||
t.Parallel()
|
t.Parallel()
|
||||||
|
|
||||||
r := newTestResolver(t)
|
r := newTestResolver(t)
|
||||||
ctx := testContext(t)
|
ns := findOneNSForDomain(t, r, "google.com")
|
||||||
ns := findOneNSForDomain(t, r, ctx, "google.com")
|
resp := liveQueryNameserver(t, r, ns, "google.com")
|
||||||
|
|
||||||
resp, err := r.QueryNameserver(
|
|
||||||
ctx, ns, "google.com",
|
|
||||||
)
|
|
||||||
require.NoError(t, err)
|
|
||||||
|
|
||||||
txtRecords := resp.Records["TXT"]
|
txtRecords := resp.Records["TXT"]
|
||||||
require.NotEmpty(t, txtRecords,
|
require.NotEmpty(t, txtRecords,
|
||||||
@@ -297,14 +224,10 @@ func TestQueryNameserver_NXDomain(t *testing.T) {
|
|||||||
t.Parallel()
|
t.Parallel()
|
||||||
|
|
||||||
r := newTestResolver(t)
|
r := newTestResolver(t)
|
||||||
ctx := testContext(t)
|
ns := findOneNSForDomain(t, r, "google.com")
|
||||||
ns := findOneNSForDomain(t, r, ctx, "google.com")
|
resp := liveQueryNameserver(
|
||||||
|
t, r, ns, "this-surely-does-not-exist-xyz.google.com",
|
||||||
resp, err := r.QueryNameserver(
|
|
||||||
ctx, ns,
|
|
||||||
"this-surely-does-not-exist-xyz.google.com",
|
|
||||||
)
|
)
|
||||||
require.NoError(t, err)
|
|
||||||
|
|
||||||
assert.Equal(t, resolver.StatusNXDomain, resp.Status)
|
assert.Equal(t, resolver.StatusNXDomain, resp.Status)
|
||||||
}
|
}
|
||||||
@@ -313,13 +236,8 @@ func TestQueryNameserver_RecordsSorted(t *testing.T) {
|
|||||||
t.Parallel()
|
t.Parallel()
|
||||||
|
|
||||||
r := newTestResolver(t)
|
r := newTestResolver(t)
|
||||||
ctx := testContext(t)
|
ns := findOneNSForDomain(t, r, "google.com")
|
||||||
ns := findOneNSForDomain(t, r, ctx, "google.com")
|
resp := liveQueryNameserver(t, r, ns, "google.com")
|
||||||
|
|
||||||
resp, err := r.QueryNameserver(
|
|
||||||
ctx, ns, "google.com",
|
|
||||||
)
|
|
||||||
require.NoError(t, err)
|
|
||||||
|
|
||||||
for recordType, values := range resp.Records {
|
for recordType, values := range resp.Records {
|
||||||
assert.True(
|
assert.True(
|
||||||
@@ -336,13 +254,8 @@ func TestQueryNameserver_ResponseIncludesNameserver(
|
|||||||
t.Parallel()
|
t.Parallel()
|
||||||
|
|
||||||
r := newTestResolver(t)
|
r := newTestResolver(t)
|
||||||
ctx := testContext(t)
|
ns := findOneNSForDomain(t, r, "cloudflare.com")
|
||||||
ns := findOneNSForDomain(t, r, ctx, "cloudflare.com")
|
resp := liveQueryNameserver(t, r, ns, "cloudflare.com")
|
||||||
|
|
||||||
resp, err := r.QueryNameserver(
|
|
||||||
ctx, ns, "cloudflare.com",
|
|
||||||
)
|
|
||||||
require.NoError(t, err)
|
|
||||||
|
|
||||||
assert.Equal(t, ns, resp.Nameserver)
|
assert.Equal(t, ns, resp.Nameserver)
|
||||||
}
|
}
|
||||||
@@ -353,14 +266,10 @@ func TestQueryNameserver_EmptyRecordsOnNXDomain(
|
|||||||
t.Parallel()
|
t.Parallel()
|
||||||
|
|
||||||
r := newTestResolver(t)
|
r := newTestResolver(t)
|
||||||
ctx := testContext(t)
|
ns := findOneNSForDomain(t, r, "google.com")
|
||||||
ns := findOneNSForDomain(t, r, ctx, "google.com")
|
resp := liveQueryNameserver(
|
||||||
|
t, r, ns, "this-surely-does-not-exist-xyz.google.com",
|
||||||
resp, err := r.QueryNameserver(
|
|
||||||
ctx, ns,
|
|
||||||
"this-surely-does-not-exist-xyz.google.com",
|
|
||||||
)
|
)
|
||||||
require.NoError(t, err)
|
|
||||||
|
|
||||||
totalRecords := 0
|
totalRecords := 0
|
||||||
for _, values := range resp.Records {
|
for _, values := range resp.Records {
|
||||||
@@ -374,18 +283,9 @@ func TestQueryNameserver_TrailingDotHandling(t *testing.T) {
|
|||||||
t.Parallel()
|
t.Parallel()
|
||||||
|
|
||||||
r := newTestResolver(t)
|
r := newTestResolver(t)
|
||||||
ctx := testContext(t)
|
ns := findOneNSForDomain(t, r, "google.com")
|
||||||
ns := findOneNSForDomain(t, r, ctx, "google.com")
|
resp1 := liveQueryNameserver(t, r, ns, "google.com")
|
||||||
|
resp2 := liveQueryNameserver(t, r, ns, "google.com.")
|
||||||
resp1, err := r.QueryNameserver(
|
|
||||||
ctx, ns, "google.com",
|
|
||||||
)
|
|
||||||
require.NoError(t, err)
|
|
||||||
|
|
||||||
resp2, err := r.QueryNameserver(
|
|
||||||
ctx, ns, "google.com.",
|
|
||||||
)
|
|
||||||
require.NoError(t, err)
|
|
||||||
|
|
||||||
assert.Equal(t, resp1.Status, resp2.Status)
|
assert.Equal(t, resp1.Status, resp2.Status)
|
||||||
}
|
}
|
||||||
@@ -398,15 +298,9 @@ func TestQueryAllNameservers_ReturnsAllNS(t *testing.T) {
|
|||||||
t.Parallel()
|
t.Parallel()
|
||||||
|
|
||||||
r := newTestResolver(t)
|
r := newTestResolver(t)
|
||||||
ctx := testContext(t)
|
results := liveQueryAllNameservers(t, r, "google.com")
|
||||||
|
|
||||||
results, err := r.QueryAllNameservers(
|
assert.GreaterOrEqual(t, len(results), minNameservers)
|
||||||
ctx, "google.com",
|
|
||||||
)
|
|
||||||
require.NoError(t, err)
|
|
||||||
require.NotEmpty(t, results)
|
|
||||||
|
|
||||||
assert.GreaterOrEqual(t, len(results), 2)
|
|
||||||
|
|
||||||
for ns, resp := range results {
|
for ns, resp := range results {
|
||||||
assert.Equal(t, ns, resp.Nameserver)
|
assert.Equal(t, ns, resp.Nameserver)
|
||||||
@@ -417,19 +311,36 @@ func TestQueryAllNameservers_AllReturnOK(t *testing.T) {
|
|||||||
t.Parallel()
|
t.Parallel()
|
||||||
|
|
||||||
r := newTestResolver(t)
|
r := newTestResolver(t)
|
||||||
ctx := testContext(t)
|
results := liveQueryAllNameservers(t, r, "google.com")
|
||||||
|
|
||||||
results, err := r.QueryAllNameservers(
|
// A quorum, not unanimity: one authoritative server being
|
||||||
ctx, "google.com",
|
// slow or rate-limiting us is a property of the live
|
||||||
|
// internet, not a resolver defect.
|
||||||
|
assert.GreaterOrEqual(
|
||||||
|
t,
|
||||||
|
countStatus(results, resolver.StatusOK),
|
||||||
|
liveQuorum(len(results)),
|
||||||
|
"a quorum of nameservers should answer OK: %s",
|
||||||
|
describeStatuses(results),
|
||||||
)
|
)
|
||||||
require.NoError(t, err)
|
|
||||||
|
|
||||||
for ns, resp := range results {
|
// Quorum tolerates SILENCE only. Every individual result must
|
||||||
assert.Equal(
|
// be either the expected answer or a non-answer: ok, timeout
|
||||||
t, resolver.StatusOK, resp.Status,
|
// or error, and nothing else. Stated as a closed allowlist so
|
||||||
"NS %s should return OK", ns,
|
// that a wrong answer no one thought to ban — nxdomain and
|
||||||
)
|
// nodata today, any status added later — fails here rather
|
||||||
}
|
// than sliding through under the quorum.
|
||||||
|
assert.Empty(
|
||||||
|
t,
|
||||||
|
unsanctionedStatuses(
|
||||||
|
results,
|
||||||
|
resolver.StatusOK,
|
||||||
|
resolver.StatusTimeout,
|
||||||
|
resolver.StatusError,
|
||||||
|
),
|
||||||
|
"every nameserver must answer OK or not answer at all: %s",
|
||||||
|
describeStatuses(results),
|
||||||
|
)
|
||||||
}
|
}
|
||||||
|
|
||||||
func TestQueryAllNameservers_NXDomainFromAllNS(
|
func TestQueryAllNameservers_NXDomainFromAllNS(
|
||||||
@@ -438,20 +349,34 @@ func TestQueryAllNameservers_NXDomainFromAllNS(
|
|||||||
t.Parallel()
|
t.Parallel()
|
||||||
|
|
||||||
r := newTestResolver(t)
|
r := newTestResolver(t)
|
||||||
ctx := testContext(t)
|
results := liveQueryAllNameservers(
|
||||||
|
t, r, "this-surely-does-not-exist-xyz.google.com",
|
||||||
results, err := r.QueryAllNameservers(
|
|
||||||
ctx,
|
|
||||||
"this-surely-does-not-exist-xyz.google.com",
|
|
||||||
)
|
)
|
||||||
require.NoError(t, err)
|
|
||||||
|
|
||||||
for ns, resp := range results {
|
assert.GreaterOrEqual(
|
||||||
assert.Equal(
|
t,
|
||||||
t, resolver.StatusNXDomain, resp.Status,
|
countStatus(results, resolver.StatusNXDomain),
|
||||||
"NS %s should return nxdomain", ns,
|
liveQuorum(len(results)),
|
||||||
)
|
"a quorum of nameservers should report NXDOMAIN: %s",
|
||||||
}
|
describeStatuses(results),
|
||||||
|
)
|
||||||
|
|
||||||
|
// Silence is tolerated; any actual answer other than NXDOMAIN
|
||||||
|
// is not. Closed allowlist for the same reason as above: a
|
||||||
|
// server answering `ok` or `nodata` for a name that must not
|
||||||
|
// exist is a wrong answer, not a slow one.
|
||||||
|
assert.Empty(
|
||||||
|
t,
|
||||||
|
unsanctionedStatuses(
|
||||||
|
results,
|
||||||
|
resolver.StatusNXDomain,
|
||||||
|
resolver.StatusTimeout,
|
||||||
|
resolver.StatusError,
|
||||||
|
),
|
||||||
|
"every nameserver must report NXDOMAIN or not answer "+
|
||||||
|
"at all: %s",
|
||||||
|
describeStatuses(results),
|
||||||
|
)
|
||||||
}
|
}
|
||||||
|
|
||||||
// ----------------------------------------------------------------
|
// ----------------------------------------------------------------
|
||||||
@@ -462,11 +387,7 @@ func TestLookupNS_ValidDomain(t *testing.T) {
|
|||||||
t.Parallel()
|
t.Parallel()
|
||||||
|
|
||||||
r := newTestResolver(t)
|
r := newTestResolver(t)
|
||||||
ctx := testContext(t)
|
nameservers := liveLookupNS(t, r, "google.com")
|
||||||
|
|
||||||
nameservers, err := r.LookupNS(ctx, "google.com")
|
|
||||||
require.NoError(t, err)
|
|
||||||
require.NotEmpty(t, nameservers)
|
|
||||||
|
|
||||||
for _, ns := range nameservers {
|
for _, ns := range nameservers {
|
||||||
assert.True(t, strings.HasSuffix(ns, "."),
|
assert.True(t, strings.HasSuffix(ns, "."),
|
||||||
@@ -479,10 +400,7 @@ func TestLookupNS_Sorted(t *testing.T) {
|
|||||||
t.Parallel()
|
t.Parallel()
|
||||||
|
|
||||||
r := newTestResolver(t)
|
r := newTestResolver(t)
|
||||||
ctx := testContext(t)
|
nameservers := liveLookupNS(t, r, "google.com")
|
||||||
|
|
||||||
nameservers, err := r.LookupNS(ctx, "google.com")
|
|
||||||
require.NoError(t, err)
|
|
||||||
|
|
||||||
assert.True(t, sort.StringsAreSorted(nameservers))
|
assert.True(t, sort.StringsAreSorted(nameservers))
|
||||||
}
|
}
|
||||||
@@ -491,15 +409,8 @@ func TestLookupNS_MatchesFindAuthoritative(t *testing.T) {
|
|||||||
t.Parallel()
|
t.Parallel()
|
||||||
|
|
||||||
r := newTestResolver(t)
|
r := newTestResolver(t)
|
||||||
ctx := testContext(t)
|
fromLookup := liveLookupNS(t, r, "google.com")
|
||||||
|
fromFind := liveFindAuthoritative(t, r, "google.com")
|
||||||
fromLookup, err := r.LookupNS(ctx, "google.com")
|
|
||||||
require.NoError(t, err)
|
|
||||||
|
|
||||||
fromFind, err := r.FindAuthoritativeNameservers(
|
|
||||||
ctx, "google.com",
|
|
||||||
)
|
|
||||||
require.NoError(t, err)
|
|
||||||
|
|
||||||
assert.Equal(t, fromFind, fromLookup)
|
assert.Equal(t, fromFind, fromLookup)
|
||||||
}
|
}
|
||||||
@@ -512,11 +423,7 @@ func TestResolveIPAddresses_ReturnsIPs(t *testing.T) {
|
|||||||
t.Parallel()
|
t.Parallel()
|
||||||
|
|
||||||
r := newTestResolver(t)
|
r := newTestResolver(t)
|
||||||
ctx := testContext(t)
|
ips := liveResolveIPs(t, r, "google.com")
|
||||||
|
|
||||||
ips, err := r.ResolveIPAddresses(ctx, "google.com")
|
|
||||||
require.NoError(t, err)
|
|
||||||
require.NotEmpty(t, ips)
|
|
||||||
|
|
||||||
for _, ip := range ips {
|
for _, ip := range ips {
|
||||||
parsed := net.ParseIP(ip)
|
parsed := net.ParseIP(ip)
|
||||||
@@ -530,10 +437,7 @@ func TestResolveIPAddresses_Deduplicated(t *testing.T) {
|
|||||||
t.Parallel()
|
t.Parallel()
|
||||||
|
|
||||||
r := newTestResolver(t)
|
r := newTestResolver(t)
|
||||||
ctx := testContext(t)
|
ips := liveResolveIPs(t, r, "google.com")
|
||||||
|
|
||||||
ips, err := r.ResolveIPAddresses(ctx, "google.com")
|
|
||||||
require.NoError(t, err)
|
|
||||||
|
|
||||||
seen := make(map[string]bool)
|
seen := make(map[string]bool)
|
||||||
|
|
||||||
@@ -547,10 +451,7 @@ func TestResolveIPAddresses_Sorted(t *testing.T) {
|
|||||||
t.Parallel()
|
t.Parallel()
|
||||||
|
|
||||||
r := newTestResolver(t)
|
r := newTestResolver(t)
|
||||||
ctx := testContext(t)
|
ips := liveResolveIPs(t, r, "google.com")
|
||||||
|
|
||||||
ips, err := r.ResolveIPAddresses(ctx, "google.com")
|
|
||||||
require.NoError(t, err)
|
|
||||||
|
|
||||||
assert.True(t, sort.StringsAreSorted(ips))
|
assert.True(t, sort.StringsAreSorted(ips))
|
||||||
}
|
}
|
||||||
@@ -561,13 +462,10 @@ func TestResolveIPAddresses_NXDomainReturnsEmpty(
|
|||||||
t.Parallel()
|
t.Parallel()
|
||||||
|
|
||||||
r := newTestResolver(t)
|
r := newTestResolver(t)
|
||||||
ctx := testContext(t)
|
ips := liveResolveIPsAllowingEmpty(
|
||||||
|
t, r, "this-surely-does-not-exist-xyz.google.com",
|
||||||
ips, err := r.ResolveIPAddresses(
|
|
||||||
ctx,
|
|
||||||
"this-surely-does-not-exist-xyz.google.com",
|
|
||||||
)
|
)
|
||||||
require.NoError(t, err)
|
|
||||||
assert.Empty(t, ips)
|
assert.Empty(t, ips)
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -575,11 +473,9 @@ func TestResolveIPAddresses_CloudflareDomain(t *testing.T) {
|
|||||||
t.Parallel()
|
t.Parallel()
|
||||||
|
|
||||||
r := newTestResolver(t)
|
r := newTestResolver(t)
|
||||||
ctx := testContext(t)
|
ips := liveResolveIPs(t, r, "cloudflare.com")
|
||||||
|
|
||||||
ips, err := r.ResolveIPAddresses(ctx, "cloudflare.com")
|
assert.NotEmpty(t, ips)
|
||||||
require.NoError(t, err)
|
|
||||||
require.NotEmpty(t, ips)
|
|
||||||
}
|
}
|
||||||
|
|
||||||
// ----------------------------------------------------------------
|
// ----------------------------------------------------------------
|
||||||
|
|||||||
+15
-7
@@ -3,15 +3,17 @@
|
|||||||
# this repo. Idempotent: every install is guarded by a check so already
|
# this repo. Idempotent: every install is guarded by a check so already
|
||||||
# installed tools are skipped. Base tooling comes from nix, apt, brew,
|
# installed tools are skipped. Base tooling comes from nix, apt, brew,
|
||||||
# or apk (detected in that order); assumes nothing is present.
|
# or apk (detected in that order); assumes nothing is present.
|
||||||
# golangci-lint and goimports are installed via `go install` at the same
|
# goimports is installed via `go install` at a pinned commit (never
|
||||||
# pinned commits the Dockerfile uses (never "latest").
|
# "latest") because script/fmt runs it on the host; script/fmt-check
|
||||||
|
# does not (it runs gofmt only).
|
||||||
|
# The linter is NOT installed here: golangci-lint runs via docker only
|
||||||
|
# (script/lint), pinned by image digest, so its only prerequisite is a
|
||||||
|
# working docker.
|
||||||
set -eu
|
set -eu
|
||||||
|
|
||||||
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
|
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
|
||||||
|
|
||||||
# Pinned versions, 2026-08-07 (same pins as the Dockerfile)
|
# Pinned version, 2026-08-07 (same pin as the Dockerfile)
|
||||||
# golangci-lint v2.12.2
|
|
||||||
GOLANGCI_LINT_REF="github.com/golangci/golangci-lint/v2/cmd/golangci-lint@c0d3ddc9cf3faa61a4e378e879ece580256d76e5"
|
|
||||||
# goimports v0.42.0
|
# goimports v0.42.0
|
||||||
GOIMPORTS_REF="golang.org/x/tools/cmd/goimports@009367f5c17a8d4c45a961a3a509277190a9a6f0"
|
GOIMPORTS_REF="golang.org/x/tools/cmd/goimports@009367f5c17a8d4c45a961a3a509277190a9a6f0"
|
||||||
|
|
||||||
@@ -69,11 +71,17 @@ main() {
|
|||||||
if missing make; then pkg_install gnumake make make make; fi
|
if missing make; then pkg_install gnumake make make make; fi
|
||||||
if missing go; then pkg_install go golang go go; fi
|
if missing go; then pkg_install go golang go go; fi
|
||||||
|
|
||||||
# Lint/format tools, pinned via go install (installs into
|
# Format tools, pinned via go install (installs into
|
||||||
# "$(go env GOPATH)/bin"; ensure that is on your PATH).
|
# "$(go env GOPATH)/bin"; ensure that is on your PATH).
|
||||||
if missing golangci-lint; then go install "$GOLANGCI_LINT_REF"; fi
|
|
||||||
if missing goimports; then go install "$GOIMPORTS_REF"; fi
|
if missing goimports; then go install "$GOIMPORTS_REF"; fi
|
||||||
|
|
||||||
|
# Linting runs via docker only (script/lint). Warn, don't fail:
|
||||||
|
# everything except `make lint` works without it.
|
||||||
|
if missing docker; then
|
||||||
|
echo "bootstrap: WARNING: docker not found; install it to" \
|
||||||
|
"run make lint and make docker." >&2
|
||||||
|
fi
|
||||||
|
|
||||||
go mod download
|
go mod download
|
||||||
|
|
||||||
echo "bootstrap complete"
|
echo "bootstrap complete"
|
||||||
|
|||||||
+3
-2
@@ -1,6 +1,7 @@
|
|||||||
#!/bin/sh
|
#!/bin/sh
|
||||||
# script/cibuild: run the CI build. The Dockerfile runs make check, so
|
# script/cibuild: run the CI build. The Dockerfile's lint stage runs
|
||||||
# a successful build implies all checks pass.
|
# make fmt-check and golangci-lint; its builder stage runs make test
|
||||||
|
# and make build. A successful build implies all of those passed.
|
||||||
set -eu
|
set -eu
|
||||||
|
|
||||||
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
|
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
|
||||||
|
|||||||
+18
-2
@@ -1,12 +1,28 @@
|
|||||||
#!/bin/sh
|
#!/bin/sh
|
||||||
# script/lint: run the linter.
|
# script/lint: run the linter. golangci-lint is never installed or run
|
||||||
|
# on the host: it runs via docker only, one way, everywhere. This
|
||||||
|
# builds Dockerfile.lint, which COPYs the repo into the digest-pinned
|
||||||
|
# golangci-lint image and lints as a build step, so a successful build
|
||||||
|
# means a clean lint.
|
||||||
|
#
|
||||||
|
# --no-cache-filter=lint forces the lint stage (source copy + linter
|
||||||
|
# run) to execute on every invocation. Without it an unchanged tree
|
||||||
|
# returns success in well under a second having linted nothing. The
|
||||||
|
# deps stage (base image + go mod download) stays cached, and no global
|
||||||
|
# cache invalidation is performed. --progress=plain keeps the linter's
|
||||||
|
# own output visible.
|
||||||
set -eu
|
set -eu
|
||||||
|
|
||||||
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
|
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
|
||||||
|
|
||||||
main() {
|
main() {
|
||||||
cd "$ROOT"
|
cd "$ROOT"
|
||||||
golangci-lint run --config .golangci.yml ./...
|
docker build \
|
||||||
|
--progress=plain \
|
||||||
|
--no-cache-filter=lint \
|
||||||
|
--target lint \
|
||||||
|
-f Dockerfile.lint \
|
||||||
|
.
|
||||||
}
|
}
|
||||||
|
|
||||||
main "$@"
|
main "$@"
|
||||||
|
|||||||
+24
-1
@@ -1,12 +1,35 @@
|
|||||||
#!/bin/sh
|
#!/bin/sh
|
||||||
# script/test: run the test suite.
|
# script/test: run the test suite.
|
||||||
|
#
|
||||||
|
# -count=1 disables Go's test cache, and is load-bearing here. This
|
||||||
|
# suite queries live DNS on every run by policy (TESTING.md); a cached
|
||||||
|
# result is a replay of an earlier run's output with no query made at
|
||||||
|
# all. On an unchanged tree the whole suite would return success in
|
||||||
|
# under a second having resolved nothing, which makes the repeated-run
|
||||||
|
# green that is used as evidence for flakiness fixes worthless. Do not
|
||||||
|
# remove it.
|
||||||
|
#
|
||||||
|
# Conditional verbose rerun per REPO_POLICIES.md: run quiet first so
|
||||||
|
# CI and docker build logs stay readable, and rerun with -v only on
|
||||||
|
# failure. The rerun also carries -count=1 (a cached replay of the
|
||||||
|
# failure would show nothing new), and the exit status is forced to 1
|
||||||
|
# no matter how the rerun ends: the first failure already proved the
|
||||||
|
# suite broken, so a flaky test that passes the second time must not
|
||||||
|
# turn the build green.
|
||||||
|
#
|
||||||
|
# -timeout 90s is a deliberate backstop above the 60s hard cap on
|
||||||
|
# suite duration. Do not lower it.
|
||||||
set -eu
|
set -eu
|
||||||
|
|
||||||
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
|
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
|
||||||
|
|
||||||
main() {
|
main() {
|
||||||
cd "$ROOT"
|
cd "$ROOT"
|
||||||
go test -v -race -timeout 30s -cover ./...
|
go test -count=1 -race -timeout 90s -cover ./... || {
|
||||||
|
echo "--- Rerunning with -v for details ---" >&2
|
||||||
|
go test -count=1 -race -timeout 90s -v ./... || true
|
||||||
|
exit 1
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
main "$@"
|
main "$@"
|
||||||
|
|||||||
Reference in New Issue
Block a user