1 Commits
Author SHA1 Message Date
clawbotandsneak e77e206fc4 next (#136)
check / check (push) Successful in 55s
Long-lived integration branch. One commit per work unit lands here; this PR accumulates them until it is merged to `main`.

## Landed units

- **Run all linting in Docker via `Dockerfile.lint` + `script/lint`** — #134

    golangci-lint is no longer installed or run on the host. New root `Dockerfile.lint` COPYs the repo into the digest-pinned `golangci/golangci-lint:v2.12.2` image and lints as a build step, so a successful build IS a clean lint; `script/lint` is reduced to a thin wrapper that builds it. This also works where the docker daemon is remote and bind mounts are impossible.

    **Pinned digest and how it was verified.** `golangci/golangci-lint:v2.12.2@sha256:5cceeef04e53efe1470638d4b4b4f5ceefd574955ab3941b2d9a68a8c9ad5240`, exactly as quoted in the issue. It resolves, and it is genuinely v2.12.2:

    ```
    $ docker buildx imagetools inspect golangci/golangci-lint:v2.12.2
    Name:      docker.io/golangci/golangci-lint:v2.12.2
    MediaType: application/vnd.oci.image.index.v1+json
    Digest:    sha256:5cceeef04e53efe1470638d4b4b4f5ceefd574955ab3941b2d9a68a8c9ad5240

    $ docker run --rm golangci/golangci-lint@sha256:5cceeef04e...ad5240 golangci-lint --version
    golangci-lint has version 2.12.2 built with go1.26.2 from c0d3ddc9 on 2026-05-06T11:07:58Z
    ```

    The tag's index digest is the quoted digest, and the binary inside reports commit `c0d3ddc9`, matching the org's canonical pin `c0d3ddc9cf3faa61a4e378e879ece580256d76e5`.

    **Forcing the linter to actually run.** `Dockerfile.lint` is split into a `deps` stage (base image + `go mod download`) and a `lint` stage (source copy + linter run). `script/lint` runs:

    ```
    docker build --progress=plain --no-cache-filter=lint --target lint -f Dockerfile.lint .
    ```

    Caching is explicitly waived for linting, and a cached build lints nothing, so the `lint` stage is invalidated on every invocation. The invalidation is scoped: the `deps` stage stays cached and no global cache wipe is performed. `--progress=plain` keeps the linter's own output visible.

    **`golangci-lint config verify`: deliberately NOT included.** It fetches its JSON schema over a live, unpinned HTTPS call, which would make linting network-dependent and defeat hash-pinning. Omitted for that reason, and the reason is recorded in a comment at the top of `Dockerfile.lint`.

    **`script/bootstrap`** no longer installs golangci-lint (and its pinned ref is gone); it warns non-fatally when `docker` is absent instead. The `goimports` install stays, because `script/fmt` and `script/fmt-check` still run on the host. Header comment updated accordingly.

    **Root `Dockerfile`** — required consequence, not scope creep. Its builder stage ran `make check`, which now calls `script/lint`, which shells out to `docker build`; there is no docker daemon inside a docker build, so `script/cibuild` and `script/docker` would have broken. It gains its own lint stage on the same pinned image (linter invoked directly, with a comment explaining why not `make lint`), with the builder stage depending on it via `COPY --from=lint /src/go.sum /dev/null` and running `make fmt-check`, `make test`, `make build`. The now-unneeded golangci-lint install is gone from the builder stage.

    **README** `Entrypoints` and `Building` sections now describe linting as a docker-only operation. `TODO.md` updated in the same commit.

    ### Verification

    All runs via `make` / `script/` entrypoints only.

    Two consecutive `make lint` runs on an unchanged tree, both executing the linter:

    ```
    # run 1
    #10 [lint 2/2] RUN golangci-lint run --config .golangci.yml ./...
    #10 10.98 0 issues.
    #10 DONE 12.0s

    # run 2, tree untouched
    #10 [lint 2/2] RUN golangci-lint run --config .golangci.yml ./...
    #10 14.63 0 issues.
    ```

    A third run shows the cache scoping is working as intended — `deps` served from cache, `lint` re-executed:

    ```
    #6 [deps 2/4] WORKDIR /src
    #6 CACHED
    #7 [deps 3/4] COPY go.mod go.sum ./
    #7 CACHED
    #8 [deps 4/4] RUN go mod download
    #8 CACHED
    #9 [lint 1/2] COPY . .
    #10 [lint 2/2] RUN golangci-lint run --config .golangci.yml ./...
    ```

    **Negative control.** A deliberate violation (an unused function containing an ineffectual assignment) was added to `internal/config/config.go`:

    ```
    #10 11.26 internal/config/config.go:29:2: ineffectual assignment to x (ineffassign)
    #10 11.26 internal/config/config.go:28:6: func negativeControlUnused is unused (unused)
    #10 11.26 2 issues:
    #10 ERROR: process "/bin/sh -c golangci-lint run --config .golangci.yml ./..." did not complete successfully: exit code: 1
    ERROR: failed to build: failed to solve: process "/bin/sh -c golangci-lint run --config .golangci.yml ./..." did not complete successfully: exit code: 1
    make: *** [Makefile:26: lint] Error 1
    ```

    `make lint` exited non-zero naming both findings and their exact lines. After reverting the file, `make lint` was clean again (`0 issues.`).

    **`make check`** green end to end (test, lint, fmt-check), exit 0.

    **`script/cibuild`** green, confirming the `Dockerfile` restructure does not recurse: the lint stage ran (`#16 12.56 0 issues.`), then `#22 [builder 8/9] RUN make test` with `PASS` lines, then `#23 [builder 9/9] RUN make build`.

    ### Notes for the owner

    This supersedes two PRs you still have queued for merge, both of which tune host linting that no longer exists after this change: #128 (isolates the host golangci-lint cache and lock) and #131 (always installs the pinned lint tools in `script/bootstrap`). Neither was merged or incorporated here.

    Made moot by this change: #121 and #130.

    ### Review outcome

    Reviewed at #136 (comment) — **PASS**. Three non-blocking comment-accuracy findings were recorded there for the next touch of those files; the reviewer independently confirmed the lint gate is live by negative control from a warm cache.

- **Live DNS tests made robust instead of gated; test caps moved to the org-wide 60s/20s/90s values** — #93

    The resolver's live-DNS tests failed nondeterministically, a different subset each run. Fixed by engineering the nondeterminism out, not by routing around the network. **Nothing is mocked, faked, stubbed, recorded or replayed; there is no `-short` flag, no build tag, no skip, and no environment-tolerance for restricted egress.** Production resolver behaviour is unchanged.

    ### Root causes, all test-side

    1. **Burst fan-out at one root server.** Every test in `internal/resolver` calls `t.Parallel()` and the build hosts have many cores (48 here), so all ~35 iterative resolutions started within milliseconds of each other, and because `queryServers` walks `rootServerList()` in fixed order they all aimed their first query at `198.41.0.4`. Root servers rate-limit that, which fits the reported symptom of a different arbitrary subset failing each run.
    2. **No retry anywhere.** One dropped UDP packet in a delegation chain failed a test outright.
    3. **Unanimity assertions.** `TestQueryAllNameservers_AllReturnOK` and `_NXDomainFromAllNS` required *every* nameserver of a domain to answer — four independent chances to fail per run, with no tolerance for one being slow.

    ### What was built

    New `internal/resolver/livedns_test.go` holds all the live-DNS machinery, so `resolver_test.go` itself takes only call-site edits:

    - **Bounded live concurrency.** A package-wide semaphore (`liveConcurrency = 6`) caps how many live resolutions are in flight at once. Tests keep `t.Parallel()`; only their network work is throttled. This is the direct fix for cause 1, and the 60s budget is what makes it affordable.
    - **Retry with exponential backoff.** Three attempts per live operation, 8s deadline each, 500ms base backoff doubling. The retry predicate is deliberately **transport-level** — "did a nameserver answer at all" — and never the assertion the test is making, so a resolver that answers *incorrectly* still fails on the first attempt rather than being retried into a false green.
    - **Quorum instead of unanimity, tolerating SILENCE ONLY.** A strict majority of the discovered nameservers must answer as expected, and every individual result must additionally fall inside a closed **allowlist** of statuses the test explicitly sanctions: `ok`/`timeout`/`error` for the all-OK test, `nxdomain`/`timeout`/`error` for the NXDOMAIN test. A nameserver that stays silent is tolerated; one that answers **wrongly** is not, at any count. The allowlist is the load-bearing part — see the rework note below for why a blocklist was not enough.

    New `internal/resolver/livedns_harness_test.go` tests that machinery directly — quorum arithmetic, status counting, the allowlist, the gate's concurrency bound, per-attempt deadlines, and recovery from a transient failure. It performs no DNS resolution of any kind, so it neither mocks DNS nor depends on it.

    ### Rework after review — the quorum could not fail on a wrong answer

    The review at #136 (comment) returned **FAIL** on `9cb2c2b`, correctly. Fixed in `87bce43`.

    **The defect.** The claim above was, as first written, false for `resolver.StatusNoData`. Each test banned exactly one wrong status — `_AllReturnOK` banned only `nxdomain`, `_NXDomainFromAllNS` banned only `ok` — and `nodata` is neither. It is a **wrong answer, not silence**: `answeredCount` counted it as answered, so it did not even trigger a retry, and with a quorum of 3-of-4 a single wrong nameserver slid through undetected. The pre-change unanimity assertions would have caught it. That is robustness work quietly becoming assertion-loosening, which is exactly what this repo cannot afford.

    **The fix.** Tolerance is now an allowlist, not a blocklist of one status. New `unsanctionedStatuses()` returns every per-nameserver result whose status the caller did not explicitly sanction, and each test asserts that list is empty in addition to its quorum. A blocklist bans the one wrong answer its author thought of and silently admits everything else, including any status added to the resolver later; an allowlist fails on anything nobody sanctioned. `answeredCount` was reframed the same way — it now counts the closed set `ok`/`nxdomain`/`nodata`, so an unfamiliar status is treated as silence and can only ever cause a retry and then a loud failure, never a quiet pass.

    **Evidence — the reviewer's exact probe, re-run.** `queryEachNS` in `internal/resolver/iterative.go` was patched to force one of `google.com`'s four nameservers to return `StatusNoData` with empty records. `make test` now goes **red**, naming the offending nameserver and status:

    ```
    exit=2
    --- FAIL: TestQueryAllNameservers_AllReturnOK (1.12s)
        Error:    Should be empty, but was [ns1.google.com.=nodata]
        Messages: every nameserver must answer OK or not answer at all:
                  ns1.google.com.=nodata ns2.google.com.=ok
                  ns3.google.com.=ok ns4.google.com.=ok
    --- FAIL: TestQueryAllNameservers_NXDomainFromAllNS (1.34s)
        Error:    Should be empty, but was [ns1.google.com.=nodata]
        Messages: every nameserver must report NXDOMAIN or not answer at all:
                  ns1.google.com.=nodata ns2.google.com.=nxdomain
                  ns3.google.com.=nxdomain ns4.google.com.=nxdomain
    FAIL  sneak.berlin/go/dnswatcher/internal/resolver  1.704s
    ```

    That is the same input that returned `exit=0` with both tests **passing** under the old assertions. Probe reverted, tree clean, suite green again:

    ```
    exit=0
    ok  sneak.berlin/go/dnswatcher/internal/resolver  4.185s  coverage: 77.1% of statements
    (zero `(cached)` lines)
    ```

    Two harness tests lock the regression in without any probe: three OK plus one `nodata` (quorum satisfied, no NXDOMAIN present — the exact input that used to pass) is reported as unsanctioned, and an unknown status is neither counted as answered nor tolerated.

    Also fixed from the review: the per-attempt deadline assertion in `livedns_harness_test.go` had no lower bound, so it passed for a deadline far shorter than intended. It now asserts the remaining time exceeds `liveAttemptTimeout/2` as well.

    Deliberately **not** done in this rework, per the review and the owner: no `-count=1` in `script/test` (the test-cache issue is real but pre-existing and repo-wide, filed separately); the remaining non-DNS mocks stay for #97; `queryServers` root-ordering stays untouched under #138.

    ### The `-timeout` backstop value: 90s

    Per the ruling at sneak/prompts#41 (comment) the cap is org-wide with two tiers: **60s hard cap for CI green, 20s target, and anything between the two must be filed as an improvement bug.** The backstop is **`90s`**, matching sneak/prompts#42 and preserving the 1.5x backstop-to-cap ratio the old 20s/30s pair had. It must strictly exceed the 60s cap or the cap is unreachable — the old `-timeout 30s` would have killed a 60s-capped suite at half its allowance. Applied to `script/test`; nothing else in the repo carried the old `30s`.

    Worst case for one live operation is 3 attempts x 8s plus ~1.5s of backoff, about 26s — comfortably inside the 90s backstop even if several operations exhaust their attempts at once.

    ### `REPO_POLICIES.md` is re-vendored, not hand-edited

    The file is org-canonical, so it was **copied byte-for-byte** from `prompts/REPO_POLICIES.md` on `sneak/prompts` branch `org-wide-60s-test-cap` (commit `52b5192`) rather than reworded to approximately the same thing. Verified:

    ```
    $ cmp prompts/REPO_POLICIES.md dnswatcher/REPO_POLICIES.md && echo identical
    identical
    $ sha256sum REPO_POLICIES.md
    bcf11c312a1bee18a0e937eb412b51914411c1ab23308b8362409f3f88379ff7
    ```

    **What that byte-identity does and does not certify.** The source branch `org-wide-60s-test-cap` is an **unmerged proposal** — sneak/prompts#42 — not `prompts` `main`. So, precisely:

    - The **60s hard cap and 20s improvement-bug tier ARE the owner's ruling** (sneak/prompts#41 (comment)).
    - The **`90s` backstop is our own proposed number and is NOT ratified** (sneak/prompts#41 (comment)).
    - The vendored text is therefore **the proposed canonical text, pending** sneak/prompts#42. If that PR lands with different numbers, this file must be re-vendored to match; it should not be hand-edited here either way.

    **Known mismatch with this repo's actual state, recorded not papered over.** Re-vendoring also picked up the paragraph at `REPO_POLICIES.md:266-271` mandating that canonical golangci-lint be installed commit-pinned via `go install ...@c0d3ddc9...`. This repo does **not** comply with that mechanism: `cc86473` in this same PR made linting Docker-only, and `script/bootstrap` now installs golangci-lint nowhere. The **version and commit match** (`v2.12.2` / `c0d3ddc9`); the **installation mechanism does not**. The vendored file is org-canonical and must not be edited downstream, so this is being raised upstream for the org text to accommodate Docker-only linting rather than patched here.

    `TESTING.md`'s stale "within the 30-second target" follows to 60. That edit is deliberately a single line so it merges cleanly when #97 lands.

    ### Verification

    All runs through `make` / `script/` entrypoints only; lint runs in Docker.

    **Ten consecutive `make check` runs, all green, none served from cache.** Go's test cache will happily report `ok pkg (cached)` without executing anything, which proves nothing about nondeterminism, so every run was forced to actually execute and each log was checked for zero `(cached)` lines:

    ```
    check#1  exit=0 wall=32s resolver=2.895s fails=0 cached=0
    check#2  exit=0 wall=27s resolver=2.872s fails=0 cached=0
    check#3  exit=0 wall=47s resolver=2.729s fails=0 cached=0
    check#4  exit=0 wall=36s resolver=2.885s fails=0 cached=0
    check#5  exit=0 wall=32s resolver=2.899s fails=0 cached=0
    check#6  exit=0 (harness tests added)      fails=0 cached=0
    check#7  exit=0 wall=41s resolver=2.891s fails=0 cached=0
    check#8  exit=0 wall=36s resolver=2.871s fails=0 cached=0
    check#9  exit=0 wall=26s resolver=2.846s fails=0 cached=0
    check#10 exit=0 wall=46s resolver=2.833s fails=0 cached=0
    ```

    After the rework commit `87bce43`, `make check` green again end to end, zero `(cached)` test lines, Docker lint stage demonstrably executed rather than served from cache:

    ```
    exit=0  cached=0
    #8 [deps 4/4] RUN go mod download
    #8 CACHED
    #10 [lint 2/2] RUN golangci-lint run --config .golangci.yml ./...
    #10 30.72 0 issues.
    ```

    The Docker lint stage was confirmed to execute rather than cache on each run:

    ```
    #8 [deps 4/4] RUN go mod download
    #8 CACHED
    #9 [lint 1/2] COPY . .
    #9 DONE 0.6s
    #10 [lint 2/2] RUN golangci-lint run --config .golangci.yml ./...
    #10 DONE 39.7s
    ```

    **`make test` wall time: 3.6-4.1s** across three timed uncached runs (`4093ms`, `3649ms`, `3729ms`). `internal/resolver` went from 2.0s to ~2.9s — the concurrency gate's cost. That is inside the 20s target, so no improvement bug is owed under the new two-tier rule.

    **Honest note on what these runs do and do not prove.** Live DNS was healthy throughout: **no live-DNS retry fired even once**, and no flake was observed either before or after the change (six pre-change baseline runs were also clean). So these runs demonstrate the change is not itself flaky and does not slow the suite; they do **not** demonstrate recovery from a real DNS failure, because no real DNS failure occurred. The original flakiness is *not reproduced* rather than *shown fixed*. The retry path is instead proven by `TestRetryLiveRecoversFromTransientFailure`, the only source of the single `retrying in 500ms` line in each log:

    ```
    livedns_harness_test.go:90: transient: attempt 1 of 3 failed
        (no answer from live DNS), retrying in 500ms
    ```

    ### Observation, not acted on

    The single most effective remaining lever against root-server rate limiting would be to stop `queryServers` always trying `rootServerList()` in the same order, so that load spreads across all thirteen roots instead of concentrating on `a.root-servers.net`. That is **production** code and this issue scopes the work as test-side, so it was left alone rather than changed quietly. It is now tracked for the owner's decision at #138.

    ### Interaction with #97

    `TESTING.md` and `internal/resolver/resolver_test.go` auto-merge — that PR touches `resolver_test.go` only at the import block and the final timeout-test section, while this change touches the body of the file and adds two new files, and it leaves the mock-`DNSClient` timeout test at the tail of `resolver_test.go` entirely alone since removing it is that PR's job. `TODO.md` does conflict; that PR is already labelled `needs-rebase`, so this adds nothing material to its rebase.

- **Go's test cache disabled, so every `make test` actually queries live DNS** — #139

    `script/test` did not pass `-count=1`, so on an unchanged tree Go served the whole suite from cache: exit 0 in ~0.2s, every package marked `(cached)`, and not one DNS query made. This repo's suite exists to exercise live resolution on every run (`TESTING.md`), so that green asserted nothing — and it is exactly the green used as evidence that a flakiness fix works, since "run it a few times" stops being runs after the first. It had already misled two agents, each of whom forced uncached runs by hand.

    `-count=1` now disables caching on every invocation.

    **The conditional verbose rerun was missing and is added here.** `REPO_POLICIES.md` mandates it; the primary run had been unconditionally `-v`, which is the failure mode the policy exists to prevent (unreadable CI and `docker build` logs on success). Tests now run quiet, and only a failure triggers the `-v` rerun. Two properties matter and both are covered: the rerun carries `-count=1` too, so it cannot replay a cached copy of the failure it is meant to diagnose; and its exit status is discarded in favour of a forced `1`, so a flake that passes the second time cannot turn the build green — the first failure already proved the suite broken.

    **`-timeout 90s` untouched.** It is a deliberate backstop that must strictly exceed the 60s hard cap.

    **No special-casing for the Docker build**, which reaches the same script via `RUN make test`: a fresh container's test cache is empty, so `-count=1` is a no-op there, and carving out an exception would only create a second code path that could drift.

    ### Verification

    **Proven from a warm cache, not a cold one.** The suite was run first to populate the cache, and the pre-change state confirmed:

    ```
    ok  sneak.berlin/go/dnswatcher/internal/config    (cached)  coverage: 92.6% of statements
    ok  sneak.berlin/go/dnswatcher/internal/resolver  (cached)  coverage: 77.1% of statements
    ...8 of 8 packages (cached)...
    real 0m0.203s
    ```

    With the change applied to that same warm cache, three back-to-back runs on an unchanged tree, **zero `(cached)` markers** in all three:

    ```
    # run 1                                          # run 2
    ok  .../internal/config     1.045s               ok  .../internal/config     1.040s
    ok  .../internal/handlers   1.029s               ok  .../internal/handlers   1.021s
    ok  .../internal/notify     1.148s               ok  .../internal/notify     1.252s
    ok  .../internal/portcheck  1.026s               ok  .../internal/portcheck  1.025s
    ok  .../internal/resolver   3.005s               ok  .../internal/resolver   2.846s
    ok  .../internal/state      1.078s               ok  .../internal/state      1.097s
    ok  .../internal/tlscheck   1.067s               ok  .../internal/tlscheck   1.079s
    ok  .../internal/watcher    1.566s               ok  .../internal/watcher    1.544s
    real 0m4.174s                                    real 0m4.015s

    # run 3: grep -c '(cached)' => 0                 real 0m4.519s
    ```

    **Measured uncached wall time: 4.0-4.5s** (was ~0.2s served from cache). Inside the 20s target, so no improvement bug is owed under the two-tier rule at sneak/prompts#41 (comment).

    **`-count=1` composes with `-race` and `-cover`**: both still present in the primary run, and the per-package coverage percentages above are identical to the pre-change values.

    **The failure path was exercised, not assumed.** A purpose-built flaky test that fails on its first run and passes every run after (marker file kept outside the module, so the tree stays byte-identical and a cached result would be served if caching were on) was run through the script:

    ```
    exit code: 1
    --- FAIL: TestFlaky (0.00s)
    FAIL    flakeproof  0.013s
    --- Rerunning with -v for details ---
    --- PASS: TestFlaky (0.00s)
    ok      flakeproof  1.014s
    ```

    Quiet failure, verbose rerun that genuinely re-executed (it passed, so it did not replay the cached `FAIL`), and exit `1` regardless of the rerun passing. Scratch module removed afterwards.

    **`make check` green**, exit 0, with the Docker lint stage demonstrably executed rather than served from cache:

    ```
    #8  [deps 4/4] RUN go mod download
    #8  CACHED
    #10 [lint 2/2] RUN golangci-lint run --config .golangci.yml ./...
    #10 21.36 0 issues.
    #10 DONE 26.2s
    ```

    `README.md` and `TESTING.md` record why the cache is waived, and `TODO.md` is updated in the same commit.

    ### Question for the owner, not filed as a defect

    The Docker lint run emits `The linter 'gomodguard' is deprecated (since v2.12.0) due to: new major version. Replaced by gomodguard_v2.` It is pre-existing and out of this issue's scope. It is not filed as an issue here because `.golangci.yml` tracks the org-canonical config, so switching to `gomodguard_v2` looks like an upstream `sneak/prompts` decision rather than a per-repo fix. Say the word and it gets filed in whichever place you consider canonical.

- **MIT `LICENSE` added; README states the licence** — #102

    The repo had no licence file at all, so publicly readable code was all-rights-reserved by default and nobody could legally use it. `LICENSE` was also the last file missing from `REPO_POLICIES.md`'s required minimum.

    MIT, by standing org policy rather than a per-repo call: any public repo lacking a licence gets MIT, and a private one with no licence is already all-rights-reserved. `sneak/dnswatcher` is public (`private: false` on the Gitea repo record).

    `README.md`'s first line now names the licence, per the Description requirement, and the License section states MIT and points at the file instead of recording the decision as pending. `TODO.md` updated in the same commit.

    ### Verification

    `LICENSE` is the canonical MIT text byte-for-byte with only the copyright line filled in (`Copyright (c) 2026 sneak`) — no clauses added, removed, reworded, or reflowed. It was not typed from memory: the file was copied from an existing verbatim MIT template on disk and only the copyright line edited (`diff` against that template shows that one line and nothing else), then the result was word-diffed against SPDX `MIT.txt` fetched from `spdx/license-list-data`, ignoring only line wrapping and the placeholder — identical.

    `make fmt` did **not** touch `LICENSE`, and cannot: `script/fmt` runs `gofmt -s -w .` and `goimports -w .` only, with no prettier or markdown step in the repo, so no exclusion was needed.

    `make check` green, exit 0. Tests executed rather than replayed (zero `(cached)` lines, `internal/resolver 3.098s`), and the Docker lint stage ran rather than cached:

    ```
    #10 [deps 4/4] RUN go mod download
    #10 CACHED
    #12 [lint 2/2] RUN golangci-lint run --config .golangci.yml ./...
    #12 36.44 0 issues.
    #12 DONE 36.7s
    ```

- **Comment-only corrections in `script/bootstrap`, `script/cibuild`, and `Dockerfile.lint`** — #137

    Follow-up to this PR's own review. Nothing executable changed: the diff touches comment lines, one warning string, and `TODO.md`.

    - `script/bootstrap`'s header justified the pinned `goimports` install by claiming `script/fmt-check` runs it on the host. Verified against the script: `script/fmt-check` runs `gofmt -l .` and nothing else. The header now credits `script/fmt` alone. That `fmt-check` does not verify goimports at all is #119 and was deliberately left alone.
    - `script/cibuild`'s header still said the `Dockerfile` runs `make check`. It now describes the current file: lint stage runs `make fmt-check` and `golangci-lint`, builder stage runs `make test` and `make build`.
    - The `docker`-missing warning was three fragments, each re-prefixed with `bootstrap:` mid-clause. Now one sentence: `bootstrap: WARNING: docker not found; install it to run make lint and make docker.`
    - `Dockerfile.lint`'s comment explained why `golangci-lint config verify` is omitted but read as though the omission were free. It now states the residual risk: unknown top-level keys in `.golangci.yml` are silently ignored, so a mistyped or wrong-schema key lints clean while applying nothing. `config verify` was **not** added — the network-dependence reasoning stands.

    ### Verification

    `make check` green, exit 0. Tests executed rather than replayed (zero `(cached)` lines, `internal/resolver 2.820s`), Docker lint stage executed rather than cached:

    ```
    #8 [deps 4/4] RUN go mod download
    #8 CACHED
    #10 [lint 2/2] RUN golangci-lint run --config .golangci.yml ./...
    #10 13.30 0 issues.
    ```

    Comment-only confirmed by reading the whole diff: no statement, flag, or command changed anywhere.

---

## Issues closed by this merge

The commits on `next` each carry a bare `(closes #N)` in their subject, but the
references in the prose above are full URLs, which Gitea's auto-close parser
does not act on. Listing them here in bare form so the merge to `main`
definitively closes them rather than leaving them open to be re-picked up as
idle work:

Closes #93
Closes #102
Closes #134
Closes #137
Closes #139

Co-authored-by: sneak <sneak@sneak.berlin>
Reviewed-on: #136
Co-authored-by: clawbot <clawbot@noreply.example.org>
Co-committed-by: clawbot <clawbot@noreply.example.org>
2026-09-09 14:57:55 +02:00
15 changed files with 1367 additions and 1099 deletions
+10 -3
View File
@@ -218,7 +218,8 @@ internal/
- **Structured logging**: All logs use `log/slog` with JSON output in
production (TTY detection for development).
- **Graceful shutdown**: All background goroutines respect context
cancellation and the fx lifecycle.
cancellation and the fx lifecycle. In-flight notification deliveries
are drained on shutdown, bounded by the shutdown timeout.
---
@@ -463,8 +464,14 @@ docker run -d \
from a previous cycle.
4. **On change detection**: Send notifications to all configured
endpoints, update in-memory state, persist to disk.
5. **Shutdown**: Persist final state to disk, complete in-flight
notifications, stop gracefully.
5. **Shutdown**: Persist final state to disk, wait for in-flight
notification deliveries to complete, stop gracefully. The wait is
bounded by the fx shutdown timeout (15s by default): deliveries still
retrying against an unreachable endpoint when that expires are
abandoned, and the number abandoned is logged at warn level rather
than dropped silently. Notifications generated after shutdown has
begun are refused and logged, so a late burst cannot extend the
shutdown.
---
+4 -58
View File
@@ -2,10 +2,8 @@
## DNS Resolution Tests
All tests that involve DNS resolution — in every package, including
consumers of the resolver such as the watcher — **MUST** use live
queries against real DNS servers. No mocking, faking, or stubbing of
DNS at any layer is permitted.
All resolver tests **MUST** use live queries against real DNS servers.
No mocking of the DNS client layer is permitted.
### Rationale
@@ -14,8 +12,6 @@ the full delegation chain. Mocked responses cannot faithfully represent
the variety of real-world DNS behavior (truncation, referrals, glue
records, DNSSEC, varied response times, EDNS, etc.). Testing against
real servers ensures the resolver works correctly in production.
Robustness comes from handling real-world DNS behavior with tolerant
assertions and sensible timeouts, not from mocks.
### Constraints
@@ -28,64 +24,14 @@ assertions and sensible timeouts, not from mocks.
- Flaky failures from transient network issues are acceptable and
should be investigated as potential resolver bugs, not papered over
with mocks or skip flags
- Watcher change-detection tests seed a synthetic *previous state*
and compare it against fresh live lookups; the DNS side is never
faked
- Live query concurrency is bounded per package (`liveGate` in
`internal/resolver`, `liveWatcherGate` in `internal/watcher`) so
parallel tests do not burst at the root servers
- Those gates are package-scoped and therefore per test binary, so
`script/test` also passes `-p 1`: with both live-DNS packages
running at once the gates sum instead of holding, and the resolver's
per-attempt deadlines start expiring
### Transport failures: loopback nameservers, not mocks
The resolver classifies a nameserver that stays silent as
`StatusTimeout` and one that answers SERVFAIL as `StatusError`. The
public network cannot be made to produce either on demand — a
black-holed address is only black-holed on some networks, and build
environments that transparently intercept UDP/53 answer it locally —
so a test built on a chosen remote address asserts on the network it
happens to run on rather than on the resolver.
`internal/resolver/transport_test.go` binds a real nameserver on
`127.0.0.1` instead and points the query at it.
`nameserverAddr` dials an address that already carries a port as
written, so no production behaviour is bypassed to arrange this.
**This is permitted, and it is not a mock.** The rule above bans
substituting `DNSClient` or any other DNS abstraction, which lets the
code under test skip DNS and hands it a manufactured verdict. A
loopback nameserver does the opposite: the resolver dials a real
socket, writes a real query with the real `miekg/dns` client, and
applies its real deadline and its real classification logic to what
comes back. Choosing which nameserver a live query is sent to is not
faking DNS — the resolver is aimed at a nameserver of the caller's
choosing in production too.
The distinction to hold on to: **substituting the client is banned;
choosing the server is not.** A test that reaches for a fake
`DNSClient` to force a classification is still forbidden, no matter
how awkward the alternative looks.
Such a test must stay cheap. The resolver asks for eight record
types and retries each once, so a nameserver silent on every type
costs sixteen query timeouts. `TestQueryNameserverIP_Timeout` is
silent on `A` alone and answers the rest, which is all the resolver
needs to classify the response and keeps the test to two.
### What NOT to do
- **Do not mock `DNSClient`**, the watcher's `DNSResolver` interface,
or any other DNS abstraction — in any package, for any reason
- **Do not mock `DNSClient`** for resolver tests (the mock constructor
exists for unit-testing other packages that consume the resolver)
- **Do not add `-short` flags** to skip slow tests
- **Do not increase `-timeout`** to hide hanging queries
- **Do not remove `-count=1` from `script/test`** — Go's test cache
replays a previous run's output without querying anything, so a
cached pass is not evidence that live resolution works
- **Do not remove `-p 1` from `script/test`** — running the live-DNS
packages in parallel oversubscribes live DNS past what their gates
bound, which surfaces as unrelated resolver tests failing on
expired deadlines
- **Do not modify linter configuration** to suppress findings
+16 -33
View File
@@ -14,10 +14,7 @@ pre-1.0. No git tags. Core resolver work in flight on feature/resolver
(dirty: internal/resolver/resolver_test.go). Local checkout has diverged
from origin: origin/main is 8 commits ahead (watcher orchestrator,
unified TARGETS) and origin/feature/resolver already contains the full
iterative resolver implementation. DNS mocking is banned in this repo
(see `TESTING.md`): all tests use live DNS only. The hermetic mocked
tests previously noted on `feature/resolver` are gone from its current
tip, which carries a live-DNS suite against `*.dns.sneak.cloud`.
iterative resolver implementation with hermetic mocked tests.
# Next Step
@@ -26,22 +23,6 @@ Rationale, Design, TODO, License, Author) if any are still missing.
# Completed Steps
- 2026-09-03: restored the transport-failure classification coverage
the DNS-mock removal had dropped, and tightened two over-tolerant
watcher assertions. `internal/resolver/transport_test.go` covers
`StatusTimeout`, `StatusError` and the connection-refused path by
binding real nameservers on loopback rather than by mocking
`DNSClient` or by aiming a query at a remote address and hoping the
network black-holes it; `queryDNS` now honours a port already
present in a nameserver address, which is what lets a query be
aimed at one. The watcher's `assertStatePopulated` and
`TestDomainPortAndTLSChecks` now assert that port and certificate
state, and the arguments the port and TLS checkers were called
with, match the addresses live DNS returned — previously they
asserted only that those sets were non-empty, which would not have
caught resolving the wrong addresses. `TESTING.md` records why a
loopback nameserver is not a mock.
- 2026-08-10: comment-only corrections to `script/bootstrap`,
`script/cibuild`, and `Dockerfile.lint`. The `goimports` pin in
`script/bootstrap` was justified by a claim that `script/fmt-check`
@@ -111,10 +92,15 @@ Rationale, Design, TODO, License, Author) if any are still missing.
`make check` into `script/lint`. `golangci-lint config verify` is
deliberately omitted: it fetches its schema over an unpinned live
HTTPS call
- 2026-08-07: DNS mocking removed from the entire test suite; watcher
tests now drive the real iterative resolver against live DNS and
`TESTING.md` bans DNS mocks in every package (`remove-dns-mocking`
branch)
- 2026-08-09: in-flight notification deliveries are now drained at
shutdown (#106): `notify.New` registers an fx `OnStop` hook that waits
on a `sync.WaitGroup` of tracked delivery goroutines, bounded by the
`OnStop` context; on expiry the outstanding count is logged at warn
level and parked retry backoffs are released instead of being dropped
silently, and deliveries submitted after the drain begins are refused
so shutdown cannot be extended indefinitely; an `OnStop` context that
is already expired on entry with nothing outstanding drains quietly
rather than warning about deliveries that were never abandoned
- 2026-08-07: golangci-lint bumped to v2.12.2 (commit-pinned installs
in `Dockerfile` and `script/bootstrap`); `.golangci.yml` set to the
org-standard v2-schema config used across the org's repos
@@ -126,8 +112,7 @@ Rationale, Design, TODO, License, Author) if any are still missing.
- 2026-07-07 Adopted scripts-to-rule-them-all: `script/` entrypoints,
Makefile shims, README Entrypoints section
- 2026-02-20: iterative DNS resolver implemented; tests made hermetic
with mocked DNS (origin/feature/resolver, unmerged; superseded — DNS
mocking is banned, see `TESTING.md`)
with mocked DNS (origin/feature/resolver, unmerged)
- 2026-02-20: CI actions and go install refs pinned to commit SHAs;
Gitea Actions workflow for make check (origin/ci/make-check, unmerged)
- 2026-02-20: watcher monitoring orchestrator merged to main (#8)
@@ -152,9 +137,8 @@ Branch reconciliation:
- Sync local checkout with origin: local main is 8 commits behind
origin/main; local feature/resolver has diverged from
origin/feature/resolver, which already implements the resolver
- Merge in-flight branches to main once green: feature/resolver
(confirm its tests remain live-DNS — DNS mocking is banned, see
`TESTING.md`), ci/make-check, feature/portcheck-implementation,
- Merge in-flight branches to main once green: feature/resolver,
ci/make-check, feature/portcheck-implementation,
feature/tlscheck-implementation
Resolver (plan from untracked TODO.md; largely implemented on
@@ -238,7 +222,6 @@ Infrastructure notes (from untracked TODO.md):
- Module path sneak.berlin/go/dnswatcher differs from the git.eeqj.de
remote intentionally; do not "fix" it
- Dependencies: github.com/miekg/dns, golang.org/x/net/publicsuffix
- Resolver tests originally used live DNS against `*.dns.sneak.cloud`
(required records documented in the test file header); `main` now
tests against live public DNS. DNS mocking is banned (see
`TESTING.md`); never reintroduce hermetic mocked DNS tests
- Resolver tests originally used live DNS against *.dns.sneak.cloud
(required records documented in the test file header); origin now has
mocked hermetic tests, keep them hermetic
+21 -5
View File
@@ -32,11 +32,27 @@ func NewRequestForTest(
// NewTestService creates a Service suitable for unit testing.
// It discards log output and uses the given transport.
func NewTestService(transport http.RoundTripper) *Service {
return &Service{
log: slog.New(slog.DiscardHandler),
transport: transport,
history: NewAlertHistory(),
}
return newService(slog.New(slog.DiscardHandler), transport)
}
// NewTestServiceWithLogger creates a Service that writes to the
// given handler, so tests can assert on emitted log records.
func NewTestServiceWithLogger(
transport http.RoundTripper,
handler slog.Handler,
) *Service {
return newService(slog.New(handler), transport)
}
// Drain exports drain for testing.
func (svc *Service) Drain(ctx context.Context) {
svc.drain(ctx)
}
// OutstandingDeliveries reports how many delivery goroutines
// are currently tracked as in flight.
func (svc *Service) OutstandingDeliveries() int64 {
return svc.outstanding.Load()
}
// SetNtfyURL sets the ntfy URL on a Service for testing.
+73 -56
View File
@@ -12,6 +12,8 @@ import (
"log/slog"
"net/http"
"net/url"
"sync"
"sync/atomic"
"time"
"go.uber.org/fx"
@@ -115,19 +117,41 @@ type Service struct {
history *AlertHistory
retryConfig RetryConfig
sleepFn func(time.Duration) <-chan time.Time
// Shutdown draining state. drainMu guards draining and
// serialises it against the counter increment in
// startDelivery; inFlight tracks the delivery goroutines
// themselves and outstanding mirrors its count so a timed
// out drain can report how many were abandoned.
drainMu sync.Mutex
draining bool
inFlight sync.WaitGroup
outstanding atomic.Int64
abandon chan struct{}
abandonOnce sync.Once
}
// newService builds a Service with the fields every Service
// needs regardless of how it was constructed.
func newService(
log *slog.Logger,
transport http.RoundTripper,
) *Service {
return &Service{
log: log,
transport: transport,
history: NewAlertHistory(),
abandon: make(chan struct{}),
}
}
// New creates a new notify Service.
func New(
_ fx.Lifecycle,
lifecycle fx.Lifecycle,
params Params,
) (*Service, error) {
svc := &Service{
log: params.Logger.Get(),
transport: http.DefaultTransport,
config: params.Config,
history: NewAlertHistory(),
}
svc := newService(params.Logger.Get(), http.DefaultTransport)
svc.config = params.Config
if params.Config.NtfyTopic != "" {
u, err := ValidateWebhookURL(
@@ -168,6 +192,14 @@ func New(
svc.mattermostWebhookURL = u
}
lifecycle.Append(fx.Hook{
OnStop: func(ctx context.Context) error {
svc.drain(ctx)
return nil
},
})
return svc, nil
}
@@ -194,6 +226,32 @@ func (svc *Service) SendNotification(
svc.dispatchMattermost(ctx, title, message, priority)
}
// dispatch delivers a notification to one endpoint on a
// tracked background goroutine.
//
// The delivery context is detached from ctx with
// context.WithoutCancel so that a cancelled caller does not
// kill a delivery already under way; the shutdown drain, not
// the caller, decides how long deliveries may keep running.
func (svc *Service) dispatch(
ctx context.Context,
endpoint string,
send func(context.Context) error,
) {
notifyCtx := context.WithoutCancel(ctx)
svc.startDelivery(endpoint, func() {
err := svc.deliverWithRetry(notifyCtx, endpoint, send)
if err != nil {
svc.log.Error(
"failed to send notification after retries",
"endpoint", endpoint,
"error", err,
)
}
})
}
func (svc *Service) dispatchNtfy(
ctx context.Context,
title, message, priority string,
@@ -202,26 +260,11 @@ func (svc *Service) dispatchNtfy(
return
}
go func() {
notifyCtx := context.WithoutCancel(ctx)
err := svc.deliverWithRetry(
notifyCtx, "ntfy",
func(c context.Context) error {
svc.dispatch(ctx, "ntfy", func(c context.Context) error {
return svc.sendNtfy(
c, svc.ntfyURL,
title, message, priority,
c, svc.ntfyURL, title, message, priority,
)
},
)
if err != nil {
svc.log.Error(
"failed to send ntfy notification "+
"after retries",
"error", err,
)
}
}()
})
}
func (svc *Service) dispatchSlack(
@@ -232,26 +275,11 @@ func (svc *Service) dispatchSlack(
return
}
go func() {
notifyCtx := context.WithoutCancel(ctx)
err := svc.deliverWithRetry(
notifyCtx, "slack",
func(c context.Context) error {
svc.dispatch(ctx, "slack", func(c context.Context) error {
return svc.sendSlack(
c, svc.slackWebhookURL,
title, message, priority,
c, svc.slackWebhookURL, title, message, priority,
)
},
)
if err != nil {
svc.log.Error(
"failed to send slack notification "+
"after retries",
"error", err,
)
}
}()
})
}
func (svc *Service) dispatchMattermost(
@@ -262,11 +290,8 @@ func (svc *Service) dispatchMattermost(
return
}
go func() {
notifyCtx := context.WithoutCancel(ctx)
err := svc.deliverWithRetry(
notifyCtx, "mattermost",
svc.dispatch(
ctx, "mattermost",
func(c context.Context) error {
return svc.sendSlack(
c, svc.mattermostWebhookURL,
@@ -274,14 +299,6 @@ func (svc *Service) dispatchMattermost(
)
},
)
if err != nil {
svc.log.Error(
"failed to send mattermost notification "+
"after retries",
"error", err,
)
}
}()
}
func (svc *Service) sendNtfy(
+9
View File
@@ -2,6 +2,7 @@ package notify
import (
"context"
"fmt"
"math"
"math/rand/v2"
"time"
@@ -121,6 +122,14 @@ func (svc *Service) deliverWithRetry(
select {
case <-ctx.Done():
return ctx.Err()
case <-svc.abandon:
// Shutdown drained past its deadline; stop
// sleeping rather than outlive the process.
// A nil channel (Service built without a
// constructor) simply never fires.
return fmt.Errorf(
"%w: %s", ErrDeliveryAbandoned, endpoint,
)
case <-svc.sleepFunc(delay):
}
}
+119
View File
@@ -0,0 +1,119 @@
package notify
import (
"context"
"errors"
)
// ErrDeliveryAbandoned is returned by a retry loop that was
// cut short because shutdown drained past its deadline.
var ErrDeliveryAbandoned = errors.New(
"notification delivery abandoned at shutdown",
)
// startDelivery runs fn on its own goroutine while tracking it,
// so that drain can wait for it during shutdown.
//
// The WaitGroup counter is incremented here, on the caller's
// goroutine, before the worker exists: incrementing it inside
// the worker would race with drain's Wait and could let
// shutdown sail past a delivery that had not started yet.
//
// Once draining has begun the delivery is refused outright
// rather than queued, so a steady stream of newly submitted
// notifications cannot keep extending the drain.
func (svc *Service) startDelivery(endpoint string, fn func()) {
svc.drainMu.Lock()
if svc.draining {
svc.drainMu.Unlock()
svc.log.Warn(
"notification not dispatched: shutdown in progress",
"endpoint", endpoint,
)
return
}
svc.outstanding.Add(1)
// WaitGroup.Go increments the counter synchronously, here,
// and only then starts the goroutine.
svc.inFlight.Go(func() {
// Runs before the WaitGroup counter is decremented, so
// a drain that times out reports an accurate count.
defer svc.outstanding.Add(-1)
fn()
})
svc.drainMu.Unlock()
}
// drain waits for in-flight notification deliveries to finish.
//
// It first stops accepting new deliveries, then waits until
// either every outstanding delivery has completed or ctx
// expires — whichever comes first. ctx is the context fx
// passes to the OnStop hook, so a permanently dead webhook
// cannot hang shutdown indefinitely.
//
// When the deadline arrives with deliveries still outstanding,
// the count is logged at warn level and the abandon channel is
// closed, which releases any retry loop sleeping in backoff.
// Deliveries already inside an HTTP round trip are bounded by
// the existing httpClientTimeout instead.
//
// A ctx that is already expired on entry is not by itself cause
// for alarm: if nothing is outstanding there is nothing to
// abandon, and the drain says so at debug level rather than
// warning about deliveries that do not exist.
func (svc *Service) drain(ctx context.Context) {
svc.drainMu.Lock()
svc.draining = true
svc.drainMu.Unlock()
done := make(chan struct{})
go func() {
svc.inFlight.Wait()
close(done)
}()
select {
case <-done:
svc.log.Debug(
"all in-flight notifications completed",
)
case <-ctx.Done():
// outstanding is decremented before the WaitGroup
// counter, and startDelivery can no longer add to it
// now that draining is set, so a zero here means every
// delivery really did finish. ctx expiring in that
// state (an OnStop context that was already cancelled
// on entry is the usual way) abandons nothing, so it
// must not close abandon or warn about it.
abandoned := svc.outstanding.Load()
if abandoned == 0 {
svc.log.Debug(
"all in-flight notifications completed",
)
return
}
svc.abandonOnce.Do(func() {
if svc.abandon != nil {
close(svc.abandon)
}
})
svc.log.Warn(
"shutdown deadline reached with notifications "+
"still in flight; abandoning them",
"abandoned", abandoned,
"error", ctx.Err(),
)
}
}
+531
View File
@@ -0,0 +1,531 @@
package notify_test
import (
"bytes"
"context"
"log/slog"
"net/http"
"net/http/httptest"
"net/url"
"strings"
"sync"
"sync/atomic"
"testing"
"time"
"go.uber.org/fx"
"sneak.berlin/go/dnswatcher/internal/config"
"sneak.berlin/go/dnswatcher/internal/globals"
"sneak.berlin/go/dnswatcher/internal/logger"
"sneak.berlin/go/dnswatcher/internal/notify"
)
// Timings used by the drain tests. They stay in the same
// 10-100ms band as the retry tests so the suite never waits on
// a real backoff delay.
const (
// inFlightHold is how long a delivery is kept mid-request
// before the handler is released.
inFlightHold = 30 * time.Millisecond
// drainDeadline bounds a drain that is expected to time
// out.
drainDeadline = 50 * time.Millisecond
// drainSlack is the upper bound on how long a bounded
// drain may take; generous enough for a loaded CI box,
// still far below the 20s test ceiling.
drainSlack = 2 * time.Second
// settleDelay is how long to wait before asserting that
// something did *not* happen.
settleDelay = 50 * time.Millisecond
// idleDrainBound is the upper bound on a drain that has
// nothing in flight. It is deliberately far above the cost
// of the goroutine hop through inFlight.Wait() — which
// reached 57ms on a loaded box under -race with the package's
// parallel tests — and far below drainSlack, the deadline
// such a drain is given. A drain that blocked until its
// deadline instead of returning on the WaitGroup therefore
// still fails this bound, but scheduling delay alone cannot.
idleDrainBound = 500 * time.Millisecond
)
// syncBuffer is an io.Writer safe for concurrent use, so log
// output written from delivery goroutines can be inspected.
type syncBuffer struct {
mu sync.Mutex
buf bytes.Buffer
}
func (sb *syncBuffer) Write(p []byte) (int, error) {
sb.mu.Lock()
defer sb.mu.Unlock()
return sb.buf.Write(p) //nolint:wrapcheck // test helper
}
func (sb *syncBuffer) String() string {
sb.mu.Lock()
defer sb.mu.Unlock()
return sb.buf.String()
}
// newLoggingService returns a Service writing JSON logs into
// the returned buffer.
func newLoggingService(
transport http.RoundTripper,
) (*notify.Service, *syncBuffer) {
logs := &syncBuffer{}
handler := slog.NewJSONHandler(logs, nil)
return notify.NewTestServiceWithLogger(transport, handler),
logs
}
// blockingNtfyServer returns a server whose handler signals on
// entered, waits for release, and then responds 200.
func blockingNtfyServer(
entered chan<- struct{},
release <-chan struct{},
served *atomic.Bool,
) *httptest.Server {
var once sync.Once
return httptest.NewServer(
http.HandlerFunc(
func(w http.ResponseWriter, _ *http.Request) {
once.Do(func() { close(entered) })
<-release
served.Store(true)
w.WriteHeader(http.StatusOK)
}),
)
}
// TestDrainWaitsForInFlightDelivery verifies that a delivery
// already under way when shutdown starts is allowed to finish.
func TestDrainWaitsForInFlightDelivery(t *testing.T) {
t.Parallel()
var served atomic.Bool
entered := make(chan struct{})
release := make(chan struct{})
srv := blockingNtfyServer(entered, release, &served)
defer srv.Close()
topicURL, _ := url.Parse(srv.URL)
svc := notify.NewTestService(http.DefaultTransport)
svc.SetNtfyURL(topicURL)
svc.SendNotification(
context.Background(), "t", "m", prioInfo,
)
// Make sure the delivery really is mid-request before the
// drain begins.
select {
case <-entered:
case <-time.After(drainSlack):
t.Fatal("delivery never reached the endpoint")
}
// As in TestDrainBoundedByContextDeadline: start is captured
// before the clock it is compared against, here the timer
// holding the delivery open, so elapsed covers the whole hold
// and the lower bound cannot come out short from scheduling
// delay alone.
start := time.Now()
timer := time.AfterFunc(inFlightHold, func() {
close(release)
})
defer timer.Stop()
ctx, cancel := context.WithTimeout(
context.Background(), drainSlack,
)
defer cancel()
svc.Drain(ctx)
elapsed := time.Since(start)
if !served.Load() {
t.Error(
"drain returned before the in-flight delivery " +
"completed",
)
}
if elapsed < inFlightHold {
t.Errorf(
"drain took %v, want at least %v",
elapsed, inFlightHold,
)
}
if got := svc.OutstandingDeliveries(); got != 0 {
t.Errorf("outstanding deliveries = %d, want 0", got)
}
}
// neverFires returns a channel that never delivers, standing in
// for a long backoff sleep without actually sleeping.
func neverFires(_ time.Duration) <-chan time.Time {
return make(chan time.Time)
}
// TestDrainBoundedByContextDeadline verifies that a delivery
// stuck retrying against a dead endpoint does not hold shutdown
// past the OnStop context deadline, and that the abandoned
// deliveries are logged at warn level rather than dropped
// silently.
func TestDrainBoundedByContextDeadline(t *testing.T) {
t.Parallel()
var requests atomic.Int64
srv := httptest.NewServer(
http.HandlerFunc(
func(w http.ResponseWriter, _ *http.Request) {
requests.Add(1)
w.WriteHeader(http.StatusInternalServerError)
}),
)
defer srv.Close()
topicURL, _ := url.Parse(srv.URL)
svc, logs := newLoggingService(http.DefaultTransport)
svc.SetNtfyURL(topicURL)
// Never let the backoff sleep complete: the delivery is
// parked in its retry wait until shutdown releases it.
svc.SetSleepFunc(neverFires)
svc.SetRetryConfig(notify.RetryConfig{
MaxRetries: 5,
BaseDelay: time.Hour,
MaxDelay: time.Hour,
})
svc.SendNotification(
context.Background(), "t", "m", prioError,
)
waitForCondition(t, func() bool {
return requests.Load() >= 1 &&
svc.OutstandingDeliveries() == 1
})
// start must be captured *before* the deadline clock starts,
// so that the measured interval is a superset of the deadline
// interval. Capturing it after context.WithTimeout would
// make elapsed structurally smaller than drainDeadline and
// the lower bound below unfalsifiable-by-luck: it would fail
// whenever the two statements were separated by any
// scheduling delay, and pass otherwise, regardless of what
// the drain did.
start := time.Now()
ctx, cancel := context.WithTimeout(
context.Background(), drainDeadline,
)
defer cancel()
// The upper bound is enforced by a watchdog rather than by
// measuring after the fact: a drain that is not bounded at
// all never returns here (the delivery is parked in a backoff
// that never fires), so an unbounded drain must fail this
// test promptly instead of hanging the package until the test
// binary's 30s timeout.
returned := make(chan struct{})
go func() {
defer close(returned)
svc.Drain(ctx)
}()
select {
case <-returned:
case <-time.After(drainSlack):
t.Fatalf(
"drain did not return within %v; its %v deadline "+
"did not bound it",
drainSlack, drainDeadline,
)
}
// The lower bound is the real assertion: the drain must have
// waited for its whole deadline rather than giving up on the
// outstanding delivery early. With start captured above, an
// early return is the only thing that can make it fail.
if elapsed := time.Since(start); elapsed < drainDeadline {
t.Errorf(
"drain returned after %v, before its %v deadline",
elapsed, drainDeadline,
)
}
assertAbandonLogged(t, logs.String())
// The abandoned delivery must stop retrying rather than
// outlive the drain.
waitForCondition(t, func() bool {
return svc.OutstandingDeliveries() == 0
})
}
// assertAbandonLogged checks that the drain logged the
// abandoned deliveries at warn level with a count.
func assertAbandonLogged(t *testing.T, output string) {
t.Helper()
if !strings.Contains(output, `"level":"WARN"`) {
t.Errorf(
"abandoned deliveries not logged at warn level; "+
"log output: %s",
output,
)
}
if !strings.Contains(output, `"abandoned":1`) {
t.Errorf(
"abandoned delivery count not logged; "+
"log output: %s",
output,
)
}
}
// TestDrainRefusesNewDeliveries verifies that notifications
// submitted after the drain has begun are refused and logged,
// so a stream of new work cannot extend shutdown indefinitely.
func TestDrainRefusesNewDeliveries(t *testing.T) {
t.Parallel()
var requests atomic.Int64
srv := httptest.NewServer(
http.HandlerFunc(
func(w http.ResponseWriter, _ *http.Request) {
requests.Add(1)
w.WriteHeader(http.StatusOK)
}),
)
defer srv.Close()
target, _ := url.Parse(srv.URL)
svc, logs := newLoggingService(http.DefaultTransport)
svc.SetNtfyURL(target)
svc.SetSlackWebhookURL(target)
svc.SetMattermostWebhookURL(target)
ctx, cancel := context.WithTimeout(
context.Background(), drainSlack,
)
defer cancel()
// Nothing is in flight, so this returns immediately and
// leaves the service refusing further deliveries.
svc.Drain(ctx)
for range 3 {
svc.SendNotification(
context.Background(), "t", "m", prioInfo,
)
}
time.Sleep(settleDelay)
if got := requests.Load(); got != 0 {
t.Errorf(
"%d requests reached the endpoint after drain, "+
"want 0",
got,
)
}
if got := svc.OutstandingDeliveries(); got != 0 {
t.Errorf("outstanding deliveries = %d, want 0", got)
}
output := logs.String()
if !strings.Contains(output, "shutdown in progress") {
t.Errorf(
"refused deliveries not logged; log output: %s",
output,
)
}
}
// recordingLifecycle is a minimal fx.Lifecycle that records the
// hooks appended to it, so the wiring done by notify.New can be
// inspected without standing up a whole fx application.
type recordingLifecycle struct {
hooks []fx.Hook
}
func (l *recordingLifecycle) Append(hook fx.Hook) {
l.hooks = append(l.hooks, hook)
}
// newNotifyService builds a Service through the real
// constructor, wired to the given lifecycle.
func newNotifyService(
t *testing.T,
lifecycle fx.Lifecycle,
ntfyTopic string,
) *notify.Service {
t.Helper()
g, err := globals.New(nil)
if err != nil {
t.Fatalf("globals.New: %v", err)
}
log, err := logger.New(nil, logger.Params{Globals: g})
if err != nil {
t.Fatalf("logger.New: %v", err)
}
svc, err := notify.New(lifecycle, notify.Params{
Logger: log,
Config: &config.Config{NtfyTopic: ntfyTopic},
})
if err != nil {
t.Fatalf("notify.New: %v", err)
}
return svc
}
// TestNewRegistersDrainingStopHook verifies that notify.New
// wires an OnStop hook into the fx lifecycle and that the hook
// waits for in-flight deliveries.
func TestNewRegistersDrainingStopHook(t *testing.T) {
t.Parallel()
var served atomic.Bool
entered := make(chan struct{})
release := make(chan struct{})
srv := blockingNtfyServer(entered, release, &served)
defer srv.Close()
lifecycle := &recordingLifecycle{}
svc := newNotifyService(t, lifecycle, srv.URL)
if len(lifecycle.hooks) != 1 {
t.Fatalf(
"appended %d lifecycle hooks, want 1",
len(lifecycle.hooks),
)
}
stop := lifecycle.hooks[0].OnStop
if stop == nil {
t.Fatal("lifecycle hook has no OnStop function")
}
svc.SendNotification(
context.Background(), "t", "m", prioInfo,
)
select {
case <-entered:
case <-time.After(drainSlack):
t.Fatal("delivery never reached the endpoint")
}
timer := time.AfterFunc(inFlightHold, func() {
close(release)
})
defer timer.Stop()
ctx, cancel := context.WithTimeout(
context.Background(), drainSlack,
)
defer cancel()
err := stop(ctx)
if err != nil {
t.Fatalf("OnStop returned error: %v", err)
}
if !served.Load() {
t.Error(
"OnStop returned before the in-flight delivery " +
"completed",
)
}
}
// TestDrainWithoutDeliveriesReturnsImmediately verifies the
// common case: nothing in flight, shutdown is not delayed.
func TestDrainWithoutDeliveriesReturnsImmediately(t *testing.T) {
t.Parallel()
svc := notify.NewTestService(http.DefaultTransport)
// Captured before the deadline clock, as elsewhere in this
// file; for an upper bound that is the conservative
// direction, since the measured interval can then only be
// longer than the drain itself.
start := time.Now()
ctx, cancel := context.WithTimeout(
context.Background(), drainSlack,
)
defer cancel()
svc.Drain(ctx)
if elapsed := time.Since(start); elapsed > idleDrainBound {
t.Errorf(
"drain of an idle service took %v, want well "+
"under its %v deadline",
elapsed, drainSlack,
)
}
}
// TestDrainWithCancelledContextDoesNotWarn verifies that an
// OnStop context that is already dead on entry does not produce
// an "abandoning them" warning when there was nothing in flight
// to abandon. The expired context wins the select immediately,
// so only the outstanding count can tell the difference between
// a genuine timeout and a shutdown that had simply already run
// out of time with no work left.
func TestDrainWithCancelledContextDoesNotWarn(t *testing.T) {
t.Parallel()
svc, logs := newLoggingService(http.DefaultTransport)
ctx, cancel := context.WithCancel(context.Background())
cancel()
svc.Drain(ctx)
if output := logs.String(); strings.Contains(
output, `"level":"WARN"`,
) {
t.Errorf(
"drain with nothing in flight warned about "+
"abandoned deliveries; log output: %s",
output,
)
}
}
+2 -2
View File
@@ -7,8 +7,8 @@ import (
"github.com/miekg/dns"
)
// DNSClient abstracts DNS wire-protocol exchanges over a single
// transport, letting the resolver switch between UDP and TCP.
// DNSClient abstracts DNS wire-protocol exchanges so the resolver
// can be tested without hitting real nameservers.
type DNSClient interface {
ExchangeContext(
ctx context.Context,
+1 -19
View File
@@ -20,10 +20,6 @@ const (
minDomainLabels = 2
)
// defaultDNSPort is the port a nameserver is assumed to listen on
// when its address does not carry one.
const defaultDNSPort = "53"
// ErrRefused is returned when a DNS server refuses a query.
var ErrRefused = errors.New("dns query refused")
@@ -110,20 +106,6 @@ func (r *Resolver) retryTCP(
return resp
}
// nameserverAddr renders a nameserver address for dialling. A bare
// address — the normal case, and what a delegation's glue records
// carry — is given the default DNS port. An address that already
// specifies a port is dialled as written, which is what makes a
// nameserver listening somewhere other than 53 reachable.
func nameserverAddr(nsIP string) string {
_, _, err := net.SplitHostPort(nsIP)
if err == nil {
return nsIP
}
return net.JoinHostPort(nsIP, defaultDNSPort)
}
// queryDNS sends a DNS query to a specific server IP.
// Tries non-recursive first, falls back to recursive on
// REFUSED (handles DNS interception environments).
@@ -138,7 +120,7 @@ func (r *Resolver) queryDNS(
}
name = dns.Fqdn(name)
addr := nameserverAddr(serverIP)
addr := net.JoinHostPort(serverIP, "53")
msg := new(dns.Msg)
msg.SetQuestion(name, qtype)
+13
View File
@@ -67,4 +67,17 @@ func NewFromLogger(log *slog.Logger) *Resolver {
}
}
// NewFromLoggerWithClient creates a Resolver with a custom DNS
// client, useful for testing with mock DNS responses.
func NewFromLoggerWithClient(
log *slog.Logger,
client DNSClient,
) *Resolver {
return &Resolver{
log: log,
client: client,
tcp: client,
}
}
// Method implementations are in iterative.go.
+54 -3
View File
@@ -8,7 +8,9 @@ import (
"sort"
"strings"
"testing"
"time"
"github.com/miekg/dns"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
@@ -517,9 +519,58 @@ func TestQueryAllNameservers_ContextCanceled(t *testing.T) {
assert.Error(t, err)
}
// Transport-failure classification (StatusTimeout / StatusError)
// is covered in transport_test.go, against real nameservers bound on
// loopback.
// ----------------------------------------------------------------
// Timeout tests
// ----------------------------------------------------------------
func TestQueryNameserverIP_Timeout(t *testing.T) {
t.Parallel()
log := slog.New(slog.NewTextHandler(
os.Stderr,
&slog.HandlerOptions{Level: slog.LevelDebug},
))
r := resolver.NewFromLoggerWithClient(
log, &timeoutClient{},
)
ctx, cancel := context.WithTimeout(
context.Background(), 10*time.Second,
)
t.Cleanup(cancel)
// Query any IP — the client always returns a timeout error.
resp, err := r.QueryNameserverIP(
ctx, "unreachable.test.", "192.0.2.1",
"example.com",
)
require.NoError(t, err)
assert.Equal(t, resolver.StatusTimeout, resp.Status)
assert.NotEmpty(t, resp.Error)
}
// timeoutClient simulates DNS timeout errors for testing.
type timeoutClient struct{}
func (c *timeoutClient) ExchangeContext(
_ context.Context,
_ *dns.Msg,
_ string,
) (*dns.Msg, time.Duration, error) {
return nil, 0, &net.OpError{
Op: "read",
Net: "udp",
Err: &timeoutError{},
}
}
type timeoutError struct{}
func (e *timeoutError) Error() string { return "i/o timeout" }
func (e *timeoutError) Timeout() bool { return true }
func (e *timeoutError) Temporary() bool { return true }
func TestResolveIPAddresses_ContextCanceled(t *testing.T) {
t.Parallel()
-259
View File
@@ -1,259 +0,0 @@
package resolver_test
// Transport-failure classification tests.
//
// These are live tests, not mocks. Nothing here substitutes the
// resolver's DNSClient: the resolver dials a real UDP socket, writes
// a real DNS query with the real miekg/dns client, and applies its
// real deadline and its real classification logic to what comes
// back. The only thing under test control is which address the query
// is sent to, and what — if anything — is listening there.
//
// That distinction is what the no-DNS-mocks rule in TESTING.md is
// about. A fake DNSClient lets the code under test skip DNS entirely
// and hands it a manufactured verdict; a nameserver bound on
// loopback makes it speak DNS for real and earn one. Pointing a live
// query at a nameserver of the test's choosing is no more a mock
// than pointing it at a.root-servers.net.
//
// The public network cannot produce these outcomes on demand. A
// black-holed address is not black-holed everywhere — build
// environments that intercept UDP/53 answer it locally — so a test
// built on one asserts on the network it happens to run on rather
// than on the resolver. A loopback nameserver is deterministic
// everywhere, and it is fast, because the test picks the deadline.
import (
"context"
"net"
"testing"
"time"
"github.com/miekg/dns"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
"sneak.berlin/go/dnswatcher/internal/resolver"
)
const (
// transportBudget is the wall time the timeout test must stay
// under. The resolver asks a nameserver for eight record types
// in turn and retries each one once, so a nameserver silent on
// every type would cost sixteen query timeouts. The test's
// nameserver is silent on exactly one type, which costs two,
// and this budget fails loudly if that ever stops being true.
transportBudget = 8 * time.Second
// transportDeadline is the caller deadline the tests run
// under. It is generous on purpose: these tests are about the
// resolver classifying a nameserver's behaviour, so the
// caller's deadline must never be the thing that expires.
transportDeadline = 30 * time.Second
// silentNS and failingNS are the nameserver names reported
// back in NameserverResponse.Nameserver. They are .test names
// (RFC 6761) and are never resolved: the tests address the
// nameserver by its socket address.
silentNS = "silent.ns.test."
failingNS = "servfail.ns.test."
// transportHostname is the name queried. Nothing resolves it;
// the point is entirely how the nameserver behaves.
transportHostname = "example.com"
)
// startNameserver binds a real UDP nameserver on loopback and serves
// every datagram it receives with handle, which returns the reply to
// send or nil to stay silent. It returns the "host:port" address to
// aim a query at, and stops the server when the test ends.
func startNameserver(
t *testing.T,
handle func(query *dns.Msg) *dns.Msg,
) string {
t.Helper()
var lc net.ListenConfig
conn, err := lc.ListenPacket(t.Context(), "udp", "127.0.0.1:0")
require.NoError(t, err, "binding loopback nameserver")
stopped := make(chan struct{})
t.Cleanup(func() {
_ = conn.Close()
<-stopped
})
go serveNameserver(conn, handle, stopped)
return conn.LocalAddr().String()
}
// serveNameserver reads queries until conn is closed, replying with
// whatever handle produces.
func serveNameserver(
conn net.PacketConn,
handle func(query *dns.Msg) *dns.Msg,
stopped chan<- struct{},
) {
defer close(stopped)
buf := make([]byte, dns.MaxMsgSize)
for {
n, from, err := conn.ReadFrom(buf)
if err != nil {
return
}
query := new(dns.Msg)
if query.Unpack(buf[:n]) != nil {
continue
}
reply := handle(query)
if reply == nil {
continue
}
wire, err := reply.Pack()
if err != nil {
continue
}
_, err = conn.WriteTo(wire, from)
if err != nil {
return
}
}
}
// unservedAddr returns a loopback address with nothing listening on
// it, by binding a port and releasing it again.
func unservedAddr(t *testing.T) string {
t.Helper()
var lc net.ListenConfig
conn, err := lc.ListenPacket(t.Context(), "udp", "127.0.0.1:0")
require.NoError(t, err, "binding loopback port")
addr := conn.LocalAddr().String()
require.NoError(t, conn.Close(), "releasing loopback port")
return addr
}
// TestQueryNameserverIP_Timeout covers the StatusTimeout branch: a
// nameserver that takes the query and never answers.
func TestQueryNameserverIP_Timeout(t *testing.T) {
t.Parallel()
// A real nameserver that drops A queries and answers every
// other type. Silence on one type is all the resolver needs to
// classify the response as a timeout, and it keeps the test
// two query timeouts long instead of sixteen.
addr := startNameserver(t, func(query *dns.Msg) *dns.Msg {
if len(query.Question) > 0 &&
query.Question[0].Qtype == dns.TypeA {
return nil
}
reply := new(dns.Msg)
reply.SetReply(query)
return reply
})
r := newTestResolver(t)
ctx, cancel := context.WithTimeout(
context.Background(), transportDeadline,
)
defer cancel()
start := time.Now()
resp, err := r.QueryNameserverIP(
ctx, silentNS, addr, transportHostname,
)
elapsed := time.Since(start)
require.NoError(t, err)
require.NotNil(t, resp)
assert.Equal(t, resolver.StatusTimeout, resp.Status)
assert.Equal(t, "all queries timed out", resp.Error)
assert.Empty(t, resp.Records)
assert.Equal(t, silentNS, resp.Nameserver)
assert.Less(
t, elapsed, transportBudget,
"one silent record type must cost one query's retries, "+
"not every record type's",
)
}
// TestQueryNameserverIP_ServFail covers the StatusError branch: a
// nameserver that answers, and answers SERVFAIL.
func TestQueryNameserverIP_ServFail(t *testing.T) {
t.Parallel()
addr := startNameserver(t, func(query *dns.Msg) *dns.Msg {
reply := new(dns.Msg)
reply.SetRcode(query, dns.RcodeServerFailure)
return reply
})
r := newTestResolver(t)
ctx, cancel := context.WithTimeout(
context.Background(), transportDeadline,
)
defer cancel()
resp, err := r.QueryNameserverIP(
ctx, failingNS, addr, transportHostname,
)
require.NoError(t, err)
require.NotNil(t, resp)
assert.Equal(t, resolver.StatusError, resp.Status)
assert.Equal(t, "server returned SERVFAIL", resp.Error)
assert.Empty(t, resp.Records)
assert.Equal(t, failingNS, resp.Nameserver)
}
// TestQueryNameserverIP_NoListener pins the third transport outcome:
// a refused datagram is not a timeout. The socket fails immediately
// with ECONNREFUSED rather than going quiet, so isTimeout is false,
// no failure flag is set, and the response classifies as NoData.
// Asserting it here is what stops that path being mistaken for the
// timeout path, in either direction.
func TestQueryNameserverIP_NoListener(t *testing.T) {
t.Parallel()
addr := unservedAddr(t)
r := newTestResolver(t)
ctx, cancel := context.WithTimeout(
context.Background(), transportDeadline,
)
defer cancel()
resp, err := r.QueryNameserverIP(
ctx, silentNS, addr, transportHostname,
)
require.NoError(t, err)
require.NotNil(t, resp)
assert.Equal(t, resolver.StatusNoData, resp.Status)
assert.Empty(t, resp.Records)
}
File diff suppressed because it is too large Load Diff
+2 -13
View File
@@ -19,26 +19,15 @@
#
# -timeout 90s is a deliberate backstop above the 60s hard cap on
# suite duration. Do not lower it.
#
# -p 1 runs one test package at a time, and is load-bearing. Live DNS
# is a resource outside the process: the concurrency gates that keep
# this suite from bursting at the root servers
# (internal/resolver/livedns_test.go, internal/watcher/watcher_test.go)
# are package-scoped, so each one only bounds its own test binary. Go
# runs package binaries in parallel by default, so with both live-DNS
# packages in flight at once their gates sum instead of holding, the
# root and TLD servers rate-limit the excess, and the resolver
# package's per-attempt deadlines expire. Serialising packages is what
# makes each gate authoritative while its package runs.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
main() {
cd "$ROOT"
go test -count=1 -p 1 -race -timeout 90s -cover ./... || {
go test -count=1 -race -timeout 90s -cover ./... || {
echo "--- Rerunning with -v for details ---" >&2
go test -count=1 -p 1 -race -timeout 90s -v ./... || true
go test -count=1 -race -timeout 90s -v ./... || true
exit 1
}
}