Long-lived integration branch. One commit per work unit lands here; this PR accumulates them until it is merged to `main`. ## Landed units - **Run all linting in Docker via `Dockerfile.lint` + `script/lint`** — #134 golangci-lint is no longer installed or run on the host. New root `Dockerfile.lint` COPYs the repo into the digest-pinned `golangci/golangci-lint:v2.12.2` image and lints as a build step, so a successful build IS a clean lint; `script/lint` is reduced to a thin wrapper that builds it. This also works where the docker daemon is remote and bind mounts are impossible. **Pinned digest and how it was verified.** `golangci/golangci-lint:v2.12.2@sha256:5cceeef04e53efe1470638d4b4b4f5ceefd574955ab3941b2d9a68a8c9ad5240`, exactly as quoted in the issue. It resolves, and it is genuinely v2.12.2: ``` $ docker buildx imagetools inspect golangci/golangci-lint:v2.12.2 Name: docker.io/golangci/golangci-lint:v2.12.2 MediaType: application/vnd.oci.image.index.v1+json Digest: sha256:5cceeef04e53efe1470638d4b4b4f5ceefd574955ab3941b2d9a68a8c9ad5240 $ docker run --rm golangci/golangci-lint@sha256:5cceeef04e...ad5240 golangci-lint --version golangci-lint has version 2.12.2 built with go1.26.2 from c0d3ddc9 on 2026-05-06T11:07:58Z ``` The tag's index digest is the quoted digest, and the binary inside reports commit `c0d3ddc9`, matching the org's canonical pin `c0d3ddc9cf3faa61a4e378e879ece580256d76e5`. **Forcing the linter to actually run.** `Dockerfile.lint` is split into a `deps` stage (base image + `go mod download`) and a `lint` stage (source copy + linter run). `script/lint` runs: ``` docker build --progress=plain --no-cache-filter=lint --target lint -f Dockerfile.lint . ``` Caching is explicitly waived for linting, and a cached build lints nothing, so the `lint` stage is invalidated on every invocation. The invalidation is scoped: the `deps` stage stays cached and no global cache wipe is performed. `--progress=plain` keeps the linter's own output visible. **`golangci-lint config verify`: deliberately NOT included.** It fetches its JSON schema over a live, unpinned HTTPS call, which would make linting network-dependent and defeat hash-pinning. Omitted for that reason, and the reason is recorded in a comment at the top of `Dockerfile.lint`. **`script/bootstrap`** no longer installs golangci-lint (and its pinned ref is gone); it warns non-fatally when `docker` is absent instead. The `goimports` install stays, because `script/fmt` and `script/fmt-check` still run on the host. Header comment updated accordingly. **Root `Dockerfile`** — required consequence, not scope creep. Its builder stage ran `make check`, which now calls `script/lint`, which shells out to `docker build`; there is no docker daemon inside a docker build, so `script/cibuild` and `script/docker` would have broken. It gains its own lint stage on the same pinned image (linter invoked directly, with a comment explaining why not `make lint`), with the builder stage depending on it via `COPY --from=lint /src/go.sum /dev/null` and running `make fmt-check`, `make test`, `make build`. The now-unneeded golangci-lint install is gone from the builder stage. **README** `Entrypoints` and `Building` sections now describe linting as a docker-only operation. `TODO.md` updated in the same commit. ### Verification All runs via `make` / `script/` entrypoints only. Two consecutive `make lint` runs on an unchanged tree, both executing the linter: ``` # run 1 #10 [lint 2/2] RUN golangci-lint run --config .golangci.yml ./... #10 10.98 0 issues. #10 DONE 12.0s # run 2, tree untouched #10 [lint 2/2] RUN golangci-lint run --config .golangci.yml ./... #10 14.63 0 issues. ``` A third run shows the cache scoping is working as intended — `deps` served from cache, `lint` re-executed: ``` #6 [deps 2/4] WORKDIR /src #6 CACHED #7 [deps 3/4] COPY go.mod go.sum ./ #7 CACHED #8 [deps 4/4] RUN go mod download #8 CACHED #9 [lint 1/2] COPY . . #10 [lint 2/2] RUN golangci-lint run --config .golangci.yml ./... ``` **Negative control.** A deliberate violation (an unused function containing an ineffectual assignment) was added to `internal/config/config.go`: ``` #10 11.26 internal/config/config.go:29:2: ineffectual assignment to x (ineffassign) #10 11.26 internal/config/config.go:28:6: func negativeControlUnused is unused (unused) #10 11.26 2 issues: #10 ERROR: process "/bin/sh -c golangci-lint run --config .golangci.yml ./..." did not complete successfully: exit code: 1 ERROR: failed to build: failed to solve: process "/bin/sh -c golangci-lint run --config .golangci.yml ./..." did not complete successfully: exit code: 1 make: *** [Makefile:26: lint] Error 1 ``` `make lint` exited non-zero naming both findings and their exact lines. After reverting the file, `make lint` was clean again (`0 issues.`). **`make check`** green end to end (test, lint, fmt-check), exit 0. **`script/cibuild`** green, confirming the `Dockerfile` restructure does not recurse: the lint stage ran (`#16 12.56 0 issues.`), then `#22 [builder 8/9] RUN make test` with `PASS` lines, then `#23 [builder 9/9] RUN make build`. ### Notes for the owner This supersedes two PRs you still have queued for merge, both of which tune host linting that no longer exists after this change: #128 (isolates the host golangci-lint cache and lock) and #131 (always installs the pinned lint tools in `script/bootstrap`). Neither was merged or incorporated here. Made moot by this change: #121 and #130. ### Review outcome Reviewed at #136 (comment) — **PASS**. Three non-blocking comment-accuracy findings were recorded there for the next touch of those files; the reviewer independently confirmed the lint gate is live by negative control from a warm cache. - **Live DNS tests made robust instead of gated; test caps moved to the org-wide 60s/20s/90s values** — #93 The resolver's live-DNS tests failed nondeterministically, a different subset each run. Fixed by engineering the nondeterminism out, not by routing around the network. **Nothing is mocked, faked, stubbed, recorded or replayed; there is no `-short` flag, no build tag, no skip, and no environment-tolerance for restricted egress.** Production resolver behaviour is unchanged. ### Root causes, all test-side 1. **Burst fan-out at one root server.** Every test in `internal/resolver` calls `t.Parallel()` and the build hosts have many cores (48 here), so all ~35 iterative resolutions started within milliseconds of each other, and because `queryServers` walks `rootServerList()` in fixed order they all aimed their first query at `198.41.0.4`. Root servers rate-limit that, which fits the reported symptom of a different arbitrary subset failing each run. 2. **No retry anywhere.** One dropped UDP packet in a delegation chain failed a test outright. 3. **Unanimity assertions.** `TestQueryAllNameservers_AllReturnOK` and `_NXDomainFromAllNS` required *every* nameserver of a domain to answer — four independent chances to fail per run, with no tolerance for one being slow. ### What was built New `internal/resolver/livedns_test.go` holds all the live-DNS machinery, so `resolver_test.go` itself takes only call-site edits: - **Bounded live concurrency.** A package-wide semaphore (`liveConcurrency = 6`) caps how many live resolutions are in flight at once. Tests keep `t.Parallel()`; only their network work is throttled. This is the direct fix for cause 1, and the 60s budget is what makes it affordable. - **Retry with exponential backoff.** Three attempts per live operation, 8s deadline each, 500ms base backoff doubling. The retry predicate is deliberately **transport-level** — "did a nameserver answer at all" — and never the assertion the test is making, so a resolver that answers *incorrectly* still fails on the first attempt rather than being retried into a false green. - **Quorum instead of unanimity, tolerating SILENCE ONLY.** A strict majority of the discovered nameservers must answer as expected, and every individual result must additionally fall inside a closed **allowlist** of statuses the test explicitly sanctions: `ok`/`timeout`/`error` for the all-OK test, `nxdomain`/`timeout`/`error` for the NXDOMAIN test. A nameserver that stays silent is tolerated; one that answers **wrongly** is not, at any count. The allowlist is the load-bearing part — see the rework note below for why a blocklist was not enough. New `internal/resolver/livedns_harness_test.go` tests that machinery directly — quorum arithmetic, status counting, the allowlist, the gate's concurrency bound, per-attempt deadlines, and recovery from a transient failure. It performs no DNS resolution of any kind, so it neither mocks DNS nor depends on it. ### Rework after review — the quorum could not fail on a wrong answer The review at #136 (comment) returned **FAIL** on `9cb2c2b`, correctly. Fixed in `87bce43`. **The defect.** The claim above was, as first written, false for `resolver.StatusNoData`. Each test banned exactly one wrong status — `_AllReturnOK` banned only `nxdomain`, `_NXDomainFromAllNS` banned only `ok` — and `nodata` is neither. It is a **wrong answer, not silence**: `answeredCount` counted it as answered, so it did not even trigger a retry, and with a quorum of 3-of-4 a single wrong nameserver slid through undetected. The pre-change unanimity assertions would have caught it. That is robustness work quietly becoming assertion-loosening, which is exactly what this repo cannot afford. **The fix.** Tolerance is now an allowlist, not a blocklist of one status. New `unsanctionedStatuses()` returns every per-nameserver result whose status the caller did not explicitly sanction, and each test asserts that list is empty in addition to its quorum. A blocklist bans the one wrong answer its author thought of and silently admits everything else, including any status added to the resolver later; an allowlist fails on anything nobody sanctioned. `answeredCount` was reframed the same way — it now counts the closed set `ok`/`nxdomain`/`nodata`, so an unfamiliar status is treated as silence and can only ever cause a retry and then a loud failure, never a quiet pass. **Evidence — the reviewer's exact probe, re-run.** `queryEachNS` in `internal/resolver/iterative.go` was patched to force one of `google.com`'s four nameservers to return `StatusNoData` with empty records. `make test` now goes **red**, naming the offending nameserver and status: ``` exit=2 --- FAIL: TestQueryAllNameservers_AllReturnOK (1.12s) Error: Should be empty, but was [ns1.google.com.=nodata] Messages: every nameserver must answer OK or not answer at all: ns1.google.com.=nodata ns2.google.com.=ok ns3.google.com.=ok ns4.google.com.=ok --- FAIL: TestQueryAllNameservers_NXDomainFromAllNS (1.34s) Error: Should be empty, but was [ns1.google.com.=nodata] Messages: every nameserver must report NXDOMAIN or not answer at all: ns1.google.com.=nodata ns2.google.com.=nxdomain ns3.google.com.=nxdomain ns4.google.com.=nxdomain FAIL sneak.berlin/go/dnswatcher/internal/resolver 1.704s ``` That is the same input that returned `exit=0` with both tests **passing** under the old assertions. Probe reverted, tree clean, suite green again: ``` exit=0 ok sneak.berlin/go/dnswatcher/internal/resolver 4.185s coverage: 77.1% of statements (zero `(cached)` lines) ``` Two harness tests lock the regression in without any probe: three OK plus one `nodata` (quorum satisfied, no NXDOMAIN present — the exact input that used to pass) is reported as unsanctioned, and an unknown status is neither counted as answered nor tolerated. Also fixed from the review: the per-attempt deadline assertion in `livedns_harness_test.go` had no lower bound, so it passed for a deadline far shorter than intended. It now asserts the remaining time exceeds `liveAttemptTimeout/2` as well. Deliberately **not** done in this rework, per the review and the owner: no `-count=1` in `script/test` (the test-cache issue is real but pre-existing and repo-wide, filed separately); the remaining non-DNS mocks stay for #97; `queryServers` root-ordering stays untouched under #138. ### The `-timeout` backstop value: 90s Per the ruling at sneak/prompts#41 (comment) the cap is org-wide with two tiers: **60s hard cap for CI green, 20s target, and anything between the two must be filed as an improvement bug.** The backstop is **`90s`**, matching sneak/prompts#42 and preserving the 1.5x backstop-to-cap ratio the old 20s/30s pair had. It must strictly exceed the 60s cap or the cap is unreachable — the old `-timeout 30s` would have killed a 60s-capped suite at half its allowance. Applied to `script/test`; nothing else in the repo carried the old `30s`. Worst case for one live operation is 3 attempts x 8s plus ~1.5s of backoff, about 26s — comfortably inside the 90s backstop even if several operations exhaust their attempts at once. ### `REPO_POLICIES.md` is re-vendored, not hand-edited The file is org-canonical, so it was **copied byte-for-byte** from `prompts/REPO_POLICIES.md` on `sneak/prompts` branch `org-wide-60s-test-cap` (commit `52b5192`) rather than reworded to approximately the same thing. Verified: ``` $ cmp prompts/REPO_POLICIES.md dnswatcher/REPO_POLICIES.md && echo identical identical $ sha256sum REPO_POLICIES.md bcf11c312a1bee18a0e937eb412b51914411c1ab23308b8362409f3f88379ff7 ``` **What that byte-identity does and does not certify.** The source branch `org-wide-60s-test-cap` is an **unmerged proposal** — sneak/prompts#42 — not `prompts` `main`. So, precisely: - The **60s hard cap and 20s improvement-bug tier ARE the owner's ruling** (sneak/prompts#41 (comment)). - The **`90s` backstop is our own proposed number and is NOT ratified** (sneak/prompts#41 (comment)). - The vendored text is therefore **the proposed canonical text, pending** sneak/prompts#42. If that PR lands with different numbers, this file must be re-vendored to match; it should not be hand-edited here either way. **Known mismatch with this repo's actual state, recorded not papered over.** Re-vendoring also picked up the paragraph at `REPO_POLICIES.md:266-271` mandating that canonical golangci-lint be installed commit-pinned via `go install ...@c0d3ddc9...`. This repo does **not** comply with that mechanism: `cc86473` in this same PR made linting Docker-only, and `script/bootstrap` now installs golangci-lint nowhere. The **version and commit match** (`v2.12.2` / `c0d3ddc9`); the **installation mechanism does not**. The vendored file is org-canonical and must not be edited downstream, so this is being raised upstream for the org text to accommodate Docker-only linting rather than patched here. `TESTING.md`'s stale "within the 30-second target" follows to 60. That edit is deliberately a single line so it merges cleanly when #97 lands. ### Verification All runs through `make` / `script/` entrypoints only; lint runs in Docker. **Ten consecutive `make check` runs, all green, none served from cache.** Go's test cache will happily report `ok pkg (cached)` without executing anything, which proves nothing about nondeterminism, so every run was forced to actually execute and each log was checked for zero `(cached)` lines: ``` check#1 exit=0 wall=32s resolver=2.895s fails=0 cached=0 check#2 exit=0 wall=27s resolver=2.872s fails=0 cached=0 check#3 exit=0 wall=47s resolver=2.729s fails=0 cached=0 check#4 exit=0 wall=36s resolver=2.885s fails=0 cached=0 check#5 exit=0 wall=32s resolver=2.899s fails=0 cached=0 check#6 exit=0 (harness tests added) fails=0 cached=0 check#7 exit=0 wall=41s resolver=2.891s fails=0 cached=0 check#8 exit=0 wall=36s resolver=2.871s fails=0 cached=0 check#9 exit=0 wall=26s resolver=2.846s fails=0 cached=0 check#10 exit=0 wall=46s resolver=2.833s fails=0 cached=0 ``` After the rework commit `87bce43`, `make check` green again end to end, zero `(cached)` test lines, Docker lint stage demonstrably executed rather than served from cache: ``` exit=0 cached=0 #8 [deps 4/4] RUN go mod download #8 CACHED #10 [lint 2/2] RUN golangci-lint run --config .golangci.yml ./... #10 30.72 0 issues. ``` The Docker lint stage was confirmed to execute rather than cache on each run: ``` #8 [deps 4/4] RUN go mod download #8 CACHED #9 [lint 1/2] COPY . . #9 DONE 0.6s #10 [lint 2/2] RUN golangci-lint run --config .golangci.yml ./... #10 DONE 39.7s ``` **`make test` wall time: 3.6-4.1s** across three timed uncached runs (`4093ms`, `3649ms`, `3729ms`). `internal/resolver` went from 2.0s to ~2.9s — the concurrency gate's cost. That is inside the 20s target, so no improvement bug is owed under the new two-tier rule. **Honest note on what these runs do and do not prove.** Live DNS was healthy throughout: **no live-DNS retry fired even once**, and no flake was observed either before or after the change (six pre-change baseline runs were also clean). So these runs demonstrate the change is not itself flaky and does not slow the suite; they do **not** demonstrate recovery from a real DNS failure, because no real DNS failure occurred. The original flakiness is *not reproduced* rather than *shown fixed*. The retry path is instead proven by `TestRetryLiveRecoversFromTransientFailure`, the only source of the single `retrying in 500ms` line in each log: ``` livedns_harness_test.go:90: transient: attempt 1 of 3 failed (no answer from live DNS), retrying in 500ms ``` ### Observation, not acted on The single most effective remaining lever against root-server rate limiting would be to stop `queryServers` always trying `rootServerList()` in the same order, so that load spreads across all thirteen roots instead of concentrating on `a.root-servers.net`. That is **production** code and this issue scopes the work as test-side, so it was left alone rather than changed quietly. It is now tracked for the owner's decision at #138. ### Interaction with #97 `TESTING.md` and `internal/resolver/resolver_test.go` auto-merge — that PR touches `resolver_test.go` only at the import block and the final timeout-test section, while this change touches the body of the file and adds two new files, and it leaves the mock-`DNSClient` timeout test at the tail of `resolver_test.go` entirely alone since removing it is that PR's job. `TODO.md` does conflict; that PR is already labelled `needs-rebase`, so this adds nothing material to its rebase. - **Go's test cache disabled, so every `make test` actually queries live DNS** — #139 `script/test` did not pass `-count=1`, so on an unchanged tree Go served the whole suite from cache: exit 0 in ~0.2s, every package marked `(cached)`, and not one DNS query made. This repo's suite exists to exercise live resolution on every run (`TESTING.md`), so that green asserted nothing — and it is exactly the green used as evidence that a flakiness fix works, since "run it a few times" stops being runs after the first. It had already misled two agents, each of whom forced uncached runs by hand. `-count=1` now disables caching on every invocation. **The conditional verbose rerun was missing and is added here.** `REPO_POLICIES.md` mandates it; the primary run had been unconditionally `-v`, which is the failure mode the policy exists to prevent (unreadable CI and `docker build` logs on success). Tests now run quiet, and only a failure triggers the `-v` rerun. Two properties matter and both are covered: the rerun carries `-count=1` too, so it cannot replay a cached copy of the failure it is meant to diagnose; and its exit status is discarded in favour of a forced `1`, so a flake that passes the second time cannot turn the build green — the first failure already proved the suite broken. **`-timeout 90s` untouched.** It is a deliberate backstop that must strictly exceed the 60s hard cap. **No special-casing for the Docker build**, which reaches the same script via `RUN make test`: a fresh container's test cache is empty, so `-count=1` is a no-op there, and carving out an exception would only create a second code path that could drift. ### Verification **Proven from a warm cache, not a cold one.** The suite was run first to populate the cache, and the pre-change state confirmed: ``` ok sneak.berlin/go/dnswatcher/internal/config (cached) coverage: 92.6% of statements ok sneak.berlin/go/dnswatcher/internal/resolver (cached) coverage: 77.1% of statements ...8 of 8 packages (cached)... real 0m0.203s ``` With the change applied to that same warm cache, three back-to-back runs on an unchanged tree, **zero `(cached)` markers** in all three: ``` # run 1 # run 2 ok .../internal/config 1.045s ok .../internal/config 1.040s ok .../internal/handlers 1.029s ok .../internal/handlers 1.021s ok .../internal/notify 1.148s ok .../internal/notify 1.252s ok .../internal/portcheck 1.026s ok .../internal/portcheck 1.025s ok .../internal/resolver 3.005s ok .../internal/resolver 2.846s ok .../internal/state 1.078s ok .../internal/state 1.097s ok .../internal/tlscheck 1.067s ok .../internal/tlscheck 1.079s ok .../internal/watcher 1.566s ok .../internal/watcher 1.544s real 0m4.174s real 0m4.015s # run 3: grep -c '(cached)' => 0 real 0m4.519s ``` **Measured uncached wall time: 4.0-4.5s** (was ~0.2s served from cache). Inside the 20s target, so no improvement bug is owed under the two-tier rule at sneak/prompts#41 (comment). **`-count=1` composes with `-race` and `-cover`**: both still present in the primary run, and the per-package coverage percentages above are identical to the pre-change values. **The failure path was exercised, not assumed.** A purpose-built flaky test that fails on its first run and passes every run after (marker file kept outside the module, so the tree stays byte-identical and a cached result would be served if caching were on) was run through the script: ``` exit code: 1 --- FAIL: TestFlaky (0.00s) FAIL flakeproof 0.013s --- Rerunning with -v for details --- --- PASS: TestFlaky (0.00s) ok flakeproof 1.014s ``` Quiet failure, verbose rerun that genuinely re-executed (it passed, so it did not replay the cached `FAIL`), and exit `1` regardless of the rerun passing. Scratch module removed afterwards. **`make check` green**, exit 0, with the Docker lint stage demonstrably executed rather than served from cache: ``` #8 [deps 4/4] RUN go mod download #8 CACHED #10 [lint 2/2] RUN golangci-lint run --config .golangci.yml ./... #10 21.36 0 issues. #10 DONE 26.2s ``` `README.md` and `TESTING.md` record why the cache is waived, and `TODO.md` is updated in the same commit. ### Question for the owner, not filed as a defect The Docker lint run emits `The linter 'gomodguard' is deprecated (since v2.12.0) due to: new major version. Replaced by gomodguard_v2.` It is pre-existing and out of this issue's scope. It is not filed as an issue here because `.golangci.yml` tracks the org-canonical config, so switching to `gomodguard_v2` looks like an upstream `sneak/prompts` decision rather than a per-repo fix. Say the word and it gets filed in whichever place you consider canonical. - **MIT `LICENSE` added; README states the licence** — #102 The repo had no licence file at all, so publicly readable code was all-rights-reserved by default and nobody could legally use it. `LICENSE` was also the last file missing from `REPO_POLICIES.md`'s required minimum. MIT, by standing org policy rather than a per-repo call: any public repo lacking a licence gets MIT, and a private one with no licence is already all-rights-reserved. `sneak/dnswatcher` is public (`private: false` on the Gitea repo record). `README.md`'s first line now names the licence, per the Description requirement, and the License section states MIT and points at the file instead of recording the decision as pending. `TODO.md` updated in the same commit. ### Verification `LICENSE` is the canonical MIT text byte-for-byte with only the copyright line filled in (`Copyright (c) 2026 sneak`) — no clauses added, removed, reworded, or reflowed. It was not typed from memory: the file was copied from an existing verbatim MIT template on disk and only the copyright line edited (`diff` against that template shows that one line and nothing else), then the result was word-diffed against SPDX `MIT.txt` fetched from `spdx/license-list-data`, ignoring only line wrapping and the placeholder — identical. `make fmt` did **not** touch `LICENSE`, and cannot: `script/fmt` runs `gofmt -s -w .` and `goimports -w .` only, with no prettier or markdown step in the repo, so no exclusion was needed. `make check` green, exit 0. Tests executed rather than replayed (zero `(cached)` lines, `internal/resolver 3.098s`), and the Docker lint stage ran rather than cached: ``` #10 [deps 4/4] RUN go mod download #10 CACHED #12 [lint 2/2] RUN golangci-lint run --config .golangci.yml ./... #12 36.44 0 issues. #12 DONE 36.7s ``` - **Comment-only corrections in `script/bootstrap`, `script/cibuild`, and `Dockerfile.lint`** — #137 Follow-up to this PR's own review. Nothing executable changed: the diff touches comment lines, one warning string, and `TODO.md`. - `script/bootstrap`'s header justified the pinned `goimports` install by claiming `script/fmt-check` runs it on the host. Verified against the script: `script/fmt-check` runs `gofmt -l .` and nothing else. The header now credits `script/fmt` alone. That `fmt-check` does not verify goimports at all is #119 and was deliberately left alone. - `script/cibuild`'s header still said the `Dockerfile` runs `make check`. It now describes the current file: lint stage runs `make fmt-check` and `golangci-lint`, builder stage runs `make test` and `make build`. - The `docker`-missing warning was three fragments, each re-prefixed with `bootstrap:` mid-clause. Now one sentence: `bootstrap: WARNING: docker not found; install it to run make lint and make docker.` - `Dockerfile.lint`'s comment explained why `golangci-lint config verify` is omitted but read as though the omission were free. It now states the residual risk: unknown top-level keys in `.golangci.yml` are silently ignored, so a mistyped or wrong-schema key lints clean while applying nothing. `config verify` was **not** added — the network-dependence reasoning stands. ### Verification `make check` green, exit 0. Tests executed rather than replayed (zero `(cached)` lines, `internal/resolver 2.820s`), Docker lint stage executed rather than cached: ``` #8 [deps 4/4] RUN go mod download #8 CACHED #10 [lint 2/2] RUN golangci-lint run --config .golangci.yml ./... #10 13.30 0 issues. ``` Comment-only confirmed by reading the whole diff: no statement, flag, or command changed anywhere. --- ## Issues closed by this merge The commits on `next` each carry a bare `(closes #N)` in their subject, but the references in the prose above are full URLs, which Gitea's auto-close parser does not act on. Listing them here in bare form so the merge to `main` definitively closes them rather than leaving them open to be re-picked up as idle work: Closes #93 Closes #102 Closes #134 Closes #137 Closes #139 Co-authored-by: sneak <sneak@sneak.berlin> Reviewed-on: #136 Co-authored-by: clawbot <clawbot@noreply.example.org> Co-committed-by: clawbot <clawbot@noreply.example.org>
dnswatcher
dnswatcher is an MIT-licensed, pre-1.0 Go daemon by @sneak that monitors DNS records, TCP port availability, and TLS certificates, delivering real-time change notifications via Slack, Mattermost, and ntfy webhooks.
⚠️ Pre-1.0 software. APIs, configuration, and behavior may change without notice.
dnswatcher watches configured DNS domains and hostnames for changes, monitors TCP port availability, tracks TLS certificate expiry, and delivers real-time notifications via Slack, Mattermost, and/or ntfy webhooks.
It performs all DNS resolution itself via iterative (non-recursive) queries, tracing from root nameservers to authoritative servers directly—never relying on upstream recursive resolvers.
State is persisted to a local JSON file so that monitoring survives restarts without requiring an external database.
No DNS mocking. Ever.
DNS is never mocked in this project — not in tests, not anywhere else. No mock resolvers, no fake DNS servers, no stubbed lookups.
dnswatcher's entire purpose is correct behavior against the real DNS. Tests exercise real iterative resolution against live nameservers by design; a test suite that passes against a mock proves nothing about the one thing this program exists to do.
When live tests are flaky, that is a robustness problem, and it gets fixed with robustness: retries with backoff, querying multiple independent nameservers, longer timeouts — or explicit opt-in gating decided by the project owner. Never with mocks.
Contributions that introduce mocked, faked, or stubbed DNS will be rejected.
Features
DNS Domain Monitoring (Apex Domains)
- Accepts a list of DNS domain names (apex domains, identified via the Public Suffix List).
- Every 1 hour, performs a full iterative trace from root servers to discover all authoritative nameservers (NS records) for each domain.
- Queries every discovered authoritative nameserver independently.
- Stores the NS record set as observed by the delegation chain.
- Any change triggers a notification:
- NS added to or removed from the delegation.
- NS IP address changed (glue record change).
DNS Hostname Monitoring (Subdomains)
- Accepts a list of DNS hostnames (subdomains, distinguished from apex domains via the Public Suffix List).
- Every 1 hour, performs a full iterative trace to discover the authoritative nameservers for the hostname's parent domain.
- Queries each authoritative nameserver independently for all record types: A, AAAA, CNAME, MX, TXT, SRV, CAA, NS.
- Stores results per nameserver. The state for a hostname is not a merged view — it is a map from nameserver to record set.
- Any observable change in any nameserver's response triggers a
notification. This includes:
- Record change: A nameserver returns different records than it did on the previous check (additions, removals, value changes).
- NS query failure: A nameserver that previously responded becomes unreachable (timeout, SERVFAIL, REFUSED, network error). This is distinct from "responded with no records."
- NS recovery: A previously-unreachable nameserver starts responding again.
- Inconsistency detected: Two nameservers that previously agreed now return different record sets for the same hostname.
TCP Port Monitoring
- For every configured domain and hostname, constructs a deduplicated list of all IPv4 and IPv6 addresses resolved via A, AAAA, and CNAME chain resolution across all authoritative nameservers.
- Checks TCP connectivity on ports 80 and 443 for each IP address.
- Every 1 hour, re-checks all ports.
- Any change in port availability triggers a notification:
- Port transitioned from open to closed (or vice versa).
- New IP appeared (from DNS change) and its port state was recorded.
- IP disappeared (from DNS change) — noted in the DNS change notification; port state for that IP is removed.
TLS Certificate Monitoring
- Every 12 hours, for each IP address listening on port 443, connects via TLS using the correct SNI hostname.
- Records the certificate's Subject CN, SANs, issuer, and expiry date.
- Any change triggers a notification:
- Certificate is expiring within 7 days (warning, repeated each check until renewed or expired).
- Certificate CN, issuer, or SANs changed (replacement detected, reports old and new values).
- TLS connection failure to a previously-reachable IP:443 (handshake error, timeout, connection refused after previously succeeding).
- TLS recovery: a previously-failing IP:443 now completes a handshake again.
Notifications
Every observable state change produces a notification. dnswatcher is designed as a real-time change feed — degradations, failures, recoveries, and routine changes are all reported equally.
Supported notification backends:
| Backend | Configuration | Payload Format |
|---|---|---|
| Slack | Incoming Webhook URL | Attachments with color |
| Mattermost | Incoming Webhook URL | Slack-compatible attachments |
| ntfy | Topic URL (e.g. https://ntfy.sh/mytopic) |
Title + body + priority |
All configured endpoints receive every notification. Notification content includes:
- DNS record changes: Which hostname, which nameserver, what record type, old values, new values.
- DNS NS changes: Which domain, which nameservers were added/removed.
- NS query failures: Which nameserver failed, error type (timeout, SERVFAIL, REFUSED, network error), which hostname/domain affected.
- NS recoveries: Which nameserver recovered, which hostname/domain.
- NS inconsistencies: Which nameservers disagree, what each one returned, which hostname affected.
- Port changes: Which IP:port, old state, new state, all associated hostnames.
- TLS expiry warnings: Which certificate, days remaining, CN, issuer, associated hostname and IP.
- TLS certificate changes: Old and new CN/issuer/SANs, associated hostname and IP.
- TLS connection failures/recoveries: Which IP:port, error details, associated hostname.
State Management
- All monitoring state is kept in memory and persisted to a JSON file on
disk (
DATA_DIR/state.json). - State is loaded on startup to resume monitoring without triggering false-positive change notifications.
- State is written atomically (write to temp file, then rename) to prevent corruption.
Web Dashboard
dnswatcher includes an unauthenticated, read-only web dashboard at the
root URL (/). It displays:
- Summary counts for monitored domains, hostnames, ports, and certificates.
- Domains with their discovered nameservers.
- Hostnames with per-nameserver DNS records and status.
- Ports with open/closed state and associated hostnames.
- TLS certificates with CN, issuer, expiry, and status.
- Recent alerts (last 100 notifications sent since the process started), displayed in reverse chronological order.
Every data point shows its age (e.g. "5m ago") so you can tell at a glance how fresh the information is. The page auto-refreshes every 30 seconds.
The dashboard intentionally does not expose any configuration details such as webhook URLs, notification endpoints, or API tokens.
All assets (CSS) are embedded in the binary and served from the application itself. The dashboard makes zero external HTTP requests — no CDN dependencies or third-party resources are loaded at runtime.
HTTP API
dnswatcher exposes a lightweight HTTP API for operational visibility:
| Endpoint | Description |
|---|---|
GET / |
Web dashboard (HTML) |
GET /s/... |
Static assets (embedded CSS) |
GET /.well-known/healthcheck |
Health check (JSON) |
GET /health |
Health check (JSON, legacy) |
GET /api/v1/status |
Current monitoring state |
GET /metrics |
Prometheus metrics (optional) |
Architecture
cmd/dnswatcher/main.go Entry point (uber/fx bootstrap)
internal/
config/config.go Viper-based configuration
globals/globals.go Build-time variables (version)
logger/logger.go slog structured logging (TTY detection)
healthcheck/healthcheck.go Health check service
middleware/middleware.go HTTP middleware (logging, CORS, metrics auth)
handlers/handlers.go HTTP request handlers
server/
server.go HTTP server lifecycle
routes.go Route definitions
state/state.go JSON file state persistence
resolver/resolver.go Iterative DNS resolution engine
portcheck/portcheck.go TCP port connectivity checker
tlscheck/tlscheck.go TLS certificate inspector
notify/notify.go Notification service (Slack, Mattermost, ntfy)
watcher/watcher.go Main monitoring orchestrator and scheduler
Design Principles
- No recursive resolvers: All DNS resolution is performed iteratively, tracing from root nameservers through the delegation chain to authoritative servers.
- No external database: State is persisted as a single JSON file.
- Dependency injection: All components are wired via uber/fx.
- Structured logging: All logs use
log/slogwith JSON output in production (TTY detection for development). - Graceful shutdown: All background goroutines respect context cancellation and the fx lifecycle. In-flight notification deliveries are drained on shutdown, bounded by the shutdown timeout.
Configuration
Configuration is loaded via Viper with the following precedence (highest to lowest):
- Environment variables (prefixed with
DNSWATCHER_) .envfile (loaded via godotenv)- Config file:
/etc/dnswatcher/dnswatcher.yaml,~/.config/dnswatcher/dnswatcher.yaml, or./dnswatcher.yaml - Defaults
Environment Variables
| Variable | Description | Default |
|---|---|---|
PORT |
HTTP listen port | 8080 |
DNSWATCHER_DEBUG |
Enable debug logging | false |
DNSWATCHER_DATA_DIR |
Directory for state file | /var/lib/dnswatcher |
DNSWATCHER_TARGETS |
Comma-separated DNS names (auto-classified via PSL) | "" |
DNSWATCHER_SLACK_WEBHOOK |
Slack incoming webhook URL | "" |
DNSWATCHER_MATTERMOST_WEBHOOK |
Mattermost incoming webhook URL | "" |
DNSWATCHER_NTFY_TOPIC |
ntfy topic URL | "" |
DNSWATCHER_DNS_INTERVAL |
DNS check interval | 1h |
DNSWATCHER_TLS_INTERVAL |
TLS check interval | 12h |
DNSWATCHER_TLS_EXPIRY_WARNING |
Days before expiry to warn | 7 |
DNSWATCHER_SENTRY_DSN |
Sentry DSN for error reporting | "" |
DNSWATCHER_MAINTENANCE_MODE |
Enable maintenance mode | false |
DNSWATCHER_METRICS_USERNAME |
Basic auth username for /metrics | "" |
DNSWATCHER_METRICS_PASSWORD |
Basic auth password for /metrics | "" |
DNSWATCHER_SEND_TEST_NOTIFICATION |
Send a test notification after first scan completes | false |
DNSWATCHER_TARGETS is required. dnswatcher will refuse to start if no
monitoring targets are configured. A monitoring daemon with nothing to monitor
is a misconfiguration, so dnswatcher fails fast with a clear error message
rather than running silently. Set DNSWATCHER_TARGETS to a comma-separated
list of DNS names before starting.
Example .env
PORT=8080
DNSWATCHER_DEBUG=false
DNSWATCHER_DATA_DIR=/var/lib/dnswatcher
DNSWATCHER_TARGETS=example.com,example.org,www.example.com,api.example.com,mail.example.org
DNSWATCHER_SLACK_WEBHOOK=https://hooks.slack.com/services/T.../B.../xxx
DNSWATCHER_MATTERMOST_WEBHOOK=https://mattermost.example.com/hooks/xxx
DNSWATCHER_NTFY_TOPIC=https://ntfy.sh/my-dns-alerts
DNSWATCHER_SEND_TEST_NOTIFICATION=true
DNS Resolution Strategy
dnswatcher never uses the system's configured recursive resolver. Instead, it performs full iterative resolution:
- Root servers: Starts from the IANA root nameserver list (hardcoded, with periodic refresh).
- TLD delegation: Queries root servers for the TLD NS records.
- Domain delegation: Queries TLD nameservers for the domain's NS records.
- Authoritative query: Queries all discovered authoritative nameservers directly for the requested records.
This approach ensures:
- Independence from any upstream resolver's cache or filtering.
- Ability to detect split-horizon or inconsistent responses across authoritative servers.
- Visibility into the full delegation chain.
For hostname monitoring, the resolver follows CNAME chains (with a depth limit to prevent loops) before collecting terminal A/AAAA records.
State File Format
The state file (DATA_DIR/state.json) contains the complete monitoring
snapshot. Hostname records are stored per authoritative nameserver,
not as a merged view, to enable inconsistency detection.
{
"version": 1,
"lastUpdated": "2026-02-19T12:00:00Z",
"domains": {
"example.com": {
"nameservers": ["ns1.example.com.", "ns2.example.com."],
"lastChecked": "2026-02-19T12:00:00Z"
}
},
"hostnames": {
"www.example.com": {
"recordsByNameserver": {
"ns1.example.com.": {
"records": {
"A": ["93.184.216.34"],
"AAAA": ["2606:2800:220:1:248:1893:25c8:1946"]
},
"status": "ok",
"lastChecked": "2026-02-19T12:00:00Z"
},
"ns2.example.com.": {
"records": {
"A": ["93.184.216.34"],
"AAAA": ["2606:2800:220:1:248:1893:25c8:1946"]
},
"status": "ok",
"lastChecked": "2026-02-19T12:00:00Z"
}
},
"lastChecked": "2026-02-19T12:00:00Z"
}
},
"ports": {
"93.184.216.34:80": {
"open": true,
"hostnames": ["www.example.com"],
"lastChecked": "2026-02-19T12:00:00Z"
},
"93.184.216.34:443": {
"open": true,
"hostnames": ["www.example.com"],
"lastChecked": "2026-02-19T12:00:00Z"
}
},
"certificates": {
"93.184.216.34:443:www.example.com": {
"commonName": "www.example.com",
"issuer": "DigiCert TLS RSA SHA256 2020 CA1",
"notAfter": "2027-01-15T23:59:59Z",
"subjectAlternativeNames": ["www.example.com"],
"status": "ok",
"lastChecked": "2026-02-19T06:00:00Z"
}
}
}
The status field for each per-nameserver entry and certificate entry
tracks reachability:
| Status | Meaning |
|---|---|
ok |
Query succeeded, records are current |
error |
Query failed (timeout, SERVFAIL, network error) |
Entrypoints
This repository adheres to the
Scripts to Rule Them All
standard: normalized scripts in script/ are the entrypoints for the
development workflow, and the Makefile targets are thin shims that call
them. We provide:
script/bootstrap— install all dependencies (go, pinned goimports,go mod download). It does not install golangci-lint: seescript/lintbelow.script/setup— make a fresh clone ready for development: bootstrap plus the git pre-commit hookscript/projectname— print the project name (used for the Docker image tag)script/test— run the test suite (race detector, coverage). Caching is waived for testing, exactly as it is for linting:-count=1forces every invocation to execute, because the suite queries live DNS and a cached pass queries nothing. Failures are rerun with-vautomatically, and the build fails even if that rerun passes.script/lint— run golangci-lint, always inside Docker: it buildsDockerfile.lint, which COPYs the repo into the digest-pinnedgolangci-lintimage and lints as a build step, so a successful build is a clean lint. The linter is never installed or run on the host, and Docker is the only prerequisite. Caching is waived for linting: the lint stage is forced to execute on every run with--no-cache-filter, because a cached build lints nothing.script/fmt— format all code (gofmt -s, goimports)script/fmt-check— check formatting (read-only)script/check— run test, lint, and fmt-checkscript/docker— build the Docker image tagged viascript/projectnamescript/cibuild— CI entrypoint: plaindocker build .script/precommit— run by the git pre-commit hook;go mod tidyguard, thenscript/checkscript/install-precommit— install the git pre-commit hook
Building
make build # Build binary to bin/dnswatcher
make test # Run tests with race detector
make lint # Run golangci-lint in Docker (requires docker)
make fmt # Format code
make check # Run all checks (test, lint, fmt-check)
make clean # Remove build artifacts
Build-Time Variables
Version is injected via -ldflags:
go build -ldflags "-X main.Version=$(git describe --tags --always)" ./cmd/dnswatcher
Docker
docker build -t dnswatcher .
docker run -d \
-p 8080:8080 \
-v dnswatcher-data:/var/lib/dnswatcher \
-e DNSWATCHER_TARGETS=example.com,www.example.com \
-e DNSWATCHER_NTFY_TOPIC=https://ntfy.sh/my-alerts \
-e DNSWATCHER_SEND_TEST_NOTIFICATION=true \
dnswatcher
Monitoring Lifecycle
- Startup: Load state from disk. If no state file exists, start with empty state (first check will establish baseline without triggering change notifications).
- Initial check: Immediately perform all DNS, port, and TLS checks on startup.
- Periodic checks (DNS always runs first):
- DNS checks: every
DNSWATCHER_DNS_INTERVAL(default 1h). Also re-run before every TLS check cycle to ensure fresh IPs. - Port checks: every
DNSWATCHER_DNS_INTERVAL, after DNS completes. - TLS checks: every
DNSWATCHER_TLS_INTERVAL(default 12h), after DNS completes. - Port and TLS checks always use freshly resolved IP addresses from the DNS phase that immediately precedes them — never stale IPs from a previous cycle.
- DNS checks: every
- On change detection: Send notifications to all configured endpoints, update in-memory state, persist to disk.
- Shutdown: Persist final state to disk, wait for in-flight notification deliveries to complete, stop gracefully. The wait is bounded by the fx shutdown timeout (15s by default): deliveries still retrying against an unreachable endpoint when that expires are abandoned, and the number abandoned is logged at warn level rather than dropped silently. Notifications generated after shutdown has begun are refused and logged, so a late burst cannot extend the shutdown.
Planned Future Features (Post-1.0)
- DNSSEC validation: Validate the DNSSEC chain of trust during iterative resolution and report DNSSEC failures as notifications.
Project Structure
Follows the conventions defined in REPO_POLICIES.md, adapted from the
upaas project template. Uses uber/fx
for dependency injection, go-chi for HTTP routing, slog for logging, and
Viper for configuration.
License
dnswatcher is released under the MIT License, Copyright (c) 2026
@sneak. See the LICENSE file in the
repository root for the full text.