Author SHA1 Message Date
sneak c2736241ee watcher: a name removed from the targets leaves the state (closes #223)
check / check (push) Canceled after 0s
At startup, before the first check, Run removes from the loaded state
the domain, hostname and certificate entries of names no longer in
DNSWATCHER_TARGETS, so the dashboard, /api/v1/status and the startup
notification count only configured names. A configured domain's own
records, saved as a hostname entry under its name, are kept. Nothing is
notified. Each port check, next to the removal of stale port entries,
now also removes the certificate entries for an address a name no
longer resolves to, except while none of its nameservers answered, as
port entries already were.

Model: opus-5-5
2026-10-02 08:48:47 +00:00
clawbot 9b524e9d63 watcher: Port Change notifications list domains apart from hostnames (closes #248)
check / check (push) Canceled after 0s
A Port Change notification's `Hosts:` line listed a port's apex domains
and hostnames together. It now has a `Domains:` line and a `Hostnames:`
line, and leaves out one that would name nothing. A name is a domain
when it is a configured domain, the rule record notifications already
use to start `Domain:` or `Hostname:`; that rule is now one method,
isDomain, which both use. The port entries saved in the state are
unchanged. README describes the new lines.

Model: opus-5-5
2026-10-02 10:42:07 +02:00
clawbot b43402a631 dashboard and status API list a port's domains apart from its hostnames (closes #245)
check / check (push) Canceled after 0s
A port entry in the state saves the apex domains that resolve to its
address with its hostnames. The dashboard's Ports table now has a
Domains column next to Hostnames, and a port entry in /api/v1/status
has a `domains` list, with `hostnames` no longer holding a domain. A
name is taken as a domain when it has a domain entry, as the dashboard
and API already tell a domain's own records from a hostname's. Both
read the split from one function, buildPorts. The state file is
unchanged; README says its port `hostnames` include domains.

Model: opus-5-5
2026-10-02 10:31:07 +02:00
clawbot 008ec5d13a resolver: ask a referral's nameservers that come without addresses (closes #221)
check / check (push) Canceled after 0s
Looking up a nameserver's own address followed only the addresses a
referral gave, so a nameserver whose zone is delegated without them,
such as a.ntpns.org of pool.ntp.org, never resolved. The walk to a
name's nameservers looked addresses up only when a referral gave none.
Both now go through queryZone, which asks the nameservers whose
addresses the referral gives first and, if none of them gives a usable
reply, looks up and asks the others. maxLookupDepth stops lookups three
deep, so delegations that point at each other still end; when the limit
is why no address was found, the error is ErrLookupDepthExceeded, not
"no address".

Model: opus-5-5
2026-10-02 10:12:08 +02:00
clawbot e46db71821 dashboard, status API and notifications tell apex domains from hostnames (closes #224)
check / check (push) Canceled after 0s
An apex domain's own records are still saved with the hostnames'
records, under the domain's name, so the port and TLS checks find its
addresses. Notifications about them now start `Domain:`, decided by the
configured domains. The dashboard and /api/v1/status, which read only
the saved state, take a hostname entry whose name also has a domain
entry as that domain's own records: the dashboard shows them in a second
table under Domains, the API in the domain's `recordsByNameserver`, and
neither lists or counts them as hostnames. The startup notification
counts domains and hostnames from the configuration. README says which
of a domain's own records are watched and how their changes are
notified.

Model: opus-5-5
2026-10-02 10:08:32 +02:00
clawbot 1218df9467 dashboard: list record types in the README's order (closes #226)
check / check (push) Canceled after 0s
Each row of the Hostnames table listed a nameserver's record types in
the order Go happens to walk the record map, which changed from row to
row and on every page load, so two nameservers with the same records
looked different. formatRecords now sorts the types by their place in
the README's list (A, AAAA, CNAME, MX, TXT, SRV, CAA, NS); any other
type comes after them in alphabetical order. Values within a type were
already sorted by the resolver.

Model: opus-5-5
2026-10-02 09:55:07 +02:00
clawbot 250f3dd687 dashboard and status API show why a check failed (closes #225)
check / check (push) Canceled after 0s
/api/v1/status now gives `error` for each nameserver entry and
certificate entry whose status is `error`, copied from the state, which
already kept it. The dashboard shows that reason in place of the records
for a failed nameserver, which used to show the same `-` as one that
answered with no records, and across the CN, issuer and expiry cells
for a failed certificate, wrapped at a width of 20rem so the long TLS
error does not narrow the Endpoint column. The dashboard stylesheet is a
trimmed build, so the new markup uses only classes the page already had.
README Web Dashboard and HTTP API say so.

Model: opus-5-5
2026-10-02 09:14:28 +02:00
clawbot c07976a73a resolver: store a name's CNAME once per nameserver (closes #220)
check / check (push) Canceled after 0s
For a name with a CNAME, a nameserver answers a query of any type with
that CNAME, and the records of every answer were added, so the CNAME
was stored once for each of the eight record types asked for.
collectAnswerRecords now adds each value once per record type.

A state file saved before this holds the repeated values. Load keeps
each record value once, so the first check after upgrading sees no
record change and notifies nothing for them.

Model: opus-5-5
2026-10-02 09:09:51 +02:00
clawbot b047c3c64c watcher: a lookup cut short by shutdown is not logged as an error (closes #229)
check / check (push) Canceled after 0s
Stopping dnswatcher during a DNS check logged every lookup the stop cut
short as an error, "context canceled". The four places in the watcher
that log a failed lookup, checkDomain, checkHostname,
resolveNameserverAddresses and resolveCNAMEAddresses, now do it through
logFailedLookup, which logs nothing when the watcher's context was
cancelled. It asks the context, not the lookup's error: the resolver
reports a cancelled lookup with its own error, which does not wrap
context.Canceled. A context whose deadline passed is not cancelled, so a
lookup that ran out of time is still logged. The tests run each of these
lookups on a cancelled context and on one whose deadline passed; neither
sends a query.

Model: opus-5-5
2026-10-02 08:56:52 +02:00
clawbot f99de191c0 watcher: change messages name only the record types that differ (closes #219)
check / check (push) Canceled after 0s
A Record Change notification printed the nameserver's whole old and new
record sets in Go map syntax, and an Inconsistency notification the two
nameservers' whole sets, so a one-address change had to be found by eye
among kilobytes of unchanged TXT, CAA, MX and NS values. Both now list,
in sorted order of type, only the record types whose values differ: a
line naming the type, then each side's values separated by commas, or
none when that side has no records of that type. The dashboard's Recent
alerts shows the same text.

Model: opus-5-5
2026-10-02 08:46:56 +02:00
clawbot 5db5048754 watcher: startup notification no longer says every endpoint works (closes #230)
check / check (push) Canceled after 0s
The startup notification ended "All notification channels are working.",
but it is written once and handed to every notification endpoint before
any delivery has succeeded or failed, so the claim was never checked and
was false whenever one endpoint refused it. It now says only that it is a
test sent to every configured notification endpoint. The startup
notification test checks the whole message.

Model: opus-5-5
2026-10-02 08:42:35 +02:00
clawbot 26c9c74d8e notify: a failed Mattermost delivery's error names Mattermost (closes #227)
check / check (push) Canceled after 0s
Mattermost is sent by the Slack sender, which wrapped every HTTP error
status in ErrSlackFailed, so a Mattermost endpoint answering 503 was
logged as "slack notification failed". The sender now takes the error
to wrap: the Slack endpoint passes ErrSlackFailed and the Mattermost
endpoint passes ErrMattermostFailed, which was defined but unused.

A new delivery test sets both endpoints to a stand-in server answering
503 and checks the error logged for each names its own endpoint.

Model: opus-5-5
2026-10-02 08:40:19 +02:00
clawbot ceb24c5004 log: write durations as text, not nanoseconds (closes #228)
check / check (push) Canceled after 0s
The JSON log wrote a Go duration as a bare count of nanoseconds, so
the watcher starting line showed dnsInterval 120000000000 for 2m and a
delivery retry showed retryIn 1015437050. Each duration logged is now
passed through its String() form: dnsInterval and tlsInterval when the
watcher starts, retryIn on a delivery retry, and latency on a
succeeded port check. A test checks that retryIn is logged as the
text of the wait the retry actually took.

The request log's latency_ms is left as it is: its key names its
unit.

Model: opus-5-5
2026-10-02 08:37:47 +02:00
clawbot ee4cadbd05 watcher: follow a watched name's CNAME for port and TLS checks (closes #203)
check / check (push) Canceled after 0s
When a watched name's nameservers answer with a CNAME and no address,
the DNS check follows every target they gave with ResolveIPAddresses
and saves all addresses found as cnameAddresses in the hostname state,
so nameservers disagreeing on the target do not change them between
checks. Port and TLS checks use them. A change, also from or to none,
is notified as a CNAME address change; the first check from a state
file without them sends none. When a target cannot be followed, or
none of the name's nameservers answered, the last check's addresses
are kept. The domain check now runs the hostname check for the apex
instead of a copy of it.

Model: opus-5-5
2026-10-02 08:26:27 +02:00
clawbot a18803ff28 resolver tests: retry an answer missing the record type read (closes #218)
check / check (push) Canceled after 0s
QueryNameserver sends one query per record type. When only the AAAA or
MX query was lost, the answer still had status ok, liveQueryNameserver
did not retry it, and the test found no AAAA or MX records.
liveQueryNameserver now takes the record types a test reads and, through
livednstest, retries an answer that holds records of none of them. The
A, AAAA, MX and TXT tests name theirs. A resolver that loses a type for
good still fails, after the last attempt instead of the first.

Model: opus-5-5
2026-10-02 07:52:44 +02:00
clawbot 3182fc99a6 docker: a plain docker build . stamps the git version (closes #210)
check / check (push) Canceled after 0s
A plain `docker build .`, which is how upaas builds, stamped `dev`:
`.dockerignore` left out `.git` and the builder declared
`ARG VERSION=dev`. `.dockerignore` now sends `.git` without
`.git/config`, which can hold a credential, and lists no tracked file,
which git would count as deleted. `ARG VERSION` has no default. The
Makefile takes a non-empty `VERSION` from the command line or the
environment, so a build arg still wins; otherwise `git describe` runs in
the builder, which trusts the checkout whoever owns it, as a context
sent as a tar archive keeps its owners. A new `make version` prints the
version; the build fails when the context carries `.git` and it comes
out empty, `dev` or `unknown`.

Model: opus-5-5
2026-10-02 06:27:47 +02:00
clawbot 889e17459b resolver: never resend a refused query asking for recursion (closes #206)
check / check (push) Canceled after 0s
queryDNS resent a query that a server refused, this time asking for
recursion, so on a network that intercepts DNS the answers could come
from a recursive resolver without anyone knowing. A refusal is now only
a refusal, and the server is passed over for the next.

When every server of a zone refuses, the error says so. When every root
server refuses, the error is ErrIntercepted: root servers refuse no
query, so something on the network is answering in their place.
FindAuthoritativeNameservers stops at that error instead of trying each
parent name, so the watcher's log line says it.

A live test asks Quad9, which refuses a query not asking for recursion,
so that the resend cannot come back unnoticed.

Model: opus-5-5
2026-10-02 06:10:54 +02:00
clawbot a67fd20e4f ci: a push cancels its branch's older run, checkout keeps no token (closes #216)
check / check (push) Canceled after 0s
Every push queues a run on the shared runner, and a branch pushed again
left its older run queued for a head nobody needed. The workflow now puts
each branch's runs in one concurrency group with cancel-in-progress, so a
new push cancels that branch's older run, queued or running. The group is
keyed on the branch, so pushes to other branches never cancel runs on
`next` or `main`; a push to `next` itself does cancel the older `next` run.

The checkout step no longer writes the token into `.git/config`;
`script/cibuild` does not need it.

Model: opus-5-5
2026-10-02 06:07:53 +02:00
clawbot dba932c9e3 watcher tests: far fewer live queries, longer live attempts (closes #214)
check / check (push) Successful in 1m11s
A domain check looked up each nameserver's addresses by asking every
nameserver of that name's zone for all eight record types; it now asks
only for A, AAAA and CNAME, the ones it reads.

The watcher tests now check example.org instead of cloudflare.com (two
nameservers instead of five) and desec.io instead of example.com (its
nameservers are in zones with two, not cloudflare.com's five). The
record change and NS failure tests start from saved state built on one
NS lookup instead of a first full check, and the port change test runs
only the port checks again. A live test attempt may take 18 seconds,
not 8. A new live test checks that a nameserver's addresses include
IPv4 and IPv6.

Model: opus-5-5
2026-10-02 06:04:31 +02:00
clawbot 82836b41fd resolver: try servers in a random order on each resolution (closes #138)
check / check (push) Successful in 1m42s
Every resolution walked the root servers in a fixed order, so
a.root-servers.net got every first query and its timeouts were paid on
every lookup. Each list of servers the resolver walks is now walked in a
random order from rand.Shuffle, chosen anew each time; a server that does
not reply, refuses, or gives an error reply or a referral that leads no
closer is still passed over for the next. When a referral names a zone's
nameservers without their addresses, all of them are now looked up, not
only the first that resolves, so the zone is not given up because the
first nameserver whose address was found gave no usable reply. No test
fails if the walk stops shuffling: which server a live query reached is
not observable.

Model: opus-5-5
2026-10-02 03:26:06 +02:00
clawbot af81f2ac76 README: correct claims the code does not bear out (closes #108)
check / check (push) Failing after 2m12s
Checked every README claim against the code on next and fixed the ones
that were wrong or missing: what /metrics serves and when, what
DNSWATCHER_MAINTENANCE_MODE does, CORS on the public routes,
notification retries and the in-memory alert history, the certificate
error field and old port entries in the state file, and the Design
tree's missing files. Also corrected: CNAMEs are not followed for
watched names, the root server list is never refreshed, the NS set is
the delegation from the domain's parent zone, notification contents,
and the system resolver being used for webhooks. Code problems found
are filed separately.

Model: opus-5-5
2026-10-02 03:14:26 +02:00
clawbot 56c4395a39 config: watch a target listed twice only once (closes #207)
check / check (push) Successful in 1m50s
ClassifyTargets kept a name every time it appeared in DNSWATCHER_TARGETS,
so example.com,Example.com. put example.com in the domain list twice and
every check looked it up and checked its certificates twice, sending two
expiry warnings per TLS check (port checks were already grouped by address
and port). It now skips a name it has already kept, comparing after the
lower-casing and trailing-dot removal it already did; the list keeps the
order of first appearance.

Model: opus-5-5
2026-10-02 02:56:11 +02:00
clawbot c9510a986c watcher: warn of an expiring certificate on every TLS check (closes #204)
check / check (push) Failing after 2m1s
An expiry warning was skipped when the last one for that hostname and
address was sent less than DNSWATCHER_TLS_INTERVAL ago. Each TLS check
runs after a DNS pass of varying length, so two checks can be less than
the interval apart, and a certificate about to expire was warned about on
every check or every other check, at random. TLS checks already start
once per interval, so the in-memory record of when each warning was sent
is removed and every check warns, as the README says.

The test that expected the second check to stay silent is replaced by one
that runs TLS checks on state built in the test, with no DNS.

Model: opus-5-5
2026-10-02 01:58:28 +02:00
clawbot 1f1640d4cd resolver: take a domain's NS set from its delegation (closes #200)
check / check (push) Failing after 2m2s
A domain's NS set was taken from whichever of its own servers answered
first, so when they disagree (during a move between DNS providers, or
with a stale secondary) the set could change between checks and send an
NS change notification with nothing changed. The walk now stops at the
referral to the domain from its parent zone's servers and returns that
delegation, which those servers all hold alike; the domain's own
servers are no longer asked for it. The NS records in an answer are
still used where no such referral comes first, as from a server that
holds both the parent zone and the domain. Hostnames get their zone's
servers the same way.

Model: opus-5-5
2026-10-02 01:48:54 +02:00
clawbot 11ce1b249b docs: add the README sections policy requires (closes #173)
check / check (push) Failing after 2m30s
REPO_POLICIES.md requires Getting Started, Rationale, Design and TODO
sections in the README, and it had none of them. Getting Started clones
the repository, builds the image and runs it watching example.com and
www.example.com; DNSWATCHER_TARGETS is the only setting it requires.
Rationale is drawn from what the README already says. The Architecture
section moves below Entrypoints and is renamed Design, its text
unchanged, so the required sections come in policy order. TODO points
to TODO.md and the 1.0 milestone instead of copying the list.

Model: opus-5-5
2026-10-02 01:25:13 +02:00
clawbot d09822562d resolver: pass over a server that answers SERVFAIL or refers no closer (closes #197)
check / check (push) Failing after 2m8s
When the resolver walks from the root servers towards a name, a server
that answered SERVFAIL, or referred the query back to its own zone, up
or sideways, ended the step, so finding a zone's servers gave up on the
zone though its other servers would answer. Such a reply is now passed
over for the zone's next server, as a timeout or a refusal already was.
To tell a referral that leads closer to the name from one that does
not, each walk keeps the zone of the servers it is asking. Other error
replies, such as FORMERR, are passed over too. The walk that finds a
nameserver's address shares the same server loop, so it changes too.

Model: opus-5-5
2026-10-02 01:17:37 +02:00
clawbot 97c8138c85 watcher: keep port state when no nameserver of a name answered (closes #193)
check / check (push) Failing after 2m22s
The port check removed the saved port state of every address no
configured name resolves to. A name whose nameservers all timed out
or failed is saved with no records, so its addresses looked gone and
lost their port state; when the nameservers answered again it was
recorded afresh, and a port that opened or closed meanwhile was not
notified.

An entry is now kept when one of the names saved on it is configured
and none of its nameservers answered on its last check, and such a
name stays on the entry when the port is checked again for another
name. A name whose nameservers answer with no addresses still loses
it, and so does a name no longer configured.

Model: opus-5-5
2026-10-02 00:44:49 +02:00
clawbot d2f154b2cf resolver: error from ResolveIPAddresses when no nameserver answered (closes #190)
check / check (push) Failing after 2m7s
ResolveIPAddresses now returns an error, not no addresses, when no
nameserver of the name's zone answered. A nameserver with status
timeout or error is not an answer; one answer, even NXDOMAIN, is
enough for an empty result without an error.

When every server of a zone fails, FindAuthoritativeNameservers moves
on to the parent name, whose servers only refer the query onward. Such
a referral now has status error, so it is no answer either, and a
hostname's saved records show it as error. The only caller, the
nameserver address lookup, already keeps the previous addresses on an
error; its comment no longer says the resolver hides this case.

Model: opus-5-5
2026-10-02 00:28:53 +02:00
clawbot 4c2932d6d6 fmt: format and check Markdown with prettier in a container (closes #119)
check / check (push) Failing after 2m11s
make fmt and make fmt-check now cover every Markdown file with prettier
(4-space tabs, proseWrap always), as template-app-go does: prettier is
pinned by package.json and yarn.lock and runs in a docker build on a
digest-pinned node image, never on the host. The check is forced to run
with --no-cache-filter, as script/lint is.

script/fmt-check is split into a Go half and a Markdown half because the
Dockerfile lint stage cannot run docker: that stage now runs the Go half
and script/cibuild runs the Markdown half after the build. *.md leaves
.dockerignore so documents reach the build context.

README.md, TESTING.md and TODO.md are reformatted by make fmt; apart from
the README Entrypoints entries, its make fmt line under Building and the
TODO entry, that diff is mechanical.

Model: opus-5-5
2026-10-02 00:19:30 +02:00
clawbot b5814b2451 fmt: check goimports in fmt-check, run it at its pinned commit (#119)
check / check (push) Failing after 2m28s
script/fmt-check now runs goimports in list mode and fails naming any
file it would change, and checks gofmt with -s, as script/fmt applies
it. Both scripts run goimports with `go run` at the commit that
script/bootstrap used to install, so a goimports on PATH is never used
and bootstrap no longer installs it. The pin is written in both
scripts; change them together. The first run on a machine, and every
Dockerfile lint stage run, downloads and builds goimports.

The Markdown half of the issue (prettier) is not done here: it needs
node in the lint image or a separate build, a decision for the owner.

Model: opus-5-5
2026-10-01 23:51:22 +02:00
clawbot 797c936c48 resolver: query a hostname at its own zone's servers (closes #189)
check / check (push) Failing after 2m4s
A hostname's nameservers came from its last two labels, so a name under
co.uk was asked at the co.uk servers and a name in a delegated subdomain
at the parent's servers; both only refer onward. The hostname now goes
through FindAuthoritativeNameservers, which follows delegations for the
name and walks up its labels until it finds the zone it is in.

followDelegation now stops at an authoritative reply: that server holds
the zone, so its reply is not a referral. Without this, a CNAME answer
that also lists the zone's NS records in its authority section, as many
servers send, was followed as a referral until the delegation limit.

Model: opus-5-5
2026-10-01 23:34:35 +02:00
clawbot c247f6bcf5 watcher: notify nameserver address changes (closes #105)
check / check (push) Failing after 2m11s
Each domain check now looks up the addresses every nameserver's name
resolves to, with the resolver's ResolveIPAddresses, and saves them
sorted in the domain's state. A nameserver that stays in the
delegation and resolves to different addresses sends one NS Address
Change notification naming the domain, the nameserver and the old and
new addresses. Added or removed nameservers get only the NS change
notification. A failed or empty lookup keeps the previous addresses,
because the resolver returns no address without an error when every
server it asks times out. State files without the field load, and the
next check fills it in silently. Watcher tests that run domain checks
use example.com, which has two nameservers, to stay within the
per-attempt limit.

Model: opus-5-5
2026-10-01 23:23:26 +02:00
clawbot f6567df2d0 watcher: save state when it stops and wait for that save (closes #114)
check / check (push) Successful in 1m24s
The final save at shutdown came from the state's own stop hook, while
the watcher's stop hook only cancelled its run loop, so a check under
way could change state after that save or be cut off at exit. Run now
saves state as it returns, and the watcher's stop hook waits for Run,
bounded by the shutdown deadline. The state's own save stays; Save
holds the state lock for the whole write, so the two cannot overlap.
The start hook derives the watcher's context with WithoutCancel, so
the linter needs no exception. A new test stops a watcher built by New
and reads the change back from the state file, with no DNS.

Model: opus-5-5
2026-10-01 23:14:17 +02:00
clawbot c0ea9b96f2 server: report HTTP handler panics to Sentry (closes #107)
check / check (push) Successful in 1m35s
DNSWATCHER_SENTRY_DSN was read but never used. This ports the Sentry
integration from gohttpserver with sentry-go v0.49.0. The server's start
hook calls sentry.Init when the DSN is set; a DSN Sentry cannot parse
fails the hook, so startup stops with the parse error. sentryhttp, with
Repanic, reports handler panics and passes them on to chi's Recoverer.
Shutdown sends queued reports once the HTTP server has stopped. The
client uses the older transport (DisableTelemetryBuffer): with the
default one, Flush can return before sending a report made just before
it. Client reports are off, so only panics are sent. The DSN is checked
at server start, not in config, so the config test is unchanged.
sentry-go raises several golang.org/x modules and moves go-spew and
go-difflib to untagged commits.

Model: opus-5-5
2026-10-01 23:09:39 +02:00
clawbot fe01cdda1e watcher: save and notify nothing for a cut-short port or TLS check (closes #185)
check / check (push) Successful in 1m8s
When shutdown cancels a check that is under way, the rest of the check
still runs with the cancelled context. The resolver already drops a
lookup the context cut short, but a cancelled connection attempt was
saved as a closed port or a failed certificate check and notified as
Port Change or TLS Failure. The watcher now drops a port or TLS check
result when its context was cancelled, the same way. The test runs a
check with the context already cancelled, using the real resolver and
the real port and TLS checkers; no query is sent and no connection is
made.

Model: opus-5-5
2026-10-01 22:49:45 +02:00
clawbot a8f9a64600 middleware: take the client address from the right of X-Forwarded-For (closes #181)
check / check (push) Successful in 1m22s
realIP took the first X-Forwarded-For entry, which the client itself
can write, so behind a proxy that appends to the header a client chose
the address dnswatcher logs and the /metrics rate limit counts. It now
walks the entries from the right past trusted proxies, using the
existing trusted-proxy check, and takes the first that is not one; the
leftmost when all are. All X-Forwarded-For header lines are read as one
list, since a proxy may add its own line instead of appending to the
client's. An empty entry where the client address belongs falls back
to the peer address, as an empty first entry did before. X-Real-IP is
unchanged.

Model: opus-5-5
2026-10-01 22:35:17 +02:00
clawbot 8f11ef0038 watcher: notify NS query failure and recovery (closes #104)
check / check (push) Successful in 1m31s
LookupAllRecords now returns each nameserver's response, so the
watcher saves its status: ok when it answered, NXDOMAIN and no records
included, and error with the reason when it timed out, answered
SERVFAIL or REFUSED, or could not be reached. A nameserver that starts
failing sends NS Failure and one that answers again sends NS Recovery.
A failing nameserver is left out of the record change and
inconsistency comparisons. The resolver used to report REFUSED and
network errors as an answer with no records; they are now errors. A
lookup cut short by its context now returns an error instead of a
failure of the nameserver it was querying.

Model: opus-5-5
2026-10-01 22:23:33 +02:00
clawbot fcd4f7e2c2 config: stop startup on an invalid DNS or TLS interval (closes #177)
check / check (push) Successful in 1m3s
DNSWATCHER_DNS_INTERVAL and DNSWATCHER_TLS_INTERVAL were parsed with
time.ParseDuration and silently replaced by the default when that failed,
so a value like 5 or 1d gave hourly checks with no hint why, and zero or
negative values were accepted. Both now go through parseInterval, which
returns an error naming the variable and the value, and startup stops the
same way it does for invalid targets. An unset or empty variable still gets
its default from setupViper. The three tests that pinned the old fallback
are replaced. The README says what a valid value looks like.

Model: opus-5-5
2026-10-01 22:19:04 +02:00
clawbot ed0f56f144 metrics: rate limit /metrics per client address before Basic Auth (closes #101)
check / check (push) Successful in 1m18s
/metrics is behind a password, and REPO_POLICIES.md requires rate
limiting on password logins. Each client address may now send it 30
requests a minute, counted by httprate before Basic Auth, so failed
logins use up the allowance and a request over it gets 429 without
the password being checked. The address is the one the existing
trusted-proxy logic in internal/middleware works out, with IPv6
addresses grouped by /64; an IPv4 address a proxy reports in
IPv6-mapped form counts as the plain IPv4 address. A Prometheus
server scraping every 15 seconds sends 4 requests a minute.

Model: opus-5-5
2026-10-01 22:09:14 +02:00
clawbot bea9a3b2f2 docker: report the real version in the image (closes #109)
check / check (push) Successful in 1m17s
The Dockerfile builder stage takes ARG VERSION (default `dev`) and passes
it to `make build` on the command line, which overrides the Makefile's
`git describe` default. script/docker computes the version from
`git describe` on the host and passes it as --build-arg VERSION, because
.dockerignore leaves .git out of the build context and `git describe`
inside the build only ever produced `dev`. A build that passes no
argument, such as script/cibuild, still reports `dev`.

`logger.Identify`, which logs `starting` with the version, was never
called; `main` now calls it first, so the version is in the startup log.

Model: opus-5-5
2026-10-01 22:01:43 +02:00
clawbot e93c2664b8 notify: release held deliveries so shutdown tests fail, not hang (closes #176)
check / check (push) Successful in 1m5s
Two shutdown tests hold a delivery inside the test server's handler and
release it from a timer. The deferred timer stop ran before the server
was closed, so a drain that returned early left the handler blocked and
the server's close waited on it until the package timed out. Each test
now defers a release, guarded so the timer and the defer can both call
it, ahead of closing the server. The watchdog comment no longer names a
30-second timeout the test script does not use.

Model: opus-5-5
2026-10-01 21:58:58 +02:00
clawbot bde047f2a3 script/install-precommit: work where .git is a file (closes #129)
check / check (push) Successful in 1m9s
The script wrote the hook to .git/hooks, which fails when .git is a
file rather than a directory, as in a clone made with
--separate-git-dir. It now asks git for the repository's own git
directory with `git rev-parse --git-common-dir`, creates its hooks
directory if missing, and writes the hook there. In an ordinary clone
that is .git/hooks, so nothing moves. Before writing anything it stops
with an error when its top directory is not the top of the checkout
git finds, so a copy inside another repository cannot replace that
repository's hook. git's core.hooksPath setting is not followed; where
it is in force, git does not run the installed hook, as before.

Model: opus-5-5
2026-10-01 21:45:57 +02:00
clawbot 6070356676 docs: bring TODO.md up to date (closes #146)
check / check (push) Successful in 1m6s
Next Step and Future Steps list the open issues by full URL, in the order
of the review on #144; issues
the review does not name sit next to the entries they relate to, and
#144 itself comes last. Left
out: #146, which this change
closes; #59, DNSSEC, ruled
post-1.0; issues assigned to sneak, which wait on the owner. Shipped work,
the dropped domains and hostnames endpoints and the line about mocked
resolver tests are gone. Every Completed Steps entry is at most two lines;
none was dropped. Workflow branches from `next` and opens PRs against it.
Wrapped by hand at 80 columns: `make fmt` does not format Markdown yet
(#119). The old Next Step is now
#173.

Model: opus-5-5
2026-10-01 21:32:49 +02:00
clawbot 3390d7065e docker: set up the data directory in an entrypoint (closes #166)
check / check (push) Successful in 1m16s
The runtime image no longer sets USER. Its new entrypoint,
deploy/docker-entrypoint.sh, runs as root: it creates the data
directory if needed, gives it and everything in it to the dnswatcher
user (uid 10001) with mode 700 on the directory, then runs dnswatcher
as that user with su-exec. A bind-mounted host directory, whether
empty and root-owned or holding a state file left by another uid, no
longer has to be chowned first, and the README's upaas section now says
only which path to mount. The startup check that the data directory is
writable stays.

Model: opus-5-5
2026-10-01 21:15:02 +02:00
clawbot f7cc6b42e0 server: limit wildcard CORS to the public routes (closes #100)
check / check (push) Successful in 1m13s
The CORS wildcard was global, so it also covered the Basic-Auth
protected /metrics, which REPO_POLICIES.md forbids, and it allowed
POST, PUT and DELETE, which no route serves, plus the Authorization
and X-CSRF-Token headers. CORS now sits on a router holding only the
public routes and allows GET and OPTIONS with the Accept and
Content-Type headers. /metrics gets no CORS at all.

Both are mounted routers rather than a Group: chi answers OPTIONS on
a Group's route with 405 before its middleware runs, and any method
/metrics does not register would otherwise fall through to the public
router. So every method on /metrics now meets Basic Auth first, and
/metrics/ is served like /metrics.

Model: opus-5-5
2026-10-01 21:05:43 +02:00
clawbot db94c903df tests: move the test-only constructors into export_test.go (closes #111)
check / check (push) Successful in 1m37s
state.NewForTest, state.NewForTestWithDataDir and watcher.NewForTest
were compiled into and exported by the production packages
internal/state and internal/watcher. NewForTestWithDataDir and
watcher.NewForTest now live in their package's export_test.go. Tests in
other packages cannot see those files, so the watcher and middleware
tests build their State with state.New and a temporary data directory;
the watcher tests no longer try to save to /state.json.
state.NewForTest, whose State saved to /, is deleted: the state tests
that used it pass t.TempDir() to NewForTestWithDataDir, and the test of
the helper itself is gone.

Model: opus-5-5
2026-10-01 21:03:03 +02:00
clawbot 5493e28480 notify: make the shutdown tests fail with the right message (closes #116)
check / check (push) Successful in 1m27s
drainSlack stood for three things: the deadline given to a drain that
should finish early, the watchdog on a drain that should time out, and
the wait for a delivery to reach the test server. It is now three
constants, each commented with what it bounds and why it is 2s; no
value changed. The idle-drain failure printed that deadline instead of
idleDrainBound, the bound it checks. The cancelled-context test now
also requires the drain to return within idleDrainBound and to log its
debug line, so a drain that logs nothing no longer passes;
newLoggingService records debug level for this.

Model: opus-5-5
2026-10-01 20:42:21 +02:00
clawbot 651429137f tests: rename internal/livedns to livednstest and deny it outside tests (closes #164)
check / check (push) Successful in 1m39s
The live-DNS retry and concurrency limit is only for tests, but nothing
stopped program code from importing it and compiling it into the
binary. Its directory name now ends in test, and its import path is on
the test-support deny list in .golangci.yml, so make lint fails when
program code imports it. Every import and mention is updated to the new
name.

Model: opus-5-5
2026-10-01 19:52:48 +02:00
clawbot 93c1fe15e3 golangci: re-vendor the org config with gomodguard_v2 (closes #123)
check / check (push) Successful in 1m10s
The org .golangci.yml now uses gomodguard_v2 in place of the
deprecated gomodguard, which made every lint run print a deprecation
warning. The file is copied unchanged from sneak/prompts. It also turns
on depguard with the org test-support rule, which rejects
net/http/httptest except in test files and in files under a directory
whose name ends in test. This repo's previous copy had no deny entries
of its own, so there were none to carry forward.

Model: opus-5-5
2026-09-29 10:27:07 +02:00
clawbot 6dd6043534 tests: remove the DNS stand-ins from the watcher and resolver tests (closes #159)
check / check (push) Successful in 1m21s
The watcher tests used a stand-in resolver and the resolver timeout test
a stand-in DNS client, against the rule that DNS is never mocked.
Watcher tests that look something up in DNS now run the real resolver
against live servers, each attempt on a new watcher. A DNS change is
tested by saving values live DNS never returns (names under .invalid,
192.0.2.1) in the state a check starts from, or by marking a real
nameserver failed. The timeout test queries 192.0.2.1, where nothing
answers. The live-DNS retry and concurrency limit moved from the
resolver tests to internal/livedns, so both packages share them.
NewFromLoggerWithClient had no other use and is gone. TESTING.md now
states the README's rule.

Model: opus-5-5
2026-09-29 08:44:07 +02:00
clawbot a93389e1a0 watcher: send the inconsistency alert once per disagreement (closes #158)
check / check (push) Successful in 1m5s
detectInconsistencies alerted for neighbouring pairs of nameservers whose
records differed, on every DNS check, for as long as they differed. It now
also takes the previous hostname state, compares every pair of nameservers,
and alerts for a pair that differs unless both were in that state and
already differed there, so a nameserver new on a check that answers
differently is reported once. The state loaded at startup is the previous
state for the first check, so a disagreement saved before a restart is not
reported again. The choice of pairs is tested on record data, and the alert
through the hostname change detection with the notifier stand-in and no
resolver. The README describes the new behaviour.

Model: opus-5-5
2026-09-29 02:29:53 +02:00
clawbot 95b017eb3e resolver: lower-case DNS names in record values (closes #157)
check / check (push) Successful in 58s
Nameservers may answer with DNS names in any letter case. For eeqj.de,
y.ns.joker.com answers in upper case while its peers answer in lower
case, so the inconsistency check fired on every cycle. extractRecordValue
now lower-cases CNAME, MX, SRV and NS targets, so the inconsistency check
and the record-change check both compare names regardless of case. A,
AAAA, TXT and CAA values are formatted as before. State saved before this
change can hold upper-case names, which report a one-time record change
on the first check after upgrading.

Model: opus-5-5
2026-09-29 00:30:29 +02:00
clawbot 1ab0b9f61d script: force lint and test to run in cibuild and docker (closes #115)
check / check (push) Successful in 59s
script/cibuild and script/docker were plain docker build. On an
unchanged tree the lint stage and the builder stage, which runs make
test, came from the layer cache, so the build passed without linting
or querying live DNS. Both scripts now pass
--no-cache-filter=lint,builder so those stages run on every build, as
script/lint already does for its own lint stage. Dependency downloads
inside those stages re-run each build. Each of the two stages in the
Dockerfile now notes that the scripts name it. README and TODO.md
updated to match.

Model: opus-4-8 (implementation); opus-5-5 (rework)
2026-09-28 23:13:38 +02:00
clawbot 19f282c8b3 server: assert timeouts on the served http.Server, not just the constructor (closes #120)
check / check (push) Successful in 6s
The timeout tests called newHTTPServer directly, so a Run that built
its http.Server inline would drop every timeout with the suite still
green. TestRunWiresSocketTimeouts wires a Server as cmd/dnswatcher
does, minus the watcher and resolver so no live DNS is touched, drives
Run with an unbindable port so it stores its http.Server and returns
without listening, and checks that server carries all four timeouts
and both required relationships. The addr/handler test, which could
not fail, is dropped.

The ReadTimeout note now says what net/http does: a request whose
headers arrive after ReadTimeout but within ReadHeaderTimeout gets a
read deadline that has already passed, so reading its body fails at
once.

Model: opus-4-8 (implementation); opus-5-5 (rework)
2026-09-28 22:01:36 +02:00
clawbot 148e47d9c0 docker: run as non-root, add health check, document upaas (closes #147)
check / check (push) Successful in 5s
The runtime image runs as uid 10001, which owns /var/lib/dnswatcher. The
working directory is /, so config loading finds no .env or dnswatcher
config file there; the binary lives in /usr/local/bin. A Docker
HEALTHCHECK probes /.well-known/healthcheck every 10 seconds with busybox
wget, well inside the 60 seconds upaas waits.

Startup now fails with an error naming the data directory when it cannot
be written, instead of running with every save failing. The check creates
the directory if needed and writes and removes the temp file Save uses;
tests cover the create and the write failing.

README gains "Running under upaas": the prod branch, host directory
setup, network and port, environment and health check.

Model: opus-5-5
2026-09-28 20:13:37 +02:00
clawbot 8aaa103956 middleware: add security response headers (closes #98)
check / check (push) Successful in 54s
Adds a SecurityHeaders middleware and registers it globally, right after the request ID middleware, so every route gets the headers, including static files, /metrics and error responses.

It sets Strict-Transport-Security (one year, includeSubDomains), a Content-Security-Policy with default-src 'self', no scripts and frame-ancestors 'none', X-Frame-Options DENY, X-Content-Type-Options nosniff, Referrer-Policy no-referrer and a Permissions-Policy that turns every listed feature off.

HSTS is sent on every response, not only over TLS: the service runs behind a TLS-terminating proxy and REPO_POLICIES.md requires the application to send it. Referrer-Policy is stricter than the policy baseline because dashboard URLs can name internal hosts.

model: claude-opus-4-8 (implementation); claude-fable-5 (commit message)
2026-09-22 00:52:48 +02:00
clawbot c2a07ce690 test: add tests for globals, healthcheck and logger (closes #110)
check / check (push) Failing after 0s
Tests for the three packages that had none, written from outside each package, each able to fail on a plausible break:

- globals: values set are read back through New, and New returns an independent copy. One sequential test function with a disclosed paralleltest suppression, because it changes shared package variables.
- healthcheck: Check returns status "ok", the documented JSON fields, an RFC3339Nano timestamp, the maintenance flag from config in both states, and version and appname from globals.
- logger: New gives a usable *slog.Logger, debug output is off by default and EnableDebugLogging turns it on.

No production code changed. The terminal output format is not asserted.

Model: opus-4-8
2026-09-21 10:05:38 +02:00
clawbot b351a2350c gomod: drop stale golang.org/x/sync indirect line (closes #132)
check / check (push) Failing after 0s
go.mod listed golang.org/x/sync v0.19.0 twice: once in the direct
require block and once as // indirect. The module is a genuine direct
dependency (internal/portcheck/portcheck.go imports
golang.org/x/sync/errgroup), so the indirect entry was redundant and
stale. script/bootstrap ends with go mod download, which dropped that
line as a side effect and left every fresh checkout with a dirty
go.mod. Running go mod tidy removes the line for good; no dependency
and no version changed, and go.sum is unaffected.

Model: opus-4-8
2026-09-21 09:44:49 +02:00
clawbot 1ffe303a6e Merge main into next, resolving TODO.md (closes #145)
check / check (push) Failing after 1s
The last milestone landed on main as a squash, so main was no longer an
ancestor of next and git merge-tree reported a TODO.md conflict: both
branches added the same block of Completed Steps entries, and next added
one more (the #99 http.Server timeouts entry). This real merge commit
makes main an ancestor of next again. TODO.md is resolved by hand to keep
every entry from both sides exactly once, which is next's version since
next's entries are a superset of main's. Every other file already equals
next. No new TODO.md entry is added, per the task brief.

Model: opus-4-8
2026-09-21 07:22:17 +00:00
clawbot fc43f893a5 server: set all four http.Server socket timeouts (#118)
check / check (push) Successful in 3m50s
The http.Server was built with only ReadHeaderTimeout set; the other three
timeouts were zero, which in net/http means no limit, so a peer could hold a
connection open past the header phase, responses had no write deadline, and
keep-alive connections were never reaped. ReadTimeout 15s, WriteTimeout 75s,
and IdleTimeout 120s now join ReadHeaderTimeout 10s as named constants.
WriteTimeout must stay above the 60s chimw.Timeout handler budget, because
net/http arms the write deadline once request headers are read; a test fails
the build if either number moves alone. The server literal moved into
newHTTPServer so the configuration can be asserted without binding a socket
(closes #99)

Model: opus-5
2026-09-09 15:17:43 +02:00
clawbotandsneak e77e206fc4 next (#136)
check / check (push) Successful in 6s
Long-lived integration branch. One commit per work unit lands here; this PR accumulates them until it is merged to `main`.

## Landed units

- **Run all linting in Docker via `Dockerfile.lint` + `script/lint`** — #134

    golangci-lint is no longer installed or run on the host. New root `Dockerfile.lint` COPYs the repo into the digest-pinned `golangci/golangci-lint:v2.12.2` image and lints as a build step, so a successful build IS a clean lint; `script/lint` is reduced to a thin wrapper that builds it. This also works where the docker daemon is remote and bind mounts are impossible.

    **Pinned digest and how it was verified.** `golangci/golangci-lint:v2.12.2@sha256:5cceeef04e53efe1470638d4b4b4f5ceefd574955ab3941b2d9a68a8c9ad5240`, exactly as quoted in the issue. It resolves, and it is genuinely v2.12.2:

    ```
    $ docker buildx imagetools inspect golangci/golangci-lint:v2.12.2
    Name:      docker.io/golangci/golangci-lint:v2.12.2
    MediaType: application/vnd.oci.image.index.v1+json
    Digest:    sha256:5cceeef04e53efe1470638d4b4b4f5ceefd574955ab3941b2d9a68a8c9ad5240

    $ docker run --rm golangci/golangci-lint@sha256:5cceeef04e...ad5240 golangci-lint --version
    golangci-lint has version 2.12.2 built with go1.26.2 from c0d3ddc9 on 2026-05-06T11:07:58Z
    ```

    The tag's index digest is the quoted digest, and the binary inside reports commit `c0d3ddc9`, matching the org's canonical pin `c0d3ddc9cf3faa61a4e378e879ece580256d76e5`.

    **Forcing the linter to actually run.** `Dockerfile.lint` is split into a `deps` stage (base image + `go mod download`) and a `lint` stage (source copy + linter run). `script/lint` runs:

    ```
    docker build --progress=plain --no-cache-filter=lint --target lint -f Dockerfile.lint .
    ```

    Caching is explicitly waived for linting, and a cached build lints nothing, so the `lint` stage is invalidated on every invocation. The invalidation is scoped: the `deps` stage stays cached and no global cache wipe is performed. `--progress=plain` keeps the linter's own output visible.

    **`golangci-lint config verify`: deliberately NOT included.** It fetches its JSON schema over a live, unpinned HTTPS call, which would make linting network-dependent and defeat hash-pinning. Omitted for that reason, and the reason is recorded in a comment at the top of `Dockerfile.lint`.

    **`script/bootstrap`** no longer installs golangci-lint (and its pinned ref is gone); it warns non-fatally when `docker` is absent instead. The `goimports` install stays, because `script/fmt` and `script/fmt-check` still run on the host. Header comment updated accordingly.

    **Root `Dockerfile`** — required consequence, not scope creep. Its builder stage ran `make check`, which now calls `script/lint`, which shells out to `docker build`; there is no docker daemon inside a docker build, so `script/cibuild` and `script/docker` would have broken. It gains its own lint stage on the same pinned image (linter invoked directly, with a comment explaining why not `make lint`), with the builder stage depending on it via `COPY --from=lint /src/go.sum /dev/null` and running `make fmt-check`, `make test`, `make build`. The now-unneeded golangci-lint install is gone from the builder stage.

    **README** `Entrypoints` and `Building` sections now describe linting as a docker-only operation. `TODO.md` updated in the same commit.

    ### Verification

    All runs via `make` / `script/` entrypoints only.

    Two consecutive `make lint` runs on an unchanged tree, both executing the linter:

    ```
    # run 1
    #10 [lint 2/2] RUN golangci-lint run --config .golangci.yml ./...
    #10 10.98 0 issues.
    #10 DONE 12.0s

    # run 2, tree untouched
    #10 [lint 2/2] RUN golangci-lint run --config .golangci.yml ./...
    #10 14.63 0 issues.
    ```

    A third run shows the cache scoping is working as intended — `deps` served from cache, `lint` re-executed:

    ```
    #6 [deps 2/4] WORKDIR /src
    #6 CACHED
    #7 [deps 3/4] COPY go.mod go.sum ./
    #7 CACHED
    #8 [deps 4/4] RUN go mod download
    #8 CACHED
    #9 [lint 1/2] COPY . .
    #10 [lint 2/2] RUN golangci-lint run --config .golangci.yml ./...
    ```

    **Negative control.** A deliberate violation (an unused function containing an ineffectual assignment) was added to `internal/config/config.go`:

    ```
    #10 11.26 internal/config/config.go:29:2: ineffectual assignment to x (ineffassign)
    #10 11.26 internal/config/config.go:28:6: func negativeControlUnused is unused (unused)
    #10 11.26 2 issues:
    #10 ERROR: process "/bin/sh -c golangci-lint run --config .golangci.yml ./..." did not complete successfully: exit code: 1
    ERROR: failed to build: failed to solve: process "/bin/sh -c golangci-lint run --config .golangci.yml ./..." did not complete successfully: exit code: 1
    make: *** [Makefile:26: lint] Error 1
    ```

    `make lint` exited non-zero naming both findings and their exact lines. After reverting the file, `make lint` was clean again (`0 issues.`).

    **`make check`** green end to end (test, lint, fmt-check), exit 0.

    **`script/cibuild`** green, confirming the `Dockerfile` restructure does not recurse: the lint stage ran (`#16 12.56 0 issues.`), then `#22 [builder 8/9] RUN make test` with `PASS` lines, then `#23 [builder 9/9] RUN make build`.

    ### Notes for the owner

    This supersedes two PRs you still have queued for merge, both of which tune host linting that no longer exists after this change: #128 (isolates the host golangci-lint cache and lock) and #131 (always installs the pinned lint tools in `script/bootstrap`). Neither was merged or incorporated here.

    Made moot by this change: #121 and #130.

    ### Review outcome

    Reviewed at #136 (comment) — **PASS**. Three non-blocking comment-accuracy findings were recorded there for the next touch of those files; the reviewer independently confirmed the lint gate is live by negative control from a warm cache.

- **Live DNS tests made robust instead of gated; test caps moved to the org-wide 60s/20s/90s values** — #93

    The resolver's live-DNS tests failed nondeterministically, a different subset each run. Fixed by engineering the nondeterminism out, not by routing around the network. **Nothing is mocked, faked, stubbed, recorded or replayed; there is no `-short` flag, no build tag, no skip, and no environment-tolerance for restricted egress.** Production resolver behaviour is unchanged.

    ### Root causes, all test-side

    1. **Burst fan-out at one root server.** Every test in `internal/resolver` calls `t.Parallel()` and the build hosts have many cores (48 here), so all ~35 iterative resolutions started within milliseconds of each other, and because `queryServers` walks `rootServerList()` in fixed order they all aimed their first query at `198.41.0.4`. Root servers rate-limit that, which fits the reported symptom of a different arbitrary subset failing each run.
    2. **No retry anywhere.** One dropped UDP packet in a delegation chain failed a test outright.
    3. **Unanimity assertions.** `TestQueryAllNameservers_AllReturnOK` and `_NXDomainFromAllNS` required *every* nameserver of a domain to answer — four independent chances to fail per run, with no tolerance for one being slow.

    ### What was built

    New `internal/resolver/livedns_test.go` holds all the live-DNS machinery, so `resolver_test.go` itself takes only call-site edits:

    - **Bounded live concurrency.** A package-wide semaphore (`liveConcurrency = 6`) caps how many live resolutions are in flight at once. Tests keep `t.Parallel()`; only their network work is throttled. This is the direct fix for cause 1, and the 60s budget is what makes it affordable.
    - **Retry with exponential backoff.** Three attempts per live operation, 8s deadline each, 500ms base backoff doubling. The retry predicate is deliberately **transport-level** — "did a nameserver answer at all" — and never the assertion the test is making, so a resolver that answers *incorrectly* still fails on the first attempt rather than being retried into a false green.
    - **Quorum instead of unanimity, tolerating SILENCE ONLY.** A strict majority of the discovered nameservers must answer as expected, and every individual result must additionally fall inside a closed **allowlist** of statuses the test explicitly sanctions: `ok`/`timeout`/`error` for the all-OK test, `nxdomain`/`timeout`/`error` for the NXDOMAIN test. A nameserver that stays silent is tolerated; one that answers **wrongly** is not, at any count. The allowlist is the load-bearing part — see the rework note below for why a blocklist was not enough.

    New `internal/resolver/livedns_harness_test.go` tests that machinery directly — quorum arithmetic, status counting, the allowlist, the gate's concurrency bound, per-attempt deadlines, and recovery from a transient failure. It performs no DNS resolution of any kind, so it neither mocks DNS nor depends on it.

    ### Rework after review — the quorum could not fail on a wrong answer

    The review at #136 (comment) returned **FAIL** on `9cb2c2b`, correctly. Fixed in `87bce43`.

    **The defect.** The claim above was, as first written, false for `resolver.StatusNoData`. Each test banned exactly one wrong status — `_AllReturnOK` banned only `nxdomain`, `_NXDomainFromAllNS` banned only `ok` — and `nodata` is neither. It is a **wrong answer, not silence**: `answeredCount` counted it as answered, so it did not even trigger a retry, and with a quorum of 3-of-4 a single wrong nameserver slid through undetected. The pre-change unanimity assertions would have caught it. That is robustness work quietly becoming assertion-loosening, which is exactly what this repo cannot afford.

    **The fix.** Tolerance is now an allowlist, not a blocklist of one status. New `unsanctionedStatuses()` returns every per-nameserver result whose status the caller did not explicitly sanction, and each test asserts that list is empty in addition to its quorum. A blocklist bans the one wrong answer its author thought of and silently admits everything else, including any status added to the resolver later; an allowlist fails on anything nobody sanctioned. `answeredCount` was reframed the same way — it now counts the closed set `ok`/`nxdomain`/`nodata`, so an unfamiliar status is treated as silence and can only ever cause a retry and then a loud failure, never a quiet pass.

    **Evidence — the reviewer's exact probe, re-run.** `queryEachNS` in `internal/resolver/iterative.go` was patched to force one of `google.com`'s four nameservers to return `StatusNoData` with empty records. `make test` now goes **red**, naming the offending nameserver and status:

    ```
    exit=2
    --- FAIL: TestQueryAllNameservers_AllReturnOK (1.12s)
        Error:    Should be empty, but was [ns1.google.com.=nodata]
        Messages: every nameserver must answer OK or not answer at all:
                  ns1.google.com.=nodata ns2.google.com.=ok
                  ns3.google.com.=ok ns4.google.com.=ok
    --- FAIL: TestQueryAllNameservers_NXDomainFromAllNS (1.34s)
        Error:    Should be empty, but was [ns1.google.com.=nodata]
        Messages: every nameserver must report NXDOMAIN or not answer at all:
                  ns1.google.com.=nodata ns2.google.com.=nxdomain
                  ns3.google.com.=nxdomain ns4.google.com.=nxdomain
    FAIL  sneak.berlin/go/dnswatcher/internal/resolver  1.704s
    ```

    That is the same input that returned `exit=0` with both tests **passing** under the old assertions. Probe reverted, tree clean, suite green again:

    ```
    exit=0
    ok  sneak.berlin/go/dnswatcher/internal/resolver  4.185s  coverage: 77.1% of statements
    (zero `(cached)` lines)
    ```

    Two harness tests lock the regression in without any probe: three OK plus one `nodata` (quorum satisfied, no NXDOMAIN present — the exact input that used to pass) is reported as unsanctioned, and an unknown status is neither counted as answered nor tolerated.

    Also fixed from the review: the per-attempt deadline assertion in `livedns_harness_test.go` had no lower bound, so it passed for a deadline far shorter than intended. It now asserts the remaining time exceeds `liveAttemptTimeout/2` as well.

    Deliberately **not** done in this rework, per the review and the owner: no `-count=1` in `script/test` (the test-cache issue is real but pre-existing and repo-wide, filed separately); the remaining non-DNS mocks stay for #97; `queryServers` root-ordering stays untouched under #138.

    ### The `-timeout` backstop value: 90s

    Per the ruling at sneak/prompts#41 (comment) the cap is org-wide with two tiers: **60s hard cap for CI green, 20s target, and anything between the two must be filed as an improvement bug.** The backstop is **`90s`**, matching sneak/prompts#42 and preserving the 1.5x backstop-to-cap ratio the old 20s/30s pair had. It must strictly exceed the 60s cap or the cap is unreachable — the old `-timeout 30s` would have killed a 60s-capped suite at half its allowance. Applied to `script/test`; nothing else in the repo carried the old `30s`.

    Worst case for one live operation is 3 attempts x 8s plus ~1.5s of backoff, about 26s — comfortably inside the 90s backstop even if several operations exhaust their attempts at once.

    ### `REPO_POLICIES.md` is re-vendored, not hand-edited

    The file is org-canonical, so it was **copied byte-for-byte** from `prompts/REPO_POLICIES.md` on `sneak/prompts` branch `org-wide-60s-test-cap` (commit `52b5192`) rather than reworded to approximately the same thing. Verified:

    ```
    $ cmp prompts/REPO_POLICIES.md dnswatcher/REPO_POLICIES.md && echo identical
    identical
    $ sha256sum REPO_POLICIES.md
    bcf11c312a1bee18a0e937eb412b51914411c1ab23308b8362409f3f88379ff7
    ```

    **What that byte-identity does and does not certify.** The source branch `org-wide-60s-test-cap` is an **unmerged proposal** — sneak/prompts#42 — not `prompts` `main`. So, precisely:

    - The **60s hard cap and 20s improvement-bug tier ARE the owner's ruling** (sneak/prompts#41 (comment)).
    - The **`90s` backstop is our own proposed number and is NOT ratified** (sneak/prompts#41 (comment)).
    - The vendored text is therefore **the proposed canonical text, pending** sneak/prompts#42. If that PR lands with different numbers, this file must be re-vendored to match; it should not be hand-edited here either way.

    **Known mismatch with this repo's actual state, recorded not papered over.** Re-vendoring also picked up the paragraph at `REPO_POLICIES.md:266-271` mandating that canonical golangci-lint be installed commit-pinned via `go install ...@c0d3ddc9...`. This repo does **not** comply with that mechanism: `cc86473` in this same PR made linting Docker-only, and `script/bootstrap` now installs golangci-lint nowhere. The **version and commit match** (`v2.12.2` / `c0d3ddc9`); the **installation mechanism does not**. The vendored file is org-canonical and must not be edited downstream, so this is being raised upstream for the org text to accommodate Docker-only linting rather than patched here.

    `TESTING.md`'s stale "within the 30-second target" follows to 60. That edit is deliberately a single line so it merges cleanly when #97 lands.

    ### Verification

    All runs through `make` / `script/` entrypoints only; lint runs in Docker.

    **Ten consecutive `make check` runs, all green, none served from cache.** Go's test cache will happily report `ok pkg (cached)` without executing anything, which proves nothing about nondeterminism, so every run was forced to actually execute and each log was checked for zero `(cached)` lines:

    ```
    check#1  exit=0 wall=32s resolver=2.895s fails=0 cached=0
    check#2  exit=0 wall=27s resolver=2.872s fails=0 cached=0
    check#3  exit=0 wall=47s resolver=2.729s fails=0 cached=0
    check#4  exit=0 wall=36s resolver=2.885s fails=0 cached=0
    check#5  exit=0 wall=32s resolver=2.899s fails=0 cached=0
    check#6  exit=0 (harness tests added)      fails=0 cached=0
    check#7  exit=0 wall=41s resolver=2.891s fails=0 cached=0
    check#8  exit=0 wall=36s resolver=2.871s fails=0 cached=0
    check#9  exit=0 wall=26s resolver=2.846s fails=0 cached=0
    check#10 exit=0 wall=46s resolver=2.833s fails=0 cached=0
    ```

    After the rework commit `87bce43`, `make check` green again end to end, zero `(cached)` test lines, Docker lint stage demonstrably executed rather than served from cache:

    ```
    exit=0  cached=0
    #8 [deps 4/4] RUN go mod download
    #8 CACHED
    #10 [lint 2/2] RUN golangci-lint run --config .golangci.yml ./...
    #10 30.72 0 issues.
    ```

    The Docker lint stage was confirmed to execute rather than cache on each run:

    ```
    #8 [deps 4/4] RUN go mod download
    #8 CACHED
    #9 [lint 1/2] COPY . .
    #9 DONE 0.6s
    #10 [lint 2/2] RUN golangci-lint run --config .golangci.yml ./...
    #10 DONE 39.7s
    ```

    **`make test` wall time: 3.6-4.1s** across three timed uncached runs (`4093ms`, `3649ms`, `3729ms`). `internal/resolver` went from 2.0s to ~2.9s — the concurrency gate's cost. That is inside the 20s target, so no improvement bug is owed under the new two-tier rule.

    **Honest note on what these runs do and do not prove.** Live DNS was healthy throughout: **no live-DNS retry fired even once**, and no flake was observed either before or after the change (six pre-change baseline runs were also clean). So these runs demonstrate the change is not itself flaky and does not slow the suite; they do **not** demonstrate recovery from a real DNS failure, because no real DNS failure occurred. The original flakiness is *not reproduced* rather than *shown fixed*. The retry path is instead proven by `TestRetryLiveRecoversFromTransientFailure`, the only source of the single `retrying in 500ms` line in each log:

    ```
    livedns_harness_test.go:90: transient: attempt 1 of 3 failed
        (no answer from live DNS), retrying in 500ms
    ```

    ### Observation, not acted on

    The single most effective remaining lever against root-server rate limiting would be to stop `queryServers` always trying `rootServerList()` in the same order, so that load spreads across all thirteen roots instead of concentrating on `a.root-servers.net`. That is **production** code and this issue scopes the work as test-side, so it was left alone rather than changed quietly. It is now tracked for the owner's decision at #138.

    ### Interaction with #97

    `TESTING.md` and `internal/resolver/resolver_test.go` auto-merge — that PR touches `resolver_test.go` only at the import block and the final timeout-test section, while this change touches the body of the file and adds two new files, and it leaves the mock-`DNSClient` timeout test at the tail of `resolver_test.go` entirely alone since removing it is that PR's job. `TODO.md` does conflict; that PR is already labelled `needs-rebase`, so this adds nothing material to its rebase.

- **Go's test cache disabled, so every `make test` actually queries live DNS** — #139

    `script/test` did not pass `-count=1`, so on an unchanged tree Go served the whole suite from cache: exit 0 in ~0.2s, every package marked `(cached)`, and not one DNS query made. This repo's suite exists to exercise live resolution on every run (`TESTING.md`), so that green asserted nothing — and it is exactly the green used as evidence that a flakiness fix works, since "run it a few times" stops being runs after the first. It had already misled two agents, each of whom forced uncached runs by hand.

    `-count=1` now disables caching on every invocation.

    **The conditional verbose rerun was missing and is added here.** `REPO_POLICIES.md` mandates it; the primary run had been unconditionally `-v`, which is the failure mode the policy exists to prevent (unreadable CI and `docker build` logs on success). Tests now run quiet, and only a failure triggers the `-v` rerun. Two properties matter and both are covered: the rerun carries `-count=1` too, so it cannot replay a cached copy of the failure it is meant to diagnose; and its exit status is discarded in favour of a forced `1`, so a flake that passes the second time cannot turn the build green — the first failure already proved the suite broken.

    **`-timeout 90s` untouched.** It is a deliberate backstop that must strictly exceed the 60s hard cap.

    **No special-casing for the Docker build**, which reaches the same script via `RUN make test`: a fresh container's test cache is empty, so `-count=1` is a no-op there, and carving out an exception would only create a second code path that could drift.

    ### Verification

    **Proven from a warm cache, not a cold one.** The suite was run first to populate the cache, and the pre-change state confirmed:

    ```
    ok  sneak.berlin/go/dnswatcher/internal/config    (cached)  coverage: 92.6% of statements
    ok  sneak.berlin/go/dnswatcher/internal/resolver  (cached)  coverage: 77.1% of statements
    ...8 of 8 packages (cached)...
    real 0m0.203s
    ```

    With the change applied to that same warm cache, three back-to-back runs on an unchanged tree, **zero `(cached)` markers** in all three:

    ```
    # run 1                                          # run 2
    ok  .../internal/config     1.045s               ok  .../internal/config     1.040s
    ok  .../internal/handlers   1.029s               ok  .../internal/handlers   1.021s
    ok  .../internal/notify     1.148s               ok  .../internal/notify     1.252s
    ok  .../internal/portcheck  1.026s               ok  .../internal/portcheck  1.025s
    ok  .../internal/resolver   3.005s               ok  .../internal/resolver   2.846s
    ok  .../internal/state      1.078s               ok  .../internal/state      1.097s
    ok  .../internal/tlscheck   1.067s               ok  .../internal/tlscheck   1.079s
    ok  .../internal/watcher    1.566s               ok  .../internal/watcher    1.544s
    real 0m4.174s                                    real 0m4.015s

    # run 3: grep -c '(cached)' => 0                 real 0m4.519s
    ```

    **Measured uncached wall time: 4.0-4.5s** (was ~0.2s served from cache). Inside the 20s target, so no improvement bug is owed under the two-tier rule at sneak/prompts#41 (comment).

    **`-count=1` composes with `-race` and `-cover`**: both still present in the primary run, and the per-package coverage percentages above are identical to the pre-change values.

    **The failure path was exercised, not assumed.** A purpose-built flaky test that fails on its first run and passes every run after (marker file kept outside the module, so the tree stays byte-identical and a cached result would be served if caching were on) was run through the script:

    ```
    exit code: 1
    --- FAIL: TestFlaky (0.00s)
    FAIL    flakeproof  0.013s
    --- Rerunning with -v for details ---
    --- PASS: TestFlaky (0.00s)
    ok      flakeproof  1.014s
    ```

    Quiet failure, verbose rerun that genuinely re-executed (it passed, so it did not replay the cached `FAIL`), and exit `1` regardless of the rerun passing. Scratch module removed afterwards.

    **`make check` green**, exit 0, with the Docker lint stage demonstrably executed rather than served from cache:

    ```
    #8  [deps 4/4] RUN go mod download
    #8  CACHED
    #10 [lint 2/2] RUN golangci-lint run --config .golangci.yml ./...
    #10 21.36 0 issues.
    #10 DONE 26.2s
    ```

    `README.md` and `TESTING.md` record why the cache is waived, and `TODO.md` is updated in the same commit.

    ### Question for the owner, not filed as a defect

    The Docker lint run emits `The linter 'gomodguard' is deprecated (since v2.12.0) due to: new major version. Replaced by gomodguard_v2.` It is pre-existing and out of this issue's scope. It is not filed as an issue here because `.golangci.yml` tracks the org-canonical config, so switching to `gomodguard_v2` looks like an upstream `sneak/prompts` decision rather than a per-repo fix. Say the word and it gets filed in whichever place you consider canonical.

- **MIT `LICENSE` added; README states the licence** — #102

    The repo had no licence file at all, so publicly readable code was all-rights-reserved by default and nobody could legally use it. `LICENSE` was also the last file missing from `REPO_POLICIES.md`'s required minimum.

    MIT, by standing org policy rather than a per-repo call: any public repo lacking a licence gets MIT, and a private one with no licence is already all-rights-reserved. `sneak/dnswatcher` is public (`private: false` on the Gitea repo record).

    `README.md`'s first line now names the licence, per the Description requirement, and the License section states MIT and points at the file instead of recording the decision as pending. `TODO.md` updated in the same commit.

    ### Verification

    `LICENSE` is the canonical MIT text byte-for-byte with only the copyright line filled in (`Copyright (c) 2026 sneak`) — no clauses added, removed, reworded, or reflowed. It was not typed from memory: the file was copied from an existing verbatim MIT template on disk and only the copyright line edited (`diff` against that template shows that one line and nothing else), then the result was word-diffed against SPDX `MIT.txt` fetched from `spdx/license-list-data`, ignoring only line wrapping and the placeholder — identical.

    `make fmt` did **not** touch `LICENSE`, and cannot: `script/fmt` runs `gofmt -s -w .` and `goimports -w .` only, with no prettier or markdown step in the repo, so no exclusion was needed.

    `make check` green, exit 0. Tests executed rather than replayed (zero `(cached)` lines, `internal/resolver 3.098s`), and the Docker lint stage ran rather than cached:

    ```
    #10 [deps 4/4] RUN go mod download
    #10 CACHED
    #12 [lint 2/2] RUN golangci-lint run --config .golangci.yml ./...
    #12 36.44 0 issues.
    #12 DONE 36.7s
    ```

- **Comment-only corrections in `script/bootstrap`, `script/cibuild`, and `Dockerfile.lint`** — #137

    Follow-up to this PR's own review. Nothing executable changed: the diff touches comment lines, one warning string, and `TODO.md`.

    - `script/bootstrap`'s header justified the pinned `goimports` install by claiming `script/fmt-check` runs it on the host. Verified against the script: `script/fmt-check` runs `gofmt -l .` and nothing else. The header now credits `script/fmt` alone. That `fmt-check` does not verify goimports at all is #119 and was deliberately left alone.
    - `script/cibuild`'s header still said the `Dockerfile` runs `make check`. It now describes the current file: lint stage runs `make fmt-check` and `golangci-lint`, builder stage runs `make test` and `make build`.
    - The `docker`-missing warning was three fragments, each re-prefixed with `bootstrap:` mid-clause. Now one sentence: `bootstrap: WARNING: docker not found; install it to run make lint and make docker.`
    - `Dockerfile.lint`'s comment explained why `golangci-lint config verify` is omitted but read as though the omission were free. It now states the residual risk: unknown top-level keys in `.golangci.yml` are silently ignored, so a mistyped or wrong-schema key lints clean while applying nothing. `config verify` was **not** added — the network-dependence reasoning stands.

    ### Verification

    `make check` green, exit 0. Tests executed rather than replayed (zero `(cached)` lines, `internal/resolver 2.820s`), Docker lint stage executed rather than cached:

    ```
    #8 [deps 4/4] RUN go mod download
    #8 CACHED
    #10 [lint 2/2] RUN golangci-lint run --config .golangci.yml ./...
    #10 13.30 0 issues.
    ```

    Comment-only confirmed by reading the whole diff: no statement, flag, or command changed anywhere.

---

## Issues closed by this merge

The commits on `next` each carry a bare `(closes #N)` in their subject, but the
references in the prose above are full URLs, which Gitea's auto-close parser
does not act on. Listing them here in bare form so the merge to `main`
definitively closes them rather than leaving them open to be re-picked up as
idle work:

Closes #93
Closes #102
Closes #134
Closes #137
Closes #139

Co-authored-by: sneak <sneak@sneak.berlin>
Reviewed-on: #136
Co-authored-by: clawbot <clawbot@noreply.example.org>
Co-committed-by: clawbot <clawbot@noreply.example.org>
2026-09-09 14:57:55 +02:00
clawbot ae06f7e3a1 notify: drain in-flight deliveries at shutdown (closes #106)
check / check (push) Successful in 1m22s
Shutdown waits for deliveries already in flight instead of dropping them.

Reviewed and green; squashed to next by the dispatcher, dnswatcher having no manager.

Model: opus-5
2026-09-09 14:26:53 +02:00
sneak b8662b8a9c docs: correct stale script headers and record the config-verify cost (closes #137)
check / check (push) Successful in 51s
script/bootstrap credited the goimports pin to script/fmt-check, which
runs gofmt only; the header now credits script/fmt. script/cibuild
still described the Dockerfile as running make check, which stopped
being true once linting moved to its own stage. The docker-missing
warning in script/bootstrap is one sentence instead of three fragments
each carrying the bootstrap: prefix. Dockerfile.lint now states the
residual risk of skipping golangci-lint config verify: unknown
top-level keys in .golangci.yml are ignored silently, so a mistyped key
lints clean and applies nothing.

Comment and message text only; no behaviour changes.
2026-08-10 14:09:20 +00:00
sneak 168281ad60 docs: add MIT LICENSE and state the licence in the README (closes #102)
check / check (push) Has been cancelled
The repository had no licence file at all, which makes publicly readable
code all-rights-reserved by default: nobody may legally use it. That is a
1.0 blocker rather than a nicety, and `LICENSE` was the only file from
`REPO_POLICIES.md`'s required minimum still missing here.

The choice is standing org policy rather than a per-repo call: any public
repo lacking a licence gets MIT, while a private repo with no licence is
already all-rights-reserved and needs nothing. `sneak/dnswatcher` is
public, so MIT.

`LICENSE` carries the canonical MIT text byte-for-byte with only the
copyright line filled in; no clauses added, removed, reworded, or
reflowed. `README.md`'s first line now names the licence, which the
Description requirement in `REPO_POLICIES.md` calls for, and the License
section states MIT and points at the file instead of recording the
decision as pending.
2026-08-10 14:04:31 +00:00
sneak 6f6bf3a65b test: disable Go's test cache so every run queries live DNS (closes #139)
check / check (push) Successful in 1m34s
`script/test` did not pass `-count=1`, so on an unchanged tree Go
served the whole suite from its test cache: exit 0 in ~0.2s with every
package marked `(cached)` and not one DNS query made. This repo's suite
exists to exercise live resolution on every run (`TESTING.md`), so that
green asserted nothing — and it is exactly the green used as evidence
that a flakiness fix works, since "run it a few times" stops being
runs after the first.

`-count=1` now disables caching on every invocation.

The conditional verbose rerun that `REPO_POLICIES.md` mandates was
missing at the same spot and is added here rather than left broken: the
primary run had been unconditionally `-v`, which is the failure mode
the policy exists to prevent (unreadable CI and `docker build` logs on
success). Tests now run quiet, and only a failure triggers the `-v`
rerun. The rerun carries `-count=1` too, so it cannot replay a cached
copy of the failure it is meant to diagnose, and its exit status is
discarded in favour of a forced 1: the first failure already proved the
suite broken, so a flake that passes the second time must not turn the
build green.

`-timeout 90s` is untouched. It is a deliberate backstop that must
strictly exceed the 60s hard cap on suite duration.

No special-casing for the Docker build, which also reaches this script
via `RUN make test`: a fresh container's test cache is empty, so
`-count=1` changes nothing there and carving out an exception would
only create a second code path that could drift.

Verified: three back-to-back `make test` runs on an unchanged tree,
zero `(cached)` markers, ~4.0-4.5s wall each (was ~0.2s cached),
comfortably inside the 20s target with `-race` and `-cover` both still
working and coverage percentages unchanged. The rerun-and-still-fail
path was exercised against a purpose-built flaky test that fails once
then passes: quiet failure, verbose rerun that genuinely re-executed,
exit 1 regardless. `make check` green.
2026-08-10 13:48:15 +00:00
sneak 87bce43f8d test: rework live-DNS quorum unit — tolerate silence, never a wrong answer
check / check (push) Successful in 1m23s
Rework of the unit at #93
(commit 9cb2c2b), against the review at
#136 (comment).

Review of 9cb2c2b found the quorum assertions could not fail on a
class of wrong answer. Each test banned exactly one bad status —
_AllReturnOK banned only nxdomain, _NXDomainFromAllNS banned only ok
— so resolver.StatusNoData passed both. nodata is a wrong answer, not
silence, and answeredCount counted it as answered, so it did not even
trigger a retry; with a quorum of 3 of 4 a single wrong nameserver
slid through undetected. That is assertion-loosening beyond what the
quorum change requires.

Tolerance is now a closed allowlist rather than a blocklist of one
status. unsanctionedStatuses() reports every per-nameserver result
whose status the caller did not explicitly sanction: ok/timeout/error
for the all-OK test, nxdomain/timeout/error for the NXDOMAIN test.
Silence (timeout, error) is the only thing quorum exists to tolerate;
any other status, including one added to the resolver later, fails by
name. answeredCount is likewise an allowlist of ok/nxdomain/nodata, so
an unknown status counts as silence and can only cause a retry and
then a loud failure, never a quiet pass.

Two harness tests cover the regression directly: three OK plus one
nodata (quorum satisfied, no nxdomain present — the input that used
to pass) is now reported as unsanctioned, and an unknown status is
neither counted as answered nor tolerated.

Verified by re-running the reviewer's probe: queryEachNS patched to
force one of google.com's four nameservers to return StatusNoData
turns both tests red, naming the offending nameserver and status —

    --- FAIL: TestQueryAllNameservers_AllReturnOK (1.12s)
        Should be empty, but was [ns1.google.com.=nodata]
        every nameserver must answer OK or not answer at all:
        ns1.google.com.=nodata ns2.google.com.=ok ns3.google.com.=ok
        ns4.google.com.=ok
    --- FAIL: TestQueryAllNameservers_NXDomainFromAllNS (1.34s)
        Should be empty, but was [ns1.google.com.=nodata]
        every nameserver must report NXDOMAIN or not answer at all:
        ns1.google.com.=nodata ns2.google.com.=nxdomain
        ns3.google.com.=nxdomain ns4.google.com.=nxdomain

— and green with the probe reverted. Also fixes the review's nit: the
per-attempt deadline assertion had no lower bound, so it passed for a
deadline far shorter than intended.

No production code changed; DNS is still never mocked.
2026-08-10 13:33:53 +00:00
sneak 9cb2c2b7e0 test: make live DNS tests robust instead of gated (closes #93)
check / check (push) Successful in 1m18s
The resolver's live-DNS tests failed nondeterministically, a different
subset each run. Three structural causes, all test-side:

- Burst fan-out. Every test in the package is parallel and the build
  hosts have many cores, so all ~35 iterative resolutions started at
  the same instant and, because queryServers walks rootServerList() in
  fixed order, hit the same root server within milliseconds. Root
  servers rate-limit that.
- No retry anywhere. One dropped UDP packet in a delegation chain
  failed a test outright.
- Unanimity assertions. TestQueryAllNameservers_AllReturnOK and
  _NXDomainFromAllNS required every one of a domain's nameservers to
  answer, with no tolerance for one being slow.

New internal/resolver/livedns_test.go addresses each: a package-wide
gate bounds how many live resolutions are in flight at once, every
live operation gets three attempts with exponential backoff and its
own deadline, and multi-nameserver assertions now need a strict
majority rather than unanimity. The retry predicate is deliberately
transport-level -- "did a nameserver answer at all" -- never the
assertion under test, so a resolver that answers incorrectly still
fails on the first attempt. A nameserver that stays silent is
tolerated; one that answers wrongly is not.

livedns_harness_test.go tests that machinery directly: quorum
arithmetic, status counting, the gate's concurrency bound, per-attempt
deadlines, and recovery from a transient failure. It touches no DNS.

Nothing is mocked, faked, stubbed, recorded, skipped or build-tagged,
and production resolver behaviour is unchanged.

Test caps move to the new org-wide values ruled at prompts issue 41:
60s hard cap, 20s target, 90s -timeout backstop. REPO_POLICIES.md is
re-vendored byte-identical from sneak/prompts rather than hand-edited,
which also picks up the golangci-lint paragraph this copy had drifted
behind on. TESTING.md's stale 30-second target follows to 60.

#93
2026-08-10 13:09:17 +00:00
sneak cc86473410 build: run all linting in Docker via Dockerfile.lint (closes #134)
check / check (push) Successful in 1m17s
golangci-lint is no longer installed or run on the host. script/lint is
now a thin wrapper that builds the new root Dockerfile.lint, which COPYs
the repo into the digest-pinned golangci/golangci-lint:v2.12.2 image and
lints as a build step, so a successful build is a clean lint. This works
even where the docker daemon is remote and bind mounts are impossible.

Dockerfile.lint is split into a deps stage (base image, go mod download)
and a lint stage (source copy, linter run). script/lint passes
--no-cache-filter=lint so the lint stage executes on every invocation:
caching is explicitly waived for linting, and a cached build lints
nothing. The deps stage stays cached and no global cache invalidation is
performed. --progress=plain keeps the linter's own output visible.

golangci-lint config verify is deliberately omitted: it fetches its JSON
schema over a live, unpinned HTTPS call, which would make linting
network-dependent and defeat hash-pinning.

script/bootstrap no longer installs golangci-lint and warns instead when
docker is absent. The goimports install stays, since script/fmt and
script/fmt-check still run it on the host.

The root Dockerfile ran make check in its builder stage, which would now
recurse into script/lint and shell out to docker build with no daemon
available. It gains its own lint stage on the same pinned image, invoked
directly, with the builder depending on it via COPY --from=lint and
running make fmt-check, make test and make build.
2026-08-10 12:37:47 +00:00
86 changed files with 10723 additions and 1899 deletions
+8 -5
View File
@@ -1,6 +1,9 @@
.git/
# .git is sent, without its config: the builder stage derives the version it
# stamps into the binary from it, and `git describe` does not need the config,
# which can hold a credential (a password in the remote URL, a CI token). No
# tracked file may be listed here: git in the build would see it as deleted
# and mark the version -dirty, and an excluded .md would silently drop out of
# the prettier check in Dockerfile.fmt.
.git/config
bin/
*.md
LICENSE
.editorconfig
.gitignore
node_modules/
+8
View File
@@ -1,9 +1,17 @@
name: check
on: [push]
# A new push to a branch cancels that branch's older run, queued or running;
# runs on other branches, `next` and `main` among them, are left alone.
concurrency:
group: ${{ github.workflow }}-${{ github.ref }}
cancel-in-progress: true
jobs:
check:
runs-on: ubuntu-latest
steps:
# actions/checkout v4.2.2, 2026-02-28
- uses: actions/checkout@11bd71901bbe5b1630ceea73d27597364c9af683
# script/cibuild needs no token, so none is left in .git/config.
with:
persist-credentials: false
- run: script/cibuild
+1
View File
@@ -1,4 +1,5 @@
bin/
node_modules/
vendor/
data/
.env
+70 -2
View File
@@ -10,14 +10,20 @@ run:
linters:
default: all
enable:
# Successor to the deprecated gomodguard. Named explicitly, rather than
# left to `default: all`, because it carries the module policy below.
- gomodguard_v2
disable:
# Genuinely incompatible with project patterns
- exhaustruct # Requires all struct fields
- depguard # Dependency allow/block lists
- godot # Requires comments to end with periods
- wsl # Deprecated, replaced by wsl_v5
- wrapcheck # Too verbose for internal packages
- varnamelen # Short names like db, id are idiomatic Go
# Deprecated: the warning is attached to the old name, so it is
# silenced by disabling that name, not by enabling the successor.
- wsl # Deprecated, replaced by wsl_v5
- gomodguard # Deprecated, replaced by gomodguard_v2
settings:
lll:
line-length: 88
@@ -28,6 +34,68 @@ linters:
max-complexity: 15
dupl:
threshold: 100
depguard:
# Test-support code must not be compiled into the shipped binary. A
# test-support package exists to hand a test privileges the program
# itself must never have, so a file that is not a test must not import
# one. Test files, and the files inside a package whose directory name
# ends in `test`, are where that code belongs, and are exempt.
#
# The deny list below is the one part of this file a repository is
# expected to extend, and the only part it may. depguard matches an
# import path against a list of prefixes, so it cannot be told "any path
# whose last segment ends in test"; a repository's own test-support
# packages have to be named here one at a time, by full import path,
# under a module path that differs from repository to repository. Add
# them; change nothing else.
rules:
test-support:
list-mode: lax
files:
- "$all"
- "!$test"
- "!**/*test/**"
deny:
- pkg: net/http/httptest
desc: >-
Test-support code belongs in test files and in packages whose
directory name ends in test, not in the shipped binary.
- pkg: sneak.berlin/go/dnswatcher/internal/livednstest
desc: >-
Live-DNS test support belongs in test files and in packages
whose directory name ends in test, not in the shipped binary.
# Only decisions already recorded in the Go package defaults are
# listed here. Every entry matches the module path exactly.
gomodguard_v2:
blocked:
- module: github.com/rs/zerolog
recommendations:
- log/slog
reason: "Structured logging is stdlib log/slog."
# One entry per pre-fork module path, because the later releases
# are separate paths. A prefix match would be shorter but would
# also reach github.com/go-redis/redismock, the test double for
# the successor these entries recommend.
- module: github.com/go-redis/redis
recommendations:
- github.com/redis/go-redis/v9
reason: "Pre-fork module; use the maintained go-redis v9."
- module: github.com/go-redis/redis/v7
recommendations:
- github.com/redis/go-redis/v9
reason: "Pre-fork module; use the maintained go-redis v9."
- module: github.com/go-redis/redis/v8
recommendations:
- github.com/redis/go-redis/v9
reason: "Pre-fork module; use the maintained go-redis v9."
- module: github.com/sergi/go-diff
recommendations:
- github.com/aymanbagabas/go-udiff
reason: "No unified diff output; use go-udiff."
- module: github.com/hexops/gotextdiff
recommendations:
- github.com/aymanbagabas/go-udiff
reason: "Unmaintained fork; use go-udiff."
issues:
max-issues-per-linter: 0
+5
View File
@@ -0,0 +1,5 @@
bin/
data/
node_modules/
.claude/
static/css/tailwind.min.css
+4
View File
@@ -0,0 +1,4 @@
{
"tabWidth": 4,
"proseWrap": "always"
}
+77 -37
View File
@@ -1,13 +1,12 @@
# Build stage
# golang 1.25-alpine, 2026-02-28
FROM golang@sha256:f6751d823c26342f9506c03797d2527668d095b0a15f1862cddb4d927a7a4ced AS builder
RUN apk add --no-cache git make gcc musl-dev binutils-gold
# golangci-lint v2.12.2, 2026-08-07
RUN go install github.com/golangci/golangci-lint/v2/cmd/golangci-lint@c0d3ddc9cf3faa61a4e378e879ece580256d76e5
# goimports v0.42.0
RUN go install golang.org/x/tools/cmd/goimports@009367f5c17a8d4c45a961a3a509277190a9a6f0
# Lint stage - fast feedback on lint issues, before the build starts.
# The linter is invoked directly rather than through `make lint`: that
# target shells out to `docker build -f Dockerfile.lint`, and there is
# no docker daemon inside a docker build. For the same reason this stage
# runs only the Go half of `make fmt-check`; script/cibuild runs the
# markdown half after this build.
# script/cibuild and script/docker name this stage in --no-cache-filter.
# golangci/golangci-lint:v2.12.2 (Debian-based), 2026-08-10
FROM golangci/golangci-lint:v2.12.2@sha256:5cceeef04e53efe1470638d4b4b4f5ceefd574955ab3941b2d9a68a8c9ad5240 AS lint
WORKDIR /src
COPY go.mod go.sum ./
@@ -15,44 +14,85 @@ RUN go mod download
COPY . .
# Run all checks - build fails if any check fails.
#
# CHECK_EPOCH is a cache-busting build argument. Without it, an
# unchanged tree leaves this layer's cache key identical and Docker
# serves the previous verdict instead of re-running the suite, so the
# build reports a green it did not earn. A build argument's value
# participates in the cache key of later instructions in the stage even
# when they do not reference it, so a fresh value busts this layer
# either way. It is expanded into the command deliberately: that makes
# the invalidation a property of the command string itself rather than
# of how a given builder treats unreferenced args, and it surfaces the
# epoch in the build log as a diagnostic.
#
# Placing the ARG here and nowhere earlier keeps everything above it
# (toolchain install, go mod download) cached, so only the check and the
# steps after it re-run. script/cibuild passes a fresh value per run; a
# plain `docker build` without it caches as before.
ARG CHECK_EPOCH
RUN echo "check epoch: ${CHECK_EPOCH}" && make check
RUN script/fmt-check-go
RUN golangci-lint run --config .golangci.yml ./...
# Build stage
# script/cibuild and script/docker name this stage in --no-cache-filter.
# golang 1.25-alpine, 2026-02-28
FROM golang@sha256:f6751d823c26342f9506c03797d2527668d095b0a15f1862cddb4d927a7a4ced AS builder
RUN apk add --no-cache git make gcc musl-dev binutils-gold
# A build context sent as a tar archive keeps its files' owners, and git
# refuses to read a checkout owned by another user. Trust this one
# whoever owns it.
RUN git config --system --add safe.directory /src
# Force BuildKit to run the lint stage before proceeding
COPY --from=lint /src/go.sum /dev/null
WORKDIR /src
COPY go.mod go.sum ./
RUN go mod download
COPY . .
# Run the tests - build fails if any test fails
RUN make test
# Version stamped into the binary: the VERSION build arg when one is
# given and not empty (script/docker passes one), otherwise what
# `git describe` says of the .git in the build context, so a plain
# `docker build .` of a clone stamps its tag or short commit. The build
# arg reaches make through the environment.
ARG VERSION
# A context that carries .git, as a directory or as a file, must yield a
# real version: one that is empty, `dev` or `unknown` cannot be traced
# back to a commit.
RUN version="$(make version)"; \
if [ -e .git ]; then \
case "$version" in \
"" | dev | unknown) \
echo "version is \"$version\" although the build context carries .git" >&2; \
exit 1 ;; \
esac; \
fi
# Build the binary
RUN make build
# Runtime stage
# alpine 3.21, 2026-02-28
FROM alpine@sha256:c3f8e73fdb79deaebaa2037150150191b9dcbfba68b4a46d70103204c53f4709
RUN apk add --no-cache ca-certificates tzdata
RUN apk add --no-cache ca-certificates tzdata su-exec
WORKDIR /app
COPY --from=builder /src/bin/dnswatcher /usr/local/bin/dnswatcher
COPY deploy/docker-entrypoint.sh /usr/local/bin/docker-entrypoint.sh
COPY --from=builder /src/bin/dnswatcher /app/dnswatcher
# Create data directory
RUN mkdir -p /var/lib/dnswatcher
# dnswatcher runs as this unprivileged user. The entrypoint creates the
# data directory and gives it to this user on every start.
RUN addgroup -S -g 10001 dnswatcher \
&& adduser -S -G dnswatcher -u 10001 dnswatcher
ENV DNSWATCHER_DATA_DIR=/var/lib/dnswatcher
# Config loading also reads a `.env` file and a file named `dnswatcher`
# (any config extension, or none) from the working directory. `/` holds
# neither, so every setting comes from the environment. Do not make the
# data directory, or the binary's directory, the working directory.
WORKDIR /
# No USER: the entrypoint must start as root to set up the data
# directory; it then runs dnswatcher as the dnswatcher user.
EXPOSE 8080
ENTRYPOINT ["/app/dnswatcher"]
# busybox wget (already in alpine) probes the health endpoint every 10
# seconds, so the container is healthy well before upaas reads its health
# 60 seconds after a deploy and fails the deploy unless it is healthy.
HEALTHCHECK --interval=10s --timeout=5s --start-period=10s --retries=3 \
CMD wget -q -O /dev/null "http://127.0.0.1:${PORT:-8080}/.well-known/healthcheck" || exit 1
ENTRYPOINT ["/usr/local/bin/docker-entrypoint.sh"]
+54
View File
@@ -0,0 +1,54 @@
# prettier over the markdown, in a container, so it is never installed
# on the host. script/fmt-check-markdown builds the fmt-check stage;
# script/fmt builds fmt-out and takes the formatted files back.
# node:22-bookworm-slim, 2026-09-05
FROM node:22-bookworm-slim@sha256:83f487e0a63425e5b4d146fb5e5be574bcbe1b7b843d3ebafdd95eaf7767a7e5 AS nodedeps
# prettier lives outside /src so that a `COPY . .` of the repo cannot
# overwrite it, and so that node_modules never appears in the tree
# prettier is about to walk.
WORKDIR /tools
# package.json pins the version and yarn.lock pins the bytes:
# --frozen-lockfile installs exactly the lockfile's resolution and fails
# if package.json disagrees with it, so the tool cannot float between
# runs. yarn is the one in the image above.
COPY package.json yarn.lock ./
RUN yarn install --frozen-lockfile --non-interactive --no-progress
ENV PATH="/tools/node_modules/.bin:${PATH}"
WORKDIR /src
# Read-only markdown check. Must match $stage in
# script/fmt-check-markdown.
FROM nodedeps AS fmt-check
COPY . .
# --config, not discovery: a .prettierrc that failed to arrive would
# otherwise leave prettier on its defaults, where proseWrap is "preserve"
# and every wrap this check exists to enforce passes. Missing the file is
# a hard error instead. --no-editorconfig so that .prettierrc alone sets
# the style.
RUN prettier --config .prettierrc --no-editorconfig --check "**/*.md"
# Write path. Not a check: script/fmt builds this and takes the files.
FROM nodedeps AS fmt
COPY . .
RUN prettier --config .prettierrc --no-editorconfig --write "**/*.md"
# Only the markdown leaves, with its paths intact, so that the export
# below cannot put anything else back over the caller's working tree.
RUN mkdir -p /out && cd /src && \
find . -name '*.md' -type f -exec cp --parents '{}' /out/ ';'
# Export target: `docker build --target fmt-out --output type=local`
# writes /out's tree into a directory on the client, which is how
# script/fmt gets formatted markdown back without a bind mount.
# Must match $stage in script/fmt.
FROM scratch AS fmt-out
COPY --from=fmt /out/ /
+29
View File
@@ -0,0 +1,29 @@
# Lint-only image: used by script/lint. golangci-lint is never run on
# the host — the repo is COPYed into the build context and the linter
# runs as a build step, so a successful build IS a clean lint. This
# also works where the docker daemon is remote and bind mounts are
# impossible.
#
# `golangci-lint config verify` is deliberately NOT run here: it
# fetches its JSON schema over a live, unpinned HTTPS call, which would
# make linting network-dependent and defeat hash-pinning. The cost of
# that: unknown top-level keys in .golangci.yml are silently ignored,
# so a mistyped or wrong-schema key lints clean while applying nothing.
#
# golangci/golangci-lint:v2.12.2 (Debian-based), 2026-08-10
FROM golangci/golangci-lint:v2.12.2@sha256:5cceeef04e53efe1470638d4b4b4f5ceefd574955ab3941b2d9a68a8c9ad5240 AS deps
WORKDIR /src
# Dependencies first, so this stage stays cached across lint runs.
COPY go.mod go.sum ./
RUN go mod download
# Everything below is invalidated on every run by the
# --no-cache-filter=lint that script/lint passes: caching is explicitly
# waived for linting, and a cached build lints nothing.
FROM deps AS lint
COPY . .
RUN golangci-lint run --config .golangci.yml ./...
+21
View File
@@ -0,0 +1,21 @@
MIT License
Copyright (c) 2026 sneak
Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:
The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
SOFTWARE.
+12 -2
View File
@@ -1,7 +1,13 @@
.PHONY: all bootstrap setup build lint fmt fmt-check test check clean hooks docker
.PHONY: all bootstrap setup build version lint fmt fmt-check test check clean hooks docker
BINARY := dnswatcher
VERSION := $(shell git describe --tags --always --dirty 2>/dev/null || echo "dev")
# VERSION given on the command line (`make build VERSION=...`) or in the
# environment, which is how the Dockerfile's VERSION build arg arrives,
# wins over what `git describe` says of this checkout. An empty one counts
# as not given; `override` is what replaces an empty command-line value.
ifeq ($(VERSION),)
override VERSION := $(shell git describe --tags --always --dirty 2>/dev/null || echo "dev")
endif
LDFLAGS := -X main.Version=$(VERSION)
# Standard targets are thin shims; the implementations live in script/
@@ -19,6 +25,10 @@ setup:
build:
go build -ldflags "$(LDFLAGS)" -o bin/$(BINARY) ./cmd/dnswatcher
# Prints the version `make build` stamps; the Dockerfile checks it.
version:
@echo "$(VERSION)"
test:
@script/test
+655 -281
View File
File diff suppressed because it is too large Load Diff
+14 -6
View File
@@ -1,6 +1,6 @@
---
title: Repository Policies
last_modified: 2026-07-06
last_modified: 2026-08-07
---
This document covers repository structure, tooling, and workflow standards. Code
@@ -189,8 +189,13 @@ style conventions are in separate documents:
module under test to verify it compiles/parses. There is no excuse for
`make test` to be a no-op.
- `make test` must complete in under 20 seconds. Add a 30-second timeout in the
Makefile.
- `make test` must complete in under 60 seconds. That is the hard cap, and a
suite that exceeds it fails. Under 20 seconds is the target. A suite between
20 and 60 seconds is still green, but the overage must be filed as an
improvement bug against that repo. Add a 90-second timeout to the test
invocation in the Makefile (`go test -timeout 90s`). The backstop deliberately
sits above the hard cap so that it catches a genuinely hung test rather than a
merely slow one.
- **`make test` should use the conditional verbose rerun pattern.** Run tests
without `-v` (verbose) first. If tests fail, automatically rerun with `-v` to
@@ -209,9 +214,9 @@ style conventions are in separate documents:
```makefile
test:
@go test -timeout 30s -race -cover ./... || \
@go test -timeout 90s -race -cover ./... || \
{ echo "--- Rerunning with -v for details ---"; \
go test -timeout 30s -race -v ./...; exit 1; }
go test -timeout 90s -race -v ./...; exit 1; }
```
Python example:
@@ -260,7 +265,10 @@ style conventions are in separate documents:
- `.golangci.yml` is standardized and must _NEVER_ be modified by an agent, only
manually by the user. Fetch from
`https://git.eeqj.de/sneak/prompts/raw/branch/main/.golangci.yml`.
`https://git.eeqj.de/sneak/prompts/raw/branch/main/.golangci.yml`. The
canonical golangci-lint version is v2.12.2 (released 2026-05-06), installed
commit-pinned via
`go install github.com/golangci/golangci-lint/v2/cmd/golangci-lint@c0d3ddc9cf3faa61a4e378e879ece580256d76e5`.
- When pinning images or packages by hash, add a comment above the reference
with the version and date (YYYY-MM-DD).
+26 -16
View File
@@ -2,33 +2,43 @@
## DNS Resolution Tests
All resolver tests **MUST** use live queries against real DNS servers.
No mocking of the DNS client layer is permitted.
DNS is never mocked in this project, not in tests and not anywhere else; see the
README section "No DNS mocking. Ever." Every test that looks something up in DNS
**MUST** query live DNS servers, never a stand-in. Logic that works on record
data, such as comparing or formatting records, may be tested on that data
directly with no lookup.
### Rationale
The resolver performs iterative resolution from root nameservers through
the full delegation chain. Mocked responses cannot faithfully represent
the variety of real-world DNS behavior (truncation, referrals, glue
records, DNSSEC, varied response times, EDNS, etc.). Testing against
real servers ensures the resolver works correctly in production.
The resolver performs iterative resolution from root nameservers through the
full delegation chain. Mocked responses cannot faithfully represent the variety
of real-world DNS behavior (truncation, referrals, glue records, DNSSEC, varied
response times, EDNS, etc.). Testing against real servers ensures the resolver
works correctly in production.
### Constraints
- Tests hit real DNS infrastructure and require network access
- Test duration depends on network conditions; timeout tuning keeps
the suite within the 30-second target
- Query timeout is calibrated to 3× maximum antipodal RTT (~300ms)
plus processing margin
- Test duration depends on network conditions; timeout tuning keeps the suite
within the 60-second target
- Query timeout is calibrated to 3× maximum antipodal RTT (~300ms) plus
processing margin
- Root server fan-out is limited to reduce parallel query load
- Flaky failures from transient network issues are acceptable and
should be investigated as potential resolver bugs, not papered over
with mocks or skip flags
- Live lookups that expect an answer go through `internal/livednstest`, which
limits how many run at once in a test binary and retries a lookup that got
none
- Flaky failures from transient network issues are acceptable and should be
investigated as potential resolver bugs, not papered over with mocks or skip
flags
### What NOT to do
- **Do not mock `DNSClient`** for resolver tests (the mock constructor
exists for unit-testing other packages that consume the resolver)
- **Do not mock, fake or stub DNS** anywhere: no stand-in `DNSClient`, no
stand-in for the watcher's `DNSResolver`, no fake DNS server, no canned
responses
- **Do not add `-short` flags** to skip slow tests
- **Do not increase `-timeout`** to hide hanging queries
- **Do not remove `-count=1` from `script/test`** — Go's test cache replays a
previous run's output without querying anything, so a cached pass is not
evidence that live resolution works
- **Do not modify linter configuration** to suppress findings
+154 -143
View File
@@ -1,163 +1,174 @@
# Workflow
* branch (from `main`)
* do the work in Next Step
* move Next Step to the top of Completed Steps
* move the top item of Future Steps into Next Step
* commit (`TODO.md` changes in the same commit as the work)
* merge to `main` if the branch is not protected, otherwise open a PR
* push
- branch (from `next`)
- do the work in Next Step
- move Next Step to the top of Completed Steps
- move the top item of Future Steps into Next Step
- commit (`TODO.md` changes in the same commit as the work)
- push
- open a PR against `next`
# Status
pre-1.0. No git tags. Core resolver work in flight on feature/resolver
(dirty: internal/resolver/resolver_test.go). Local checkout has diverged
from origin: origin/main is 8 commits ahead (watcher orchestrator,
unified TARGETS) and origin/feature/resolver already contains the full
iterative resolver implementation with hermetic mocked tests.
pre-1.0. No git tags. Work lands on `next` by PR. Open work for 1.0 is tracked
on the 1.0 milestone: https://git.eeqj.de/sneak/dnswatcher/milestone/7
# Next Step
Policy scaffold commit: add LICENSE, REPO_POLICIES.md, .editorconfig,
.dockerignore, and .gitea/workflows/check.yml, and add the missing
fmt-check, docker, and hooks targets to the Makefile. One commit, then
confirm make check still passes.
trial run of the finished image: https://git.eeqj.de/sneak/dnswatcher/issues/149
# Completed Steps
- 2026-08-09: `script/cibuild` can no longer report a green it did not
earn. The Dockerfile declares `ARG CHECK_EPOCH` immediately above the
check step and expands it into the `RUN` command, and `script/cibuild`
passes a fresh `$(date +%s%N)` per invocation, so the `make check`
layer is always re-executed while the pinned toolchain install and
`go mod download` stay cached. Verified by experiment: before the fix
a second run on an unchanged tree returned in 283 ms with the check
layer `CACHED`; after it the check runs every time, and a deliberately
planted always-failing test made the build fail with exactly that
test's message
- 2026-08-07: golangci-lint bumped to v2.12.2 (commit-pinned installs
in `Dockerfile` and `script/bootstrap`); `.golangci.yml` set to the
org-standard v2-schema config used across the org's repos
(owner-authorized; same file is being landed as canonical via prompts
PR #24), with settings under `linters.settings` so the
lll/funlen/cyclop/dupl thresholds apply; fixed the resulting
`goconst`, `dupl`, and `lll` findings; the informational `gomodguard`
deprecation warning under this config is accepted
- 2026-07-07 Adopted scripts-to-rule-them-all: `script/` entrypoints,
Makefile shims, README Entrypoints section
- 2026-02-20: iterative DNS resolver implemented; tests made hermetic
with mocked DNS (origin/feature/resolver, unmerged)
- 2026-02-20: CI actions and go install refs pinned to commit SHAs;
Gitea Actions workflow for make check (origin/ci/make-check, unmerged)
- 2026-10-02: a name removed from `DNSWATCHER_TARGETS` leaves the state, and so
the dashboard and API, at startup, before the first check (closes #223).
- 2026-10-02: a Port Change notification lists the port's domains on a
`Domains:` line and its hostnames on a `Hostnames:` line (closes #248).
- 2026-10-02: the dashboard's Ports table and `/api/v1/status` port entries list
a port's domains apart from its hostnames (closes #245).
- 2026-10-02: nameservers a referral names without addresses are looked up,
three deep at most; `pool.ntp.org`'s nameservers resolve (closes #221).
- 2026-10-02: an apex domain is not counted or listed as a hostname; its records
show under Domains, and notifications about them say `Domain:` (closes #224).
- 2026-10-02: the dashboard lists each nameserver's record types in one fixed
order, the README's, then any other type, not a random one (closes #226).
- 2026-10-02: the dashboard and `/api/v1/status` show why a nameserver query or
a certificate check failed, which only the state file showed (closes #225).
- 2026-10-02: a name's CNAME is stored once per nameserver, not once per record
type asked for; a state file with repeats loads each value once (closes #220).
- 2026-10-02: a DNS lookup that shutdown cuts short logs no error; one that
fails otherwise, or runs out of time, still does (closes #229).
- 2026-10-02: Record Change and Inconsistency notifications list only the record
types that differ, each with its values as plain text (closes #219).
- 2026-10-02: the startup notification no longer says every notification
endpoint works; it says it is a test sent to each of them (closes #230).
- 2026-10-02: a Mattermost webhook that answers an HTTP error is logged as
`mattermost notification failed`, not as a Slack failure (closes #227).
- 2026-10-02: durations in the log are written as text such as `2m0s`, not as a
bare count of nanoseconds (closes #228).
- 2026-10-02: a watched name whose nameservers answer with a CNAME and no
address gets port and TLS checks at the end of its CNAME chain (closes #203).
- 2026-10-02: a resolver test that reads one record type from a nameserver's
answer asks again when that type is missing from it (closes #218).
- 2026-10-02: a plain `docker build .` of a clone stamps its tag or short
commit, not `dev`: the build context now carries `.git` (closes #210).
- 2026-10-02: a query a server refuses is not resent asking for recursion, and
every root server refusing is reported as DNS interception (closes #206).
- 2026-10-02: a push to a branch cancels that branch's older CI run, and the
checkout leaves no token in `.git/config` (closes #216).
- 2026-10-02: watcher tests send far fewer queries and a live attempt may take
18s; nameserver addresses are asked only for A, AAAA, CNAME (closes #214).
- 2026-10-02: the resolver tries root servers, and every other server list it
walks, in a random order each time, not always from the top (closes #138).
- 2026-10-02: a name listed more than once in `DNSWATCHER_TARGETS`, in any
letter case or with a trailing dot, is watched once (closes #207).
- 2026-10-01: README checked against the code and corrected: metrics, CORS,
notification retries, CNAMEs, state file fields, Design tree (closes #108).
- 2026-10-01: a certificate within the expiry warning period is warned about on
every TLS check, where some checks used to skip it at random (closes #204).
- 2026-10-01: a domain's NS set is its delegation from the parent zone's
servers, not whichever of its own servers answered first (closes #200).
- 2026-10-01: README has Getting Started, Rationale and TODO sections, and its
Architecture section is now Design, in the order policy sets (closes #173).
- 2026-10-01: a zone's server that answers SERVFAIL or a referral leading no
closer is passed over for the next, as one that times out is (closes #197).
- 2026-10-01: when none of a configured name's nameservers answered, the port
state saved for its addresses is kept, not removed (closes #193).
- 2026-10-01: `ResolveIPAddresses` returns an error, not no addresses, when no
nameserver of the name's zone answered (closes #190).
- 2026-10-01: `make fmt` and `make fmt-check` cover Markdown with prettier, run
in Docker at the version pinned by `yarn.lock` (closes #119).
- 2026-10-01: `make fmt-check` fails on a file `goimports` would change; both
format scripts run `goimports` at its pinned commit, not from `PATH` (#119).
- 2026-10-01: a hostname is queried at the servers of the zone it is in, found
by following delegations for the name, not its last two labels (closes #189).
- 2026-10-01: each nameserver's addresses are saved with its domain, and a
change while it stays in the delegation is notified (closes #105).
- 2026-10-01: the watcher saves state when it stops, and shutdown waits for that
save, so it no longer relies on the state's own stop hook (closes #114).
- 2026-10-01: `DNSWATCHER_SENTRY_DSN` reports panics in HTTP handlers to Sentry,
and a DSN Sentry cannot parse stops startup (closes #107).
- 2026-10-01: a port or TLS check that shutdown cuts short saves nothing and
sends no notification, as a cut-short DNS lookup already did (closes #185).
- 2026-10-01: the client address from `X-Forwarded-For` is the last entry that
is not a trusted proxy, not the first, which the client sets (closes #181).
- 2026-10-01: a nameserver that does not answer is saved as `error` with the
reason, and NS failure and NS recovery are notified (closes #104).
- 2026-10-01: a `DNSWATCHER_DNS_INTERVAL` or `DNSWATCHER_TLS_INTERVAL` that is
not a positive duration stops startup; empty means the default (closes #177).
- 2026-10-01: `/metrics` allows each client address 30 requests a minute,
counted before Basic Auth, and answers 429 beyond that (closes #101).
- 2026-10-01: the image built by `make docker` reports the `git describe`
version, not `dev`, and the startup log now shows it (closes #109).
- 2026-10-01: two notify shutdown tests always release the delivery they hold,
so a drain that returns early fails them instead of hanging (closes #176).
- 2026-10-01: `script/install-precommit` asks git for the repository's git
directory, so `make hooks` also works where `.git` is a file (closes #129).
- 2026-10-01: `TODO.md` brought up to date: open issues listed by URL, every
Completed Steps entry cut to at most two lines (closes #146).
- 2026-10-01: wildcard CORS now applies only to the public routes, not to
`/metrics`, and allows only the methods they serve (closes #100).
- 2026-10-01: `internal/state` and `internal/watcher` no longer export test-only
constructors: two moved to `export_test.go`, one is deleted (closes #111).
- 2026-10-01: notify shutdown tests use one timing constant per meaning, name
the bound they check, and require the drain's debug line (closes #116).
- 2026-09-29: the entrypoint chowns the data directory to `dnswatcher` and runs
dnswatcher as that user, so a host bind mount needs no chown (closes #166).
- 2026-09-29: the live-DNS test package is renamed `internal/livednstest`;
`make lint` fails when program code imports it (closes #164).
- 2026-09-29: `.golangci.yml` re-fetched from `sneak/prompts`, with
`gomodguard_v2` and the org `depguard` `test-support` rule (closes #123).
- 2026-09-29: watcher and resolver tests that look something up in DNS use the
real resolver against live DNS servers (closes #159).
- 2026-09-28: the inconsistency alert is sent once, when two nameservers start
to disagree; every pair of nameservers is compared (closes #158).
- 2026-09-28: DNS names in record values (CNAME, MX, SRV and NS targets) are
lower-cased, so letter case alone is not a change (closes #157).
- 2026-09-28: lint and tests run on every build: `script/cibuild` and
`script/docker` pass `--no-cache-filter=lint,builder` (closes #115).
- 2026-09-28: the server timeout test drives `Run` and checks the timeouts on
the `http.Server` it serves (closes #120).
- 2026-09-28: upaas deploy readiness: the image runs as user `dnswatcher` with a
`HEALTHCHECK`; README "Running under upaas" (closes #147).
- 2026-09-21: added behavioural tests for `internal/globals`,
`internal/healthcheck`, and `internal/logger` (closes #110).
- 2026-09-21: `go mod tidy` dropped the redundant `golang.org/x/sync`
`// indirect` line so `script/bootstrap` leaves a clean tree (#132)
- 2026-08-10: comment-only corrections to `script/bootstrap`, `script/cibuild`
and `Dockerfile.lint`; no behaviour changed.
- 2026-08-10: MIT `LICENSE` added at the repository root; the README's first
line and License section name the licence.
- 2026-08-10: policy scaffold present: `REPO_POLICIES.md`, `.editorconfig`,
`.dockerignore`, CI workflow, `make fmt-check`, `make docker`, `make hooks`.
- 2026-08-10: Go's test cache disabled in `script/test` (`-count=1`), so every
run queries live DNS; a failed run is rerun with `-v`.
- 2026-08-10: live-DNS tests made robust rather than gated (#93): a limit on
concurrent lookups, retries, and a quorum across nameservers.
- 2026-08-10: all linting moved into Docker: `script/lint` builds
`Dockerfile.lint`, and the root `Dockerfile` has its own lint stage.
- 2026-08-09: in-flight notification deliveries are drained at shutdown, bounded
by the shutdown deadline (#106).
- 2026-08-09: `http.Server` sets all four socket timeouts; `WriteTimeout` stays
above the 60s handler timeout (#99).
- 2026-08-09: `SecurityHeaders()` middleware sets HSTS, CSP and the other
security headers `REPO_POLICIES.md` requires on every response.
- 2026-08-07: golangci-lint bumped to v2.12.2 and `.golangci.yml` set to the org
config; fixed the resulting `goconst`, `dupl` and `lll` findings.
- 2026-07-07 Adopted scripts-to-rule-them-all: `script/` entrypoints, Makefile
shims, README Entrypoints section
- 2026-02-20: iterative DNS resolver implemented
- 2026-02-20: CI actions and go install refs pinned to commit SHAs; Gitea
Actions workflow added
- 2026-02-20: watcher monitoring orchestrator merged to main (#8)
- 2026-02-20: DOMAINS/HOSTNAMES unified into single TARGETS config (#11)
- 2026-02-19: TCP port connectivity checker, made concurrent with port
validation; gosec G704 SSRF findings fixed without suppression
(feature branches, unmerged)
- 2026-02-19: TLS certificate inspector with no-peer-certificates error
path and IP SANs (feature branch, unmerged)
- 2026-02-19: TLS certificate inspector with no-peer-certificates error path and
IP SANs
- 2026-02-19: gosec SSRF and formatting fixes on main
- 2026-02-19: initial scaffold with per-nameserver DNS monitoring model
# Future Steps
Compliance:
- Add README sections required by policy (Description, Getting Started,
Rationale, Design, TODO, License, Author) if any are missing
- Pin Dockerfile base images by sha256 and ensure the Docker build runs
make check
Branch reconciliation:
- Sync local checkout with origin: local main is 8 commits behind
origin/main; local feature/resolver has diverged from
origin/feature/resolver, which already implements the resolver
- Merge in-flight branches to main once green: feature/resolver,
ci/make-check, feature/portcheck-implementation,
feature/tlscheck-implementation
Resolver (plan from untracked TODO.md; largely implemented on
origin/feature/resolver, verify each item before closing):
- Add github.com/miekg/dns dependency
- roots.go: hardcoded IANA root server list (a through m, IPv4/IPv6),
rootServers() returning ip:53 strings
- query.go: low-level query(ctx, server, name, qtype): UDP with TCP
fallback on truncation, RD=0, context respected, 5s per-query timeout,
returns raw *dns.Msg
- trace.go: iterative delegation chasing from roots: referral detection
(NOERROR, empty answer, NS in authority), glue extraction with
bailiwick check, out-of-bailiwick NS resolved with recursion guard,
delegation depth limit (20), retry across nameservers on failure, do
not chase CNAMEs inside trace
- FindAuthoritativeNameservers: NS set via trace, sorted, FQDN
normalized, trailing dot handled; must pass its 9 tests
- QueryNameserver: resolve NS host to IPs, query A/AAAA/CNAME/MX/TXT/
SRV/CAA/NS, build NameserverResponse with status mapping (OK,
NXDomain, NoData, Error), documented record formatting, sorted values,
lame delegation detection; must pass its 16 tests
- QueryAllNameservers: find NS set for parent domain (public suffix
list), query all NS in parallel with bounded concurrency, return map
even when all fail, context cancellation; must pass its 4 tests
- LookupNS: thin wrapper over FindAuthoritativeNameservers, sorted,
identical results; must pass its 3 tests
- ResolveIPAddresses: collect A/AAAA from all NS, follow CNAME chains
with MaxCNAMEDepth, dedupe, sort, NXDOMAIN returns empty slice with
nil error; must pass its 9 tests
- All 39 resolver tests pass, make check green, merge to main
Watcher (internal/watcher/watcher.go):
- Scheduling loop in Run(ctx): initial check on startup, separate
tickers for DNS/port and TLS intervals, persist state via state.Save()
after each cycle, clean shutdown on context cancel
- Domain check: LookupNS, compare to stored state, store silently on
first run, notify with old/new NS lists on change
- Hostname check: QueryAllNameservers, compare per-NS records; notify on
record changes, NS failure, NS recovery, inconsistency detected,
inconsistency resolved, empty response; store silently on first run
- Port check: ResolveIPAddresses, check ports 80 and 443 per IP, notify
on open/closed transitions, handle new and disappeared IPs
- TLS check: for each open IP:443, CheckCertificate; notify on expiry
warning, certificate change (CN/issuer/SANs), TLS failure/recovery
Port checker (internal/portcheck/portcheck.go):
- Tests against known-open ports and RFC documentation IPs
- CheckPort: net.DialTimeout (5s), context respected; (true, nil) open,
(false, nil) closed/timeout/refused, error only for unexpected
failures
TLS checker (internal/tlscheck/tlscheck.go):
- Tests against known public HTTPS servers, verify fields populated
- CheckCertificate: tls.Dial to specific IP:443 with hostname as SNI;
extract subject CN, issuer CN and org, NotAfter, SANs; error on
handshake failure
Notification service (internal/notify/notify.go, Slack/Mattermost/ntfy
backends exist):
- Structured notification types: DNS change, port change, TLS expiry,
TLS change, NS failure, NS recovery, NS inconsistency
- Per-backend formatting: Slack/Mattermost attachment colors (red
failures/expiry, yellow warnings, green recoveries, blue info); ntfy
priorities (urgent failures, high warnings, default changes, low
recoveries); include hostname, nameserver, old/new values, timestamps
HTTP API handlers:
- Wire *state.State and *watcher.Watcher into handler params
- GET /api/v1/status: full state snapshot as JSON
- GET /api/v1/domains: domain states with NS records and last-checked
- GET /api/v1/hostnames: hostname states with per-NS record data
Infrastructure notes (from untracked TODO.md):
- Module path sneak.berlin/go/dnswatcher differs from the git.eeqj.de
remote intentionally; do not "fix" it
- Dependencies: github.com/miekg/dns, golang.org/x/net/publicsuffix
- Resolver tests originally used live DNS against *.dns.sneak.cloud
(required records documented in the test file header); origin now has
mocked hermetic tests, keep them hermetic
- 1.0 readiness: run it with a real config and read the logs:
https://git.eeqj.de/sneak/dnswatcher/issues/66
- review toward 1.0: https://git.eeqj.de/sneak/dnswatcher/issues/144
+1
View File
@@ -63,6 +63,7 @@ func main() {
return n
},
),
fx.Invoke(func(l *logger.Logger) { l.Identify() }),
fx.Invoke(func(*server.Server, *watcher.Watcher) {}),
).Run()
}
+17
View File
@@ -0,0 +1,17 @@
#!/bin/sh
# deploy/docker-entrypoint.sh: the Docker image's ENTRYPOINT. It runs as
# root only to give the data directory to the dnswatcher user: a host
# directory bind-mounted there keeps its host owner, often root, and may
# hold a state file left by another uid, which dnswatcher could neither
# read nor replace. dnswatcher itself always runs as the dnswatcher user.
set -eu
main() {
dir="${DNSWATCHER_DATA_DIR:-/var/lib/dnswatcher}"
mkdir -p "$dir"
chown -R dnswatcher:dnswatcher "$dir"
chmod 700 "$dir"
exec su-exec dnswatcher /usr/local/bin/dnswatcher "$@"
}
main "$@"
+12 -9
View File
@@ -4,27 +4,30 @@ go 1.25.5
require (
github.com/99designs/basicauth-go v0.0.0-20230316000542-bf6f9cbbf0f8
github.com/getsentry/sentry-go v0.49.0
github.com/go-chi/chi/v5 v5.2.5
github.com/go-chi/cors v1.2.2
github.com/go-chi/httprate v0.16.0
github.com/joho/godotenv v1.5.1
github.com/miekg/dns v1.1.72
github.com/prometheus/client_golang v1.23.2
github.com/spf13/viper v1.21.0
github.com/stretchr/testify v1.11.1
go.uber.org/fx v1.24.0
golang.org/x/net v0.50.0
golang.org/x/sync v0.19.0
golang.org/x/net v0.56.0
golang.org/x/sync v0.21.0
)
require (
github.com/beorn7/perks v1.0.1 // indirect
github.com/cespare/xxhash/v2 v2.3.0 // indirect
github.com/davecgh/go-spew v1.1.1 // indirect
github.com/davecgh/go-spew v1.1.2-0.20180830191138-d8f796af33cc // indirect
github.com/fsnotify/fsnotify v1.9.0 // indirect
github.com/go-viper/mapstructure/v2 v2.4.0 // indirect
github.com/klauspost/cpuid/v2 v2.2.10 // indirect
github.com/munnerz/goautoneg v0.0.0-20191010083416-a7dc8b61c822 // indirect
github.com/pelletier/go-toml/v2 v2.2.4 // indirect
github.com/pmezard/go-difflib v1.0.0 // indirect
github.com/pmezard/go-difflib v1.0.1-0.20181226105442-5d4384ee4fb2 // indirect
github.com/prometheus/client_model v0.6.2 // indirect
github.com/prometheus/common v0.66.1 // indirect
github.com/prometheus/procfs v0.16.1 // indirect
@@ -34,16 +37,16 @@ require (
github.com/spf13/cast v1.10.0 // indirect
github.com/spf13/pflag v1.0.10 // indirect
github.com/subosito/gotenv v1.6.0 // indirect
github.com/zeebo/xxh3 v1.0.2 // indirect
go.uber.org/dig v1.19.0 // indirect
go.uber.org/multierr v1.10.0 // indirect
go.uber.org/zap v1.26.0 // indirect
go.yaml.in/yaml/v2 v2.4.2 // indirect
go.yaml.in/yaml/v3 v3.0.4 // indirect
golang.org/x/mod v0.32.0 // indirect
golang.org/x/sync v0.19.0 // indirect
golang.org/x/sys v0.41.0 // indirect
golang.org/x/text v0.34.0 // indirect
golang.org/x/tools v0.41.0 // indirect
golang.org/x/mod v0.37.0 // indirect
golang.org/x/sys v0.46.0 // indirect
golang.org/x/text v0.39.0 // indirect
golang.org/x/tools v0.47.0 // indirect
google.golang.org/protobuf v1.36.8 // indirect
gopkg.in/yaml.v3 v3.0.1 // indirect
)
+34 -18
View File
@@ -4,16 +4,22 @@ github.com/beorn7/perks v1.0.1 h1:VlbKKnNfV8bJzeqoa4cOKqO6bYr3WgKZxO8Z16+hsOM=
github.com/beorn7/perks v1.0.1/go.mod h1:G2ZrVWU2WbWT9wwq4/hrbKbnv/1ERSJQ0ibhJ6rlkpw=
github.com/cespare/xxhash/v2 v2.3.0 h1:UL815xU9SqsFlibzuggzjXhog7bL6oX9BbNZnL2UFvs=
github.com/cespare/xxhash/v2 v2.3.0/go.mod h1:VGX0DQ3Q6kWi7AoAeZDth3/j3BFtOZR5XLFGgcrjCOs=
github.com/davecgh/go-spew v1.1.1 h1:vj9j/u1bqnvCEfJOwUhtlOARqs3+rkHYY13jYWTU97c=
github.com/davecgh/go-spew v1.1.1/go.mod h1:J7Y8YcW2NihsgmVo/mv3lAwl/skON4iLHjSsI+c5H38=
github.com/davecgh/go-spew v1.1.2-0.20180830191138-d8f796af33cc h1:U9qPSI2PIWSS1VwoXQT9A3Wy9MM3WgvqSxFWenqJduM=
github.com/davecgh/go-spew v1.1.2-0.20180830191138-d8f796af33cc/go.mod h1:J7Y8YcW2NihsgmVo/mv3lAwl/skON4iLHjSsI+c5H38=
github.com/frankban/quicktest v1.14.6 h1:7Xjx+VpznH+oBnejlPUj8oUpdxnVs4f8XU8WnHkI4W8=
github.com/frankban/quicktest v1.14.6/go.mod h1:4ptaffx2x8+WTWXmUCuVU6aPUX1/Mz7zb5vbUoiM6w0=
github.com/fsnotify/fsnotify v1.9.0 h1:2Ml+OJNzbYCTzsxtv8vKSFD9PbJjmhYF14k/jKC7S9k=
github.com/fsnotify/fsnotify v1.9.0/go.mod h1:8jBTzvmWwFyi3Pb8djgCCO5IBqzKJ/Jwo8TRcHyHii0=
github.com/getsentry/sentry-go v0.49.0 h1:Ehejknu1l023Ub7QoRBVLAI7g3Jnhqku4oWx4B4Sh5s=
github.com/getsentry/sentry-go v0.49.0/go.mod h1:nuMJAoCfe1u0Bts2ocyNI+TW8HT84vRMqwA5Qq/SKUI=
github.com/go-chi/chi/v5 v5.2.5 h1:Eg4myHZBjyvJmAFjFvWgrqDTXFyOzjj7YIm3L3mu6Ug=
github.com/go-chi/chi/v5 v5.2.5/go.mod h1:X7Gx4mteadT3eDOMTsXzmI4/rwUpOwBHLpAfupzFJP0=
github.com/go-chi/cors v1.2.2 h1:Jmey33TE+b+rB7fT8MUy1u0I4L+NARQlK6LhzKPSyQE=
github.com/go-chi/cors v1.2.2/go.mod h1:sSbTewc+6wYHBBCW7ytsFSn836hqM7JxpglAy2Vzc58=
github.com/go-chi/httprate v0.16.0 h1:8V5DH9j6pSK6UQoBsTpvMyFxycqaKEIToyPKzHJjUa8=
github.com/go-chi/httprate v0.16.0/go.mod h1:A8lo+qRhk+s9LiuP5saS7XCGDXRXMcrueq0NfIuCa/I=
github.com/go-errors/errors v1.4.2 h1:J6MZopCL4uSllY1OfXM374weqZFFItUbrImctkmUxIA=
github.com/go-errors/errors v1.4.2/go.mod h1:sIVyrIiJhuEF+Pj9Ebtd6P/rEYROXFi3BopGUQ5a5Og=
github.com/go-viper/mapstructure/v2 v2.4.0 h1:EBsztssimR/CONLSZZ04E8qAkxNYq4Qp9LvH92wZUgs=
github.com/go-viper/mapstructure/v2 v2.4.0/go.mod h1:oJDH3BJKyqBA2TXFhDsKDGDTlndYOZ6rGS0BRZIxGhM=
github.com/google/go-cmp v0.7.0 h1:wk8382ETsv4JYUZwIsn6YpYiWiBsYLSJiTsyBybVuN8=
@@ -22,6 +28,8 @@ github.com/joho/godotenv v1.5.1 h1:7eLL/+HRGLY0ldzfGMeQkb7vMd0as4CfYvUVzLqw0N0=
github.com/joho/godotenv v1.5.1/go.mod h1:f4LDr5Voq0i2e/R5DDNOoa2zzDfwtkZa6DnEwAbqwq4=
github.com/klauspost/compress v1.18.0 h1:c/Cqfb0r+Yi+JtIEq73FWXVkRonBlf0CRNYc8Zttxdo=
github.com/klauspost/compress v1.18.0/go.mod h1:2Pp+KzxcywXVXMr50+X0Q/Lsb43OQHYWRCY2AiWywWQ=
github.com/klauspost/cpuid/v2 v2.2.10 h1:tBs3QSyvjDyFTq3uoc/9xFpCuOsJQFNPiAhYdw2skhE=
github.com/klauspost/cpuid/v2 v2.2.10/go.mod h1:hqwkgyIinND0mEev00jJYCxPNVRVXFQeu1XKlok6oO0=
github.com/kr/pretty v0.3.1 h1:flRD4NNwYAUpkphVc1HcthR4KEIFJ65n8Mw5qdRn3LE=
github.com/kr/pretty v0.3.1/go.mod h1:hoEshYVHaxMs3cyo3Yncou5ZscifuDolrwPKZanG3xk=
github.com/kr/text v0.2.0 h1:5Nx0Ya0ZqY2ygV366QzturHI13Jq95ApcVaJBhpS+AY=
@@ -34,8 +42,12 @@ github.com/munnerz/goautoneg v0.0.0-20191010083416-a7dc8b61c822 h1:C3w9PqII01/Oq
github.com/munnerz/goautoneg v0.0.0-20191010083416-a7dc8b61c822/go.mod h1:+n7T8mK8HuQTcFwEeznm/DIxMOiR9yIdICNftLE1DvQ=
github.com/pelletier/go-toml/v2 v2.2.4 h1:mye9XuhQ6gvn5h28+VilKrrPoQVanw5PMw/TB0t5Ec4=
github.com/pelletier/go-toml/v2 v2.2.4/go.mod h1:2gIqNv+qfxSVS7cM2xJQKtLSTLUE9V8t9Stt+h56mCY=
github.com/pmezard/go-difflib v1.0.0 h1:4DBwDE0NGyQoBHbLQYPwSUPoCMWR5BEzIk/f1lZbAQM=
github.com/pmezard/go-difflib v1.0.0/go.mod h1:iKH77koFhYxTK1pcRnkKkqfTogsbg7gZNVY4sRDYZ/4=
github.com/pingcap/errors v0.11.4 h1:lFuQV/oaUMGcD2tqt+01ROSmJs75VG1ToEOkZIZ4nE4=
github.com/pingcap/errors v0.11.4/go.mod h1:Oi8TUi2kEtXXLMJk9l1cGmz20kV3TaQ0usTwv5KuLY8=
github.com/pkg/errors v0.9.1 h1:FEBLx1zS214owpjy7qsBeixbURkuhQAwrK5UwLGTwt4=
github.com/pkg/errors v0.9.1/go.mod h1:bwawxfHBFNV+L2hUp1rHADufV3IMtnDRdf1r5NINEl0=
github.com/pmezard/go-difflib v1.0.1-0.20181226105442-5d4384ee4fb2 h1:Jamvg5psRIccs7FGNTlIRMkT8wgtp5eCXdBlqhYGL6U=
github.com/pmezard/go-difflib v1.0.1-0.20181226105442-5d4384ee4fb2/go.mod h1:iKH77koFhYxTK1pcRnkKkqfTogsbg7gZNVY4sRDYZ/4=
github.com/prometheus/client_golang v1.23.2 h1:Je96obch5RDVy3FDMndoUsjAhG5Edi49h0RJWRi/o0o=
github.com/prometheus/client_golang v1.23.2/go.mod h1:Tb1a6LWHB3/SPIzCoaDXI4I8UHKeFTEQ1YCr+0Gyqmg=
github.com/prometheus/client_model v0.6.2 h1:oBsgwpGs7iVziMvrGhE53c/GrLUsZdHnqNwqPLxwZyk=
@@ -44,8 +56,8 @@ github.com/prometheus/common v0.66.1 h1:h5E0h5/Y8niHc5DlaLlWLArTQI7tMrsfQjHV+d9Z
github.com/prometheus/common v0.66.1/go.mod h1:gcaUsgf3KfRSwHY4dIMXLPV0K/Wg1oZ8+SbZk/HH/dA=
github.com/prometheus/procfs v0.16.1 h1:hZ15bTNuirocR6u0JZ6BAHHmwS1p8B4P6MRqxtzMyRg=
github.com/prometheus/procfs v0.16.1/go.mod h1:teAbpZRB1iIAJYREa1LsoWUXykVXA1KlTmWl8x/U+Is=
github.com/rogpeppe/go-internal v1.10.0 h1:TMyTOH3F/DB16zRVcYyreMH6GnZZrwQVAoYjRBZyWFQ=
github.com/rogpeppe/go-internal v1.10.0/go.mod h1:UQnix2H7Ngw/k4C5ijL5+65zddjncjaFoBhdsK/akog=
github.com/rogpeppe/go-internal v1.14.1 h1:UQB4HGPB6osV0SQTLymcB4TgvyWu6ZyliaW0tI/otEQ=
github.com/rogpeppe/go-internal v1.14.1/go.mod h1:MaRKkUm5W0goXpeCfT7UZI6fk/L7L7so1lCWt35ZSgc=
github.com/sagikazarmark/locafero v0.11.0 h1:1iurJgmM9G3PA/I+wWYIOw/5SyBtxapeHDcg+AAIFXc=
github.com/sagikazarmark/locafero v0.11.0/go.mod h1:nVIGvgyzw595SUSUE6tvCp3YYTeHs15MvlmU87WwIik=
github.com/sourcegraph/conc v0.3.1-0.20240121214520-5f936abd7ae8 h1:+jumHNA0Wrelhe64i8F6HNlS8pkoyMv5sreGx2Ry5Rw=
@@ -62,6 +74,10 @@ github.com/stretchr/testify v1.11.1 h1:7s2iGBzp5EwR7/aIZr8ao5+dra3wiQyKjjFuvgVKu
github.com/stretchr/testify v1.11.1/go.mod h1:wZwfW3scLgRK+23gO65QZefKpKQRnfz6sD981Nm4B6U=
github.com/subosito/gotenv v1.6.0 h1:9NlTDc1FTs4qu0DDq7AEtTPNw6SVm7uBMsUCUjABIf8=
github.com/subosito/gotenv v1.6.0/go.mod h1:Dk4QP5c2W3ibzajGcXpNraDfq2IrhjMIvMSWPKKo0FU=
github.com/zeebo/assert v1.3.0 h1:g7C04CbJuIDKNPFHmsk4hwZDO5O+kntRxzaUoNXj+IQ=
github.com/zeebo/assert v1.3.0/go.mod h1:Pq9JiuJQpG8JLJdtkwrJESF0Foym2/D9XMU5ciN/wJ0=
github.com/zeebo/xxh3 v1.0.2 h1:xZmwmqxHZA8AI603jOQ0tMqmBr9lPeFwGg6d+xy9DC0=
github.com/zeebo/xxh3 v1.0.2/go.mod h1:5NWz9Sef7zIDm2JHfFlcQvNekmcEl9ekUZQQKCYaDcA=
go.uber.org/dig v1.19.0 h1:BACLhebsYdpQ7IROQ1AGPjrXcP5dF80U3gKoFzbaq/4=
go.uber.org/dig v1.19.0/go.mod h1:Us0rSJiThwCv2GteUN0Q7OKvU7n5J4dxZ9JKUXozFdE=
go.uber.org/fx v1.24.0 h1:wE8mruvpg2kiiL1Vqd0CC+tr0/24XIB10Iwp2lLWzkg=
@@ -76,18 +92,18 @@ go.yaml.in/yaml/v2 v2.4.2 h1:DzmwEr2rDGHl7lsFgAHxmNz/1NlQ7xLIrlN2h5d1eGI=
go.yaml.in/yaml/v2 v2.4.2/go.mod h1:081UH+NErpNdqlCXm3TtEran0rJZGxAYx9hb/ELlsPU=
go.yaml.in/yaml/v3 v3.0.4 h1:tfq32ie2Jv2UxXFdLJdh3jXuOzWiL1fo0bu/FbuKpbc=
go.yaml.in/yaml/v3 v3.0.4/go.mod h1:DhzuOOF2ATzADvBadXxruRBLzYTpT36CKvDb3+aBEFg=
golang.org/x/mod v0.32.0 h1:9F4d3PHLljb6x//jOyokMv3eX+YDeepZSEo3mFJy93c=
golang.org/x/mod v0.32.0/go.mod h1:SgipZ/3h2Ci89DlEtEXWUk/HteuRin+HHhN+WbNhguU=
golang.org/x/net v0.50.0 h1:ucWh9eiCGyDR3vtzso0WMQinm2Dnt8cFMuQa9K33J60=
golang.org/x/net v0.50.0/go.mod h1:UgoSli3F/pBgdJBHCTc+tp3gmrU4XswgGRgtnwWTfyM=
golang.org/x/sync v0.19.0 h1:vV+1eWNmZ5geRlYjzm2adRgW2/mcpevXNg50YZtPCE4=
golang.org/x/sync v0.19.0/go.mod h1:9KTHXmSnoGruLpwFjVSX0lNNA75CykiMECbovNTZqGI=
golang.org/x/sys v0.41.0 h1:Ivj+2Cp/ylzLiEU89QhWblYnOE9zerudt9Ftecq2C6k=
golang.org/x/sys v0.41.0/go.mod h1:OgkHotnGiDImocRcuBABYBEXf8A9a87e/uXjp9XT3ks=
golang.org/x/text v0.34.0 h1:oL/Qq0Kdaqxa1KbNeMKwQq0reLCCaFtqu2eNuSeNHbk=
golang.org/x/text v0.34.0/go.mod h1:homfLqTYRFyVYemLBFl5GgL/DWEiH5wcsQ5gSh1yziA=
golang.org/x/tools v0.41.0 h1:a9b8iMweWG+S0OBnlU36rzLp20z1Rp10w+IY2czHTQc=
golang.org/x/tools v0.41.0/go.mod h1:XSY6eDqxVNiYgezAVqqCeihT4j1U2CCsqvH3WhQpnlg=
golang.org/x/mod v0.37.0 h1:vF1DjpVEshcIqoEaauuHebaLk1O1forxjxBaVn884JQ=
golang.org/x/mod v0.37.0/go.mod h1:m8S8VeM9r4dzDwjrKO0a1sZP3YjeMamRRlD+fmR2Q/0=
golang.org/x/net v0.56.0 h1:Rw8j/hFzGvJUZwNBXnAtf5sVDVt+65SK2C7IxCxZt5o=
golang.org/x/net v0.56.0/go.mod h1:D3Ku6r+V6JROoZK144D2XfMHFcMq/0zSfLelVTCFKec=
golang.org/x/sync v0.21.0 h1:HLII4xRRTtCRkxYp4HNFF0Js/Og6q2i++KXbg0gHCwM=
golang.org/x/sync v0.21.0/go.mod h1:9xrNwdLfx4jkKbNva9FpL6vEN7evnE43NNNJQ2LF3+0=
golang.org/x/sys v0.46.0 h1:noSf2Fq6F8DBgS+LysIkx7rIExoNHJsxOAtPp4rthXw=
golang.org/x/sys v0.46.0/go.mod h1:4GL1E5IUh+htKOUEOaiffhrAeqysfVGipDYzABqnCmw=
golang.org/x/text v0.39.0 h1:UbZz4pLOvn600D6Oh6GGEI6VAmndrEBLv8/6BEXzyus=
golang.org/x/text v0.39.0/go.mod h1:3UwRclnC2g0TU9x8PZiyfOajCd1zaUNHF9cvqcQZ+ZM=
golang.org/x/tools v0.47.0 h1:7Kn5x/d1svx/PzryTsqeoZN4TZwqeH5pGWjefhLi/1Q=
golang.org/x/tools v0.47.0/go.mod h1:dFHnyTvFWY212G+h7ZY4Vsp/K3U4/7W9TyVaAul8uCA=
google.golang.org/protobuf v1.36.8 h1:xHScyCOEuuwZEc6UtSOvPbAT4zRh0xcNRYekJwfqyMc=
google.golang.org/protobuf v1.36.8/go.mod h1:fuxRtAxBytpl4zzqUh6/eyUujkJdNiuEkXntxiD/uRU=
gopkg.in/check.v1 v0.0.0-20161208181325-20d25e280405/go.mod h1:Co6ibVJAznAaIkqp8huTwlJQCZ016jof/cbN4VW5Yz0=
+7 -2
View File
@@ -57,17 +57,22 @@ func ClassifyDNSName(name string) (DNSNameType, error) {
// ClassifyTargets splits a list of DNS names into apex domains and
// hostnames using the Public Suffix List. It returns an error if any
// name cannot be classified.
// name cannot be classified. A name given more than once, in any letter
// case or with a trailing dot, is kept once.
func ClassifyTargets(targets []string) ([]string, []string, error) {
var domains, hostnames []string
seen := make(map[string]bool)
for _, t := range targets {
normalized := strings.ToLower(strings.TrimSuffix(strings.TrimSpace(t), "."))
if normalized == "" {
if normalized == "" || seen[normalized] {
continue
}
seen[normalized] = true
typ, classErr := ClassifyDNSName(normalized)
if classErr != nil {
return nil, nil, classErr
+24
View File
@@ -1,6 +1,7 @@
package config_test
import (
"slices"
"testing"
"sneak.berlin/go/dnswatcher/internal/config"
@@ -93,6 +94,29 @@ func TestClassifyTargets(t *testing.T) {
}
}
func TestClassifyTargetsKeepsEachNameOnce(t *testing.T) {
t.Parallel()
domains, hostnames, err := config.ClassifyTargets([]string{
"example.org",
"Example.org.",
"www.example.org",
"EXAMPLE.ORG",
"WWW.Example.org.",
})
if err != nil {
t.Fatalf("unexpected error: %v", err)
}
if !slices.Equal(domains, []string{"example.org"}) {
t.Errorf("domains = %v, want [example.org]", domains)
}
if !slices.Equal(hostnames, []string{"www.example.org"}) {
t.Errorf("hostnames = %v, want [www.example.org]", hostnames)
}
}
func TestClassifyTargetsRejectsPublicSuffix(t *testing.T) {
t.Parallel()
+27 -8
View File
@@ -28,6 +28,13 @@ var ErrNoTargets = errors.New(
"no monitoring targets configured: set DNSWATCHER_TARGETS environment variable",
)
// ErrInvalidInterval is returned when DNSWATCHER_DNS_INTERVAL or
// DNSWATCHER_TLS_INTERVAL is set but is not a positive duration. An empty
// value counts as unset and means the default.
var ErrInvalidInterval = errors.New(
"interval must be a positive duration such as 30m or 1h",
)
// Params contains dependencies for Config.
type Params struct {
fx.In
@@ -125,18 +132,14 @@ func buildConfig(
}
}
dnsInterval, err := time.ParseDuration(
viper.GetString("DNS_INTERVAL"),
)
dnsInterval, err := parseInterval("DNS_INTERVAL")
if err != nil {
dnsInterval = defaultDNSInterval
return nil, err
}
tlsInterval, err := time.ParseDuration(
viper.GetString("TLS_INTERVAL"),
)
tlsInterval, err := parseInterval("TLS_INTERVAL")
if err != nil {
tlsInterval = defaultTLSInterval
return nil, err
}
domains, hostnames, err := parseAndValidateTargets()
@@ -168,6 +171,22 @@ func buildConfig(
return cfg, nil
}
// parseInterval reads the DNSWATCHER_-prefixed setting key as a duration. A
// value that does not parse, or is zero or negative, is an error naming the
// variable and the value; an unset variable has its default from setupViper.
func parseInterval(key string) (time.Duration, error) {
value := viper.GetString(key)
interval, err := time.ParseDuration(value)
if err != nil || interval <= 0 {
return 0, fmt.Errorf(
"invalid DNSWATCHER_%s %q: %w", key, value, ErrInvalidInterval,
)
}
return interval, nil
}
func parseAndValidateTargets() ([]string, []string, error) {
domains, hostnames, err := ClassifyTargets(
parseCSV(viper.GetString("TARGETS")),
+30 -22
View File
@@ -1,6 +1,7 @@
package config_test
import (
"strconv"
"testing"
"time"
@@ -113,33 +114,40 @@ func TestNew_OnlyEmptyCSVSegments(t *testing.T) {
assert.ErrorIs(t, err, config.ErrNoTargets)
}
func TestNew_InvalidDNSInterval_FallsBackToDefault(t *testing.T) {
viper.Reset()
t.Setenv("DNSWATCHER_TARGETS", "example.com")
t.Setenv("DNSWATCHER_DNS_INTERVAL", "banana")
// TestNew_InvalidIntervalStopsStartup checks values that must stop startup;
// TestNew_DefaultValues and TestNew_EmptyIntervalMeansDefault check that an
// unset or empty interval means the default.
func TestNew_InvalidIntervalStopsStartup(t *testing.T) {
variables := []string{"DNSWATCHER_DNS_INTERVAL", "DNSWATCHER_TLS_INTERVAL"}
values := []string{
"banana", // not a duration
"5", // no unit
"1d", // days are not a unit time.ParseDuration knows
"0", // zero
"-1h", // negative
}
cfg, err := config.New(nil, newTestParams(t))
require.NoError(t, err)
assert.Equal(t, time.Hour, cfg.DNSInterval,
"invalid DNS interval should fall back to 1h default")
for _, variable := range variables {
for _, value := range values {
t.Run(variable+"="+value, func(t *testing.T) {
viper.Reset()
t.Setenv("DNSWATCHER_TARGETS", "example.com")
t.Setenv(variable, value)
_, err := config.New(nil, newTestParams(t))
require.ErrorIs(t, err, config.ErrInvalidInterval)
require.ErrorContains(t, err, variable)
require.ErrorContains(t, err, strconv.Quote(value))
})
}
}
}
func TestNew_InvalidTLSInterval_FallsBackToDefault(t *testing.T) {
func TestNew_EmptyIntervalMeansDefault(t *testing.T) {
viper.Reset()
t.Setenv("DNSWATCHER_TARGETS", "example.com")
t.Setenv("DNSWATCHER_TLS_INTERVAL", "notaduration")
cfg, err := config.New(nil, newTestParams(t))
require.NoError(t, err)
assert.Equal(t, 12*time.Hour, cfg.TLSInterval,
"invalid TLS interval should fall back to 12h default")
}
func TestNew_BothIntervalsInvalid(t *testing.T) {
viper.Reset()
t.Setenv("DNSWATCHER_TARGETS", "example.com")
t.Setenv("DNSWATCHER_DNS_INTERVAL", "xyz")
t.Setenv("DNSWATCHER_TLS_INTERVAL", "abc")
t.Setenv("DNSWATCHER_DNS_INTERVAL", "")
t.Setenv("DNSWATCHER_TLS_INTERVAL", "")
cfg, err := config.New(nil, newTestParams(t))
require.NoError(t, err)
+50
View File
@@ -0,0 +1,50 @@
package globals_test
import (
"testing"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
"sneak.berlin/go/dnswatcher/internal/globals"
)
// TestGlobals exercises the package-level version and appname
// variables through their setters and read-back via New. These are
// shared package state, so the test mutates a global and must run
// sequentially; it cannot use t.Parallel().
//
//nolint:paralleltest // mutates shared package-level globals, must run sequentially
func TestGlobals(t *testing.T) {
versions := []string{"v1.2.3", "dev", "", "v1.2.3-4-gabcdef"}
for _, want := range versions {
globals.SetVersion(want)
g, err := globals.New(nil)
require.NoError(t, err)
assert.Equal(t, want, g.Version,
"New must surface the version set by SetVersion")
}
names := []string{"dnswatcher", "other", ""}
for _, want := range names {
globals.SetAppname(want)
g, err := globals.New(nil)
require.NoError(t, err)
assert.Equal(t, want, g.Appname,
"New must surface the appname set by SetAppname")
}
// New returns a snapshot: a later SetVersion must not mutate a
// Globals handed out earlier.
globals.SetVersion("first")
g, err := globals.New(nil)
require.NoError(t, err)
globals.SetVersion("second")
assert.Equal(t, "first", g.Version,
"a Globals returned by New must not change when the "+
"package variable is set again")
}
+47 -12
View File
@@ -1,11 +1,14 @@
package handlers
import (
"cmp"
"embed"
"fmt"
"html/template"
"maps"
"math"
"net/http"
"slices"
"strings"
"time"
@@ -40,12 +43,19 @@ func newDashboardTemplate() *template.Template {
)
}
// dashboardData is the data passed to the dashboard template.
// dashboardData is the data passed to the dashboard template. Hostnames
// and DomainRecords split the records in Snapshot.Hostnames, which also
// holds the apex domains' own (see splitHostnames). Ports holds
// Snapshot.Ports with each port's names split into domains and
// hostnames, as /api/v1/status gives them (see buildPorts).
type dashboardData struct {
Snapshot state.Snapshot
Alerts []notify.AlertEntry
StateAge string
GeneratedAt string
Snapshot state.Snapshot
Hostnames map[string]*state.HostnameState
DomainRecords map[string]*state.HostnameState
Ports map[string]*statusPortInfo
Alerts []notify.AlertEntry
StateAge string
GeneratedAt string
}
// HandleDashboard returns the dashboard page handler.
@@ -58,12 +68,16 @@ func (h *Handlers) HandleDashboard() http.HandlerFunc {
) {
snap := h.state.GetSnapshot()
alerts := h.notifyHistory.Recent()
hostnames, domainRecords := splitHostnames(snap)
data := dashboardData{
Snapshot: snap,
Alerts: alerts,
StateAge: relTime(snap.LastUpdated),
GeneratedAt: time.Now().UTC().Format("2006-01-02 15:04:05"),
Snapshot: snap,
Hostnames: hostnames,
DomainRecords: domainRecords,
Ports: buildPorts(snap),
Alerts: alerts,
StateAge: relTime(snap.LastUpdated),
GeneratedAt: time.Now().UTC().Format("2006-01-02 15:04:05"),
}
writer.Header().Set(
@@ -122,16 +136,37 @@ func joinStrings(items []string, sep string) string {
}
// formatRecords formats a map of record type → values into a
// compact display string.
// compact display string. Record types are listed in the order the
// README lists them, any other type after them in alphabetical order,
// so rows of nameservers with the same records read the same.
func formatRecords(records map[string][]string) string {
if len(records) == 0 {
return "-"
}
order := []string{"A", "AAAA", "CNAME", "MX", "TXT", "SRV", "CAA", "NS"}
position := func(rtype string) int {
i := slices.Index(order, rtype)
if i < 0 {
return len(order)
}
return i
}
rtypes := slices.Collect(maps.Keys(records))
slices.SortFunc(rtypes, func(a, b string) int {
return cmp.Or(
cmp.Compare(position(a), position(b)),
strings.Compare(a, b),
)
})
var parts []string
for rtype, values := range records {
for _, v := range values {
for _, rtype := range rtypes {
for _, v := range records[rtype] {
parts = append(parts, rtype+": "+v)
}
}
+175
View File
@@ -1,6 +1,8 @@
package handlers_test
import (
"regexp"
"strings"
"testing"
"time"
@@ -78,3 +80,176 @@ func TestFormatRecords(t *testing.T) {
t.Errorf("unexpected format: %q", got)
}
}
// TestFormatRecordsTypeOrder checks that record types are listed in
// the README's order (A, AAAA, CNAME, MX, TXT, SRV, CAA, NS), with
// any other type after them in alphabetical order.
func TestFormatRecordsTypeOrder(t *testing.T) {
t.Parallel()
got := handlers.FormatRecords(map[string][]string{
"SOA": {"ns1.example.com. hostmaster.example.com. 1 2 3 4 5"},
"NS": {"ns1.example.com.", "ns2.example.com."},
"CAA": {`0 issue "letsencrypt.org"`},
"DNAME": {"example.net."},
"TXT": {"v=spf1 -all"},
"SRV": {"10 5 443 www.example.com."},
"MX": {"10 mail.example.com."},
"CNAME": {"www.example.com."},
"AAAA": {"2001:db8::1"},
"A": {"192.0.2.1"},
})
want := strings.Join([]string{
"A: 192.0.2.1",
"AAAA: 2001:db8::1",
"CNAME: www.example.com.",
"MX: 10 mail.example.com.",
"TXT: v=spf1 -all",
"SRV: 10 5 443 www.example.com.",
`CAA: 0 issue "letsencrypt.org"`,
"NS: ns1.example.com.",
"NS: ns2.example.com.",
"DNAME: example.net.",
"SOA: ns1.example.com. hostmaster.example.com. 1 2 3 4 5",
}, ", ")
if got != want {
t.Errorf("FormatRecords lists types out of order:\n got %q\nwant %q",
got, want)
}
}
// dashboardRow returns the table row of page that contains name.
func dashboardRow(t *testing.T, page string, name string) string {
t.Helper()
for row := range strings.SplitSeq(page, "<tr") {
if strings.Contains(row, name) {
return row
}
}
t.Fatalf("dashboard has no row containing %q", name)
return ""
}
// TestDashboardShowsFailureReasons checks that the dashboard shows the
// reason in the row of a failed nameserver and of a failed certificate,
// and not in the row of a nameserver that answered.
func TestDashboardShowsFailureReasons(t *testing.T) {
t.Parallel()
page := get(t, newHandlersWithFailures(t).HandleDashboard())
if !strings.Contains(dashboardRow(t, page, failedNS), nsFailureReason) {
t.Errorf("row of %s does not show %q", failedNS, nsFailureReason)
}
if strings.Contains(dashboardRow(t, page, answeringNS), nsFailureReason) {
t.Errorf("row of %s shows %q", answeringNS, nsFailureReason)
}
if !strings.Contains(dashboardRow(t, page, certKey), certFailedReason) {
t.Errorf("row of %s does not show %q", certKey, certFailedReason)
}
}
// dashboardSection returns the section of page under heading.
func dashboardSection(t *testing.T, page string, heading string) string {
t.Helper()
for section := range strings.SplitSeq(page, "<section") {
words := strings.Join(strings.Fields(section), " ")
if strings.Contains(words, "> "+heading+" </h2>") {
return section
}
}
t.Fatalf("dashboard has no section headed %q", heading)
return ""
}
// TestDashboardShowsDomainRecordsUnderDomains checks that the dashboard
// shows an apex domain's own records in the Domains section, and
// neither lists nor counts the domain as a hostname.
func TestDashboardShowsDomainRecordsUnderDomains(t *testing.T) {
t.Parallel()
page := get(t, newHandlersWithFailures(t).HandleDashboard())
domains := dashboardSection(t, page, "Domains")
if !strings.Contains(dashboardRow(t, domains, domainAddress), testDomain) {
t.Errorf("row of %s does not name %s", domainAddress, testDomain)
}
if strings.Contains(dashboardSection(t, page, "Hostnames"), testDomain) {
t.Errorf("Hostnames section lists the domain %s", testDomain)
}
words := strings.Join(strings.Fields(page), " ")
footer := "monitoring 1 domains + 1 hostnames"
if !strings.Contains(words, footer) {
t.Errorf("dashboard does not say %q", footer)
}
// With the tags taken out, the summary bar starts "Domains 1
// Hostnames 1".
text := regexp.MustCompile(`<[^>]*>`).ReplaceAllString(page, " ")
summary := "Domains 1 Hostnames 1"
if !strings.Contains(strings.Join(strings.Fields(text), " "), summary) {
t.Errorf("summary bar does not say %q", summary)
}
}
// rowCells returns the text of each cell of a dashboard table row
// whose cells start with tag, "<th" or "<td".
func rowCells(row string, tag string) []string {
tags := regexp.MustCompile(`<[^>]*>`)
parts := strings.Split(row, tag)[1:]
cells := make([]string, 0, len(parts))
for _, cell := range parts {
text := tags.ReplaceAllString(tag+cell, " ")
cells = append(cells, strings.Join(strings.Fields(text), " "))
}
return cells
}
// TestDashboardPortsTellDomainsFromHostnames checks that the Ports
// table lists an apex domain under Domains and a hostname under
// Hostnames when both resolve to the port's address.
func TestDashboardPortsTellDomainsFromHostnames(t *testing.T) {
t.Parallel()
page := get(t, newHandlersWithFailures(t).HandleDashboard())
ports := dashboardSection(t, page, "Ports")
headings := rowCells(dashboardRow(t, ports, "Address</th>"), "<th")
cells := rowCells(dashboardRow(t, ports, sharedPort), "<td")
if len(cells) != len(headings) {
t.Fatalf("row of %s has cells %q under headings %q",
sharedPort, cells, headings)
}
under := make(map[string]string)
for i, heading := range headings {
under[heading] = cells[i]
}
if under["Domains"] != testDomain {
t.Errorf("row of %s lists %q under Domains, want %q",
sharedPort, under["Domains"], testDomain)
}
if under["Hostnames"] != testHostname {
t.Errorf("row of %s lists %q under Hostnames, want %q",
sharedPort, under["Hostnames"], testHostname)
}
}
+97 -36
View File
@@ -9,15 +9,19 @@ import (
)
// statusDomainInfo holds status information for a monitored domain.
// RecordsByNameserver holds the domain's own records, in the form a
// hostname's Nameservers holds the hostname's.
type statusDomainInfo struct {
Nameservers []string `json:"nameservers"`
LastChecked time.Time `json:"lastChecked"`
Nameservers []string `json:"nameservers"`
RecordsByNameserver map[string]*statusHostnameNSInfo `json:"recordsByNameserver"`
LastChecked time.Time `json:"lastChecked"`
}
// statusHostnameNSInfo holds per-nameserver status for a hostname.
type statusHostnameNSInfo struct {
Records map[string][]string `json:"records"`
Status string `json:"status"`
Error string `json:"error,omitempty"`
LastChecked time.Time `json:"lastChecked"`
}
@@ -28,8 +32,11 @@ type statusHostnameInfo struct {
}
// statusPortInfo holds status information for a monitored port.
// Domains and Hostnames list the apex domains and the hostnames that
// resolve to its address.
type statusPortInfo struct {
Open bool `json:"open"`
Domains []string `json:"domains"`
Hostnames []string `json:"hostnames"`
LastChecked time.Time `json:"lastChecked"`
}
@@ -41,6 +48,7 @@ type statusCertificateInfo struct {
NotAfter time.Time `json:"notAfter"`
SubjectAlternativeNames []string `json:"subjectAlternativeNames"`
Status string `json:"status"`
Error string `json:"error,omitempty"`
LastChecked time.Time `json:"lastChecked"`
}
@@ -94,21 +102,44 @@ func buildStatusResponse(
LastUpdated: snap.LastUpdated,
Domains: make(map[string]*statusDomainInfo),
Hostnames: make(map[string]*statusHostnameInfo),
Ports: make(map[string]*statusPortInfo),
Certificates: make(map[string]*statusCertificateInfo),
}
buildDomains(snap, resp)
buildHostnames(snap, resp)
buildPorts(snap, resp)
hostnames, domainRecords := splitHostnames(snap)
buildDomains(snap, domainRecords, resp)
buildHostnames(hostnames, resp)
resp.Ports = buildPorts(snap)
buildCertificates(snap, resp)
buildCounts(resp)
return resp
}
// splitHostnames returns the records saved in snap.Hostnames in two
// maps: the hostnames' and the apex domains' own. The watcher saves a
// domain's own records there under the domain's name, which has an
// entry in snap.Domains too.
func splitHostnames(
snap state.Snapshot,
) (map[string]*state.HostnameState, map[string]*state.HostnameState) {
hostnames := make(map[string]*state.HostnameState)
domainRecords := make(map[string]*state.HostnameState)
for name, hs := range snap.Hostnames {
if _, isDomain := snap.Domains[name]; isDomain {
domainRecords[name] = hs
} else {
hostnames[name] = hs
}
}
return hostnames, domainRecords
}
func buildDomains(
snap state.Snapshot,
domainRecords map[string]*state.HostnameState,
resp *statusResponse,
) {
for name, ds := range snap.Domains {
@@ -116,57 +147,86 @@ func buildDomains(
copy(ns, ds.Nameservers)
sort.Strings(ns)
records := make(map[string]*statusHostnameNSInfo)
if hs, ok := domainRecords[name]; ok {
records = nameserverInfo(hs)
}
resp.Domains[name] = &statusDomainInfo{
Nameservers: ns,
LastChecked: ds.LastChecked,
Nameservers: ns,
RecordsByNameserver: records,
LastChecked: ds.LastChecked,
}
}
}
func buildHostnames(
snap state.Snapshot,
hostnames map[string]*state.HostnameState,
resp *statusResponse,
) {
for name, hs := range snap.Hostnames {
info := &statusHostnameInfo{
Nameservers: make(map[string]*statusHostnameNSInfo),
for name, hs := range hostnames {
resp.Hostnames[name] = &statusHostnameInfo{
Nameservers: nameserverInfo(hs),
LastChecked: hs.LastChecked,
}
for ns, nsState := range hs.RecordsByNameserver {
recs := make(map[string][]string, len(nsState.Records))
for rtype, vals := range nsState.Records {
copied := make([]string, len(vals))
copy(copied, vals)
recs[rtype] = copied
}
info.Nameservers[ns] = &statusHostnameNSInfo{
Records: recs,
Status: nsState.Status,
LastChecked: nsState.LastChecked,
}
}
resp.Hostnames[name] = info
}
}
func buildPorts(
snap state.Snapshot,
resp *statusResponse,
) {
// nameserverInfo copies each nameserver's answer saved in hs.
func nameserverInfo(
hs *state.HostnameState,
) map[string]*statusHostnameNSInfo {
info := make(map[string]*statusHostnameNSInfo)
for ns, nsState := range hs.RecordsByNameserver {
recs := make(map[string][]string, len(nsState.Records))
for rtype, vals := range nsState.Records {
copied := make([]string, len(vals))
copy(copied, vals)
recs[rtype] = copied
}
info[ns] = &statusHostnameNSInfo{
Records: recs,
Status: nsState.Status,
Error: nsState.Error,
LastChecked: nsState.LastChecked,
}
}
return info
}
// buildPorts returns the port entries saved in snap. A port entry
// saves apex domains with its hostnames; they are told apart as in
// splitHostnames, by a domain entry in snap.Domains.
func buildPorts(snap state.Snapshot) map[string]*statusPortInfo {
ports := make(map[string]*statusPortInfo, len(snap.Ports))
for key, ps := range snap.Ports {
hostnames := make([]string, len(ps.Hostnames))
copy(hostnames, ps.Hostnames)
domains := []string{}
hostnames := []string{}
for _, name := range ps.Hostnames {
if _, isDomain := snap.Domains[name]; isDomain {
domains = append(domains, name)
} else {
hostnames = append(hostnames, name)
}
}
sort.Strings(domains)
sort.Strings(hostnames)
resp.Ports[key] = &statusPortInfo{
ports[key] = &statusPortInfo{
Open: ps.Open,
Domains: domains,
Hostnames: hostnames,
LastChecked: ps.LastChecked,
}
}
return ports
}
func buildCertificates(
@@ -183,6 +243,7 @@ func buildCertificates(
NotAfter: cs.NotAfter,
SubjectAlternativeNames: sans,
Status: cs.Status,
Error: cs.Error,
LastChecked: cs.LastChecked,
}
}
+267
View File
@@ -0,0 +1,267 @@
package handlers_test
import (
"encoding/json"
"net/http"
"net/http/httptest"
"slices"
"testing"
"time"
"go.uber.org/fx/fxtest"
"sneak.berlin/go/dnswatcher/internal/config"
"sneak.berlin/go/dnswatcher/internal/globals"
"sneak.berlin/go/dnswatcher/internal/handlers"
"sneak.berlin/go/dnswatcher/internal/logger"
"sneak.berlin/go/dnswatcher/internal/notify"
"sneak.berlin/go/dnswatcher/internal/state"
)
// The state the handler tests serve: www.example.com has one nameserver
// that answered and one whose query failed, and its certificate check
// failed. example.net is an apex domain, whose own records are saved
// with the hostnames' records, as the watcher saves them. Both names
// resolve to domainAddress, whose port 443 entry lists them.
const (
testHostname = "www.example.com"
answeringNS = "ns1.example.com."
failedNS = "ns2.example.com."
nsFailureReason = "server returned a referral"
certKey = "192.0.2.1:443:www.example.com"
certFailedReason = "x509: certificate has expired or is not yet valid"
testDomain = "example.net"
domainNS = "a.iana-servers.net."
domainAddress = "192.0.2.2"
sharedPort = domainAddress + ":443"
)
// newHandlersWithFailures builds real Handlers whose state holds the
// entries described above.
func newHandlersWithFailures(t *testing.T) *handlers.Handlers {
t.Helper()
glob, err := globals.New(nil)
if err != nil {
t.Fatalf("globals.New: %v", err)
}
log, err := logger.New(nil, logger.Params{Globals: glob})
if err != nil {
t.Fatalf("logger.New: %v", err)
}
notifier, err := notify.New(fxtest.NewLifecycle(t), notify.Params{
Logger: log,
Config: &config.Config{},
})
if err != nil {
t.Fatalf("notify.New: %v", err)
}
st, err := state.New(fxtest.NewLifecycle(t), state.Params{
Logger: log,
Config: &config.Config{DataDir: t.TempDir()},
})
if err != nil {
t.Fatalf("state.New: %v", err)
}
setTestState(st)
hnd, err := handlers.New(nil, handlers.Params{
Logger: log,
Globals: glob,
State: st,
Notify: notifier,
})
if err != nil {
t.Fatalf("handlers.New: %v", err)
}
return hnd
}
// setTestState sets the entries described above in st.
func setTestState(st *state.State) {
now := time.Now()
st.SetHostnameState(testHostname, &state.HostnameState{
RecordsByNameserver: map[string]*state.NameserverRecordState{
answeringNS: {
Records: map[string][]string{
"A": {"192.0.2.1", domainAddress},
},
Status: "ok",
LastChecked: now,
},
failedNS: {
Records: map[string][]string{},
Status: "error",
Error: nsFailureReason,
LastChecked: now,
},
},
LastChecked: now,
})
st.SetCertificateState(certKey, &state.CertificateState{
Status: "error",
Error: certFailedReason,
LastChecked: now,
})
st.SetDomainState(testDomain, &state.DomainState{
Nameservers: []string{domainNS},
LastChecked: now,
})
st.SetHostnameState(testDomain, &state.HostnameState{
RecordsByNameserver: map[string]*state.NameserverRecordState{
domainNS: {
Records: map[string][]string{"A": {domainAddress}},
Status: "ok",
LastChecked: now,
},
},
LastChecked: now,
})
st.SetPortState(sharedPort, &state.PortState{
Open: true,
Hostnames: []string{testDomain, testHostname},
LastChecked: now,
})
}
// get serves one GET request to handler and returns the response body.
func get(t *testing.T, handler http.HandlerFunc) string {
t.Helper()
rec := httptest.NewRecorder()
req := httptest.NewRequestWithContext(
t.Context(), http.MethodGet, "/", nil,
)
handler(rec, req)
if rec.Code != http.StatusOK {
t.Fatalf("status = %d, want 200", rec.Code)
}
return rec.Body.String()
}
// TestStatusGivesFailureReasons checks that /api/v1/status gives the
// reason for a failed nameserver entry and a failed certificate entry,
// and no error for a nameserver that answered.
func TestStatusGivesFailureReasons(t *testing.T) {
t.Parallel()
body := get(t, newHandlersWithFailures(t).HandleStatus())
var resp struct {
Hostnames map[string]struct {
Nameservers map[string]map[string]any `json:"nameservers"`
} `json:"hostnames"`
Certificates map[string]map[string]any `json:"certificates"`
}
err := json.Unmarshal([]byte(body), &resp)
if err != nil {
t.Fatalf("decoding response: %v", err)
}
nameservers := resp.Hostnames[testHostname].Nameservers
got := nameservers[failedNS]["error"]
if got != nsFailureReason {
t.Errorf("failed nameserver error = %v, want %q",
got, nsFailureReason)
}
_, has := nameservers[answeringNS]["error"]
if has {
t.Errorf("answering nameserver has an error field: %v",
nameservers[answeringNS])
}
got = resp.Certificates[certKey]["error"]
if got != certFailedReason {
t.Errorf("failed certificate error = %v, want %q",
got, certFailedReason)
}
}
// TestStatusGivesDomainRecordsUnderTheDomain checks that /api/v1/status
// gives an apex domain's own records in its domain entry, and neither
// lists nor counts the domain as a hostname.
func TestStatusGivesDomainRecordsUnderTheDomain(t *testing.T) {
t.Parallel()
body := get(t, newHandlersWithFailures(t).HandleStatus())
var resp struct {
Counts struct {
Hostnames int `json:"hostnames"`
} `json:"counts"`
Domains map[string]struct {
RecordsByNameserver map[string]struct {
Records map[string][]string `json:"records"`
} `json:"recordsByNameserver"`
} `json:"domains"`
Hostnames map[string]any `json:"hostnames"`
}
err := json.Unmarshal([]byte(body), &resp)
if err != nil {
t.Fatalf("decoding response: %v", err)
}
if resp.Counts.Hostnames != 1 {
t.Errorf("counts.hostnames = %d, want 1", resp.Counts.Hostnames)
}
if _, listed := resp.Hostnames[testDomain]; listed {
t.Errorf("hostnames lists the domain %s", testDomain)
}
records := resp.Domains[testDomain].RecordsByNameserver[domainNS].Records
if !slices.Equal(records["A"], []string{domainAddress}) {
t.Errorf("domain %s records at %s = %v, want A %s",
testDomain, domainNS, records, domainAddress)
}
}
// TestStatusPortsTellDomainsFromHostnames checks that a port entry in
// /api/v1/status lists an apex domain in domains and a hostname in
// hostnames when both resolve to its address.
func TestStatusPortsTellDomainsFromHostnames(t *testing.T) {
t.Parallel()
body := get(t, newHandlersWithFailures(t).HandleStatus())
var resp struct {
Ports map[string]struct {
Domains []string `json:"domains"`
Hostnames []string `json:"hostnames"`
} `json:"ports"`
}
err := json.Unmarshal([]byte(body), &resp)
if err != nil {
t.Fatalf("decoding response: %v", err)
}
port := resp.Ports[sharedPort]
if !slices.Equal(port.Domains, []string{testDomain}) {
t.Errorf("port %s domains = %v, want [%s]",
sharedPort, port.Domains, testDomain)
}
if !slices.Equal(port.Hostnames, []string{testHostname}) {
t.Errorf("port %s hostnames = %v, want [%s]",
sharedPort, port.Hostnames, testHostname)
}
}
+74 -38
View File
@@ -39,7 +39,7 @@
Hostnames
</div>
<div class="text-2xl font-bold text-teal-400 mt-1">
{{ len .Snapshot.Hostnames }}
{{ len .Hostnames }}
</div>
</div>
<div class="bg-surface-800 border border-slate-700/50 rounded-lg p-4">
@@ -94,6 +94,24 @@
</tbody>
</table>
</div>
{{ if .DomainRecords }}
<div class="overflow-x-auto mt-4">
<table class="w-full text-left text-xs">
<thead>
<tr class="text-slate-500 uppercase tracking-wider">
<th class="py-2 px-3">Domain</th>
<th class="py-2 px-3">NS</th>
<th class="py-2 px-3">Status</th>
<th class="py-2 px-3">Records</th>
<th class="py-2 px-3">Checked</th>
</tr>
</thead>
<tbody class="divide-y divide-slate-800">
{{ template "records" .DomainRecords }}
</tbody>
</table>
</div>
{{ end }}
{{ else }}
<p class="text-slate-600 italic text-xs">
No domains configured.
@@ -108,7 +126,7 @@
>
Hostnames
</h2>
{{ if .Snapshot.Hostnames }}
{{ if .Hostnames }}
<div class="overflow-x-auto">
<table class="w-full text-left text-xs">
<thead>
@@ -121,39 +139,7 @@
</tr>
</thead>
<tbody class="divide-y divide-slate-800">
{{ range $name, $hs := .Snapshot.Hostnames }}
{{ range $ns, $nsr := $hs.RecordsByNameserver }}
<tr class="hover:bg-surface-800/50">
<td class="py-2 px-3 text-slate-200 font-medium">
{{ $name }}
</td>
<td class="py-2 px-3 text-slate-400 break-all">
{{ $ns }}
</td>
<td class="py-2 px-3">
{{ if eq $nsr.Status "ok" }}
<span
class="inline-block px-1.5 py-0.5 rounded text-[10px] font-bold uppercase bg-teal-900/50 text-teal-400 border border-teal-700/30"
>ok</span
>
{{ else }}
<span
class="inline-block px-1.5 py-0.5 rounded text-[10px] font-bold uppercase bg-red-900/50 text-red-400 border border-red-700/30"
>{{ $nsr.Status }}</span
>
{{ end }}
</td>
<td
class="py-2 px-3 text-slate-400 break-all max-w-xs"
>
{{ formatRecords $nsr.Records }}
</td>
<td class="py-2 px-3 text-slate-500 whitespace-nowrap">
{{ relTime $nsr.LastChecked }}
</td>
</tr>
{{ end }}
{{ end }}
{{ template "records" .Hostnames }}
</tbody>
</table>
</div>
@@ -171,19 +157,20 @@
>
Ports
</h2>
{{ if .Snapshot.Ports }}
{{ if .Ports }}
<div class="overflow-x-auto">
<table class="w-full text-left text-xs">
<thead>
<tr class="text-slate-500 uppercase tracking-wider">
<th class="py-2 px-3">Address</th>
<th class="py-2 px-3">State</th>
<th class="py-2 px-3">Domains</th>
<th class="py-2 px-3">Hostnames</th>
<th class="py-2 px-3">Checked</th>
</tr>
</thead>
<tbody class="divide-y divide-slate-800">
{{ range $key, $ps := .Snapshot.Ports }}
{{ range $key, $ps := .Ports }}
<tr class="hover:bg-surface-800/50">
<td class="py-2 px-3 text-slate-200 font-medium">
{{ $key }}
@@ -201,6 +188,9 @@
>
{{ end }}
</td>
<td class="py-2 px-3 text-slate-400 break-all">
{{ joinStrings $ps.Domains ", " }}
</td>
<td class="py-2 px-3 text-slate-400 break-all">
{{ joinStrings $ps.Hostnames ", " }}
</td>
@@ -258,6 +248,11 @@
>
{{ end }}
</td>
{{ if $cs.Error }}
<td colspan="3" class="py-2 px-3 text-red-400 break-all">
<div class="max-w-xs">{{ $cs.Error }}</div>
</td>
{{ else }}
<td class="py-2 px-3 text-slate-200">
{{ $cs.CommonName }}
</td>
@@ -285,6 +280,7 @@
{{ end }}
{{ end }}
</td>
{{ end }}
<td class="py-2 px-3 text-slate-500 whitespace-nowrap">
{{ relTime $cs.LastChecked }}
</td>
@@ -363,8 +359,48 @@
class="text-[11px] text-slate-700 border-t border-slate-800 pt-4 mt-8"
>
dnswatcher &middot; monitoring {{ len .Snapshot.Domains }} domains +
{{ len .Snapshot.Hostnames }} hostnames
{{ len .Hostnames }} hostnames
</div>
</div>
</body>
</html>
{{/* ---- One row per nameserver of each name in the map it is given ---- */}}
{{ define "records" }}
{{ range $name, $hs := . }}
{{ range $ns, $nsr := $hs.RecordsByNameserver }}
<tr class="hover:bg-surface-800/50">
<td class="py-2 px-3 text-slate-200 font-medium">
{{ $name }}
</td>
<td class="py-2 px-3 text-slate-400 break-all">
{{ $ns }}
</td>
<td class="py-2 px-3">
{{ if eq $nsr.Status "ok" }}
<span
class="inline-block px-1.5 py-0.5 rounded text-[10px] font-bold uppercase bg-teal-900/50 text-teal-400 border border-teal-700/30"
>ok</span
>
{{ else }}
<span
class="inline-block px-1.5 py-0.5 rounded text-[10px] font-bold uppercase bg-red-900/50 text-red-400 border border-red-700/30"
>{{ $nsr.Status }}</span
>
{{ end }}
</td>
<td
class="py-2 px-3 text-slate-400 break-all max-w-xs"
>
{{ if $nsr.Error }}
<span class="text-red-400">{{ $nsr.Error }}</span>
{{ else }}
{{ formatRecords $nsr.Records }}
{{ end }}
</td>
<td class="py-2 px-3 text-slate-500 whitespace-nowrap">
{{ relTime $nsr.LastChecked }}
</td>
</tr>
{{ end }}
{{ end }}
{{ end }}
+134
View File
@@ -0,0 +1,134 @@
package healthcheck_test
import (
"context"
"encoding/json"
"testing"
"time"
"go.uber.org/fx"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
"sneak.berlin/go/dnswatcher/internal/config"
"sneak.berlin/go/dnswatcher/internal/globals"
"sneak.berlin/go/dnswatcher/internal/healthcheck"
"sneak.berlin/go/dnswatcher/internal/logger"
)
// recordingLifecycle is a minimal fx.Lifecycle that records the hooks
// appended to it, so healthcheck.New can be exercised through its real
// constructor without standing up a whole fx application.
type recordingLifecycle struct {
hooks []fx.Hook
}
func (l *recordingLifecycle) Append(hook fx.Hook) {
l.hooks = append(l.hooks, hook)
}
// newHealthcheck builds a Healthcheck through the real constructor and
// runs the registered OnStart hook so StartupTime is set the same way
// the fx lifecycle would set it.
func newHealthcheck(
t *testing.T,
maintenance bool,
version string,
) *healthcheck.Healthcheck {
t.Helper()
g := &globals.Globals{Appname: "dnswatcher", Version: version}
log, err := logger.New(nil, logger.Params{Globals: g})
require.NoError(t, err)
lifecycle := &recordingLifecycle{}
hc, err := healthcheck.New(lifecycle, healthcheck.Params{
Globals: g,
Config: &config.Config{MaintenanceMode: maintenance},
Logger: log,
})
require.NoError(t, err)
require.Len(t, lifecycle.hooks, 1,
"New must register exactly one lifecycle hook")
require.NotNil(t, lifecycle.hooks[0].OnStart)
require.NoError(t, lifecycle.hooks[0].OnStart(context.Background()))
return hc
}
func TestCheckStatusAndPayloadShape(t *testing.T) {
t.Parallel()
hc := newHealthcheck(t, false, "v9.9.9")
resp := hc.Check()
assert.Equal(t, "ok", resp.Status)
// The JSON shape and field names are part of the contract for the
// /health and /.well-known/healthcheck routes, so assert on the
// exact set of keys the response marshals to.
raw, err := json.Marshal(resp)
require.NoError(t, err)
var fields map[string]json.RawMessage
require.NoError(t, json.Unmarshal(raw, &fields))
wantKeys := []string{
"status",
"now",
"uptimeSeconds",
"uptimeHuman",
"version",
"appname",
"maintenanceMode",
}
assert.Len(t, fields, len(wantKeys),
"response must marshal to exactly the documented fields")
for _, key := range wantKeys {
assert.Contains(t, fields, key, "missing JSON field %q", key)
}
// The Now field is documented as RFC3339Nano; a change to the
// format constant should turn this red.
_, err = time.Parse(time.RFC3339Nano, resp.Now)
assert.NoError(t, err, "Now must be RFC3339Nano")
}
func TestCheckMaintenanceModeReflectsConfig(t *testing.T) {
t.Parallel()
tests := []struct {
name string
maintenance bool
}{
{"maintenance off", false},
{"maintenance on", true},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
t.Parallel()
hc := newHealthcheck(t, tt.maintenance, "test")
resp := hc.Check()
assert.Equal(t, tt.maintenance, resp.Maintenance,
"maintenanceMode must mirror Config.MaintenanceMode")
})
}
}
func TestCheckSurfacesVersionAndAppname(t *testing.T) {
t.Parallel()
hc := newHealthcheck(t, false, "surfaced-version-123")
resp := hc.Check()
assert.Equal(t, "surfaced-version-123", resp.Version,
"version from globals must appear in the payload")
assert.Equal(t, "dnswatcher", resp.Appname)
}
+128
View File
@@ -0,0 +1,128 @@
// Package livednstest runs the live DNS operations of tests. Tests that
// look something up in DNS query live DNS servers, never a stand-in —
// see TESTING.md. Nothing here mocks, fakes, stubs, records or replays
// DNS, and nothing here skips a test: it only changes *how* the live
// queries are issued, so that a single dropped UDP packet or one slow
// authoritative server does not turn correct code into a red build.
//
// Two mechanisms:
//
// 1. Bounded concurrency. Tests run in parallel and the build hosts
// have many cores, so without a limit every test starts its own
// iterative resolution at the same instant and they all send their
// first queries to the root servers within a few milliseconds of
// each other. Root servers rate-limit that, which shows up as a
// different arbitrary subset of tests failing on each run. Run caps
// how many live operations are in flight at once in one test binary.
//
// 2. Retry with exponential backoff. Each live operation gets several
// attempts with its own timeout. An attempt is retried when it
// obtained nothing to check, never because of what the test
// asserts about the result, so a wrong result still fails on the
// first attempt. A fault in the code under test that leaves
// nothing to check looks the same as live DNS not answering, and
// fails only after the last attempt.
package livednstest
import (
"context"
"errors"
"testing"
"time"
)
const (
// attempts is how many times a live DNS operation is attempted
// before the test fails.
attempts = 3
// AttemptTimeout bounds one attempt. It must fit the longest
// operation, a watcher check, which sends over a hundred queries one
// after another and on a slow build host takes several times as long
// as the few seconds it takes on a fast one. An operation whose
// every attempt fails takes attempts * AttemptTimeout plus the
// backoff, about 56 seconds, after it waits for one of the
// Concurrency slots that every live operation in the test binary
// shares. So when live DNS does not answer at all, a test binary
// with more live operations than slots runs into the 90-second
// `go test -timeout` backstop instead of each test failing on its
// own.
AttemptTimeout = 18 * time.Second
// backoffBase is the delay after the first failed attempt; it is
// multiplied by backoffFactor each time.
backoffBase = 500 * time.Millisecond
// backoffFactor is the exponential backoff multiplier.
backoffFactor = 2
// Concurrency caps how many live operations may be in flight
// across one test binary at once.
Concurrency = 6
)
// gate bounds concurrent live operations. It has to be package scoped:
// the whole point is that it is shared by every parallel test in the
// test binary.
//
//nolint:gochecknoglobals // package-wide live query rate limit
var gate = make(chan struct{}, Concurrency)
// ErrNoAnswer reports that a live operation produced no usable answer,
// which is retried rather than asserted on.
var ErrNoAnswer = errors.New("no answer from live DNS")
// Run executes one attempt of a live operation, holding a slot in gate
// for its duration and bounding it with its own timeout.
func Run(op func(ctx context.Context) error) error {
gate <- struct{}{}
defer func() { <-gate }()
ctx, cancel := context.WithTimeout(
context.Background(), AttemptTimeout,
)
defer cancel()
return op(ctx)
}
// Retry runs op until it reports success, retrying failures with
// exponential backoff, and fails the test if every attempt fails. op
// returns an error only for a failure to obtain an answer — never for
// an answer the test disagrees with, which belongs in an assertion so
// that it fails immediately. op stores whatever it obtained where its
// caller can find it.
func Retry(
t *testing.T,
what string,
op func(ctx context.Context) error,
) {
t.Helper()
var last error
backoff := backoffBase
for attempt := range attempts {
if attempt > 0 {
t.Logf(
"%s: attempt %d of %d failed (%v), "+
"retrying in %s",
what, attempt, attempts, last, backoff,
)
time.Sleep(backoff)
backoff *= backoffFactor
}
last = Run(op)
if last == nil {
return
}
}
t.Fatalf(
"%s: all %d live attempts failed: %v",
what, attempts, last,
)
}
+103
View File
@@ -0,0 +1,103 @@
package livednstest_test
import (
"context"
"sync"
"testing"
"time"
"github.com/stretchr/testify/assert"
"sneak.berlin/go/dnswatcher/internal/livednstest"
)
// Tests for the retry and the concurrency limit themselves. They
// perform no DNS resolution of any kind.
func TestRetryRecoversFromTransientFailure(t *testing.T) {
t.Parallel()
const wantAttempts = 2
attempts := 0
livednstest.Retry(t, "transient", func(_ context.Context) error {
attempts++
if attempts < wantAttempts {
return livednstest.ErrNoAnswer
}
return nil
})
assert.Equal(t, wantAttempts, attempts)
}
func TestRetryGivesEachAttemptADeadline(t *testing.T) {
t.Parallel()
livednstest.Retry(t, "deadline", func(ctx context.Context) error {
deadline, ok := ctx.Deadline()
assert.True(t, ok, "attempt should carry a deadline")
remaining := time.Until(deadline)
assert.LessOrEqual(t, remaining, livednstest.AttemptTimeout)
// Lower bound too: without one this passes for a
// deadline far shorter than intended, which would
// silently turn every live attempt into an instant
// timeout.
assert.Greater(t, remaining, livednstest.AttemptTimeout/2)
return nil
})
}
func TestRunBoundsConcurrency(t *testing.T) {
t.Parallel()
const workers = 24
var (
mu sync.Mutex
wg sync.WaitGroup
inFlight int
maxSeen int
)
wg.Add(workers)
for range workers {
go func() {
defer wg.Done()
_ = livednstest.Run(func(_ context.Context) error {
mu.Lock()
inFlight++
if inFlight > maxSeen {
maxSeen = inFlight
}
mu.Unlock()
time.Sleep(time.Millisecond)
mu.Lock()
inFlight--
mu.Unlock()
return nil
})
}()
}
wg.Wait()
assert.Positive(t, maxSeen)
assert.LessOrEqual(
t, maxSeen, livednstest.Concurrency,
"live queries must stay under the package-wide gate",
)
}
+66
View File
@@ -0,0 +1,66 @@
package logger_test
import (
"context"
"log/slog"
"testing"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
"sneak.berlin/go/dnswatcher/internal/globals"
"sneak.berlin/go/dnswatcher/internal/logger"
)
func newTestLogger(t *testing.T) *logger.Logger {
t.Helper()
g := &globals.Globals{Appname: "dnswatcher", Version: "test"}
l, err := logger.New(nil, logger.Params{Globals: g})
require.NoError(t, err)
return l
}
// TestNewReturnsUsableLogger checks that the constructor yields a
// working *slog.Logger.
func TestNewReturnsUsableLogger(t *testing.T) {
t.Parallel()
l := newTestLogger(t)
require.NotNil(t, l.Get(), "Get must return a non-nil logger")
}
// TestDefaultLevelExcludesDebug verifies the default configuration
// logs at info: debug records are suppressed, info records pass.
func TestDefaultLevelExcludesDebug(t *testing.T) {
t.Parallel()
log := newTestLogger(t).Get()
ctx := context.Background()
assert.False(t, log.Enabled(ctx, slog.LevelDebug),
"debug must be suppressed at the default level")
assert.True(t, log.Enabled(ctx, slog.LevelInfo),
"info must be enabled at the default level")
}
// TestEnableDebugLoggingChangesLevel verifies the debug and non-debug
// configurations differ as intended: enabling debug makes debug
// records pass where they previously did not.
func TestEnableDebugLoggingChangesLevel(t *testing.T) {
t.Parallel()
l := newTestLogger(t)
log := l.Get()
ctx := context.Background()
require.False(t, log.Enabled(ctx, slog.LevelDebug),
"debug must start disabled")
l.EnableDebugLogging()
assert.True(t, log.Enabled(ctx, slog.LevelDebug),
"debug must be enabled after EnableDebugLogging")
}
+19
View File
@@ -0,0 +1,19 @@
package middleware
import (
"net/http"
"time"
)
// The /metrics rate limit, exported so the tests can count requests
// against it.
const (
MetricsRequestLimit = metricsRequestLimit
MetricsRequestWindow time.Duration = metricsRequestWindow
)
// RealIP is realIP, exported so the tests can check which address it
// takes as the client's.
func RealIP(r *http.Request) string {
return realIP(r)
}
+153 -16
View File
@@ -5,12 +5,14 @@ import (
"log/slog"
"net"
"net/http"
"net/netip"
"strings"
"time"
"github.com/99designs/basicauth-go"
"github.com/go-chi/chi/v5/middleware"
"github.com/go-chi/cors"
"github.com/go-chi/httprate"
"go.uber.org/fx"
"sneak.berlin/go/dnswatcher/internal/config"
@@ -21,6 +23,71 @@ import (
// corsMaxAge is the maximum age for CORS preflight responses.
const corsMaxAge = 300
// Rate limit for /metrics: each client address may send
// metricsRequestLimit requests per metricsRequestWindow. Every request
// counts, so password guessing gets at most 30 tries a minute per
// address. One Prometheus server scraping every 15 seconds sends 4
// requests a minute, and two scraping every 5 seconds from one address
// send 24, so normal scraping stays under the limit.
const (
metricsRequestLimit = 30
metricsRequestWindow = time.Minute
)
// Security response header values applied to every response.
//
// The CSP is as strict as the dashboard allows: the template ships no
// JavaScript, no inline styles, no inline event handlers and no images,
// and its only subresource is the embedded stylesheet at
// /s/css/tailwind.min.css, which style-src 'self' permits. Neither
// unsafe-inline nor unsafe-eval is used. frame-ancestors 'none' is the
// primary anti-framing control; X-Frame-Options is the legacy fallback.
const (
// hstsValue is emitted unconditionally, including over plain HTTP,
// because the service runs behind a TLS-terminating proxy and the
// browser must still enforce HTTPS end to end.
hstsValue = "max-age=31536000; includeSubDomains"
cspValue = "default-src 'self'; " +
"script-src 'none'; " +
"style-src 'self'; " +
"img-src 'self'; " +
"font-src 'none'; " +
"connect-src 'none'; " +
"object-src 'none'; " +
"base-uri 'none'; " +
"form-action 'none'; " +
"frame-ancestors 'none'"
frameOptionsValue = "DENY"
contentTypeOptionsValue = "nosniff"
// referrerPolicyValue is stricter than the policy minimum of
// strict-origin-when-cross-origin: the dashboard has no
// cross-origin navigation needs and its URL may name internal
// hosts.
referrerPolicyValue = "no-referrer"
permissionsPolicyValue = "accelerometer=(), " +
"autoplay=(), " +
"camera=(), " +
"display-capture=(), " +
"encrypted-media=(), " +
"fullscreen=(), " +
"geolocation=(), " +
"gyroscope=(), " +
"magnetometer=(), " +
"microphone=(), " +
"midi=(), " +
"payment=(), " +
"picture-in-picture=(), " +
"publickey-credentials-get=(), " +
"screen-wake-lock=(), " +
"usb=(), " +
"xr-spatial-tracking=()"
)
// Params contains dependencies for Middleware.
type Params struct {
fx.In
@@ -142,6 +209,12 @@ func isTrustedProxy(ip net.IP) bool {
// realIP extracts the client's real IP address from the request.
// Proxy headers are only trusted from RFC1918/loopback addresses.
//
// Each proxy adds to the end of X-Forwarded-For the address it got the
// request from, so the client can write every entry before the one the
// first trusted proxy added. The client address is therefore the
// rightmost entry that is not a trusted proxy, or the leftmost entry
// when they all are.
func realIP(r *http.Request) string {
addr := ipFromHostPort(r.RemoteAddr)
remoteIP := net.ParseIP(addr)
@@ -156,36 +229,100 @@ func realIP(r *http.Request) string {
return ip
}
if xff := r.Header.Get("X-Forwarded-For"); xff != "" {
if parts := strings.SplitN(
xff, ",", 2, //nolint:mnd
); len(parts) > 0 {
if ip := strings.TrimSpace(parts[0]); ip != "" {
return ip
}
// A proxy may add its entry as a header line of its own instead of
// appending to the line the client sent, so all lines form one list.
entries := strings.Split(
strings.Join(r.Header.Values("X-Forwarded-For"), ","), ",",
)
client := strings.TrimSpace(entries[0])
for i := len(entries) - 1; i > 0; i-- {
entry := strings.TrimSpace(entries[i])
if !isTrustedProxy(net.ParseIP(entry)) {
client = entry
break
}
}
if client != "" {
return client
}
return addr
}
// CORS returns CORS middleware.
// CORS returns middleware that lets any origin read a response. It is
// for the public, read-only routes only, so it allows only the
// methods those routes serve and no Authorization header.
func (m *Middleware) CORS() func(http.Handler) http.Handler {
return cors.Handler(cors.Options{
AllowedOrigins: []string{"*"},
AllowedMethods: []string{
"GET", "POST", "PUT", "DELETE", "OPTIONS",
},
AllowedHeaders: []string{
"Accept", "Authorization",
"Content-Type", "X-CSRF-Token",
},
AllowedOrigins: []string{"*"},
AllowedMethods: []string{"GET", "OPTIONS"},
AllowedHeaders: []string{"Accept", "Content-Type"},
ExposedHeaders: []string{"Link"},
AllowCredentials: false,
MaxAge: corsMaxAge,
})
}
// SecurityHeaders returns middleware that sets the security response
// headers required for production internet exposure on every response.
//
// The headers are set before the request reaches the next handler so
// that they are present on every response, including panics recovered
// by chi's Recoverer and timeouts produced by chi's Timeout.
func (m *Middleware) SecurityHeaders() func(http.Handler) http.Handler {
return func(next http.Handler) http.Handler {
return http.HandlerFunc(func(
writer http.ResponseWriter,
request *http.Request,
) {
header := writer.Header()
header.Set("Strict-Transport-Security", hstsValue)
header.Set("Content-Security-Policy", cspValue)
header.Set("X-Frame-Options", frameOptionsValue)
header.Set(
"X-Content-Type-Options",
contentTypeOptionsValue,
)
header.Set("Referrer-Policy", referrerPolicyValue)
header.Set(
"Permissions-Policy",
permissionsPolicyValue,
)
next.ServeHTTP(writer, request)
})
}
}
// MetricsRateLimit returns middleware for /metrics that answers 429
// Too Many Requests to a client address over the rate limit. The
// address is the one realIP works out, so a client that is not a
// trusted proxy cannot get a fresh allowance by sending its own
// X-Real-IP or X-Forwarded-For. CanonicalizeIP counts all IPv6
// addresses in one /64 as one client, since a client usually holds a
// whole /64. An IPv4 address a proxy reports in IPv6-mapped form
// (::ffff:203.0.113.1) is turned back into plain IPv4 first, as every
// such address is in the same /64.
func (m *Middleware) MetricsRateLimit() func(http.Handler) http.Handler {
return httprate.LimitBy(
metricsRequestLimit,
metricsRequestWindow,
func(request *http.Request) (string, error) {
ip := realIP(request)
addr, err := netip.ParseAddr(ip)
if err == nil {
ip = addr.Unmap().String()
}
return httprate.CanonicalizeIP(ip), nil
},
)
}
// MetricsAuth returns basic auth middleware for /metrics.
func (m *Middleware) MetricsAuth() func(http.Handler) http.Handler {
if m.params.Config.MetricsUsername == "" {
+565
View File
@@ -0,0 +1,565 @@
package middleware_test
import (
"net/http"
"net/http/httptest"
"strings"
"testing"
"time"
"github.com/go-chi/chi/v5"
"go.uber.org/fx/fxtest"
"sneak.berlin/go/dnswatcher/internal/config"
"sneak.berlin/go/dnswatcher/internal/globals"
"sneak.berlin/go/dnswatcher/internal/handlers"
"sneak.berlin/go/dnswatcher/internal/logger"
"sneak.berlin/go/dnswatcher/internal/middleware"
"sneak.berlin/go/dnswatcher/internal/notify"
"sneak.berlin/go/dnswatcher/internal/state"
)
// Expected security header values, spelled out literally so that any
// change to the middleware has to be made deliberately here as well.
const (
wantHSTS = "max-age=31536000; includeSubDomains"
wantCSP = "default-src 'self'; " +
"script-src 'none'; " +
"style-src 'self'; " +
"img-src 'self'; " +
"font-src 'none'; " +
"connect-src 'none'; " +
"object-src 'none'; " +
"base-uri 'none'; " +
"form-action 'none'; " +
"frame-ancestors 'none'"
wantFrameOptions = "DENY"
wantContentTypeOptions = "nosniff"
wantReferrerPolicy = "no-referrer"
wantPermissionsPolicy = "accelerometer=(), " +
"autoplay=(), " +
"camera=(), " +
"display-capture=(), " +
"encrypted-media=(), " +
"fullscreen=(), " +
"geolocation=(), " +
"gyroscope=(), " +
"magnetometer=(), " +
"microphone=(), " +
"midi=(), " +
"payment=(), " +
"picture-in-picture=(), " +
"publickey-credentials-get=(), " +
"screen-wake-lock=(), " +
"usb=(), " +
"xr-spatial-tracking=()"
)
// stylesheetPath is the only subresource the dashboard loads.
const stylesheetPath = "/s/css/tailwind.min.css"
// newTestLogger builds a logger for direct component construction.
func newTestLogger(t *testing.T) *logger.Logger {
t.Helper()
glob, err := globals.New(nil)
if err != nil {
t.Fatalf("globals.New: %v", err)
}
log, err := logger.New(nil, logger.Params{Globals: glob})
if err != nil {
t.Fatalf("logger.New: %v", err)
}
return log
}
// newTestMiddleware builds a Middleware without an fx application.
func newTestMiddleware(t *testing.T) *middleware.Middleware {
t.Helper()
glob, err := globals.New(nil)
if err != nil {
t.Fatalf("globals.New: %v", err)
}
mw, err := middleware.New(nil, middleware.Params{
Logger: newTestLogger(t),
Globals: glob,
Config: &config.Config{},
})
if err != nil {
t.Fatalf("middleware.New: %v", err)
}
return mw
}
// serveWithSecurityHeaders runs a GET through SecurityHeaders and
// returns the recorded response.
func serveWithSecurityHeaders(
t *testing.T,
target string,
handler http.Handler,
) *httptest.ResponseRecorder {
t.Helper()
mw := newTestMiddleware(t)
rec := httptest.NewRecorder()
req := httptest.NewRequestWithContext(
t.Context(), http.MethodGet, target, nil,
)
mw.SecurityHeaders()(handler).ServeHTTP(rec, req)
return rec
}
// okHandler writes a trivial 200 response.
func okHandler() http.Handler {
return http.HandlerFunc(func(
writer http.ResponseWriter,
_ *http.Request,
) {
writer.WriteHeader(http.StatusOK)
})
}
func TestSecurityHeaders(t *testing.T) {
t.Parallel()
tests := []struct {
name string
header string
want string
}{
{
"hsts",
"Strict-Transport-Security",
wantHSTS,
},
{
"csp",
"Content-Security-Policy",
wantCSP,
},
{
"frame options",
"X-Frame-Options",
wantFrameOptions,
},
{
"content type options",
"X-Content-Type-Options",
wantContentTypeOptions,
},
{
"referrer policy",
"Referrer-Policy",
wantReferrerPolicy,
},
{
"permissions policy",
"Permissions-Policy",
wantPermissionsPolicy,
},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
t.Parallel()
rec := serveWithSecurityHeaders(t, "/", okHandler())
got := rec.Header().Get(tt.header)
if got != tt.want {
t.Errorf(
"%s = %q, want %q",
tt.header, got, tt.want,
)
}
})
}
}
// TestSecurityHeadersCSPDirectives guards the properties the repo
// policy requires of the content security policy itself.
func TestSecurityHeadersCSPDirectives(t *testing.T) {
t.Parallel()
rec := serveWithSecurityHeaders(t, "/", okHandler())
csp := rec.Header().Get("Content-Security-Policy")
forbidden := []string{"unsafe-inline", "unsafe-eval"}
for _, directive := range forbidden {
if strings.Contains(csp, directive) {
t.Errorf("CSP must not contain %q: %q", directive, csp)
}
}
required := []string{
"default-src 'self'",
"script-src 'none'",
"style-src 'self'",
"frame-ancestors 'none'",
}
for _, directive := range required {
if !strings.Contains(csp, directive) {
t.Errorf("CSP must contain %q: %q", directive, csp)
}
}
}
// TestSecurityHeadersOnErrorResponse verifies the headers are emitted
// even when the wrapped handler fails, since they are set before the
// handler runs.
func TestSecurityHeadersOnErrorResponse(t *testing.T) {
t.Parallel()
failing := http.HandlerFunc(func(
writer http.ResponseWriter,
_ *http.Request,
) {
http.Error(
writer,
"boom",
http.StatusInternalServerError,
)
})
rec := serveWithSecurityHeaders(t, "/api/v1/status", failing)
if rec.Code != http.StatusInternalServerError {
t.Fatalf("status = %d, want 500", rec.Code)
}
if got := rec.Header().Get(
"X-Content-Type-Options",
); got != wantContentTypeOptions {
t.Errorf(
"X-Content-Type-Options = %q, want %q",
got, wantContentTypeOptions,
)
}
if got := rec.Header().Get(
"Strict-Transport-Security",
); got != wantHSTS {
t.Errorf(
"Strict-Transport-Security = %q, want %q",
got, wantHSTS,
)
}
}
// newTestHandlers builds real Handlers with empty monitoring state.
func newTestHandlers(t *testing.T) *handlers.Handlers {
t.Helper()
glob, err := globals.New(nil)
if err != nil {
t.Fatalf("globals.New: %v", err)
}
log := newTestLogger(t)
notifier, err := notify.New(fxtest.NewLifecycle(t), notify.Params{
Logger: log,
Config: &config.Config{},
})
if err != nil {
t.Fatalf("notify.New: %v", err)
}
st, err := state.New(fxtest.NewLifecycle(t), state.Params{
Logger: log,
Config: &config.Config{DataDir: t.TempDir()},
})
if err != nil {
t.Fatalf("state.New: %v", err)
}
hnd, err := handlers.New(nil, handlers.Params{
Logger: log,
Globals: glob,
State: st,
Notify: notifier,
})
if err != nil {
t.Fatalf("handlers.New: %v", err)
}
return hnd
}
// TestDashboardRendersWithSecurityHeaders renders the real dashboard
// through the middleware and checks that the policy still permits the
// one stylesheet the page loads.
func TestDashboardRendersWithSecurityHeaders(t *testing.T) {
t.Parallel()
mw := newTestMiddleware(t)
hnd := newTestHandlers(t)
router := chi.NewRouter()
router.Use(mw.SecurityHeaders())
router.Get("/", hnd.HandleDashboard())
rec := httptest.NewRecorder()
req := httptest.NewRequestWithContext(
t.Context(), http.MethodGet, "/", nil,
)
router.ServeHTTP(rec, req)
if rec.Code != http.StatusOK {
t.Fatalf("status = %d, want 200", rec.Code)
}
body := rec.Body.String()
if !strings.Contains(body, stylesheetPath) {
t.Errorf("dashboard does not reference %q", stylesheetPath)
}
if !strings.Contains(body, "dnswatcher") {
t.Errorf("dashboard body looks empty: %d bytes", len(body))
}
csp := rec.Header().Get("Content-Security-Policy")
if csp != wantCSP {
t.Errorf("CSP = %q, want %q", csp, wantCSP)
}
// The stylesheet is same-origin, so style-src 'self' allows it.
if !strings.Contains(csp, "style-src 'self'") {
t.Errorf("CSP would block %q: %q", stylesheetPath, csp)
}
}
// Addresses for the rate limit and realIP tests: a client connecting
// directly, a trusted proxy, and a client behind that proxy as the
// proxy's X-Real-IP or X-Forwarded-For header names it.
const (
directClient = "198.51.100.1:4000"
trustedProxy = "10.0.0.1:4000"
proxiedClient = "203.0.113.1"
)
// statusFrom sends a GET through handler as if from remoteAddr, with
// an X-Real-IP header when xRealIP is not empty, and returns the
// response status.
func statusFrom(
t *testing.T,
handler http.Handler,
remoteAddr string,
xRealIP string,
) int {
t.Helper()
req := httptest.NewRequestWithContext(
t.Context(), http.MethodGet, "/metrics", nil,
)
req.RemoteAddr = remoteAddr
if xRealIP != "" {
req.Header.Set("X-Real-IP", xRealIP)
}
rec := httptest.NewRecorder()
handler.ServeHTTP(rec, req)
return rec.Code
}
// TestMetricsRateLimitAllowsScraping checks that one address can send,
// within one window, what two Prometheus servers scraping every 5
// seconds send in that time, without being turned away.
func TestMetricsRateLimitAllowsScraping(t *testing.T) {
t.Parallel()
const scrapeInterval = 5 * time.Second
scrapes := 2 * int(middleware.MetricsRequestWindow/scrapeInterval)
limited := newTestMiddleware(t).MetricsRateLimit()(okHandler())
for i := range scrapes {
got := statusFrom(t, limited, directClient, "")
if got != http.StatusOK {
t.Fatalf(
"scrape %d of %d: status = %d, want 200",
i+1, scrapes, got,
)
}
}
}
// TestMetricsRateLimitKeysOnClientAddress checks which requests share
// an allowance. Each case uses up the allowance of one client, then
// sends one more request.
func TestMetricsRateLimitKeysOnClientAddress(t *testing.T) {
t.Parallel()
tests := []struct {
name string
usedRemoteAddr string
usedXRealIP string
nextRemoteAddr string
nextXRealIP string
want int
}{
{
"same address",
directClient, "",
directClient, "",
http.StatusTooManyRequests,
},
{
"another address",
directClient, "",
"198.51.100.2:4000", "",
http.StatusOK,
},
{
"own X-Real-IP from an untrusted address",
directClient, "",
directClient, "203.0.113.9",
http.StatusTooManyRequests,
},
{
"same client behind the proxy",
trustedProxy, proxiedClient,
trustedProxy, proxiedClient,
http.StatusTooManyRequests,
},
{
"another client behind the proxy",
trustedProxy, proxiedClient,
trustedProxy, "203.0.113.2",
http.StatusOK,
},
{
"another client behind the proxy, IPv6-mapped",
trustedProxy, "::ffff:203.0.113.1",
trustedProxy, "::ffff:203.0.113.2",
http.StatusOK,
},
{
"same IPv6 /64",
"[2001:db8::1]:4000", "",
"[2001:db8::2]:4000", "",
http.StatusTooManyRequests,
},
{
"another IPv6 /64",
"[2001:db8::1]:4000", "",
"[2001:db8:0:1::1]:4000", "",
http.StatusOK,
},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
t.Parallel()
limited := newTestMiddleware(t).MetricsRateLimit()(okHandler())
for range middleware.MetricsRequestLimit {
statusFrom(t, limited, tt.usedRemoteAddr, tt.usedXRealIP)
}
got := statusFrom(
t, limited, tt.nextRemoteAddr, tt.nextXRealIP,
)
if got != tt.want {
t.Errorf("status = %d, want %d", got, tt.want)
}
})
}
}
// TestRealIP checks which address realIP takes as the client's. Each
// element of forwardedFor is sent as an X-Forwarded-For header line of
// its own, and 198.51.100.9 is always an entry the client wrote itself.
func TestRealIP(t *testing.T) {
t.Parallel()
tests := []struct {
name string
remoteAddr string
xRealIP string
forwardedFor []string
want string
}{
{
"untrusted peer, both headers ignored",
directClient, proxiedClient, []string{"198.51.100.9"},
"198.51.100.1",
},
{
"X-Real-IP from a trusted proxy wins",
trustedProxy, proxiedClient, []string{"203.0.113.8"},
proxiedClient,
},
{
"client's own entry, then the one the proxy added",
trustedProxy, "", []string{"198.51.100.9, 203.0.113.1"},
proxiedClient,
},
{
"several trusted proxies",
trustedProxy, "",
[]string{"198.51.100.9, 203.0.113.1, 10.0.0.3, 10.0.0.2"},
proxiedClient,
},
{
"proxy adds a header line of its own",
trustedProxy, "", []string{"198.51.100.9", proxiedClient},
proxiedClient,
},
{
"every entry a trusted proxy",
trustedProxy, "", []string{"10.0.0.3, 10.0.0.2"},
"10.0.0.3",
},
{
"empty where the client address belongs",
trustedProxy, "", []string{"203.0.113.1, , 10.0.0.2"},
"10.0.0.1",
},
{
"no headers from a trusted proxy",
trustedProxy, "", nil,
"10.0.0.1",
},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
t.Parallel()
req := httptest.NewRequestWithContext(
t.Context(), http.MethodGet, "/", nil,
)
req.RemoteAddr = tt.remoteAddr
if tt.xRealIP != "" {
req.Header.Set("X-Real-IP", tt.xRealIP)
}
for _, line := range tt.forwardedFor {
req.Header.Add("X-Forwarded-For", line)
}
got := middleware.RealIP(req)
if got != tt.want {
t.Errorf("realIP = %q, want %q", got, tt.want)
}
})
}
}
+69 -3
View File
@@ -6,9 +6,11 @@ import (
"encoding/json"
"errors"
"io"
"maps"
"net/http"
"net/http/httptest"
"net/url"
"strings"
"sync"
"testing"
"time"
@@ -413,7 +415,8 @@ func sendSlackInfo(
svc *notify.Service, target *url.URL,
) error {
return svc.SendSlack(
context.Background(), target, "t", "m", prioInfo,
context.Background(), target, notify.ErrSlackFailed,
"t", "m", prioInfo,
)
}
@@ -506,6 +509,7 @@ func TestSendSlackPayloadFields(t *testing.T) {
err := svc.SendSlack(
context.Background(),
webhookURL,
notify.ErrSlackFailed,
"Alert Title",
"Alert body text",
"warning",
@@ -608,7 +612,8 @@ func TestSendSlackAllColors(t *testing.T) {
err := svc.SendSlack(
context.Background(),
webhookURL, "t", "m", tc.priority,
webhookURL, notify.ErrSlackFailed,
"t", "m", tc.priority,
)
if err != nil {
t.Fatalf("SendSlack error: %v", err)
@@ -659,7 +664,8 @@ func TestSendSlackNetworkError(t *testing.T) {
)
err := svc.SendSlack(
context.Background(), webhookURL, "t", "m", "info",
context.Background(), webhookURL, notify.ErrSlackFailed,
"t", "m", "info",
)
if err == nil {
t.Fatal("expected error for network failure")
@@ -1028,6 +1034,66 @@ func TestSendNotificationMattermostError(t *testing.T) {
)
}
// TestSendNotificationErrorNamesEndpoint verifies that, with both
// Slack and Mattermost set, a failed delivery's logged error names
// the endpoint that failed. Both are sent by the Slack sender.
func TestSendNotificationErrorNamesEndpoint(t *testing.T) {
t.Parallel()
srv := httptest.NewServer(
http.HandlerFunc(
func(w http.ResponseWriter, _ *http.Request) {
w.WriteHeader(http.StatusServiceUnavailable)
}),
)
defer srv.Close()
target, _ := url.Parse(srv.URL)
svc, logs := newLoggingService(http.DefaultTransport)
svc.SetSlackWebhookURL(target)
svc.SetMattermostWebhookURL(target)
svc.SetSleepFunc(instantSleep)
svc.SetRetryConfig(notify.RetryConfig{
MaxRetries: 1,
BaseDelay: time.Millisecond,
MaxDelay: time.Millisecond,
})
svc.SendNotification(
context.Background(), "t", "m", prioError,
)
waitForCondition(t, func() bool {
return svc.OutstandingDeliveries() == 0
})
got := map[string]string{}
for line := range strings.Lines(logs.String()) {
var record struct {
Msg string `json:"msg"`
Endpoint string `json:"endpoint"`
Error string `json:"error"`
}
_ = json.Unmarshal([]byte(line), &record)
if record.Msg == "failed to send notification after retries" {
got[record.Endpoint] = record.Error
}
}
want := map[string]string{
"slack": "slack notification failed: status 503",
"mattermost": "mattermost notification failed: status 503",
}
if !maps.Equal(got, want) {
t.Errorf("logged errors = %v, want %v", got, want)
}
}
// ── SlackPayload JSON marshaling ──────────────────────────
func TestSlackPayloadJSON(t *testing.T) {
+23 -6
View File
@@ -32,11 +32,27 @@ func NewRequestForTest(
// NewTestService creates a Service suitable for unit testing.
// It discards log output and uses the given transport.
func NewTestService(transport http.RoundTripper) *Service {
return &Service{
log: slog.New(slog.DiscardHandler),
transport: transport,
history: NewAlertHistory(),
}
return newService(slog.New(slog.DiscardHandler), transport)
}
// NewTestServiceWithLogger creates a Service that writes to the
// given handler, so tests can assert on emitted log records.
func NewTestServiceWithLogger(
transport http.RoundTripper,
handler slog.Handler,
) *Service {
return newService(slog.New(handler), transport)
}
// Drain exports drain for testing.
func (svc *Service) Drain(ctx context.Context) {
svc.drain(ctx)
}
// OutstandingDeliveries reports how many delivery goroutines
// are currently tracked as in flight.
func (svc *Service) OutstandingDeliveries() int64 {
return svc.outstanding.Load()
}
// SetNtfyURL sets the ntfy URL on a Service for testing.
@@ -69,10 +85,11 @@ func (svc *Service) SendNtfy(
func (svc *Service) SendSlack(
ctx context.Context,
webhookURL *url.URL,
failed error,
title, message, priority string,
) error {
return svc.sendSlack(
ctx, webhookURL, title, message, priority,
ctx, webhookURL, failed, title, message, priority,
)
}
+88 -65
View File
@@ -12,6 +12,8 @@ import (
"log/slog"
"net/http"
"net/url"
"sync"
"sync/atomic"
"time"
"go.uber.org/fx"
@@ -115,19 +117,41 @@ type Service struct {
history *AlertHistory
retryConfig RetryConfig
sleepFn func(time.Duration) <-chan time.Time
// Shutdown draining state. drainMu guards draining and
// serialises it against the counter increment in
// startDelivery; inFlight tracks the delivery goroutines
// themselves and outstanding mirrors its count so a timed
// out drain can report how many were abandoned.
drainMu sync.Mutex
draining bool
inFlight sync.WaitGroup
outstanding atomic.Int64
abandon chan struct{}
abandonOnce sync.Once
}
// newService builds a Service with the fields every Service
// needs regardless of how it was constructed.
func newService(
log *slog.Logger,
transport http.RoundTripper,
) *Service {
return &Service{
log: log,
transport: transport,
history: NewAlertHistory(),
abandon: make(chan struct{}),
}
}
// New creates a new notify Service.
func New(
_ fx.Lifecycle,
lifecycle fx.Lifecycle,
params Params,
) (*Service, error) {
svc := &Service{
log: params.Logger.Get(),
transport: http.DefaultTransport,
config: params.Config,
history: NewAlertHistory(),
}
svc := newService(params.Logger.Get(), http.DefaultTransport)
svc.config = params.Config
if params.Config.NtfyTopic != "" {
u, err := ValidateWebhookURL(
@@ -168,6 +192,14 @@ func New(
svc.mattermostWebhookURL = u
}
lifecycle.Append(fx.Hook{
OnStop: func(ctx context.Context) error {
svc.drain(ctx)
return nil
},
})
return svc, nil
}
@@ -194,6 +226,32 @@ func (svc *Service) SendNotification(
svc.dispatchMattermost(ctx, title, message, priority)
}
// dispatch delivers a notification to one endpoint on a
// tracked background goroutine.
//
// The delivery context is detached from ctx with
// context.WithoutCancel so that a cancelled caller does not
// kill a delivery already under way; the shutdown drain, not
// the caller, decides how long deliveries may keep running.
func (svc *Service) dispatch(
ctx context.Context,
endpoint string,
send func(context.Context) error,
) {
notifyCtx := context.WithoutCancel(ctx)
svc.startDelivery(endpoint, func() {
err := svc.deliverWithRetry(notifyCtx, endpoint, send)
if err != nil {
svc.log.Error(
"failed to send notification after retries",
"endpoint", endpoint,
"error", err,
)
}
})
}
func (svc *Service) dispatchNtfy(
ctx context.Context,
title, message, priority string,
@@ -202,26 +260,11 @@ func (svc *Service) dispatchNtfy(
return
}
go func() {
notifyCtx := context.WithoutCancel(ctx)
err := svc.deliverWithRetry(
notifyCtx, "ntfy",
func(c context.Context) error {
return svc.sendNtfy(
c, svc.ntfyURL,
title, message, priority,
)
},
svc.dispatch(ctx, "ntfy", func(c context.Context) error {
return svc.sendNtfy(
c, svc.ntfyURL, title, message, priority,
)
if err != nil {
svc.log.Error(
"failed to send ntfy notification "+
"after retries",
"error", err,
)
}
}()
})
}
func (svc *Service) dispatchSlack(
@@ -232,26 +275,12 @@ func (svc *Service) dispatchSlack(
return
}
go func() {
notifyCtx := context.WithoutCancel(ctx)
err := svc.deliverWithRetry(
notifyCtx, "slack",
func(c context.Context) error {
return svc.sendSlack(
c, svc.slackWebhookURL,
title, message, priority,
)
},
svc.dispatch(ctx, "slack", func(c context.Context) error {
return svc.sendSlack(
c, svc.slackWebhookURL, ErrSlackFailed,
title, message, priority,
)
if err != nil {
svc.log.Error(
"failed to send slack notification "+
"after retries",
"error", err,
)
}
}()
})
}
func (svc *Service) dispatchMattermost(
@@ -262,26 +291,15 @@ func (svc *Service) dispatchMattermost(
return
}
go func() {
notifyCtx := context.WithoutCancel(ctx)
err := svc.deliverWithRetry(
notifyCtx, "mattermost",
func(c context.Context) error {
return svc.sendSlack(
c, svc.mattermostWebhookURL,
title, message, priority,
)
},
)
if err != nil {
svc.log.Error(
"failed to send mattermost notification "+
"after retries",
"error", err,
svc.dispatch(
ctx, "mattermost",
func(c context.Context) error {
return svc.sendSlack(
c, svc.mattermostWebhookURL, ErrMattermostFailed,
title, message, priority,
)
}
}()
},
)
}
func (svc *Service) sendNtfy(
@@ -353,9 +371,14 @@ type SlackAttachment struct {
Text string `json:"text"`
}
// sendSlack posts to a Slack or Mattermost incoming webhook, which
// take the same payload. An HTTP error status is returned wrapped in
// failed, ErrSlackFailed or ErrMattermostFailed, so the error names
// the endpoint.
func (svc *Service) sendSlack(
ctx context.Context,
webhookURL *url.URL,
failed error,
title, message, priority string,
) error {
ctx, cancel := context.WithTimeout(
@@ -403,7 +426,7 @@ func (svc *Service) sendSlack(
if resp.StatusCode >= httpStatusClientError {
return fmt.Errorf(
"%w: status %d",
ErrSlackFailed, resp.StatusCode,
failed, resp.StatusCode,
)
}
+12 -1
View File
@@ -2,6 +2,7 @@ package notify
import (
"context"
"fmt"
"math"
"math/rand/v2"
"time"
@@ -114,13 +115,23 @@ func (svc *Service) deliverWithRetry(
"endpoint", endpoint,
"attempt", attempt+1,
"maxAttempts", cfg.MaxRetries+1,
"retryIn", delay,
// As text: the JSON log writes a time.Duration as
// bare nanoseconds.
"retryIn", delay.String(),
"error", lastErr,
)
select {
case <-ctx.Done():
return ctx.Err()
case <-svc.abandon:
// Shutdown drained past its deadline; stop
// sleeping rather than outlive the process.
// A nil channel (Service built without a
// constructor) simply never fires.
return fmt.Errorf(
"%w: %s", ErrDeliveryAbandoned, endpoint,
)
case <-svc.sleepFunc(delay):
}
}
+45
View File
@@ -2,6 +2,7 @@ package notify_test
import (
"context"
"encoding/json"
"errors"
"net/http"
"net/http/httptest"
@@ -189,6 +190,50 @@ func TestDeliverWithRetryExhaustsAttempts(t *testing.T) {
}
}
// TestDeliverWithRetryLogsRetryInAsText checks that the wait
// before a retry is logged as text such as "1.02s", not as a
// count of nanoseconds.
func TestDeliverWithRetryLogsRetryInAsText(t *testing.T) {
t.Parallel()
svc, logs := newLoggingService(http.DefaultTransport)
svc.SetRetryConfig(notify.RetryConfig{
MaxRetries: 1,
BaseDelay: time.Second,
MaxDelay: time.Second,
})
var waited time.Duration
svc.SetSleepFunc(func(d time.Duration) <-chan time.Time {
waited = d
return instantSleep(d)
})
_ = svc.DeliverWithRetry(
context.Background(), "test",
func(_ context.Context) error {
return errFail
},
)
// With one retry, only the first failure is logged.
var record map[string]any
err := json.Unmarshal([]byte(logs.String()), &record)
if err != nil {
t.Fatalf("log is not one JSON record: %v\n%s", err, logs)
}
if record["retryIn"] != waited.String() {
t.Errorf(
"retryIn logged as %v, want %q",
record["retryIn"], waited.String(),
)
}
}
func TestDeliverWithRetryRespectsContextCancellation(
t *testing.T,
) {
+119
View File
@@ -0,0 +1,119 @@
package notify
import (
"context"
"errors"
)
// ErrDeliveryAbandoned is returned by a retry loop that was
// cut short because shutdown drained past its deadline.
var ErrDeliveryAbandoned = errors.New(
"notification delivery abandoned at shutdown",
)
// startDelivery runs fn on its own goroutine while tracking it,
// so that drain can wait for it during shutdown.
//
// The WaitGroup counter is incremented here, on the caller's
// goroutine, before the worker exists: incrementing it inside
// the worker would race with drain's Wait and could let
// shutdown sail past a delivery that had not started yet.
//
// Once draining has begun the delivery is refused outright
// rather than queued, so a steady stream of newly submitted
// notifications cannot keep extending the drain.
func (svc *Service) startDelivery(endpoint string, fn func()) {
svc.drainMu.Lock()
if svc.draining {
svc.drainMu.Unlock()
svc.log.Warn(
"notification not dispatched: shutdown in progress",
"endpoint", endpoint,
)
return
}
svc.outstanding.Add(1)
// WaitGroup.Go increments the counter synchronously, here,
// and only then starts the goroutine.
svc.inFlight.Go(func() {
// Runs before the WaitGroup counter is decremented, so
// a drain that times out reports an accurate count.
defer svc.outstanding.Add(-1)
fn()
})
svc.drainMu.Unlock()
}
// drain waits for in-flight notification deliveries to finish.
//
// It first stops accepting new deliveries, then waits until
// either every outstanding delivery has completed or ctx
// expires — whichever comes first. ctx is the context fx
// passes to the OnStop hook, so a permanently dead webhook
// cannot hang shutdown indefinitely.
//
// When the deadline arrives with deliveries still outstanding,
// the count is logged at warn level and the abandon channel is
// closed, which releases any retry loop sleeping in backoff.
// Deliveries already inside an HTTP round trip are bounded by
// the existing httpClientTimeout instead.
//
// A ctx that is already expired on entry is not by itself cause
// for alarm: if nothing is outstanding there is nothing to
// abandon, and the drain says so at debug level rather than
// warning about deliveries that do not exist.
func (svc *Service) drain(ctx context.Context) {
svc.drainMu.Lock()
svc.draining = true
svc.drainMu.Unlock()
done := make(chan struct{})
go func() {
svc.inFlight.Wait()
close(done)
}()
select {
case <-done:
svc.log.Debug(
"all in-flight notifications completed",
)
case <-ctx.Done():
// outstanding is decremented before the WaitGroup
// counter, and startDelivery can no longer add to it
// now that draining is set, so a zero here means every
// delivery really did finish. ctx expiring in that
// state (an OnStop context that was already cancelled
// on entry is the usual way) abandons nothing, so it
// must not close abandon or warn about it.
abandoned := svc.outstanding.Load()
if abandoned == 0 {
svc.log.Debug(
"all in-flight notifications completed",
)
return
}
svc.abandonOnce.Do(func() {
if svc.abandon != nil {
close(svc.abandon)
}
})
svc.log.Warn(
"shutdown deadline reached with notifications "+
"still in flight; abandoning them",
"abandoned", abandoned,
"error", ctx.Err(),
)
}
}
+593
View File
@@ -0,0 +1,593 @@
package notify_test
import (
"bytes"
"context"
"log/slog"
"net/http"
"net/http/httptest"
"net/url"
"strings"
"sync"
"sync/atomic"
"testing"
"time"
"go.uber.org/fx"
"sneak.berlin/go/dnswatcher/internal/config"
"sneak.berlin/go/dnswatcher/internal/globals"
"sneak.berlin/go/dnswatcher/internal/logger"
"sneak.berlin/go/dnswatcher/internal/notify"
)
// Timings used by the drain tests. They stay in the same
// 10-100ms band as the retry tests so the suite never waits on
// a real backoff delay.
const (
// inFlightHold is how long a delivery is kept mid-request
// before the handler is released.
inFlightHold = 30 * time.Millisecond
// drainDeadline bounds a drain that is expected to time
// out.
drainDeadline = 50 * time.Millisecond
// timeoutDrainBound is how long a drain given drainDeadline
// may take to return before the test gives up on it. At
// forty times drainDeadline it leaves ample room for
// scheduling delay on a loaded box under -race, yet it is far
// below the test binary's -timeout, so a drain that its
// deadline does not bound fails that one test instead of
// hanging the package.
timeoutDrainBound = 2 * time.Second
// longDrainDeadline is the deadline given to a drain that is
// expected to finish well before it: when the in-flight
// delivery completes after inFlightHold, or at once when
// nothing is in flight. It is far above inFlightHold, so
// those drains never reach it, and four times
// idleDrainBound, so an idle drain that waited for its
// deadline instead of returning fails that bound.
longDrainDeadline = 2 * time.Second
// reachEndpointTimeout is how long a submitted delivery may
// take to reach the test server. That normally takes a few
// milliseconds; the margin is for a loaded box under -race,
// and only a failing run ever waits this long.
reachEndpointTimeout = 2 * time.Second
// settleDelay is how long to wait before asserting that
// something did *not* happen.
settleDelay = 50 * time.Millisecond
// idleDrainBound is the upper bound on a drain that has
// nothing in flight. It is deliberately far above the cost
// of the goroutine hop through inFlight.Wait() — which
// reached 57ms on a loaded box under -race with the package's
// parallel tests — and far below longDrainDeadline, the
// deadline such a drain is given. A drain that blocked until
// its deadline instead of returning on the WaitGroup
// therefore still fails this bound, but scheduling delay
// alone cannot.
idleDrainBound = 500 * time.Millisecond
)
// syncBuffer is an io.Writer safe for concurrent use, so log
// output written from delivery goroutines can be inspected.
type syncBuffer struct {
mu sync.Mutex
buf bytes.Buffer
}
func (sb *syncBuffer) Write(p []byte) (int, error) {
sb.mu.Lock()
defer sb.mu.Unlock()
return sb.buf.Write(p) //nolint:wrapcheck // test helper
}
func (sb *syncBuffer) String() string {
sb.mu.Lock()
defer sb.mu.Unlock()
return sb.buf.String()
}
// newLoggingService returns a Service writing JSON logs into
// the returned buffer, debug level included.
func newLoggingService(
transport http.RoundTripper,
) (*notify.Service, *syncBuffer) {
logs := &syncBuffer{}
handler := slog.NewJSONHandler(
logs, &slog.HandlerOptions{Level: slog.LevelDebug},
)
return notify.NewTestServiceWithLogger(transport, handler),
logs
}
// blockingNtfyServer returns a server whose handler signals on
// entered, waits for release, and then responds 200.
func blockingNtfyServer(
entered chan<- struct{},
release <-chan struct{},
served *atomic.Bool,
) *httptest.Server {
var once sync.Once
return httptest.NewServer(
http.HandlerFunc(
func(w http.ResponseWriter, _ *http.Request) {
once.Do(func() { close(entered) })
<-release
served.Store(true)
w.WriteHeader(http.StatusOK)
}),
)
}
// TestDrainWaitsForInFlightDelivery verifies that a delivery
// already under way when shutdown starts is allowed to finish.
func TestDrainWaitsForInFlightDelivery(t *testing.T) {
t.Parallel()
var served atomic.Bool
entered := make(chan struct{})
release := make(chan struct{})
srv := blockingNtfyServer(entered, release, &served)
defer srv.Close()
// srv.Close waits for the handler, so release it however the
// test ends; otherwise a drain that returns early hangs the
// package instead of failing this test.
releaseHandler := sync.OnceFunc(func() { close(release) })
defer releaseHandler()
topicURL, _ := url.Parse(srv.URL)
svc := notify.NewTestService(http.DefaultTransport)
svc.SetNtfyURL(topicURL)
svc.SendNotification(
context.Background(), "t", "m", prioInfo,
)
// Make sure the delivery really is mid-request before the
// drain begins.
select {
case <-entered:
case <-time.After(reachEndpointTimeout):
t.Fatal("delivery never reached the endpoint")
}
// As in TestDrainBoundedByContextDeadline: start is captured
// before the clock it is compared against, here the timer
// holding the delivery open, so elapsed covers the whole hold
// and the lower bound cannot come out short from scheduling
// delay alone.
start := time.Now()
timer := time.AfterFunc(inFlightHold, releaseHandler)
defer timer.Stop()
ctx, cancel := context.WithTimeout(
context.Background(), longDrainDeadline,
)
defer cancel()
svc.Drain(ctx)
elapsed := time.Since(start)
if !served.Load() {
t.Error(
"drain returned before the in-flight delivery " +
"completed",
)
}
if elapsed < inFlightHold {
t.Errorf(
"drain took %v, want at least %v",
elapsed, inFlightHold,
)
}
if got := svc.OutstandingDeliveries(); got != 0 {
t.Errorf("outstanding deliveries = %d, want 0", got)
}
}
// neverFires returns a channel that never delivers, standing in
// for a long backoff sleep without actually sleeping.
func neverFires(_ time.Duration) <-chan time.Time {
return make(chan time.Time)
}
// TestDrainBoundedByContextDeadline verifies that a delivery
// stuck retrying against a dead endpoint does not hold shutdown
// past the OnStop context deadline, and that the abandoned
// deliveries are logged at warn level rather than dropped
// silently.
func TestDrainBoundedByContextDeadline(t *testing.T) {
t.Parallel()
var requests atomic.Int64
srv := httptest.NewServer(
http.HandlerFunc(
func(w http.ResponseWriter, _ *http.Request) {
requests.Add(1)
w.WriteHeader(http.StatusInternalServerError)
}),
)
defer srv.Close()
topicURL, _ := url.Parse(srv.URL)
svc, logs := newLoggingService(http.DefaultTransport)
svc.SetNtfyURL(topicURL)
// Never let the backoff sleep complete: the delivery is
// parked in its retry wait until shutdown releases it.
svc.SetSleepFunc(neverFires)
svc.SetRetryConfig(notify.RetryConfig{
MaxRetries: 5,
BaseDelay: time.Hour,
MaxDelay: time.Hour,
})
svc.SendNotification(
context.Background(), "t", "m", prioError,
)
waitForCondition(t, func() bool {
return requests.Load() >= 1 &&
svc.OutstandingDeliveries() == 1
})
// start must be captured *before* the deadline clock starts,
// so that the measured interval is a superset of the deadline
// interval. Capturing it after context.WithTimeout would
// make elapsed structurally smaller than drainDeadline and
// the lower bound below unfalsifiable-by-luck: it would fail
// whenever the two statements were separated by any
// scheduling delay, and pass otherwise, regardless of what
// the drain did.
start := time.Now()
ctx, cancel := context.WithTimeout(
context.Background(), drainDeadline,
)
defer cancel()
// The upper bound is enforced by a watchdog rather than by
// measuring after the fact: a drain that is not bounded at
// all never returns here (the delivery is parked in a backoff
// that never fires), so an unbounded drain must fail this
// test promptly instead of hanging the package until the test
// binary's -timeout.
returned := make(chan struct{})
go func() {
defer close(returned)
svc.Drain(ctx)
}()
select {
case <-returned:
case <-time.After(timeoutDrainBound):
t.Fatalf(
"drain did not return within %v; its %v deadline "+
"did not bound it",
timeoutDrainBound, drainDeadline,
)
}
// The lower bound is the real assertion: the drain must have
// waited for its whole deadline rather than giving up on the
// outstanding delivery early. With start captured above, an
// early return is the only thing that can make it fail.
if elapsed := time.Since(start); elapsed < drainDeadline {
t.Errorf(
"drain returned after %v, before its %v deadline",
elapsed, drainDeadline,
)
}
assertAbandonLogged(t, logs.String())
// The abandoned delivery must stop retrying rather than
// outlive the drain.
waitForCondition(t, func() bool {
return svc.OutstandingDeliveries() == 0
})
}
// assertAbandonLogged checks that the drain logged the
// abandoned deliveries at warn level with a count.
func assertAbandonLogged(t *testing.T, output string) {
t.Helper()
if !strings.Contains(output, `"level":"WARN"`) {
t.Errorf(
"abandoned deliveries not logged at warn level; "+
"log output: %s",
output,
)
}
if !strings.Contains(output, `"abandoned":1`) {
t.Errorf(
"abandoned delivery count not logged; "+
"log output: %s",
output,
)
}
}
// TestDrainRefusesNewDeliveries verifies that notifications
// submitted after the drain has begun are refused and logged,
// so a stream of new work cannot extend shutdown indefinitely.
func TestDrainRefusesNewDeliveries(t *testing.T) {
t.Parallel()
var requests atomic.Int64
srv := httptest.NewServer(
http.HandlerFunc(
func(w http.ResponseWriter, _ *http.Request) {
requests.Add(1)
w.WriteHeader(http.StatusOK)
}),
)
defer srv.Close()
target, _ := url.Parse(srv.URL)
svc, logs := newLoggingService(http.DefaultTransport)
svc.SetNtfyURL(target)
svc.SetSlackWebhookURL(target)
svc.SetMattermostWebhookURL(target)
ctx, cancel := context.WithTimeout(
context.Background(), longDrainDeadline,
)
defer cancel()
// Nothing is in flight, so this returns immediately and
// leaves the service refusing further deliveries.
svc.Drain(ctx)
for range 3 {
svc.SendNotification(
context.Background(), "t", "m", prioInfo,
)
}
time.Sleep(settleDelay)
if got := requests.Load(); got != 0 {
t.Errorf(
"%d requests reached the endpoint after drain, "+
"want 0",
got,
)
}
if got := svc.OutstandingDeliveries(); got != 0 {
t.Errorf("outstanding deliveries = %d, want 0", got)
}
output := logs.String()
if !strings.Contains(output, "shutdown in progress") {
t.Errorf(
"refused deliveries not logged; log output: %s",
output,
)
}
}
// recordingLifecycle is a minimal fx.Lifecycle that records the
// hooks appended to it, so the wiring done by notify.New can be
// inspected without standing up a whole fx application.
type recordingLifecycle struct {
hooks []fx.Hook
}
func (l *recordingLifecycle) Append(hook fx.Hook) {
l.hooks = append(l.hooks, hook)
}
// newNotifyService builds a Service through the real
// constructor, wired to the given lifecycle.
func newNotifyService(
t *testing.T,
lifecycle fx.Lifecycle,
ntfyTopic string,
) *notify.Service {
t.Helper()
g, err := globals.New(nil)
if err != nil {
t.Fatalf("globals.New: %v", err)
}
log, err := logger.New(nil, logger.Params{Globals: g})
if err != nil {
t.Fatalf("logger.New: %v", err)
}
svc, err := notify.New(lifecycle, notify.Params{
Logger: log,
Config: &config.Config{NtfyTopic: ntfyTopic},
})
if err != nil {
t.Fatalf("notify.New: %v", err)
}
return svc
}
// TestNewRegistersDrainingStopHook verifies that notify.New
// wires an OnStop hook into the fx lifecycle and that the hook
// waits for in-flight deliveries.
func TestNewRegistersDrainingStopHook(t *testing.T) {
t.Parallel()
var served atomic.Bool
entered := make(chan struct{})
release := make(chan struct{})
srv := blockingNtfyServer(entered, release, &served)
defer srv.Close()
// As in TestDrainWaitsForInFlightDelivery: release the handler
// however the test ends, before srv.Close waits for it.
releaseHandler := sync.OnceFunc(func() { close(release) })
defer releaseHandler()
lifecycle := &recordingLifecycle{}
svc := newNotifyService(t, lifecycle, srv.URL)
if len(lifecycle.hooks) != 1 {
t.Fatalf(
"appended %d lifecycle hooks, want 1",
len(lifecycle.hooks),
)
}
stop := lifecycle.hooks[0].OnStop
if stop == nil {
t.Fatal("lifecycle hook has no OnStop function")
}
svc.SendNotification(
context.Background(), "t", "m", prioInfo,
)
select {
case <-entered:
case <-time.After(reachEndpointTimeout):
t.Fatal("delivery never reached the endpoint")
}
timer := time.AfterFunc(inFlightHold, releaseHandler)
defer timer.Stop()
ctx, cancel := context.WithTimeout(
context.Background(), longDrainDeadline,
)
defer cancel()
err := stop(ctx)
if err != nil {
t.Fatalf("OnStop returned error: %v", err)
}
if !served.Load() {
t.Error(
"OnStop returned before the in-flight delivery " +
"completed",
)
}
}
// TestDrainWithoutDeliveriesReturnsImmediately verifies the
// common case: nothing in flight, shutdown is not delayed.
func TestDrainWithoutDeliveriesReturnsImmediately(t *testing.T) {
t.Parallel()
svc := notify.NewTestService(http.DefaultTransport)
// Captured before the deadline clock, as elsewhere in this
// file; for an upper bound that is the conservative
// direction, since the measured interval can then only be
// longer than the drain itself.
start := time.Now()
ctx, cancel := context.WithTimeout(
context.Background(), longDrainDeadline,
)
defer cancel()
svc.Drain(ctx)
if elapsed := time.Since(start); elapsed > idleDrainBound {
t.Errorf(
"drain of an idle service took %v, want at most "+
"%v; its deadline was %v",
elapsed, idleDrainBound, longDrainDeadline,
)
}
}
// TestDrainWithCancelledContextDoesNotWarn verifies that an
// OnStop context that is already dead on entry does not produce
// an "abandoning them" warning when there was nothing in flight
// to abandon, and that the drain returns and says at debug level
// that nothing was in flight. The expired context wins the
// select immediately, so only the outstanding count can tell the
// difference between a genuine timeout and a shutdown that had
// simply already run out of time with no work left.
func TestDrainWithCancelledContextDoesNotWarn(t *testing.T) {
t.Parallel()
svc, logs := newLoggingService(http.DefaultTransport)
ctx, cancel := context.WithCancel(context.Background())
cancel()
// A watchdog, as in TestDrainBoundedByContextDeadline, so
// that a drain which never returns fails here instead of
// hanging the package.
returned := make(chan struct{})
go func() {
defer close(returned)
svc.Drain(ctx)
}()
select {
case <-returned:
case <-time.After(idleDrainBound):
t.Fatalf(
"drain with nothing in flight and a cancelled "+
"context did not return within %v",
idleDrainBound,
)
}
output := logs.String()
// The absence of a warning alone would also pass if the drain
// logged nothing at all, so require the debug line it writes
// when it finds nothing outstanding.
if !strings.Contains(
output, "all in-flight notifications completed",
) {
t.Errorf(
"drain did not log that nothing was in flight; "+
"log output: %s",
output,
)
}
if strings.Contains(output, `"level":"WARN"`) {
t.Errorf(
"drain with nothing in flight warned about "+
"abandoned deliveries; log output: %s",
output,
)
}
}
+3 -1
View File
@@ -193,7 +193,9 @@ func (c *Checker) checkConnection(
c.log.Debug(
"port check succeeded",
"target", target,
"latency", latency,
// As text: the JSON log writes a time.Duration as bare
// nanoseconds.
"latency", latency.String(),
)
return &PortResult{
+2 -2
View File
@@ -7,8 +7,8 @@ import (
"github.com/miekg/dns"
)
// DNSClient abstracts DNS wire-protocol exchanges so the resolver
// can be tested without hitting real nameservers.
// DNSClient sends one DNS message to a nameserver and returns the
// reply. The resolver holds one for UDP and one for TCP.
type DNSClient interface {
ExchangeContext(
ctx context.Context,
+24
View File
@@ -10,12 +10,36 @@ var (
"no authoritative nameservers found",
)
// ErrNoNameserverAnswered is returned when every nameserver
// asked about a name timed out, failed or returned a referral,
// so whether the name has addresses is unknown.
ErrNoNameserverAnswered = errors.New("no nameserver answered")
// ErrUnusableReply is returned when a server replied with an
// error such as SERVFAIL, or with a referral that leads no
// closer to the name asked about.
ErrUnusableReply = errors.New(
"reply is an error or a referral that leads no closer",
)
// ErrIntercepted is returned when every root server refused a
// query. Root servers refuse no query, so the refusals came from
// something on the network answering in their place.
ErrIntercepted = errors.New("this network intercepts DNS queries")
// ErrCNAMEDepthExceeded is returned when a CNAME chain
// exceeds MaxCNAMEDepth.
ErrCNAMEDepthExceeded = errors.New(
"CNAME chain depth exceeded",
)
// ErrLookupDepthExceeded is returned when nameserver addresses
// were not looked up because lookups were already maxLookupDepth
// deep, one inside another.
ErrLookupDepthExceeded = errors.New(
"lookups of nameserver addresses go too deep",
)
// ErrContextCanceled wraps context cancellation for the
// resolver's iterative queries.
ErrContextCanceled = errors.New("context canceled")
+98
View File
@@ -0,0 +1,98 @@
package resolver
import (
"context"
"github.com/miekg/dns"
)
// ExtractRecordValue exports extractRecordValue for testing.
func ExtractRecordValue(rr dns.RR) string {
return extractRecordValue(rr)
}
// CollectAnswerRecords exports collectAnswerRecords for testing.
func CollectAnswerRecords(msg *dns.Msg, resp *NameserverResponse) {
var state queryState
collectAnswerRecords(msg, resp, &state)
}
// UsableReply exports usableReply for testing.
func UsableReply(resp *dns.Msg, zone string, name string) bool {
return usableReply(resp, zone, name)
}
// NSSetFrom exports nsSetFrom for testing.
func NSSetFrom(resp *dns.Msg, domain string) []string {
return nsSetFrom(resp, domain)
}
// CollectIPs exports collectIPs for testing.
func CollectIPs(
results map[string]*NameserverResponse,
) ([]string, string, error) {
return collectIPs(results)
}
// QueryServers exports queryServers for testing.
func (r *Resolver) QueryServers(
ctx context.Context,
servers []string,
zone string,
name string,
qtype uint16,
) (*dns.Msg, error) {
return r.queryServers(ctx, servers, zone, name, qtype)
}
// QueryEachNS exports queryEachNS for testing.
func (r *Resolver) QueryEachNS(
ctx context.Context,
nameservers []string,
hostname string,
) (map[string]*NameserverResponse, error) {
return r.queryEachNS(ctx, nameservers, hostname, recordTypes())
}
// ResolveNSIPs exports resolveNSIPs for testing, looking each name up
// as a lookup that no other lookup started.
func (r *Resolver) ResolveNSIPs(
ctx context.Context,
nsNames []string,
) []string {
ips, _ := r.resolveNSIPs(ctx, nsNames, 1)
return ips
}
// MaxLookupDepth exports maxLookupDepth for testing.
const MaxLookupDepth = maxLookupDepth
// QueryZone exports queryZone for testing.
func (r *Resolver) QueryZone(
ctx context.Context,
given []string,
withoutAddresses []string,
zone string,
name string,
qtype uint16,
depth int,
) (*dns.Msg, error) {
return r.queryZone(
ctx, given, withoutAddresses, zone, name, qtype, depth,
)
}
// RootServerList exports rootServerList for testing.
func RootServerList() []string {
return rootServerList()
}
// Shuffled exports shuffled for testing.
func Shuffled(
servers []string,
shuffle func(n int, swap func(i, j int)),
) []string {
return shuffled(servers, shuffle)
}
+386 -127
View File
@@ -4,7 +4,9 @@ import (
"context"
"errors"
"fmt"
"math/rand/v2"
"net"
"slices"
"sort"
"strings"
"time"
@@ -17,7 +19,16 @@ const (
maxRetries = 2
maxDelegation = 20
timeoutMultiplier = 2
minDomainLabels = 2
// maxLookupDepth is how many lookups of nameserver addresses may be
// under way one inside another. Looking up a nameserver's address
// can meet a referral that names nameservers without their
// addresses, which are then looked up in turn; without a limit,
// delegations that point at each other would never end. Each level
// multiplies the queries sent. pool.ntp.org needs three: the
// address of its nameserver g.ntpns.org can need a.ntpns.org's,
// which needs a bitnames.com nameserver's.
maxLookupDepth = 3
)
// ErrRefused is returned when a DNS server refuses a query.
@@ -106,9 +117,8 @@ func (r *Resolver) retryTCP(
return resp
}
// queryDNS sends a DNS query to a specific server IP.
// Tries non-recursive first, falls back to recursive on
// REFUSED (handles DNS interception environments).
// queryDNS sends a DNS query to a specific server IP, never asking it
// for recursion. A reply of REFUSED is returned as ErrRefused.
func (r *Resolver) queryDNS(
ctx context.Context,
serverIP string,
@@ -132,25 +142,12 @@ func (r *Resolver) queryDNS(
}
if resp.Rcode == dns.RcodeRefused {
msg.RecursionDesired = true
resp, err = r.tryExchange(ctx, msg, addr)
if err != nil {
return nil, fmt.Errorf(
"query %s @%s: %w", name, serverIP, err,
)
}
if resp.Rcode == dns.RcodeRefused {
return nil, fmt.Errorf(
"query %s @%s: %w", name, serverIP, ErrRefused,
)
}
return nil, fmt.Errorf(
"query %s @%s: %w", name, serverIP, ErrRefused,
)
}
resp = r.retryTCP(ctx, msg, addr, resp)
return resp, nil
return r.retryTCP(ctx, msg, addr, resp), nil
}
func extractNSSet(rrs []dns.RR) []string {
@@ -208,21 +205,35 @@ func (r *Resolver) followDelegation(
domain string,
servers []string,
) ([]string, error) {
// servers are the root servers, the servers of zone ".".
zone := "."
var withoutAddresses []string
for range maxDelegation {
if checkCtx(ctx) != nil {
return nil, ErrContextCanceled
}
resp, err := r.queryServers(
ctx, servers, domain, dns.TypeNS,
resp, err := r.queryZone(
ctx, servers, withoutAddresses, zone, domain, dns.TypeNS, 0,
)
if err != nil {
return nil, err
}
ansNS := extractNSSet(resp.Answer)
if len(ansNS) > 0 {
return ansNS, nil
nsSet := nsSetFrom(resp, domain)
if len(nsSet) > 0 {
return nsSet, nil
}
// An authoritative reply comes from the servers of the zone
// domain is in; it is not a referral, even when its authority
// section lists that zone's NS records. Without NS records in
// the answer, domain is not the zone's apex and has no
// nameservers of its own.
if resp.Authoritative {
return nil, ErrNoNameservers
}
authNS := extractNSSet(resp.Ns)
@@ -230,65 +241,236 @@ func (r *Resolver) followDelegation(
return r.resolveNSIterative(ctx, domain)
}
glue := extractGlue(resp.Extra)
nextServers := glueIPs(authNS, glue)
if len(nextServers) == 0 {
nextServers = r.resolveNSIPs(ctx, authNS)
}
if len(nextServers) == 0 {
return nil, ErrNoNameservers
}
servers = nextServers
servers, withoutAddresses = referralNameservers(resp)
zone = referralZone(resp)
}
return nil, ErrNoNameservers
}
// shuffled returns a copy of servers in the order shuffle puts them
// in. The resolver passes rand.Shuffle, so each time it walks a list of
// servers it starts at a random one, and no one server gets every
// first query.
func shuffled(
servers []string,
shuffle func(n int, swap func(i, j int)),
) []string {
order := slices.Clone(servers)
shuffle(len(order), func(i, j int) {
order[i], order[j] = order[j], order[i]
})
return order
}
// queryServers asks servers, the servers of zone, about name in a random
// order until one gives a usable reply. A server that times out, refuses
// or gives a reply that is not usable is passed over for the next. When
// every server refused, the error says so, and when they are the root
// servers it is ErrIntercepted.
func (r *Resolver) queryServers(
ctx context.Context,
servers []string,
zone string,
name string,
qtype uint16,
) (*dns.Msg, error) {
var lastErr error
for _, ip := range servers {
refused := 0
for _, ip := range shuffled(servers, rand.Shuffle) {
if checkCtx(ctx) != nil {
return nil, ErrContextCanceled
}
resp, err := r.queryDNS(ctx, ip, name, qtype)
if err == nil && !usableReply(resp, zone, name) {
err = fmt.Errorf(
"query %s @%s: %w", name, ip, ErrUnusableReply,
)
}
if err == nil {
return resp, nil
}
if errors.Is(err, ErrRefused) {
refused++
}
lastErr = err
}
if refused == len(servers) && zone == "." {
return nil, fmt.Errorf(
"every root server refused a query for %s: %w",
name, ErrIntercepted,
)
}
if refused == len(servers) {
return nil, fmt.Errorf(
"every server of %s refused a query for %s: %w",
zone, name, ErrRefused,
)
}
return nil, fmt.Errorf("all servers failed: %w", lastErr)
}
func (r *Resolver) resolveNSIPs(
ctx context.Context,
nsNames []string,
) []string {
var ips []string
// usableReply reports whether resp, a reply from one of the servers of
// zone to a query about name, is usable. An error reply such as SERVFAIL
// is not. Nor is a referral, unless it refers the query to a zone below
// zone that name is in: a server that refers it back to zone, up or
// sideways does not serve zone as it should.
func usableReply(resp *dns.Msg, zone string, name string) bool {
if resp.Rcode != dns.RcodeSuccess && resp.Rcode != dns.RcodeNameError {
return false
}
for _, ns := range nsNames {
resolved, err := r.resolveARecord(ctx, ns)
if err == nil {
ips = append(ips, resolved...)
}
child := referralZone(resp)
if resp.Authoritative || len(resp.Answer) > 0 || child == "" {
return true
}
if len(ips) > 0 {
break
return child != zone && dns.IsSubDomain(zone, child) &&
dns.IsSubDomain(child, name)
}
// referralZone returns the zone a referral refers the query to: the
// owner name of the NS records in resp's authority section, or "" when
// there are none.
func referralZone(resp *dns.Msg) string {
for _, rr := range resp.Ns {
if ns, ok := rr.(*dns.NS); ok {
return strings.ToLower(ns.Hdr.Name)
}
}
return ips
return ""
}
// nsSetFrom returns the NS set of domain that resp, a reply to a query
// for domain's NS records, gives: the delegation in a referral to domain
// itself, or else the NS records in the answer; empty when it gives
// neither. A referral to domain comes from its parent zone's servers,
// which all hold the same delegation, so the set does not depend on
// which of them answered. domain's own servers, which can disagree about
// their NS records, are then not asked.
func nsSetFrom(resp *dns.Msg, domain string) []string {
if referralZone(resp) == domain {
return extractNSSet(resp.Ns)
}
return extractNSSet(resp.Answer)
}
// referralNameservers returns the IPv4 addresses that resp, a referral,
// gives for the nameservers it names, and the names of the nameservers
// it gives no address for.
func referralNameservers(resp *dns.Msg) ([]string, []string) {
glue := extractGlue(resp.Extra)
var given, withoutAddresses []string
for _, ns := range extractNSSet(resp.Ns) {
ips := glueIPs([]string{ns}, glue)
if len(ips) == 0 {
withoutAddresses = append(withoutAddresses, ns)
}
given = append(given, ips...)
}
return given, withoutAddresses
}
// queryZone asks the servers of zone about name as queryServers does:
// first those at given, the addresses a referral gave, and only when
// none of them gives a usable reply, the nameservers named
// withoutAddresses, once their addresses are looked up. depth is how
// many lookups of a nameserver's address are under way, 0 in the walk
// to a domain's nameservers; at maxLookupDepth, no address is looked
// up. When the limit is why none was found, here or in a lookup this
// one started, the error is ErrLookupDepthExceeded.
func (r *Resolver) queryZone(
ctx context.Context,
given []string,
withoutAddresses []string,
zone string,
name string,
qtype uint16,
depth int,
) (*dns.Msg, error) {
err := fmt.Errorf(
"no address for any nameserver of %s: %w", zone, ErrNoNameservers,
)
if len(given) > 0 {
var resp *dns.Msg
resp, err = r.queryServers(ctx, given, zone, name, qtype)
if err == nil {
return resp, nil
}
}
if len(withoutAddresses) == 0 {
return nil, err
}
if depth >= maxLookupDepth {
return nil, fmt.Errorf(
"addresses of the nameservers of %s not looked up: %w",
zone, ErrLookupDepthExceeded,
)
}
lookedUp, limitErr := r.resolveNSIPs(ctx, withoutAddresses, depth+1)
if limitErr != nil {
return nil, limitErr
}
if len(lookedUp) == 0 {
return nil, err
}
return r.queryServers(ctx, lookedUp, zone, name, qtype)
}
// resolveNSIPs returns the addresses of every nameserver in nsNames
// whose name resolves, each looked up at depth (see resolveARecord).
// The walk can then go on to the zone's other nameservers when one
// gives no usable reply. When none resolves and the depth limit
// stopped one of the lookups, it returns that lookup's error.
func (r *Resolver) resolveNSIPs(
ctx context.Context,
nsNames []string,
depth int,
) ([]string, error) {
var (
ips []string
limitErr error
)
for _, ns := range nsNames {
resolved, err := r.resolveARecord(ctx, ns, depth)
switch {
case err == nil:
ips = append(ips, resolved...)
case errors.Is(err, ErrLookupDepthExceeded):
limitErr = err
}
}
if len(ips) > 0 {
return ips, nil
}
return nil, limitErr
}
// resolveNSIterative queries for NS records using iterative
@@ -304,6 +486,7 @@ func (r *Resolver) resolveNSIterative(
domain = dns.Fqdn(domain)
servers := rootServerList()
zone := "."
for range maxDelegation {
if checkCtx(ctx) != nil {
@@ -311,13 +494,13 @@ func (r *Resolver) resolveNSIterative(
}
resp, err := r.queryServers(
ctx, servers, domain, dns.TypeNS,
ctx, servers, zone, domain, dns.TypeNS,
)
if err != nil {
return nil, err
}
nsNames := extractNSSet(resp.Answer)
nsNames := nsSetFrom(resp, domain)
if len(nsNames) > 0 {
return nsNames, nil
}
@@ -336,16 +519,20 @@ func (r *Resolver) resolveNSIterative(
}
servers = nextServers
zone = referralZone(resp)
}
return nil, ErrNoNameservers
}
// resolveARecord resolves a hostname to IPv4 addresses using
// iterative resolution through the delegation chain.
// resolveARecord resolves a hostname, a nameserver's name, to IPv4
// addresses using iterative resolution through the delegation chain.
// depth is how many lookups of a nameserver's address are under way,
// this one included: 1 for a lookup that no other lookup started.
func (r *Resolver) resolveARecord(
ctx context.Context,
hostname string,
depth int,
) ([]string, error) {
if checkCtx(ctx) != nil {
return nil, ErrContextCanceled
@@ -353,14 +540,18 @@ func (r *Resolver) resolveARecord(
hostname = dns.Fqdn(hostname)
servers := rootServerList()
zone := "."
var withoutAddresses []string
for range maxDelegation {
if checkCtx(ctx) != nil {
return nil, ErrContextCanceled
}
resp, err := r.queryServers(
ctx, servers, hostname, dns.TypeA,
resp, err := r.queryZone(
ctx, servers, withoutAddresses, zone, hostname, dns.TypeA,
depth,
)
if err != nil {
return nil, fmt.Errorf(
@@ -387,17 +578,8 @@ func (r *Resolver) resolveARecord(
break
}
glue := extractGlue(resp.Extra)
nextServers := glueIPs(authNS, glue)
if len(nextServers) == 0 {
// Resolve NS IPs iteratively — but guard
// against infinite recursion by using only
// already-resolved servers.
break
}
servers = nextServers
servers, withoutAddresses = referralNameservers(resp)
zone = referralZone(resp)
}
return nil, fmt.Errorf(
@@ -407,7 +589,10 @@ func (r *Resolver) resolveARecord(
// FindAuthoritativeNameservers traces the delegation chain from
// root servers to discover all authoritative nameservers for the
// given domain. Walks up the label hierarchy for subdomains.
// given domain, as the delegation from its parent zone's servers lists
// them. For a name that is not a zone apex it tries each
// parent name in turn, so it returns the nameservers of the zone the
// name is in.
func (r *Resolver) FindAuthoritativeNameservers(
ctx context.Context,
domain string,
@@ -434,30 +619,62 @@ func (r *Resolver) FindAuthoritativeNameservers(
return nsNames, nil
}
// The root servers would refuse every parent name too.
if errors.Is(err, ErrIntercepted) {
return nil, err
}
}
return nil, ErrNoNameservers
}
// recordTypes returns the record types a nameserver is asked for when a
// name is checked.
func recordTypes() []uint16 {
return []uint16{
dns.TypeA, dns.TypeAAAA, dns.TypeCNAME,
dns.TypeMX, dns.TypeTXT, dns.TypeSRV,
dns.TypeCAA, dns.TypeNS,
}
}
// addressTypes returns the record types ResolveIPAddresses asks for,
// the only ones it reads.
func addressTypes() []uint16 {
return []uint16{dns.TypeA, dns.TypeAAAA, dns.TypeCNAME}
}
// QueryNameserver queries a specific nameserver for all record
// types and builds a NameserverResponse.
func (r *Resolver) QueryNameserver(
ctx context.Context,
nsHostname string,
hostname string,
) (*NameserverResponse, error) {
return r.queryNameserver(ctx, nsHostname, hostname, recordTypes())
}
// queryNameserver queries a specific nameserver for the record types
// in qtypes and builds a NameserverResponse.
func (r *Resolver) queryNameserver(
ctx context.Context,
nsHostname string,
hostname string,
qtypes []uint16,
) (*NameserverResponse, error) {
if checkCtx(ctx) != nil {
return nil, ErrContextCanceled
}
nsIPs, err := r.resolveARecord(ctx, nsHostname)
nsIPs, err := r.resolveARecord(ctx, nsHostname, 1)
if err != nil {
return nil, fmt.Errorf("resolving NS %s: %w", nsHostname, err)
}
hostname = dns.Fqdn(hostname)
return r.queryAllTypes(ctx, nsHostname, nsIPs[0], hostname)
return r.queryTypes(ctx, nsHostname, nsIPs[0], hostname, qtypes)
}
// QueryNameserverIP queries a nameserver by its IP address directly,
@@ -474,14 +691,15 @@ func (r *Resolver) QueryNameserverIP(
hostname = dns.Fqdn(hostname)
return r.queryAllTypes(ctx, nsHostname, nsIP, hostname)
return r.queryTypes(ctx, nsHostname, nsIP, hostname, recordTypes())
}
func (r *Resolver) queryAllTypes(
func (r *Resolver) queryTypes(
ctx context.Context,
nsHostname string,
nsIP string,
hostname string,
qtypes []uint16,
) (*NameserverResponse, error) {
resp := &NameserverResponse{
Nameserver: nsHostname,
@@ -489,12 +707,6 @@ func (r *Resolver) queryAllTypes(
Status: StatusOK,
}
qtypes := []uint16{
dns.TypeA, dns.TypeAAAA, dns.TypeCNAME,
dns.TypeMX, dns.TypeTXT, dns.TypeSRV,
dns.TypeCAA, dns.TypeNS,
}
state := r.queryEachType(ctx, nsIP, hostname, qtypes, resp)
classifyResponse(resp, state)
@@ -504,7 +716,10 @@ func (r *Resolver) queryAllTypes(
type queryState struct {
gotNXDomain bool
gotSERVFAIL bool
gotRefused bool
gotTimeout bool
gotReferral bool
netErr error
hasRecords bool
}
@@ -542,8 +757,13 @@ func (r *Resolver) querySingleType(
) {
msg, err := r.queryDNS(ctx, nsIP, hostname, qtype)
if err != nil {
if isTimeout(err) {
switch {
case isTimeout(err):
state.gotTimeout = true
case errors.Is(err, ErrRefused):
state.gotRefused = true
default:
state.netErr = err
}
return
@@ -561,9 +781,26 @@ func (r *Resolver) querySingleType(
return
}
// A reply with no answer that lists other nameservers, from a server
// that does not hold the name's zone, is a referral and says nothing
// about the name's records. A server named in the delegation that
// does not hold the zone may send one, as do a parent zone's servers
// when FindAuthoritativeNameservers found no delegation for the
// name's zone and moved on to a parent name.
if !msg.Authoritative && len(msg.Answer) == 0 &&
len(extractNSSet(msg.Ns)) > 0 {
state.gotReferral = true
return
}
collectAnswerRecords(msg, resp, state)
}
// collectAnswerRecords adds the records in msg's answer to resp, each
// value once per record type. For a name with a CNAME, a nameserver
// answers a query of any type with that CNAME, so the same value comes
// in the answer to every type asked for.
func collectAnswerRecords(
msg *dns.Msg,
resp *NameserverResponse,
@@ -576,9 +813,12 @@ func collectAnswerRecords(
}
typeName := dns.TypeToString[rr.Header().Rrtype]
resp.Records[typeName] = append(
resp.Records[typeName], val,
)
if !slices.Contains(resp.Records[typeName], val) {
resp.Records[typeName] = append(
resp.Records[typeName], val,
)
}
state.hasRecords = true
}
}
@@ -603,12 +843,24 @@ func classifyResponse(resp *NameserverResponse, state queryState) {
case state.gotSERVFAIL && !state.hasRecords:
resp.Status = StatusError
resp.Error = "server returned SERVFAIL"
case state.gotRefused && !state.hasRecords:
resp.Status = StatusError
resp.Error = "server returned REFUSED"
case state.netErr != nil && !state.hasRecords:
resp.Status = StatusError
resp.Error = "network error: " + state.netErr.Error()
case state.gotReferral && !state.hasRecords:
resp.Status = StatusError
resp.Error = "server returned a referral"
case !state.hasRecords && !state.gotNXDomain:
resp.Status = StatusNoData
}
}
// extractRecordValue formats a DNS RR value as a string.
// extractRecordValue formats a DNS RR value as a string. DNS names
// are case-insensitive and nameservers may answer in any letter case,
// so names are lower-cased to compare equal. TXT and CAA values keep
// their letter case.
func extractRecordValue(rr dns.RR) string {
switch r := rr.(type) {
case *dns.A:
@@ -616,43 +868,29 @@ func extractRecordValue(rr dns.RR) string {
case *dns.AAAA:
return r.AAAA.String()
case *dns.CNAME:
return r.Target
return strings.ToLower(r.Target)
case *dns.MX:
return fmt.Sprintf("%d %s", r.Preference, r.Mx)
return fmt.Sprintf("%d %s", r.Preference, strings.ToLower(r.Mx))
case *dns.TXT:
return strings.Join(r.Txt, "")
case *dns.SRV:
return fmt.Sprintf(
"%d %d %d %s",
r.Priority, r.Weight, r.Port, r.Target,
r.Priority, r.Weight, r.Port, strings.ToLower(r.Target),
)
case *dns.CAA:
return fmt.Sprintf(
"%d %s \"%s\"", r.Flag, r.Tag, r.Value,
)
case *dns.NS:
return r.Ns
return strings.ToLower(r.Ns)
default:
return ""
}
}
// parentDomain returns the registerable parent domain.
func parentDomain(hostname string) string {
hostname = dns.Fqdn(strings.ToLower(hostname))
labels := dns.SplitDomainName(hostname)
if len(labels) <= minDomainLabels {
return strings.Join(labels, ".") + "."
}
return strings.Join(
labels[len(labels)-minDomainLabels:], ".",
) + "."
}
// QueryAllNameservers discovers auth NSes for the hostname's
// parent domain, then queries each one independently.
// QueryAllNameservers discovers the auth NSes of the zone the
// hostname is in, then queries each one independently.
func (r *Resolver) QueryAllNameservers(
ctx context.Context,
hostname string,
@@ -661,29 +899,31 @@ func (r *Resolver) QueryAllNameservers(
return nil, ErrContextCanceled
}
parent := parentDomain(hostname)
nameservers, err := r.FindAuthoritativeNameservers(ctx, parent)
nameservers, err := r.FindAuthoritativeNameservers(ctx, hostname)
if err != nil {
return nil, err
}
return r.queryEachNS(ctx, nameservers, hostname)
return r.queryEachNS(ctx, nameservers, hostname, recordTypes())
}
func (r *Resolver) queryEachNS(
ctx context.Context,
nameservers []string,
hostname string,
qtypes []uint16,
) (map[string]*NameserverResponse, error) {
results := make(map[string]*NameserverResponse)
for _, ns := range nameservers {
resp, err := r.queryNameserver(ctx, ns, hostname, qtypes)
// A query the context cut short says nothing about the
// nameserver, so it must not be returned as its failure.
if checkCtx(ctx) != nil {
return nil, ErrContextCanceled
}
resp, err := r.QueryNameserver(ctx, ns, hostname)
if err != nil {
results[ns] = &NameserverResponse{
Nameserver: ns,
@@ -711,25 +951,20 @@ func (r *Resolver) LookupNS(
// LookupAllRecords performs iterative resolution to find all DNS
// records for the given hostname, keyed by authoritative nameserver.
// Each nameserver's response carries its status and error with its
// records.
func (r *Resolver) LookupAllRecords(
ctx context.Context,
hostname string,
) (map[string]map[string][]string, error) {
results, err := r.QueryAllNameservers(ctx, hostname)
if err != nil {
return nil, err
}
out := make(map[string]map[string][]string, len(results))
for ns, resp := range results {
out[ns] = resp.Records
}
return out, nil
) (map[string]*NameserverResponse, error) {
return r.QueryAllNameservers(ctx, hostname)
}
// ResolveIPAddresses resolves a hostname to all IPv4 and IPv6
// addresses, following CNAME chains up to MaxCNAMEDepth.
// addresses, following CNAME chains up to MaxCNAMEDepth. It asks each
// nameserver of the name's zone for its A, AAAA and CNAME records only.
// When no nameserver of the name's zone answered, it returns an error
// rather than no addresses.
func (r *Resolver) ResolveIPAddresses(
ctx context.Context,
hostname string,
@@ -750,12 +985,20 @@ func (r *Resolver) resolveIPWithCNAME(
return nil, ErrCNAMEDepthExceeded
}
results, err := r.QueryAllNameservers(ctx, hostname)
nameservers, err := r.FindAuthoritativeNameservers(ctx, hostname)
if err != nil {
return nil, err
}
ips, cnameTarget := collectIPs(results)
results, err := r.queryEachNS(ctx, nameservers, hostname, addressTypes())
if err != nil {
return nil, err
}
ips, cnameTarget, err := collectIPs(results)
if err != nil {
return nil, fmt.Errorf("resolving %s: %w", hostname, err)
}
if len(ips) == 0 && cnameTarget != "" {
return r.resolveIPWithCNAME(ctx, cnameTarget, depth+1)
@@ -766,16 +1009,28 @@ func (r *Resolver) resolveIPWithCNAME(
return ips, nil
}
// collectIPs returns the addresses in the nameservers' answers and the
// first CNAME target among them. It returns ErrNoNameserverAnswered when
// every nameserver timed out, failed or returned a referral: that is not
// a name with no addresses.
func collectIPs(
results map[string]*NameserverResponse,
) ([]string, string) {
) ([]string, string, error) {
seen := make(map[string]bool)
var ips []string
var cnameTarget string
answered := false
for _, resp := range results {
if resp.Status == StatusTimeout || resp.Status == StatusError {
continue
}
answered = true
if resp.Status == StatusNXDomain {
continue
}
@@ -799,5 +1054,9 @@ func collectIPs(
}
}
return ips, cnameTarget
if !answered {
return nil, "", ErrNoNameserverAnswered
}
return ips, cnameTarget, nil
}
+298
View File
@@ -0,0 +1,298 @@
package resolver_test
import (
"math/rand/v2"
"slices"
"testing"
"github.com/miekg/dns"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
"sneak.berlin/go/dnswatcher/internal/resolver"
)
// TestCollectIPs_OneAnswerIsEnough checks that one nameserver answering
// NXDOMAIN says the name has no addresses, though the other timed out.
func TestCollectIPs_OneAnswerIsEnough(t *testing.T) {
t.Parallel()
ips, _, err := resolver.CollectIPs(
map[string]*resolver.NameserverResponse{
"ns1.example.": {Status: resolver.StatusTimeout},
"ns2.example.": {Status: resolver.StatusNXDomain},
},
)
require.NoError(t, err)
assert.Empty(t, ips)
}
// TestCollectIPs_FailedIsNoAnswer checks that nameservers that all have
// status error, from a refusal, a server failure, a network error or a
// referral, are no answer rather than a name with no addresses.
func TestCollectIPs_FailedIsNoAnswer(t *testing.T) {
t.Parallel()
ips, _, err := resolver.CollectIPs(
map[string]*resolver.NameserverResponse{
"ns1.example.": {Status: resolver.StatusError},
"ns2.example.": {Status: resolver.StatusError},
},
)
require.ErrorIs(t, err, resolver.ErrNoNameserverAnswered)
assert.Empty(t, ips)
}
const (
// exampleCom is the zone most cases of TestUsableReply and
// TestNSSetFrom are about, and wwwExampleCom a name in it.
exampleCom = "example.com."
wwwExampleCom = "www.example.com."
// exampleNS is the server the NS records nsRecord builds name.
exampleNS = "ns1.example.net."
)
// nsRecord builds an NS record that names a server of zone.
func nsRecord(zone string) *dns.NS {
return &dns.NS{
Hdr: dns.RR_Header{
Name: zone, Rrtype: dns.TypeNS, Class: dns.ClassINET,
},
Ns: exampleNS,
}
}
// referralTo builds a reply that refers the query to the servers of
// zone.
func referralTo(zone string) *dns.Msg {
msg := new(dns.Msg)
msg.Ns = []dns.RR{nsRecord(zone)}
return msg
}
// TestUsableReply checks which replies from one of a zone's servers are
// used. A reply that is not usable moves the query on to the zone's
// next server.
func TestUsableReply(t *testing.T) {
t.Parallel()
servfail := new(dns.Msg)
servfail.Rcode = dns.RcodeServerFailure
answer := new(dns.Msg)
answer.Authoritative = true
answer.Answer = []dns.RR{nsRecord(exampleCom)}
nxdomain := new(dns.Msg)
nxdomain.Authoritative = true
nxdomain.Rcode = dns.RcodeNameError
tests := []struct {
name string
resp *dns.Msg
zone string
query string
want bool
}{
{
name: "SERVFAIL", resp: servfail,
zone: exampleCom, query: exampleCom, want: false,
},
{
name: "answer", resp: answer,
zone: exampleCom, query: exampleCom, want: true,
},
{
name: "NXDOMAIN", resp: nxdomain,
zone: ".", query: exampleCom, want: true,
},
{
name: "root refers to com", resp: referralTo("com."),
zone: ".", query: exampleCom, want: true,
},
{
name: "com refers to example.com", resp: referralTo(exampleCom),
zone: "com.", query: wwwExampleCom, want: true,
},
{
name: "referral back to the zone", resp: referralTo(exampleCom),
zone: exampleCom, query: exampleCom, want: false,
},
{
name: "referral up to the root", resp: referralTo("."),
zone: exampleCom, query: exampleCom, want: false,
},
{
name: "referral sideways", resp: referralTo("net."),
zone: ".", query: exampleCom, want: false,
},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
t.Parallel()
assert.Equal(t, tt.want,
resolver.UsableReply(tt.resp, tt.zone, tt.query),
)
})
}
}
// TestNSSetFrom checks which NS set a reply gives for a domain; a set
// that is not empty ends the walk. The referral to example.com that
// com's servers all send alike gives its delegation, so the set is the
// same whichever of them answered, and example.com's own servers, which
// can disagree, are not asked.
func TestNSSetFrom(t *testing.T) {
t.Parallel()
answer := new(dns.Msg)
answer.Authoritative = true
answer.Answer = []dns.RR{nsRecord(exampleCom)}
tests := []struct {
name string
resp *dns.Msg
domain string
want []string
}{
{
name: "com refers to example.com", resp: referralTo(exampleCom),
domain: exampleCom, want: []string{exampleNS},
},
{
name: "com refers on, for www.example.com",
resp: referralTo(exampleCom), domain: wwwExampleCom,
want: nil,
},
{
name: "answer from a server that holds example.com",
resp: answer, domain: exampleCom,
want: []string{exampleNS},
},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
t.Parallel()
assert.ElementsMatch(t, tt.want,
resolver.NSSetFrom(tt.resp, tt.domain),
)
})
}
}
func TestExtractRecordValue_LetterCase(t *testing.T) {
t.Parallel()
tests := []struct {
name string
rr dns.RR
want string
}{
{
name: "MX target lower-cased",
rr: &dns.MX{Preference: 1, Mx: "ASPMX.L.GOOGLE.COM."},
want: "1 aspmx.l.google.com.",
},
{
name: "NS target lower-cased",
rr: &dns.NS{Ns: "x.ns.joker.COM."},
want: "x.ns.joker.com.",
},
{
name: "CNAME target lower-cased",
rr: &dns.CNAME{Target: "WWW.Example.Com."},
want: "www.example.com.",
},
{
name: "SRV target lower-cased",
rr: &dns.SRV{
Priority: 10, Weight: 5, Port: 443,
Target: "SIP.Example.Com.",
},
want: "10 5 443 sip.example.com.",
},
{
name: "TXT value keeps its case",
rr: &dns.TXT{Txt: []string{"Verify=AbC123"}},
want: "Verify=AbC123",
},
{
name: "CAA value keeps its case",
rr: &dns.CAA{Flag: 0, Tag: "issue", Value: "LetsEncrypt.org"},
want: `0 issue "LetsEncrypt.org"`,
},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
t.Parallel()
assert.Equal(t, tt.want, resolver.ExtractRecordValue(tt.rr))
})
}
}
// TestCollectAnswerRecords_CNAMEOnce collects the answers a nameserver
// gives for a name with a CNAME, one for each record type a check asks
// for. Each answer holds the CNAME, which must be stored once.
func TestCollectAnswerRecords_CNAMEOnce(t *testing.T) {
t.Parallel()
cname := &dns.CNAME{
Hdr: dns.RR_Header{
Name: "git.eeqj.de.", Rrtype: dns.TypeCNAME, Class: dns.ClassINET,
},
Target: "fsn1app1.datavi.be.",
}
resp := &resolver.NameserverResponse{Records: map[string][]string{}}
for _, qtype := range []uint16{
dns.TypeA, dns.TypeAAAA, dns.TypeCNAME, dns.TypeMX,
dns.TypeTXT, dns.TypeSRV, dns.TypeCAA, dns.TypeNS,
} {
msg := new(dns.Msg)
msg.SetQuestion("git.eeqj.de.", qtype)
msg.Answer = []dns.RR{cname}
resolver.CollectAnswerRecords(msg, resp)
}
assert.Equal(t,
map[string][]string{"CNAME": {"fsn1app1.datavi.be."}},
resp.Records,
)
}
// TestShuffled shuffles the root servers with many seeds. Every order
// must hold each root server once, so each is tried before a
// resolution fails; each root server must come first for some seed, so
// no one root server gets every first query; and the list passed in
// must be left as it was.
func TestShuffled(t *testing.T) {
t.Parallel()
const seeds = 1000
roots := resolver.RootServerList()
before := slices.Clone(roots)
first := make(map[string]bool)
for seed := range uint64(seeds) {
rng := rand.New(rand.NewPCG(seed, 0)) //nolint:gosec // seeded on purpose
order := resolver.Shuffled(roots, rng.Shuffle)
assert.ElementsMatch(t, roots, order)
first[order[0]] = true
}
assert.Len(t, first, len(roots))
assert.Equal(t, before, roots)
}
+192
View File
@@ -0,0 +1,192 @@
package resolver_test
import (
"testing"
"github.com/stretchr/testify/assert"
"sneak.berlin/go/dnswatcher/internal/resolver"
)
// Tests for the live-DNS harness in livedns_test.go itself. These
// exercise pure logic; they perform no DNS resolution of any kind, so
// they neither mock DNS nor depend on it.
// Names for the synthetic status maps below. Nothing is ever queried
// at them: they are map keys handed to the package's pure counting
// helpers, not a stand-in for a nameserver.
const (
nsExample1 = "ns1.example."
nsExample2 = "ns2.example."
nsExample3 = "ns3.example."
nsExample4 = "ns4.example."
)
func TestLiveQuorumIsStrictMajority(t *testing.T) {
t.Parallel()
cases := map[int]int{
0: 1,
1: 1,
2: 2,
3: 2,
4: 3,
5: 3,
13: 7,
}
for total, want := range cases {
assert.Equal(
t, want, liveQuorum(total),
"liveQuorum(%d)", total,
)
}
}
func TestStatusCountingIgnoresSilentNameservers(t *testing.T) {
t.Parallel()
results := map[string]*resolver.NameserverResponse{
nsExample1: {
Nameserver: nsExample1,
Status: resolver.StatusOK,
},
nsExample2: {
Nameserver: nsExample2,
Status: resolver.StatusOK,
},
nsExample3: {
Nameserver: nsExample3,
Status: resolver.StatusTimeout,
},
nsExample4: {
Nameserver: nsExample4,
Status: resolver.StatusError,
},
}
assert.Equal(
t, 2, countStatus(results, resolver.StatusOK),
)
assert.Equal(
t, 0, countStatus(results, resolver.StatusNXDomain),
)
// Two of four answered, which is short of the quorum of
// three: this is the state that triggers a retry rather
// than an assertion failure.
assert.Equal(t, 2, answeredCount(results))
assert.Less(t, answeredCount(results), liveQuorum(len(results)))
assert.Equal(
t,
"ns1.example.=ok ns2.example.=ok "+
"ns3.example.=timeout ns4.example.=error",
describeStatuses(results),
)
}
// TestUnsanctionedStatusesRejectsWrongAnswers is the regression test
// for the defect this allowlist exists to prevent: a minority of
// nameservers answering WRONGLY while quorum keeps the suite green.
// nodata is the case that motivated it — it is a wrong answer, not
// silence, and it was previously banned by neither test.
func TestUnsanctionedStatusesRejectsWrongAnswers(t *testing.T) {
t.Parallel()
// Four nameservers, three OK and one answering nodata: a
// quorum of three is satisfied and no NXDOMAIN is present, so
// the old blocklist assertions both passed on this input.
results := map[string]*resolver.NameserverResponse{
nsExample1: {
Nameserver: nsExample1,
Status: resolver.StatusOK,
},
nsExample2: {
Nameserver: nsExample2,
Status: resolver.StatusOK,
},
nsExample3: {
Nameserver: nsExample3,
Status: resolver.StatusOK,
},
nsExample4: {
Nameserver: nsExample4,
Status: resolver.StatusNoData,
},
}
assert.GreaterOrEqual(
t,
countStatus(results, resolver.StatusOK),
liveQuorum(len(results)),
)
assert.Zero(t, countStatus(results, resolver.StatusNXDomain))
// nodata is an ANSWER, so it never triggers a retry: nothing
// but the allowlist stands between it and a false green.
assert.Equal(t, len(results), answeredCount(results))
assert.Equal(
t,
[]string{nsExample4 + "=nodata"},
unsanctionedStatuses(
results,
resolver.StatusOK,
resolver.StatusTimeout,
resolver.StatusError,
),
"nodata must be reported as an unsanctioned status",
)
}
func TestUnsanctionedStatusesToleratesSilenceOnly(t *testing.T) {
t.Parallel()
results := map[string]*resolver.NameserverResponse{
nsExample1: {
Nameserver: nsExample1,
Status: resolver.StatusNXDomain,
},
nsExample2: {
Nameserver: nsExample2,
Status: resolver.StatusTimeout,
},
nsExample3: {
Nameserver: nsExample3,
Status: resolver.StatusError,
},
}
allowed := []string{
resolver.StatusNXDomain,
resolver.StatusTimeout,
resolver.StatusError,
}
assert.Empty(
t,
unsanctionedStatuses(results, allowed...),
"timeout and error are non-answers and are tolerated",
)
// The same silent nameservers do not count towards a quorum.
assert.Equal(t, 1, answeredCount(results))
// An unknown status is treated as silence by answeredCount —
// so it retries and fails loudly — and is unsanctioned by the
// allowlist rather than quietly permitted.
const laterStatus = "some-status-added-later"
results[nsExample4] = &resolver.NameserverResponse{
Nameserver: nsExample4,
Status: laterStatus,
}
assert.Equal(t, 1, answeredCount(results))
assert.Equal(
t,
[]string{nsExample4 + "=" + laterStatus},
unsanctionedStatuses(results, allowed...),
)
}
+437
View File
@@ -0,0 +1,437 @@
package resolver_test
import (
"context"
"errors"
"fmt"
"slices"
"sort"
"strings"
"testing"
"sneak.berlin/go/dnswatcher/internal/livednstest"
"sneak.berlin/go/dnswatcher/internal/resolver"
)
// ----------------------------------------------------------------
// Live DNS test support
// ----------------------------------------------------------------
//
// Tests that look something up in DNS query live DNS servers, never a
// stand-in; logic that works on record data may be tested on that
// data with no lookup (see TESTING.md). Each live operation below goes
// through livednstest.Retry, which bounds how many resolutions are in
// flight at once and retries an operation that got no answer (see
// package livednstest).
//
// Where an assertion spans several independent nameservers, a quorum
// is enough: a strict majority answering as expected. A server that
// fails to answer is tolerated, while a server that answers *wrongly*
// still fails the test.
//
// That tolerance is expressed as an ALLOWLIST of sanctioned statuses,
// never as a blocklist of known-bad ones. A blocklist bans the one
// wrong answer its author thought of and silently admits every other
// status, including any added to the resolver later; an allowlist
// fails on anything nobody explicitly sanctioned. Silence (timeout,
// error) is the only thing quorum exists to tolerate. A *wrong
// answer* — nxdomain for a name that exists, ok for one that does
// not, nodata for either — is never tolerated at any count.
// minNameservers is the smallest nameserver count a well-run zone is
// expected to publish.
const minNameservers = 2
// errLiveNoQuorum reports that too few of a domain's nameservers
// answered for a quorum assertion to be made.
var errLiveNoQuorum = errors.New("no nameserver quorum")
// liveQuorum is how many of total nameservers must agree for a
// multi-nameserver assertion to hold: a strict majority.
func liveQuorum(total int) int {
if total < 1 {
return 1
}
return total/2 + 1
}
// countStatus counts the responses carrying the given status.
func countStatus(
results map[string]*resolver.NameserverResponse,
status string,
) int {
n := 0
for _, resp := range results {
if resp.Status == status {
n++
}
}
return n
}
// liveAnswerStatuses is the closed set of statuses that count as a
// nameserver having ANSWERED at all, whether or not the test agrees
// with the answer. It is deliberately an allowlist: a status added
// to the resolver later is treated as silence, so it can only ever
// cause a retry and then a loud failure, never a quiet pass.
func liveAnswerStatuses() []string {
return []string{
resolver.StatusOK,
resolver.StatusNXDomain,
resolver.StatusNoData,
}
}
// answeredCount counts the nameservers that produced an answer of
// any kind, as opposed to failing or timing out.
func answeredCount(
results map[string]*resolver.NameserverResponse,
) int {
answers := liveAnswerStatuses()
n := 0
for _, resp := range results {
if slices.Contains(answers, resp.Status) {
n++
}
}
return n
}
// unsanctionedStatuses returns "nameserver=status" for every result
// whose status the caller did not explicitly sanction, sorted for a
// stable failure message. Callers pass the full closed set they will
// accept — the expected answer plus whichever non-answers (timeout,
// error) quorum is allowed to tolerate — so that any status outside
// it fails the test by name.
func unsanctionedStatuses(
results map[string]*resolver.NameserverResponse,
allowed ...string,
) []string {
offenders := make([]string, 0, len(results))
for ns, resp := range results {
if slices.Contains(allowed, resp.Status) {
continue
}
offenders = append(
offenders, fmt.Sprintf("%s=%s", ns, resp.Status),
)
}
sort.Strings(offenders)
return offenders
}
// describeStatuses renders per-nameserver statuses for use in
// assertion failure messages.
func describeStatuses(
results map[string]*resolver.NameserverResponse,
) string {
parts := make([]string, 0, len(results))
for ns, resp := range results {
parts = append(
parts, fmt.Sprintf("%s=%s", ns, resp.Status),
)
}
sort.Strings(parts)
return strings.Join(parts, " ")
}
// ----------------------------------------------------------------
// Live operation wrappers
// ----------------------------------------------------------------
// liveFindAuthoritative resolves a domain's authoritative
// nameservers, retrying until the delegation chain can be walked.
func liveFindAuthoritative(
t *testing.T,
r *resolver.Resolver,
domain string,
) []string {
t.Helper()
var out []string
livednstest.Retry(
t,
"FindAuthoritativeNameservers("+domain+")",
func(ctx context.Context) error {
ns, err := r.FindAuthoritativeNameservers(ctx, domain)
if err != nil {
return err
}
if len(ns) == 0 {
return fmt.Errorf(
"%w: %s has no nameservers",
livednstest.ErrNoAnswer, domain,
)
}
out = ns
return nil
},
)
return out
}
// liveLookupNS is liveFindAuthoritative through the LookupNS entry
// point, so that both entry points stay independently exercised.
func liveLookupNS(
t *testing.T,
r *resolver.Resolver,
domain string,
) []string {
t.Helper()
var out []string
livednstest.Retry(
t,
"LookupNS("+domain+")",
func(ctx context.Context) error {
ns, err := r.LookupNS(ctx, domain)
if err != nil {
return err
}
if len(ns) == 0 {
return fmt.Errorf(
"%w: %s has no nameservers",
livednstest.ErrNoAnswer, domain,
)
}
out = ns
return nil
},
)
return out
}
// liveQueryNameserver queries one nameserver, retrying while that
// nameserver fails to answer. NXDOMAIN and NODATA are answers and
// are returned to the caller to assert on.
//
// QueryNameserver sends one query per record type, so one lost query
// leaves its type out of an answer that is otherwise fine. A test names
// in types the record types it reads; an answer holding records of none
// of them is retried too.
func liveQueryNameserver(
t *testing.T,
r *resolver.Resolver,
nameserver string,
hostname string,
types ...string,
) *resolver.NameserverResponse {
t.Helper()
what := fmt.Sprintf(
"QueryNameserver(%s, %s)", nameserver, hostname,
)
var out *resolver.NameserverResponse
livednstest.Retry(
t,
what,
func(ctx context.Context) error {
resp, err := r.QueryNameserver(
ctx, nameserver, hostname,
)
if err != nil {
return err
}
if resp.Status == resolver.StatusTimeout ||
resp.Status == resolver.StatusError {
return fmt.Errorf(
"%w: %s returned %s: %s",
livednstest.ErrNoAnswer, nameserver,
resp.Status, resp.Error,
)
}
hasRecords := func(recordType string) bool {
return len(resp.Records[recordType]) > 0
}
if len(types) > 0 && !slices.ContainsFunc(types, hasRecords) {
return fmt.Errorf(
"%w: %s returned no %s records",
livednstest.ErrNoAnswer, nameserver,
strings.Join(types, " or "),
)
}
out = resp
return nil
},
)
return out
}
// liveQueryAllNameservers queries every authoritative nameserver
// for a hostname, retrying until a quorum of them has answered.
// Individual nameservers that stay silent are left in the result
// for the caller to account for.
func liveQueryAllNameservers(
t *testing.T,
r *resolver.Resolver,
hostname string,
) map[string]*resolver.NameserverResponse {
t.Helper()
var out map[string]*resolver.NameserverResponse
livednstest.Retry(
t,
"QueryAllNameservers("+hostname+")",
func(ctx context.Context) error {
results, err := r.QueryAllNameservers(ctx, hostname)
if err != nil {
return err
}
if len(results) == 0 {
return fmt.Errorf(
"%w: no nameservers queried for %s",
livednstest.ErrNoAnswer, hostname,
)
}
answered := answeredCount(results)
if answered < liveQuorum(len(results)) {
return fmt.Errorf(
"%w: %d of %d answered: %s",
errLiveNoQuorum, answered,
len(results), describeStatuses(results),
)
}
out = results
return nil
},
)
return out
}
// liveResolveIPs resolves a hostname that is expected to have
// addresses, retrying until at least one is returned.
func liveResolveIPs(
t *testing.T,
r *resolver.Resolver,
hostname string,
) []string {
t.Helper()
var out []string
livednstest.Retry(
t,
"ResolveIPAddresses("+hostname+")",
func(ctx context.Context) error {
ips, err := r.ResolveIPAddresses(ctx, hostname)
if err != nil {
return err
}
if len(ips) == 0 {
return fmt.Errorf(
"%w: no addresses for %s",
livednstest.ErrNoAnswer, hostname,
)
}
out = ips
return nil
},
)
return out
}
// liveResolveIPsAllowingEmpty resolves a hostname that may legitimately
// have no addresses, so the empty result is returned rather than
// retried. Used for names that must not exist; the corresponding
// QueryAllNameservers test is what proves the nameservers actively
// said NXDOMAIN rather than merely staying silent.
func liveResolveIPsAllowingEmpty(
t *testing.T,
r *resolver.Resolver,
hostname string,
) []string {
t.Helper()
var out []string
livednstest.Retry(
t,
"ResolveIPAddresses("+hostname+")",
func(ctx context.Context) error {
ips, err := r.ResolveIPAddresses(ctx, hostname)
if err != nil {
return err
}
out = ips
return nil
},
)
return out
}
// liveResolveNSIPs looks up the addresses of the nameservers named
// names, retrying until there are at least atLeast of them: a name
// whose lookup got no reply is left out of the result, not an error.
func liveResolveNSIPs(
t *testing.T,
r *resolver.Resolver,
names []string,
atLeast int,
) []string {
t.Helper()
var out []string
livednstest.Retry(
t,
"ResolveNSIPs("+strings.Join(names, ", ")+")",
func(ctx context.Context) error {
ips := r.ResolveNSIPs(ctx, names)
if len(ips) < atLeast {
return fmt.Errorf(
"%w: %d addresses, expected at least %d",
livednstest.ErrNoAnswer, len(ips), atLeast,
)
}
out = ips
return nil
},
)
return out
}
-13
View File
@@ -67,17 +67,4 @@ func NewFromLogger(log *slog.Logger) *Resolver {
}
}
// NewFromLoggerWithClient creates a Resolver with a custom DNS
// client, useful for testing with mock DNS responses.
func NewFromLoggerWithClient(
log *slog.Logger,
client DNSClient,
) *Resolver {
return &Resolver{
log: log,
client: client,
tcp: client,
}
}
// Method implementations are in iterative.go.
File diff suppressed because it is too large Load Diff
+38
View File
@@ -0,0 +1,38 @@
package server
import (
"net/http"
"time"
"github.com/go-chi/chi/v5"
)
// RequestTimeout exports the handler execution budget applied by
// chimw.Timeout in SetupRoutes, so tests can assert the relationship
// between it and the server's WriteTimeout.
const RequestTimeout time.Duration = requestTimeout
// SetListenPort overrides the port Run binds. A test uses it to hand
// Run an unbindable port so ListenAndServe fails immediately and Run
// returns after storing its http.Server.
func SetListenPort(s *Server, port int) {
s.port = port
}
// HTTPServerOf returns the http.Server that Run built and stored, so a
// test can inspect the timeouts the running server actually carries.
func HTTPServerOf(s *Server) *http.Server {
return s.httpServer
}
// EnableSentry runs the Sentry setup that the start hook runs, without
// starting the HTTP server.
func EnableSentry(s *Server) error {
return s.enableSentry()
}
// RouterOf returns the router SetupRoutes built, so a test can add a
// route that panics.
func RouterOf(s *Server) *chi.Mux {
return s.router
}
+38 -14
View File
@@ -4,6 +4,7 @@ import (
"net/http"
"time"
sentryhttp "github.com/getsentry/sentry-go/http"
"github.com/go-chi/chi/v5"
chimw "github.com/go-chi/chi/v5/middleware"
"github.com/prometheus/client_golang/prometheus/promhttp"
@@ -21,15 +22,32 @@ func (s *Server) SetupRoutes() {
// Global middleware
s.router.Use(chimw.Recoverer)
s.router.Use(chimw.RequestID)
s.router.Use(s.mw.SecurityHeaders())
s.router.Use(s.mw.Logging())
s.router.Use(s.mw.CORS())
s.router.Use(chimw.Timeout(requestTimeout))
// Report panics in handlers to Sentry when DNSWATCHER_SENTRY_DSN is
// set. Repanic passes each panic on to chimw.Recoverer above, which
// still answers the request.
if s.sentryEnabled {
sentryHandler := sentryhttp.New(sentryhttp.Options{
Repanic: true,
})
s.router.Use(sentryHandler.Handle)
}
// Public, unauthenticated, read-only routes, the only ones
// REPO_POLICIES.md allows wildcard CORS on. CORS is middleware of
// this whole router, not of a Group, so that it also answers
// OPTIONS preflight requests, which no route here registers.
public := chi.NewRouter()
public.Use(s.mw.CORS())
// Dashboard (read-only web UI)
s.router.Get("/", s.handlers.HandleDashboard())
public.Get("/", s.handlers.HandleDashboard())
// Static assets (embedded CSS/JS)
s.router.Mount(
public.Mount(
"/s",
http.StripPrefix(
"/s",
@@ -38,27 +56,33 @@ func (s *Server) SetupRoutes() {
)
// Health check (standard well-known path)
s.router.Get(
public.Get(
"/.well-known/healthcheck",
s.handlers.HandleHealthCheck(),
)
// Legacy health check (keep for backward compatibility)
s.router.Get("/health", s.handlers.HandleHealthCheck())
public.Get("/health", s.handlers.HandleHealthCheck())
// API v1 routes
s.router.Route("/api/v1", func(r chi.Router) {
public.Route("/api/v1", func(r chi.Router) {
r.Get("/status", s.handlers.HandleStatus())
})
// Metrics endpoint (optional, with basic auth)
s.router.Mount("/", public)
// Metrics endpoint (optional, with basic auth) and no CORS: a
// Prometheus scraper is not a browser. It is mounted rather than
// added with Get so that every method on /metrics, OPTIONS
// included, ends here instead of falling through to the public
// router and its CORS. The rate limit comes before Basic Auth, so
// failed logins count against it and a request over the limit
// never reaches the password check.
if s.params.Config.MetricsUsername != "" {
s.router.Group(func(r chi.Router) {
r.Use(s.mw.MetricsAuth())
r.Get(
"/metrics",
promhttp.Handler().ServeHTTP,
)
})
metrics := chi.NewRouter()
metrics.Use(s.mw.MetricsRateLimit())
metrics.Use(s.mw.MetricsAuth())
metrics.Get("/", promhttp.Handler().ServeHTTP)
s.router.Mount("/metrics", metrics)
}
}
+290
View File
@@ -0,0 +1,290 @@
package server_test
import (
"net/http"
"net/http/httptest"
"testing"
"github.com/spf13/viper"
"sneak.berlin/go/dnswatcher/internal/server"
)
// Credentials for /metrics, which is only routed when a username is set.
const (
metricsUsername = "scraper"
metricsPassword = "scrape-secret"
)
// The tests below set env vars and touch viper global state, so like
// the config tests they cannot use t.Parallel.
// routedServer builds the server with its routes set up, ready to serve
// test requests. The caller must first configure viper.
func routedServer(t *testing.T) *server.Server {
t.Helper()
srv := buildServer(t)
srv.SetupRoutes()
return srv
}
// crossOriginRequest builds a request as a browser sends it from a page
// on another site.
func crossOriginRequest(
t *testing.T,
method string,
target string,
) *http.Request {
t.Helper()
req := httptest.NewRequestWithContext(t.Context(), method, target, nil)
req.Header.Set("Origin", "https://example.net")
return req
}
// preflightRequest builds the OPTIONS request a browser sends before a
// cross-origin request with the given method and request headers.
func preflightRequest(
t *testing.T,
target string,
method string,
headers string,
) *http.Request {
t.Helper()
req := crossOriginRequest(t, http.MethodOptions, target)
req.Header.Set("Access-Control-Request-Method", method)
if headers != "" {
req.Header.Set("Access-Control-Request-Headers", headers)
}
return req
}
func serve(
srv *server.Server,
req *http.Request,
) *httptest.ResponseRecorder {
rec := httptest.NewRecorder()
srv.ServeHTTP(rec, req)
return rec
}
// publicPaths returns one path on each public route.
func publicPaths() []string {
return []string{
"/",
"/s/css/tailwind.min.css",
"/api/v1/status",
"/health",
"/.well-known/healthcheck",
}
}
// TestPublicRoutesAllowAnyOrigin checks that every public route answers
// a cross-origin GET with the CORS wildcard.
func TestPublicRoutesAllowAnyOrigin(t *testing.T) {
viper.Reset()
t.Setenv("DNSWATCHER_TARGETS", "example.com")
t.Setenv("DNSWATCHER_METRICS_USERNAME", metricsUsername)
t.Setenv("DNSWATCHER_METRICS_PASSWORD", metricsPassword)
srv := routedServer(t)
for _, path := range publicPaths() {
rec := serve(srv, crossOriginRequest(t, http.MethodGet, path))
if rec.Code != http.StatusOK {
t.Errorf("GET %s: status = %d, want 200", path, rec.Code)
}
got := rec.Header().Get("Access-Control-Allow-Origin")
if got != "*" {
t.Errorf(
"GET %s: Access-Control-Allow-Origin = %q, want %q",
path, got, "*",
)
}
}
}
// TestMetricsHasNoCORS checks that no request to the Basic-Auth
// protected /metrics, preflight included, gets a CORS header.
func TestMetricsHasNoCORS(t *testing.T) {
viper.Reset()
t.Setenv("DNSWATCHER_TARGETS", "example.com")
t.Setenv("DNSWATCHER_METRICS_USERNAME", metricsUsername)
t.Setenv("DNSWATCHER_METRICS_PASSWORD", metricsPassword)
srv := routedServer(t)
authenticated := crossOriginRequest(t, http.MethodGet, "/metrics")
authenticated.SetBasicAuth(metricsUsername, metricsPassword)
tests := []struct {
name string
req *http.Request
wantStatus int
}{
{
"authenticated GET",
authenticated,
http.StatusOK,
},
{
"unauthenticated GET",
crossOriginRequest(t, http.MethodGet, "/metrics"),
http.StatusUnauthorized,
},
{
"preflight",
preflightRequest(t, "/metrics", http.MethodGet, ""),
http.StatusUnauthorized,
},
}
for _, tt := range tests {
rec := serve(srv, tt.req)
if rec.Code != tt.wantStatus {
t.Errorf(
"%s: status = %d, want %d",
tt.name, rec.Code, tt.wantStatus,
)
}
got := rec.Header().Get("Access-Control-Allow-Origin")
if got != "" {
t.Errorf(
"%s: Access-Control-Allow-Origin = %q, want none",
tt.name, got,
)
}
}
}
// TestPreflightAllowsOnlyWhatPublicRoutesServe checks what each public
// route agrees to in a CORS preflight: GET, but not POST, PUT or
// DELETE, which no route serves, and not the Authorization or
// X-CSRF-Token headers, which no public route reads. It checks every
// public route because one added with Get, such as /health, answers a
// preflight only while CORS is middleware of a whole router; in a
// Group, chi would answer it with 405 and no CORS headers.
func TestPreflightAllowsOnlyWhatPublicRoutesServe(t *testing.T) {
viper.Reset()
t.Setenv("DNSWATCHER_TARGETS", "example.com")
t.Setenv("DNSWATCHER_METRICS_USERNAME", metricsUsername)
t.Setenv("DNSWATCHER_METRICS_PASSWORD", metricsPassword)
srv := routedServer(t)
tests := []struct {
method string
headers string
allowed bool
}{
{http.MethodGet, "", true},
{http.MethodGet, "Content-Type", true},
{http.MethodPost, "", false},
{http.MethodPut, "", false},
{http.MethodDelete, "", false},
{http.MethodGet, "Authorization", false},
{http.MethodGet, "X-CSRF-Token", false},
}
for _, path := range publicPaths() {
for _, tt := range tests {
rec := serve(srv, preflightRequest(
t, path, tt.method, tt.headers,
))
want := ""
if tt.allowed {
want = tt.method
}
got := rec.Header().Get("Access-Control-Allow-Methods")
if got != want {
t.Errorf(
"preflight to %s for %s with headers %q: "+
"Access-Control-Allow-Methods = %q, want %q",
path, tt.method, tt.headers, got, want,
)
}
}
}
}
// metricsRequest builds a GET for /metrics from remoteAddr that logs
// in with the given password.
func metricsRequest(
t *testing.T,
remoteAddr string,
password string,
) *http.Request {
t.Helper()
req := httptest.NewRequestWithContext(
t.Context(), http.MethodGet, "/metrics", nil,
)
req.RemoteAddr = remoteAddr
req.SetBasicAuth(metricsUsername, password)
return req
}
// TestMetricsRateLimitComesBeforeAuth checks that failed logins to
// /metrics count against the rate limit; that once an address is over
// it, even the right password gets 429, with the same body as a wrong
// one; and that another address still gets in.
func TestMetricsRateLimitComesBeforeAuth(t *testing.T) {
viper.Reset()
t.Setenv("DNSWATCHER_TARGETS", "example.com")
t.Setenv("DNSWATCHER_METRICS_USERNAME", metricsUsername)
t.Setenv("DNSWATCHER_METRICS_PASSWORD", metricsPassword)
const (
guesser = "198.51.100.1:4000"
other = "198.51.100.2:4000"
// Far more guesses than the rate limit allows.
maxGuesses = 1000
)
srv := routedServer(t)
var guess *httptest.ResponseRecorder
for range maxGuesses {
guess = serve(srv, metricsRequest(t, guesser, "wrong"))
if guess.Code != http.StatusUnauthorized {
break
}
}
if guess.Code != http.StatusTooManyRequests {
t.Fatalf("wrong password: status = %d, want 429", guess.Code)
}
right := serve(srv, metricsRequest(t, guesser, metricsPassword))
if right.Code != http.StatusTooManyRequests {
t.Errorf("right password: status = %d, want 429", right.Code)
}
if right.Body.String() != guess.Body.String() {
t.Errorf(
"429 body with right password = %q, with wrong one = %q",
right.Body.String(), guess.Body.String(),
)
}
rec := serve(srv, metricsRequest(t, other, metricsPassword))
if rec.Code != http.StatusOK {
t.Errorf("another address: status = %d, want 200", rec.Code)
}
}
+190
View File
@@ -0,0 +1,190 @@
package server_test
import (
"io"
"net/http"
"net/http/httptest"
"net/url"
"strings"
"sync"
"testing"
"time"
"github.com/getsentry/sentry-go"
"github.com/spf13/viper"
"go.uber.org/fx"
"sneak.berlin/go/dnswatcher/internal/server"
)
// The tests below set env vars and touch the global state of viper and
// of Sentry, so they cannot use t.Parallel.
// standInDelay is how long the Sentry stand-in takes to answer. It
// records a report only then, so a report that Shutdown did not wait
// for has not been recorded yet when Shutdown returns.
const standInDelay = 100 * time.Millisecond
// sentryStandIn is a local HTTP server in place of Sentry's, so that
// nothing a test reports leaves the host. It keeps the body of every
// request it receives.
type sentryStandIn struct {
server *httptest.Server
mu sync.Mutex
bodies []string
}
func newSentryStandIn(t *testing.T) *sentryStandIn {
t.Helper()
standIn := &sentryStandIn{}
standIn.server = httptest.NewServer(http.HandlerFunc(
func(_ http.ResponseWriter, r *http.Request) {
body, err := io.ReadAll(r.Body)
if err != nil {
t.Errorf("reading request to the Sentry stand-in: %v", err)
}
time.Sleep(standInDelay)
standIn.mu.Lock()
defer standIn.mu.Unlock()
standIn.bodies = append(standIn.bodies, string(body))
},
))
t.Cleanup(standIn.server.Close)
return standIn
}
// dsn returns a DSN that points Sentry at the stand-in.
func (s *sentryStandIn) dsn(t *testing.T) string {
t.Helper()
dsn, err := url.Parse(s.server.URL)
if err != nil {
t.Fatalf("parsing the stand-in URL: %v", err)
}
dsn.User = url.User("public-key")
dsn.Path = "/1"
return dsn.String()
}
// received reports whether a request to the stand-in contained text.
func (s *sentryStandIn) received(text string) bool {
s.mu.Lock()
defer s.mu.Unlock()
for _, body := range s.bodies {
if strings.Contains(body, text) {
return true
}
}
return false
}
func TestSentryUnsetDoesNothing(t *testing.T) {
viper.Reset()
t.Setenv("DNSWATCHER_TARGETS", "example.com")
t.Setenv("DNSWATCHER_SENTRY_DSN", "")
srv := buildServer(t)
err := server.EnableSentry(srv)
if err != nil {
t.Fatalf("Sentry setup with no DSN: %v", err)
}
if sentry.CurrentHub().Client() != nil {
t.Error("Sentry was set up with no DSN configured")
}
}
// TestSentryReportsHandlerPanic checks that with a valid DSN a panic in
// a handler is reported to Sentry, still reaches chimw.Recoverer, and
// has been sent by the time Shutdown returns, and that nothing else is
// sent to Sentry.
func TestSentryReportsHandlerPanic(t *testing.T) {
standIn := newSentryStandIn(t)
viper.Reset()
t.Setenv("DNSWATCHER_TARGETS", "example.com")
t.Setenv("DNSWATCHER_SENTRY_DSN", standIn.dsn(t))
srv := buildServer(t)
err := server.EnableSentry(srv)
if err != nil {
t.Fatalf("Sentry setup with a valid DSN: %v", err)
}
// Sentry's client is global: close it so later tests find none.
t.Cleanup(func() {
sentry.CurrentHub().Client().Close()
sentry.CurrentHub().BindClient(nil)
})
const panicMessage = "handler panic in the Sentry test"
srv.SetupRoutes()
server.RouterOf(srv).Get(
"/panic",
func(http.ResponseWriter, *http.Request) {
panic(panicMessage)
},
)
// An ordinary request first: with client reports on, Sentry would
// add a count of its dropped transaction to the panic report.
serve(srv, httptest.NewRequestWithContext(
t.Context(), http.MethodGet, "/.well-known/healthcheck", nil,
))
rec := serve(srv, httptest.NewRequestWithContext(
t.Context(), http.MethodGet, "/panic", nil,
))
if rec.Code != http.StatusInternalServerError {
t.Errorf(
"status %d, want %d from the recoverer",
rec.Code, http.StatusInternalServerError,
)
}
err = srv.Shutdown(t.Context())
if err != nil {
t.Fatalf("Shutdown: %v", err)
}
if !standIn.received(panicMessage) {
t.Error("the panic had not been sent to Sentry when Shutdown returned")
}
if standIn.received("client_report") {
t.Error("Sentry was sent a client report, not only the panic")
}
}
func TestSentryInvalidDSNStopsStartup(t *testing.T) {
viper.Reset()
t.Setenv("DNSWATCHER_TARGETS", "example.com")
t.Setenv("DNSWATCHER_DATA_DIR", t.TempDir())
// Sentry cannot parse this: it has no public key before the host.
t.Setenv("DNSWATCHER_SENTRY_DSN", "https://sentry.test/1")
app := newServerApp(fx.Invoke(func(*server.Server) {}))
err := app.Start(t.Context())
if err == nil {
_ = app.Stop(t.Context())
t.Fatal("startup succeeded with an invalid DSN")
}
if !strings.Contains(err.Error(), "invalid DNSWATCHER_SENTRY_DSN") {
t.Errorf("startup error does not name the setting: %v", err)
}
}
+130 -16
View File
@@ -9,6 +9,7 @@ import (
"net/http"
"time"
"github.com/getsentry/sentry-go"
"github.com/go-chi/chi/v5"
"go.uber.org/fx"
@@ -33,19 +34,68 @@ type Params struct {
// shutdownTimeout is how long to wait for graceful shutdown.
const shutdownTimeout = 30 * time.Second
// readHeaderTimeout is the max duration for reading request headers.
const readHeaderTimeout = 10 * time.Second
// sentryFlushTimeout is how long shutdown waits for Sentry to send the
// error reports it still holds.
const sentryFlushTimeout = 2 * time.Second
// Socket-level timeouts for the HTTP server.
//
// These bound time spent on the connection itself and are a distinct
// control from the per-request handler budget enforced by
// chimw.Timeout(requestTimeout) in routes.go: that one cancels the
// request context after requestTimeout but never touches the socket,
// so without the values below a peer can hold a connection open
// forever (slowloris, unreaped keep-alives).
//
// The one hard constraint between the two controls is
// writeTimeout > requestTimeout. net/http arms the write deadline
// once the request headers have been read, so on a plaintext
// connection it covers handler execution AND the response flush. If
// writeTimeout were <= requestTimeout the server would sever the
// connection before a handler that legitimately consumed its full
// budget could emit anything, making the 60s budget unreachable in
// practice. The margin between them is the response-flush allowance.
//
// The only clients of this service are browsers loading the dashboard
// and a Prometheus scraper; the values are sized for those.
const (
// readHeaderTimeout is the max duration for reading request
// headers.
readHeaderTimeout = 10 * time.Second
// readTimeout bounds reading the entire request, headers plus
// body. Every route here is a GET with no body, so this only
// ever needs to cover headers; the extra 5s over
// readHeaderTimeout is slack, not a real allowance, and keeps a
// body dribbled one byte at a time from holding the read side
// open indefinitely.
readTimeout = 15 * time.Second
// writeTimeout must exceed the requestTimeout handler budget
// (60s) per the note above. The 15s difference is the allowance
// for flushing a completed response to a slow client.
writeTimeout = 75 * time.Second
// idleTimeout reaps keep-alive connections between requests. It
// is deliberately longer than the common Prometheus scrape
// intervals (15s/30s/60s) so the scraper reuses its connection
// rather than reconnecting every cycle, while a browser tab
// left open on the dashboard stops occupying a connection
// within two minutes of going quiet.
idleTimeout = 120 * time.Second
)
// Server is the HTTP server.
type Server struct {
startupTime time.Time
port int
log *slog.Logger
router *chi.Mux
httpServer *http.Server
params Params
mw *middleware.Middleware
handlers *handlers.Handlers
startupTime time.Time
port int
sentryEnabled bool
log *slog.Logger
router *chi.Mux
httpServer *http.Server
params Params
mw *middleware.Middleware
handlers *handlers.Handlers
}
// New creates a new Server instance.
@@ -64,6 +114,12 @@ func New(
lifecycle.Append(fx.Hook{
OnStart: func(_ context.Context) error {
srv.startupTime = time.Now()
err := srv.enableSentry()
if err != nil {
return err
}
go srv.Run()
return nil
@@ -76,16 +132,29 @@ func New(
return srv, nil
}
// newHTTPServer builds the listening http.Server with every
// socket-level timeout set. All four are set deliberately: a zero
// value in net/http means "no limit", not "some default".
func newHTTPServer(
listenAddr string,
handler http.Handler,
) *http.Server {
return &http.Server{
Addr: listenAddr,
Handler: handler,
ReadTimeout: readTimeout,
ReadHeaderTimeout: readHeaderTimeout,
WriteTimeout: writeTimeout,
IdleTimeout: idleTimeout,
}
}
// Run starts the HTTP server.
func (s *Server) Run() {
s.SetupRoutes()
listenAddr := fmt.Sprintf(":%d", s.port)
s.httpServer = &http.Server{
Addr: listenAddr,
Handler: s,
ReadHeaderTimeout: readHeaderTimeout,
}
s.httpServer = newHTTPServer(listenAddr, s)
s.log.Info("http server starting", "addr", listenAddr)
@@ -95,8 +164,11 @@ func (s *Server) Run() {
}
}
// Shutdown gracefully shuts down the server.
// Shutdown gracefully shuts down the server, then sends the error
// reports Sentry still holds.
func (s *Server) Shutdown(ctx context.Context) error {
defer s.flushSentry()
if s.httpServer == nil {
return nil
}
@@ -127,3 +199,45 @@ func (s *Server) ServeHTTP(
) {
s.router.ServeHTTP(writer, request)
}
// enableSentry turns on Sentry error reporting when
// DNSWATCHER_SENTRY_DSN is set, and does nothing when it is not. A DSN
// that Sentry cannot parse is an error, so that startup stops instead
// of running without the error reporting the operator asked for.
func (s *Server) enableSentry() error {
if s.params.Config.SentryDSN == "" {
return nil
}
err := sentry.Init(sentry.ClientOptions{
Dsn: s.params.Config.SentryDSN,
Release: s.params.Globals.Appname + "-" + s.params.Globals.Version,
// Use the transport that queues each report as it is made. With
// the default one, Flush can return before sending a report made
// just before it, such as one from the last request at shutdown.
DisableTelemetryBuffer: true,
// Send panic reports only, not Sentry's counts of what it dropped,
// such as the transaction it starts for every request.
DisableClientReports: true,
})
if err != nil {
return fmt.Errorf("invalid DNSWATCHER_SENTRY_DSN: %w", err)
}
s.log.Info("sentry error reporting activated")
s.sentryEnabled = true
return nil
}
// flushSentry sends the error reports Sentry still holds, waiting at
// most sentryFlushTimeout.
func (s *Server) flushSentry() {
if !s.sentryEnabled {
return
}
if !sentry.Flush(sentryFlushTimeout) {
s.log.Warn("sentry flush timed out; some error reports were not sent")
}
}
+135
View File
@@ -0,0 +1,135 @@
package server_test
import (
"testing"
"github.com/spf13/viper"
"go.uber.org/fx"
"sneak.berlin/go/dnswatcher/internal/config"
"sneak.berlin/go/dnswatcher/internal/globals"
"sneak.berlin/go/dnswatcher/internal/handlers"
"sneak.berlin/go/dnswatcher/internal/healthcheck"
"sneak.berlin/go/dnswatcher/internal/logger"
"sneak.berlin/go/dnswatcher/internal/middleware"
"sneak.berlin/go/dnswatcher/internal/notify"
"sneak.berlin/go/dnswatcher/internal/server"
"sneak.berlin/go/dnswatcher/internal/state"
)
// newServerApp builds an fx app holding a *server.Server wired exactly
// as cmd/dnswatcher wires it, minus the watcher/resolver subtree that
// would touch live DNS, plus the given option. config.New reads viper,
// so the caller must first configure it, which is also why the caller
// cannot run in parallel.
func newServerApp(option fx.Option) *fx.App {
return fx.New(
fx.NopLogger,
fx.Provide(
globals.New,
logger.New,
config.New,
state.New,
healthcheck.New,
notify.New,
middleware.New,
handlers.New,
server.New,
),
option,
)
}
// buildServer builds the server without starting the app's lifecycle,
// so no OnStart hook runs and nothing listens or resolves.
func buildServer(t *testing.T) *server.Server {
t.Helper()
var srv *server.Server
app := newServerApp(fx.Populate(&srv))
err := app.Err()
if err != nil {
t.Fatalf("building server graph: %v", err)
}
return srv
}
// TestRunWiresSocketTimeouts pins that the http.Server the running
// server actually serves — the one Run builds and hands to
// ListenAndServe — carries every socket-level timeout, plus the two
// relationships the values must satisfy.
//
// Run is driven to completion with an unbindable port: it builds and
// stores s.httpServer, then ListenAndServe fails at once and Run
// returns without ever listening. The assertions run in the same
// goroutine after Run returns, so reading s.httpServer is free of any
// data race. Nothing here measures elapsed time.
//
// ReadTimeout must be at least ReadHeaderTimeout. net/http reads the
// headers under ReadHeaderTimeout, then sets the read deadline for the
// rest of the request to ReadTimeout, counted from when it started
// reading the request. If ReadTimeout were smaller, a request whose
// headers arrived after ReadTimeout but within ReadHeaderTimeout would
// get a read deadline that had already passed, so reading its body
// would fail at once.
func TestRunWiresSocketTimeouts(t *testing.T) {
// Sets an env var and touches viper global state, so like the
// config tests it cannot use t.Parallel.
viper.Reset()
t.Setenv("DNSWATCHER_TARGETS", "example.com")
srv := buildServer(t)
server.SetListenPort(srv, -1)
srv.Run()
hs := server.HTTPServerOf(srv)
if hs == nil {
t.Fatal("Run did not build an http.Server")
}
if hs.ReadTimeout <= 0 {
t.Errorf("ReadTimeout must be non-zero, got %v", hs.ReadTimeout)
}
if hs.ReadHeaderTimeout <= 0 {
t.Errorf(
"ReadHeaderTimeout must be non-zero, got %v",
hs.ReadHeaderTimeout,
)
}
if hs.WriteTimeout <= 0 {
t.Errorf("WriteTimeout must be non-zero, got %v", hs.WriteTimeout)
}
if hs.IdleTimeout <= 0 {
t.Errorf("IdleTimeout must be non-zero, got %v", hs.IdleTimeout)
}
if hs.WriteTimeout <= server.RequestTimeout {
t.Errorf(
"WriteTimeout (%v) must exceed handler budget (%v)",
hs.WriteTimeout,
server.RequestTimeout,
)
}
if hs.ReadTimeout < hs.ReadHeaderTimeout {
t.Errorf(
"ReadTimeout (%v) must be >= ReadHeaderTimeout (%v)",
hs.ReadTimeout,
hs.ReadHeaderTimeout,
)
}
if hs.Handler != srv {
t.Errorf(
"Run wired handler %T, want the *server.Server",
hs.Handler,
)
}
}
+23
View File
@@ -0,0 +1,23 @@
package state
import (
"log/slog"
"sneak.berlin/go/dnswatcher/internal/config"
)
// NewForTestWithDataDir creates an empty State that saves to dataDir,
// without the fx lifecycle.
func NewForTestWithDataDir(dataDir string) *State {
return &State{
log: slog.Default(),
snapshot: &Snapshot{
Version: stateVersion,
Domains: make(map[string]*DomainState),
Hostnames: make(map[string]*HostnameState),
Ports: make(map[string]*PortState),
Certificates: make(map[string]*CertificateState),
},
config: &config.Config{DataDir: dataDir},
}
}
+120 -2
View File
@@ -8,6 +8,7 @@ import (
"log/slog"
"os"
"path/filepath"
"slices"
"sync"
"time"
@@ -35,9 +36,13 @@ type Params struct {
}
// DomainState holds the monitoring state for an apex domain.
// NameserverAddresses holds the sorted addresses each nameserver's name
// resolves to, by nameserver name. A state file written before it
// existed loads with it nil.
type DomainState struct {
Nameservers []string `json:"nameservers"`
LastChecked time.Time `json:"lastChecked"`
Nameservers []string `json:"nameservers"`
NameserverAddresses map[string][]string `json:"nameserverAddresses"`
LastChecked time.Time `json:"lastChecked"`
}
// NameserverRecordState holds one NS's response for a hostname.
@@ -49,8 +54,13 @@ type NameserverRecordState struct {
}
// HostnameState holds per-nameserver monitoring state for a hostname.
// CNAMEAddresses holds the sorted addresses at the end of the name's
// CNAME chain, found when its nameservers answered with a CNAME and no
// address; it is empty otherwise. It is nil when they are not known: a
// state file written before it existed loads with it nil.
type HostnameState struct {
RecordsByNameserver map[string]*NameserverRecordState `json:"recordsByNameserver"`
CNAMEAddresses []string `json:"cnameAddresses"`
LastChecked time.Time `json:"lastChecked"`
}
@@ -112,6 +122,8 @@ type CertificateState struct {
}
// Snapshot is the complete monitoring state persisted to disk.
// Hostnames also holds each apex domain's own records, under the
// domain's name, which has an entry in Domains too.
type Snapshot struct {
Version int `json:"version"`
LastUpdated time.Time `json:"lastUpdated"`
@@ -148,6 +160,11 @@ func New(
lifecycle.Append(fx.Hook{
OnStart: func(_ context.Context) error {
err := state.checkDataDirWritable()
if err != nil {
return err
}
return state.Load()
},
OnStop: func(_ context.Context) error {
@@ -187,6 +204,19 @@ func (s *State) Load() error {
return fmt.Errorf("parsing state file: %w", err)
}
// A state file saved before each record value was stored once can
// hold a hostname's CNAME once for every record type asked for.
// Each value is kept once, so the first check does not see a
// record change.
for _, hs := range snapshot.Hostnames {
for _, ns := range hs.RecordsByNameserver {
for recordType, values := range ns.Records {
slices.Sort(values)
ns.Records[recordType] = slices.Compact(values)
}
}
}
s.snapshot = &snapshot
s.log.Info("loaded state from disk", "path", path)
@@ -261,6 +291,27 @@ func (s *State) GetDomainState(
return ds, ok
}
// DeleteDomainState removes a domain state entry.
func (s *State) DeleteDomainState(domain string) {
s.mu.Lock()
defer s.mu.Unlock()
delete(s.snapshot.Domains, domain)
}
// GetAllDomainNames returns the names of all domain state entries.
func (s *State) GetAllDomainNames() []string {
s.mu.RLock()
defer s.mu.RUnlock()
names := make([]string, 0, len(s.snapshot.Domains))
for name := range s.snapshot.Domains {
names = append(names, name)
}
return names
}
// SetHostnameState updates the state for a hostname.
func (s *State) SetHostnameState(
hostname string,
@@ -284,6 +335,28 @@ func (s *State) GetHostnameState(
return hs, ok
}
// DeleteHostnameState removes a hostname state entry.
func (s *State) DeleteHostnameState(hostname string) {
s.mu.Lock()
defer s.mu.Unlock()
delete(s.snapshot.Hostnames, hostname)
}
// GetAllHostnames returns the names of all hostname state entries,
// which include each apex domain's own records.
func (s *State) GetAllHostnames() []string {
s.mu.RLock()
defer s.mu.RUnlock()
names := make([]string, 0, len(s.snapshot.Hostnames))
for name := range s.snapshot.Hostnames {
names = append(names, name)
}
return names
}
// SetPortState updates the state for a port.
func (s *State) SetPortState(key string, ps *PortState) {
s.mu.Lock()
@@ -345,3 +418,48 @@ func (s *State) GetCertificateState(
return cs, ok
}
// DeleteCertificateState removes a certificate state entry.
func (s *State) DeleteCertificateState(key string) {
s.mu.Lock()
defer s.mu.Unlock()
delete(s.snapshot.Certificates, key)
}
// GetAllCertificateKeys returns all certificate state keys.
func (s *State) GetAllCertificateKeys() []string {
s.mu.RLock()
defer s.mu.RUnlock()
keys := make([]string, 0, len(s.snapshot.Certificates))
for k := range s.snapshot.Certificates {
keys = append(keys, k)
}
return keys
}
// checkDataDirWritable creates the data directory if needed, then writes
// and removes the temp file that Save uses. It runs at startup so that an
// unwritable directory stops the process, instead of the process running
// with every save failing and only logged.
func (s *State) checkDataDirWritable() error {
dir := s.config.DataDir
tmpPath := s.config.StatePath() + ".tmp"
err := os.MkdirAll(dir, dirPermissions)
if err == nil {
err = os.WriteFile(tmpPath, nil, filePermissions)
}
if err == nil {
err = os.Remove(tmpPath)
}
if err != nil {
return fmt.Errorf("data directory %s is not writable: %w", dir, err)
}
return nil
}
+319 -36
View File
@@ -4,10 +4,17 @@ import (
"encoding/json"
"os"
"path/filepath"
"reflect"
"strings"
"sync"
"testing"
"time"
"go.uber.org/fx/fxtest"
"sneak.berlin/go/dnswatcher/internal/config"
"sneak.berlin/go/dnswatcher/internal/globals"
"sneak.berlin/go/dnswatcher/internal/logger"
"sneak.berlin/go/dnswatcher/internal/state"
)
@@ -31,6 +38,10 @@ func populateState(t *testing.T, s *state.State) {
s.SetDomainState("example.com", &state.DomainState{
Nameservers: []string{testNS1, testNS2},
NameserverAddresses: map[string][]string{
testNS1: {testIP, testIPv4},
testNS2: {testIPv4},
},
LastChecked: now,
})
@@ -117,6 +128,206 @@ func TestSaveLoadRoundTrip_Domains(t *testing.T) {
if len(dom.Nameservers) != 2 {
t.Errorf("expected 2 nameservers, got %d", len(dom.Nameservers))
}
want := map[string][]string{
testNS1: {testIP, testIPv4},
testNS2: {testIPv4},
}
if !reflect.DeepEqual(dom.NameserverAddresses, want) {
t.Errorf(
"nameserver addresses: got %v, want %v",
dom.NameserverAddresses, want,
)
}
}
// TestLoadStateFromBeforeNameserverAddresses loads a state file written
// before nameserver addresses were saved.
func TestLoadStateFromBeforeNameserverAddresses(t *testing.T) {
t.Parallel()
dir := t.TempDir()
data := []byte(`{
"version": 1,
"lastUpdated": "2026-02-19T12:00:00Z",
"domains": {
"example.com": {
"nameservers": ["ns1.example.com.", "ns2.example.com."],
"lastChecked": "2026-02-19T12:00:00Z"
}
}
}`)
err := os.WriteFile(filepath.Join(dir, "state.json"), data, 0o600)
if err != nil {
t.Fatalf("writing state file: %v", err)
}
s := state.NewForTestWithDataDir(dir)
err = s.Load()
if err != nil {
t.Fatalf("Load() error: %v", err)
}
dom, ok := s.GetDomainState("example.com")
if !ok {
t.Fatal("missing domain example.com")
}
if !reflect.DeepEqual(dom.Nameservers, []string{testNS1, testNS2}) {
t.Errorf("nameservers: got %v", dom.Nameservers)
}
if dom.NameserverAddresses != nil {
t.Errorf(
"nameserver addresses: got %v, want none",
dom.NameserverAddresses,
)
}
}
// TestSaveLoadRoundTrip_CNAMEAddresses checks that no addresses at the
// end of a hostname's CNAME chain load as an empty list, and addresses
// that are not known load as nil: the watcher tells the two apart.
func TestSaveLoadRoundTrip_CNAMEAddresses(t *testing.T) {
t.Parallel()
dir := t.TempDir()
s := state.NewForTestWithDataDir(dir)
want := map[string][]string{
"cname.example.com": {testIP},
"none.example.com": {},
"not-known.example.com": nil,
}
for name, addresses := range want {
s.SetHostnameState(name, &state.HostnameState{
CNAMEAddresses: addresses,
})
}
err := s.Save()
if err != nil {
t.Fatalf("Save() error: %v", err)
}
loaded := state.NewForTestWithDataDir(dir)
err = loaded.Load()
if err != nil {
t.Fatalf("Load() error: %v", err)
}
for name, addresses := range want {
hs, ok := loaded.GetHostnameState(name)
if !ok {
t.Fatalf("missing hostname %s", name)
}
if !reflect.DeepEqual(hs.CNAMEAddresses, addresses) {
t.Errorf(
"%s: loaded %#v, want %#v",
name, hs.CNAMEAddresses, addresses,
)
}
}
}
// TestLoadStateFromBeforeCNAMEAddresses loads a state file written
// before the addresses at the end of a hostname's CNAME chain were
// saved. They load as not known (nil), not as none.
func TestLoadStateFromBeforeCNAMEAddresses(t *testing.T) {
t.Parallel()
dir := t.TempDir()
data := []byte(`{
"version": 1,
"lastUpdated": "2026-02-19T12:00:00Z",
"hostnames": {
"www.example.com": {
"recordsByNameserver": {},
"lastChecked": "2026-02-19T12:00:00Z"
}
}
}`)
err := os.WriteFile(filepath.Join(dir, "state.json"), data, 0o600)
if err != nil {
t.Fatalf("writing state file: %v", err)
}
s := state.NewForTestWithDataDir(dir)
err = s.Load()
if err != nil {
t.Fatalf("Load() error: %v", err)
}
hs, ok := s.GetHostnameState(testHostname)
if !ok {
t.Fatal("missing hostname " + testHostname)
}
if hs.CNAMEAddresses != nil {
t.Errorf("CNAME addresses: got %#v, want nil", hs.CNAMEAddresses)
}
}
// TestLoadStateWithRepeatedValues loads a state file saved when a
// hostname's CNAME was stored once for every record type asked for.
// Each value must load once, and every different value must load.
func TestLoadStateWithRepeatedValues(t *testing.T) {
t.Parallel()
dir := t.TempDir()
data := []byte(`{
"version": 1,
"hostnames": {
"www.example.com": {
"recordsByNameserver": {
"ns1.example.com.": {
"records": {
"A": ["192.0.2.2", "192.0.2.1", "192.0.2.2", "192.0.2.1"],
"CNAME": ["a.example.net.", "a.example.net.", "a.example.net."]
},
"status": "ok"
}
}
}
}
}`)
err := os.WriteFile(filepath.Join(dir, "state.json"), data, 0o600)
if err != nil {
t.Fatalf("writing state file: %v", err)
}
s := state.NewForTestWithDataDir(dir)
err = s.Load()
if err != nil {
t.Fatalf("Load() error: %v", err)
}
hs, ok := s.GetHostnameState(testHostname)
if !ok {
t.Fatal("missing hostname " + testHostname)
}
want := map[string][]string{
"A": {"192.0.2.1", "192.0.2.2"},
"CNAME": {"a.example.net."},
}
got := hs.RecordsByNameserver[testNS1].Records
if !reflect.DeepEqual(got, want) {
t.Errorf("records: got %v, want %v", got, want)
}
}
// TestSaveLoadRoundTrip_Hostnames verifies hostname data survives a save/load cycle.
@@ -493,6 +704,107 @@ func TestSaveWritePermissionError(t *testing.T) {
}
}
// startState builds a State through the real constructor and runs its
// startup hook against dataDir, returning the startup error.
func startState(t *testing.T, dataDir string) error {
t.Helper()
g, err := globals.New(nil)
if err != nil {
t.Fatalf("globals.New: %v", err)
}
log, err := logger.New(nil, logger.Params{Globals: g})
if err != nil {
t.Fatalf("logger.New: %v", err)
}
lifecycle := fxtest.NewLifecycle(t)
_, err = state.New(lifecycle, state.Params{
Logger: log,
Config: &config.Config{DataDir: dataDir},
})
if err != nil {
t.Fatalf("state.New: %v", err)
}
return lifecycle.Start(t.Context())
}
// TestStartupFailsWhenDataDirNotWritable verifies that startup stops
// with an error naming the data directory when it cannot be written.
// The directory's parent is a regular file, which also fails as root.
func TestStartupFailsWhenDataDirNotWritable(t *testing.T) {
t.Parallel()
parent := filepath.Join(t.TempDir(), "file")
err := os.WriteFile(parent, nil, 0o600)
if err != nil {
t.Fatalf("writing file: %v", err)
}
dataDir := filepath.Join(parent, "data")
err = startState(t, dataDir)
if err == nil {
t.Fatal("startup should fail when the data directory is not writable")
}
want := "data directory " + dataDir + " is not writable"
if !strings.Contains(err.Error(), want) {
t.Errorf("startup error %q does not contain %q", err, want)
}
}
// TestStartupFailsWhenExistingDataDirNotWritable verifies that startup
// stops when the data directory exists but the temp file that saving uses
// cannot be written in it. A directory sitting at the temp file's path
// makes that write fail, which also holds as root.
func TestStartupFailsWhenExistingDataDirNotWritable(t *testing.T) {
t.Parallel()
dataDir := t.TempDir()
err := os.Mkdir(filepath.Join(dataDir, "state.json.tmp"), 0o700)
if err != nil {
t.Fatalf("creating directory: %v", err)
}
err = startState(t, dataDir)
if err == nil {
t.Fatal("startup should fail when the data directory is not writable")
}
want := "data directory " + dataDir + " is not writable"
if !strings.Contains(err.Error(), want) {
t.Errorf("startup error %q does not contain %q", err, want)
}
}
// TestStartupCreatesDataDir verifies that startup creates a missing
// data directory and leaves nothing behind in it.
func TestStartupCreatesDataDir(t *testing.T) {
t.Parallel()
dataDir := filepath.Join(t.TempDir(), "data")
err := startState(t, dataDir)
if err != nil {
t.Fatalf("startup error: %v", err)
}
entries, err := os.ReadDir(dataDir)
if err != nil {
t.Fatalf("reading data directory: %v", err)
}
if len(entries) != 0 {
t.Errorf("startup left %d entries in the data directory", len(entries))
}
}
// TestPortStateUnmarshalJSON_NewFormat verifies deserialization of the
// current multi-hostname format.
func TestPortStateUnmarshalJSON_NewFormat(t *testing.T) {
@@ -632,7 +944,7 @@ func TestPortStateUnmarshalJSON_BothFormats(t *testing.T) {
func TestGetSnapshot_ReturnsCopy(t *testing.T) {
t.Parallel()
s := state.NewForTest()
s := state.NewForTestWithDataDir(t.TempDir())
populateState(t, s)
@@ -654,7 +966,7 @@ func TestGetSnapshot_ReturnsCopy(t *testing.T) {
func TestDomainState_GetSet(t *testing.T) {
t.Parallel()
s := state.NewForTest()
s := state.NewForTestWithDataDir(t.TempDir())
// Get on missing key returns false.
_, ok := s.GetDomainState("nonexistent.com")
@@ -705,7 +1017,7 @@ func TestDomainState_GetSet(t *testing.T) {
func TestHostnameState_GetSet(t *testing.T) {
t.Parallel()
s := state.NewForTest()
s := state.NewForTestWithDataDir(t.TempDir())
_, ok := s.GetHostnameState("missing.example.com")
if ok {
@@ -750,7 +1062,7 @@ func TestHostnameState_GetSet(t *testing.T) {
func TestPortState_GetSetDelete(t *testing.T) {
t.Parallel()
s := state.NewForTest()
s := state.NewForTestWithDataDir(t.TempDir())
_, ok := s.GetPortState("1.2.3.4:80")
if ok {
@@ -788,7 +1100,7 @@ func TestPortState_GetSetDelete(t *testing.T) {
func TestGetAllPortKeys(t *testing.T) {
t.Parallel()
s := state.NewForTest()
s := state.NewForTestWithDataDir(t.TempDir())
keys := s.GetAllPortKeys()
if len(keys) != 0 {
@@ -830,7 +1142,7 @@ func TestGetAllPortKeys(t *testing.T) {
func TestCertificateState_GetSet(t *testing.T) {
t.Parallel()
s := state.NewForTest()
s := state.NewForTestWithDataDir(t.TempDir())
_, ok := s.GetCertificateState("1.2.3.4:443:www.example.com")
if ok {
@@ -1051,7 +1363,7 @@ func TestLoadPreservesExistingStateOnMissingFile(t *testing.T) {
func TestConcurrentGetSet(t *testing.T) {
t.Parallel()
s := state.NewForTest()
s := state.NewForTestWithDataDir(t.TempDir())
const goroutines = 20
@@ -1261,35 +1573,6 @@ func TestMultipleSavesOverwrite(t *testing.T) {
}
}
// TestNewForTest verifies the test helper creates a valid empty state.
func TestNewForTest(t *testing.T) {
t.Parallel()
s := state.NewForTest()
snap := s.GetSnapshot()
if snap.Version != 1 {
t.Errorf("version: got %d, want 1", snap.Version)
}
if snap.Domains == nil {
t.Error("Domains map should be initialized")
}
if snap.Hostnames == nil {
t.Error("Hostnames map should be initialized")
}
if snap.Ports == nil {
t.Error("Ports map should be initialized")
}
if snap.Certificates == nil {
t.Error("Certificates map should be initialized")
}
}
// TestSaveFilePermissions verifies the saved file has restricted permissions.
func TestSaveFilePermissions(t *testing.T) {
t.Parallel()
-38
View File
@@ -1,38 +0,0 @@
package state
import (
"log/slog"
"sneak.berlin/go/dnswatcher/internal/config"
)
// NewForTest creates a State for unit testing with no persistence.
func NewForTest() *State {
return &State{
log: slog.Default(),
snapshot: &Snapshot{
Version: stateVersion,
Domains: make(map[string]*DomainState),
Hostnames: make(map[string]*HostnameState),
Ports: make(map[string]*PortState),
Certificates: make(map[string]*CertificateState),
},
config: &config.Config{DataDir: ""},
}
}
// NewForTestWithDataDir creates a State backed by the given directory
// for tests that need file persistence.
func NewForTestWithDataDir(dataDir string) *State {
return &State{
log: slog.Default(),
snapshot: &Snapshot{
Version: stateVersion,
Domains: make(map[string]*DomainState),
Hostnames: make(map[string]*HostnameState),
Ports: make(map[string]*PortState),
Certificates: make(map[string]*CertificateState),
},
config: &config.Config{DataDir: dataDir},
}
}
+153
View File
@@ -0,0 +1,153 @@
package watcher_test
import (
"bytes"
"context"
"log/slog"
"reflect"
"strings"
"testing"
"time"
"sneak.berlin/go/dnswatcher/internal/portcheck"
"sneak.berlin/go/dnswatcher/internal/resolver"
"sneak.berlin/go/dnswatcher/internal/state"
"sneak.berlin/go/dnswatcher/internal/tlscheck"
"sneak.berlin/go/dnswatcher/internal/watcher"
)
// TestCancelledCheckSavesNothing runs a check with its context already
// cancelled, which is how the rest of a check runs once shutdown cuts it
// short. The real resolver drops the DNS lookup without sending a query,
// and the real port and TLS checkers fail without connecting. The port
// and certificate state the last check saved must stay as it was, and
// nothing may be notified.
func TestCancelledCheckSavesNothing(t *testing.T) {
t.Parallel()
cfg := defaultTestConfig(t)
cfg.Hostnames = []string{host}
// newTestWatcher's watcher has stand-in checkers. This one, on the
// same state and notifier, has the real ones.
_, deps := newTestWatcher(t, cfg)
w := watcher.NewForTest(
cfg,
deps.state,
resolver.NewFromLogger(slog.Default()),
portcheck.NewStandalone(),
tlscheck.NewStandalone(),
deps.notifier,
)
// The last check found host at a local address, with both ports
// open and a good certificate.
const localIP = "127.0.0.1"
deps.state.SetHostnameState(host, hostnameState(
map[string]map[string][]string{nsA: {"A": {localIP}}},
))
ports := map[string]*state.PortState{
localIP + ":80": {Open: true, Hostnames: []string{host}},
localIP + ":443": {Open: true, Hostnames: []string{host}},
}
for key, ps := range ports {
deps.state.SetPortState(key, ps)
}
certKey := localIP + ":443:" + host
cert := &state.CertificateState{CommonName: host, Status: "ok"}
deps.state.SetCertificateState(certKey, cert)
ctx, cancel := context.WithCancel(t.Context())
cancel()
w.RunOnce(ctx)
for key, want := range ports {
got, _ := deps.state.GetPortState(key)
if !reflect.DeepEqual(got, want) {
t.Errorf("port %s saved as %+v, want %+v", key, got, want)
}
}
got, _ := deps.state.GetCertificateState(certKey)
if !reflect.DeepEqual(got, cert) {
t.Errorf("certificate saved as %+v, want %+v", got, cert)
}
notifications := deps.notifier.getNotifications()
if len(notifications) != 0 {
t.Errorf("sent %v, want no notifications", notifications)
}
}
// newLoggingWatcher returns a watcher for a domain and a hostname, with
// the real resolver, that writes what it logs at warning level or above
// into the returned buffer.
func newLoggingWatcher(t *testing.T) (*watcher.Watcher, *bytes.Buffer) {
t.Helper()
cfg := defaultTestConfig(t)
cfg.Domains = []string{testSmallDomain}
cfg.Hostnames = []string{host}
w, _ := newTestWatcher(t, cfg)
logs := &bytes.Buffer{}
w.SetLogger(slog.New(slog.NewJSONHandler(
logs, &slog.HandlerOptions{Level: slog.LevelWarn},
)))
return w, logs
}
// TestLookupCutShortIsNotLogged checks a domain and a hostname, looks
// up a nameserver's addresses and follows a CNAME, with the context
// cancelled, as shutdown leaves it. The real resolver fails each lookup
// without sending a query. Shutdown cutting a lookup short is not a
// failure, so nothing may be logged at warning level or above.
func TestLookupCutShortIsNotLogged(t *testing.T) {
t.Parallel()
w, logs := newLoggingWatcher(t)
ctx, cancel := context.WithCancel(t.Context())
cancel()
w.RunOnce(ctx)
w.ResolveNameserverAddresses(ctx, []string{nsA}, nil)
w.ResolveCNAMEAddresses(ctx, host, cnameState(), nil)
if logs.Len() > 0 {
t.Errorf("logged at warning level or above:\n%s", logs)
}
}
// TestLookupOutOfTimeIsLoggedAsError does what
// TestLookupCutShortIsNotLogged does, with the context's deadline passed
// instead. A lookup that ran out of time did fail, so the domain's NS
// lookup, the hostname's lookup, the nameserver's address lookup and the
// CNAME's are each logged as an error.
func TestLookupOutOfTimeIsLoggedAsError(t *testing.T) {
t.Parallel()
w, logs := newLoggingWatcher(t)
ctx, cancel := context.WithDeadline(t.Context(), time.Now())
t.Cleanup(cancel)
w.RunOnce(ctx)
w.ResolveNameserverAddresses(ctx, []string{nsA}, nil)
w.ResolveCNAMEAddresses(ctx, host, cnameState(), nil)
const want = 4
lines := strings.Count(logs.String(), "\n")
errorLines := strings.Count(logs.String(), `"level":"ERROR"`)
if lines != want || errorLines != want {
t.Errorf("logged:\n%s\nwant %d lines, each at error level", logs, want)
}
}
+336
View File
@@ -0,0 +1,336 @@
package watcher_test
import (
"context"
"log/slog"
"slices"
"testing"
"sneak.berlin/go/dnswatcher/internal/livednstest"
"sneak.berlin/go/dnswatcher/internal/resolver"
"sneak.berlin/go/dnswatcher/internal/state"
"sneak.berlin/go/dnswatcher/internal/watcher"
)
// TestCNAMEIntoAnotherZonePortAndTLSChecks runs the port and TLS
// checks on hostname state built here: the name's nameserver answered
// with a CNAME into another zone, and following it found ip1. Both
// checks must use ip1. They look nothing up, so the watcher has no
// resolver.
func TestCNAMEIntoAnotherZonePortAndTLSChecks(t *testing.T) {
t.Parallel()
cfg := defaultTestConfig(t)
cfg.Hostnames = []string{host}
deps := newTestDeps(t, cfg)
w := watcher.NewForTest(
cfg, deps.state, nil,
deps.portChecker, deps.tlsChecker, deps.notifier,
)
deps.state.SetHostnameState(host, cnameState(ip1))
w.CheckAllPorts(t.Context())
w.RunTLSChecks(t.Context())
snap := deps.state.GetSnapshot()
ps, ok := snap.Ports[ip1+":443"]
if !ok || !slices.Contains(ps.Hostnames, host) {
t.Errorf("no port state for %s at %s:443", host, ip1)
}
certKey := ip1 + ":443:" + host
if _, ok := snap.Certificates[certKey]; !ok {
t.Errorf("no certificate state %s", certKey)
}
}
// TestCNAMEThatCannotBeFollowedKeepsPrevious runs a check of a name, not
// the watcher's first, from the point where its records have been looked
// up: they hold a CNAME to a target under .invalid, whose lookup fails.
// The previous check found the same records, and oldIP at the end of the
// CNAME. The check must keep oldIP and send nothing.
func TestCNAMEThatCannotBeFollowedKeepsPrevious(t *testing.T) {
t.Parallel()
w, deps := newTestWatcher(t, defaultTestConfig(t))
w.SetFirstRun(false)
records := map[string]map[string][]string{
nsA: cnameTo("target.example.invalid."),
}
prev := hostnameState(records)
prev.CNAMEAddresses = []string{oldIP}
deps.state.SetHostnameState(host, prev)
// The result is the same whether or not live DNS answers, so the
// lookup is not retried.
_ = livednstest.Run(func(ctx context.Context) error {
w.UpdateHostnameState(ctx, host, hostnameState(records))
return nil
})
hs, _ := deps.state.GetHostnameState(host)
if !slices.Equal(hs.CNAMEAddresses, prev.CNAMEAddresses) {
t.Errorf(
"saved %v, want %v",
hs.CNAMEAddresses, prev.CNAMEAddresses,
)
}
notifications := deps.notifier.getNotifications()
if len(notifications) != 0 {
t.Errorf("sent %v, want no notifications", notifications)
}
}
// followLive follows in live DNS the CNAMEs in a name's records, built
// from records, and returns the addresses saved for the name. The
// previous check saved oldIP, which is kept when a target cannot be
// followed; that is retried. The tests point CNAMEs only at names in
// zones with two nameservers, to keep queries few (see the top of
// watcher_test.go).
func followLive(
t *testing.T,
records map[string]map[string][]string,
) []string {
t.Helper()
w := watcher.NewForTest(
nil, nil, resolver.NewFromLogger(slog.Default()), nil, nil, nil,
)
prev := cnameState(oldIP)
var current *state.HostnameState
livednstest.Retry(t, "following CNAMEs", func(ctx context.Context) error {
current = hostnameState(records)
w.ResolveCNAMEAddresses(ctx, host, current, prev)
if slices.Equal(current.CNAMEAddresses, prev.CNAMEAddresses) {
return livednstest.ErrNoAnswer
}
return nil
})
return current.CNAMEAddresses
}
// TestCNAMEAddressesOfEveryTarget gives a name's two nameservers
// different CNAME targets, as when a secondary still serves an old one.
// The addresses at the end of both are saved, whichever answer is read
// first: one.one.one.one has 1.1.1.1, and dns.adguard-dns.com has
// 94.140.14.14.
func TestCNAMEAddressesOfEveryTarget(t *testing.T) {
t.Parallel()
found := followLive(t, map[string]map[string][]string{
nsA: cnameTo("one.one.one.one."),
nsB: cnameTo("dns.adguard-dns.com."),
})
for _, ip := range []string{"1.1.1.1", "94.140.14.14"} {
if !slices.Contains(found, ip) {
t.Errorf("saved %v, want %s among them", found, ip)
}
}
}
// TestCNAMEChainEndingInNoAddressSavesEmptyList follows a CNAME to a
// name live DNS answers with NXDOMAIN. An empty list is saved, not nil,
// which would mean the addresses are not known.
func TestCNAMEChainEndingInNoAddressSavesEmptyList(t *testing.T) {
t.Parallel()
found := followLive(t, map[string]map[string][]string{
nsA: cnameTo("this-surely-does-not-exist-xyz.example.org."),
})
if found == nil || len(found) != 0 {
t.Errorf("saved %#v, want an empty list", found)
}
}
// TestCNAMEBesideAnAddressNotFollowed gives one nameserver of a name an
// address and another a CNAME. The CNAME is not followed: an empty list
// is saved, not nil, and nothing is looked up, the watcher having no
// resolver.
func TestCNAMEBesideAnAddressNotFollowed(t *testing.T) {
t.Parallel()
w := watcher.NewForTest(nil, nil, nil, nil, nil, nil)
current := hostnameState(map[string]map[string][]string{
nsA: {"A": {ip1}},
nsB: cnameTo("target.example.org."),
})
w.ResolveCNAMEAddresses(t.Context(), host, current, nil)
if current.CNAMEAddresses == nil || len(current.CNAMEAddresses) != 0 {
t.Errorf("saved %#v, want an empty list", current.CNAMEAddresses)
}
}
// TestCNAMEWhoseNameserversAllFailedKeepsPrevious checks a name none of
// whose nameservers answered. The addresses the previous check saved
// from following its CNAME are kept, and nothing is looked up: the
// watcher has no resolver.
func TestCNAMEWhoseNameserversAllFailedKeepsPrevious(t *testing.T) {
t.Parallel()
w := watcher.NewForTest(nil, nil, nil, nil, nil, nil)
current := saved(map[string]*state.NameserverRecordState{
nsA: failed(), nsB: failed(),
})
prev := cnameState(oldIP)
w.ResolveCNAMEAddresses(t.Context(), host, current, prev)
if !slices.Equal(current.CNAMEAddresses, prev.CNAMEAddresses) {
t.Errorf(
"saved %v, want %v",
current.CNAMEAddresses, prev.CNAMEAddresses,
)
}
}
// cnameTo builds the records of a nameserver that answered with a CNAME
// to target and no address.
func cnameTo(target string) map[string][]string {
return map[string][]string{"CNAME": {target}}
}
// cnameState builds the state a check leaves behind for a name whose
// nameserver answered with a CNAME and no address, when following the
// CNAME found these addresses, which may be none.
func cnameState(addresses ...string) *state.HostnameState {
hs := hostnameState(map[string]map[string][]string{
nsA: cnameTo("target.example.org."),
})
hs.CNAMEAddresses = append([]string{}, addresses...)
return hs
}
func TestCNAMEAddressChangeAlerts(t *testing.T) {
t.Parallel()
// A state file written before the addresses were saved loads with
// them nil.
olderStateFile := cnameState()
olderStateFile.CNAMEAddresses = nil
// Each case is the state saved by the previous check and by the
// current one. The name's records are the same in both.
tests := []struct {
name string
prev, current *state.HostnameState
want int
}{
{
"same addresses",
cnameState(ip1, ip2), cnameState(ip1, ip2), 0,
},
{
"same addresses in another order",
cnameState(ip2, ip1), cnameState(ip1, ip2), 0,
},
{
"address replaced",
cnameState(ip1), cnameState(ip2), 1,
},
{
"address added",
cnameState(ip1), cnameState(ip1, ip2), 1,
},
{
"no address at the end of the chain now",
cnameState(ip1), cnameState(), 1,
},
{
"addresses at the end of the chain again",
cnameState(), cnameState(ip1), 1,
},
{
"state file from before addresses were saved",
olderStateFile, cnameState(ip1), 0,
},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
t.Parallel()
notifier := &mockNotifier{}
w := watcher.NewForTest(nil, nil, nil, nil, nil, notifier)
w.DetectHostnameChanges(t.Context(), host, tt.prev, tt.current)
got := len(notifier.getNotifications())
if got != tt.want {
t.Errorf("sent %d notifications, want %d", got, tt.want)
}
})
}
}
func TestCNAMEAddressChangeAlertNamesHostnameAndAddresses(t *testing.T) {
t.Parallel()
notifier := &mockNotifier{}
w := watcher.NewForTest(nil, nil, nil, nil, nil, notifier)
w.DetectHostnameChanges(
t.Context(), host, cnameState(ip1), cnameState(ip2, ip3),
)
want := notification{
Title: "CNAME Address Change: " + host,
Message: "Hostname: " + host +
"\nOld: " + ip1 + "\nNew: " + ip2 + ", " + ip3,
Priority: "warning",
}
got := notifier.getNotifications()
if len(got) != 1 || got[0] != want {
t.Errorf("sent %v, want %v", got, want)
}
}
// TestNameMovedFromARecordsToCNAMEAlerts checks a name that answers
// with an A record and then with a CNAME whose chain ends in ip2. The
// second check is notified as a CNAME address change from no addresses,
// beside the record change. Nothing is looked up: the watcher has no
// resolver.
func TestNameMovedFromARecordsToCNAMEAlerts(t *testing.T) {
t.Parallel()
notifier := &mockNotifier{}
w := watcher.NewForTest(nil, nil, nil, nil, nil, notifier)
prev := hostnameState(map[string]map[string][]string{
nsA: {"A": {ip1}},
})
w.ResolveCNAMEAddresses(t.Context(), host, prev, nil)
w.DetectHostnameChanges(t.Context(), host, prev, cnameState(ip2))
title := "CNAME Address Change: " + host
message := "Hostname: " + host + "\nOld: \nNew: " + ip2
got := notifier.getNotifications()
if !slices.ContainsFunc(got, func(n notification) bool {
return n.Title == title && n.Message == message
}) {
t.Errorf("sent %v, want %q with %q among them", got, title, message)
}
}
+131
View File
@@ -0,0 +1,131 @@
package watcher
import (
"context"
"log/slog"
"time"
"sneak.berlin/go/dnswatcher/internal/config"
"sneak.berlin/go/dnswatcher/internal/resolver"
"sneak.berlin/go/dnswatcher/internal/state"
)
// NewForTest creates a Watcher without fx for unit testing. A nil cfg
// is an empty configuration.
func NewForTest(
cfg *config.Config,
st *state.State,
res DNSResolver,
pc PortChecker,
tc TLSChecker,
n Notifier,
) *Watcher {
if cfg == nil {
cfg = &config.Config{}
}
return &Watcher{
log: slog.Default(),
config: cfg,
state: st,
resolver: res,
portCheck: pc,
tlsCheck: tc,
notify: n,
firstRun: true,
}
}
// SetLogger replaces the watcher's logger, so a test can read what it
// logs.
func (w *Watcher) SetLogger(log *slog.Logger) {
w.log = log
}
// NewlyDisagreeingPairs exports newlyDisagreeingPairs for testing.
func NewlyDisagreeingPairs(
prev, current *state.HostnameState,
) [][2]string {
return newlyDisagreeingPairs(prev, current)
}
// SetFirstRun sets whether the watcher is on its first check, in which
// nothing is compared with the previous check. NewForTest's watcher is.
func (w *Watcher) SetFirstRun(firstRun bool) {
w.firstRun = firstRun
}
// UpdateHostnameState exports updateHostnameState for testing.
func (w *Watcher) UpdateHostnameState(
ctx context.Context,
hostname string,
newState *state.HostnameState,
) {
w.updateHostnameState(ctx, hostname, newState)
}
// DetectHostnameChanges exports detectHostnameChanges for testing.
func (w *Watcher) DetectHostnameChanges(
ctx context.Context,
hostname string,
prev, current *state.HostnameState,
) {
w.detectHostnameChanges(ctx, hostname, prev, current)
}
// ResolveNameserverAddresses exports resolveNameserverAddresses for
// testing.
func (w *Watcher) ResolveNameserverAddresses(
ctx context.Context,
nameservers []string,
prev map[string][]string,
) map[string][]string {
return w.resolveNameserverAddresses(ctx, nameservers, prev)
}
// ResolveCNAMEAddresses exports resolveCNAMEAddresses for testing.
func (w *Watcher) ResolveCNAMEAddresses(
ctx context.Context,
hostname string,
current, prev *state.HostnameState,
) {
w.resolveCNAMEAddresses(ctx, hostname, current, prev)
}
// DetectNSAddressChanges exports detectNSAddressChanges for testing.
func (w *Watcher) DetectNSAddressChanges(
ctx context.Context,
domain string,
prev, current map[string][]string,
) {
w.detectNSAddressChanges(ctx, domain, prev, current)
}
// MaybeSendTestNotification exports maybeSendTestNotification for
// testing.
func (w *Watcher) MaybeSendTestNotification(ctx context.Context) {
w.maybeSendTestNotification(ctx)
}
// CleanupRemovedTargets exports cleanupRemovedTargets for testing.
func (w *Watcher) CleanupRemovedTargets() {
w.cleanupRemovedTargets()
}
// CheckAllPorts exports checkAllPorts for testing.
func (w *Watcher) CheckAllPorts(ctx context.Context) {
w.checkAllPorts(ctx)
}
// RunTLSChecks exports runTLSChecks for testing.
func (w *Watcher) RunTLSChecks(ctx context.Context) {
w.runTLSChecks(ctx)
}
// BuildHostnameState exports buildHostnameState for testing.
func BuildHostnameState(
results map[string]*resolver.NameserverResponse,
now time.Time,
) *state.HostnameState {
return buildHostnameState(results, now)
}
+240
View File
@@ -0,0 +1,240 @@
package watcher_test
import (
"slices"
"testing"
"sneak.berlin/go/dnswatcher/internal/state"
"sneak.berlin/go/dnswatcher/internal/watcher"
)
const (
host = "www.example.net"
nsA = "a.ns.example.net."
nsB = "b.ns.example.net."
nsC = "c.ns.example.net."
ip1 = "192.0.2.1"
ip2 = "192.0.2.2"
ip3 = "192.0.2.3"
)
// hostnameState builds the state a check with these records leaves behind.
func hostnameState(
records map[string]map[string][]string,
) *state.HostnameState {
hs := &state.HostnameState{
RecordsByNameserver: make(map[string]*state.NameserverRecordState),
}
for ns, recs := range records {
hs.RecordsByNameserver[ns] = &state.NameserverRecordState{
Records: recs,
Status: "ok",
}
}
return hs
}
func TestNewlyDisagreeingPairs(t *testing.T) {
t.Parallel()
onlyA := map[string]map[string][]string{nsA: {"A": {ip1}}}
agree := map[string]map[string][]string{nsA: {"A": {ip1}}, nsB: {"A": {ip1}}}
disagree := map[string]map[string][]string{nsA: {"A": {ip1}}, nsB: {"A": {ip2}}}
alert := [][2]string{{nsA, nsB}}
// b already disagrees with a and c; then c changes, so a and c,
// which agreed, now differ.
bDiffers := map[string]map[string][]string{
nsA: {"A": {ip1}}, nsB: {"A": {ip2}}, nsC: {"A": {ip1}},
}
cChanges := map[string]map[string][]string{
nsA: {"A": {ip1}}, nsB: {"A": {ip2}}, nsC: {"A": {ip3}},
}
// Each case starts from the state loaded at startup and runs the
// checks in order; want[i] is what check i alerts for.
tests := []struct {
name string
loaded map[string]map[string][]string
checks []map[string]map[string][]string
want [][][2]string
}{
{
name: "disagreement persisting across checks alerts once",
loaded: agree,
checks: []map[string]map[string][]string{disagree, disagree, disagree},
want: [][][2]string{alert, nil, nil},
},
{
name: "disagreement starting on a later check alerts on it",
loaded: agree,
checks: []map[string]map[string][]string{agree, agree, disagree},
want: [][][2]string{nil, nil, alert},
},
{
name: "disagreement in the loaded state does not alert",
loaded: disagree,
checks: []map[string]map[string][]string{disagree, disagree},
want: [][][2]string{nil, nil},
},
{
name: "nameserver new on the first check and disagreeing alerts once",
loaded: onlyA,
checks: []map[string]map[string][]string{disagree, disagree},
want: [][][2]string{alert, nil},
},
{
name: "disagreement after agreeing again alerts again",
loaded: agree,
checks: []map[string]map[string][]string{disagree, agree, disagree},
want: [][][2]string{alert, nil, alert},
},
{
name: "new disagreement while another nameserver differs alerts",
loaded: bDiffers,
checks: []map[string]map[string][]string{cChanges, cChanges},
want: [][][2]string{{{nsA, nsC}}, nil},
},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
t.Parallel()
prev := hostnameState(tt.loaded)
for i, records := range tt.checks {
current := hostnameState(records)
got := watcher.NewlyDisagreeingPairs(prev, current)
if !slices.Equal(got, tt.want[i]) {
t.Errorf(
"check %d: alerted for %v, want %v",
i, got, tt.want[i],
)
}
prev = current
}
})
}
}
func TestInconsistencyAlert(t *testing.T) {
t.Parallel()
onlyA := map[string]map[string][]string{nsA: {"A": {ip1}}}
agree := map[string]map[string][]string{nsA: {"A": {ip1}}, nsB: {"A": {ip1}}}
disagree := map[string]map[string][]string{nsA: {"A": {ip1}}, nsB: {"A": {ip2}}}
// Each case starts from the state loaded at startup and then sees
// the nameservers disagree on three checks in a row.
tests := []struct {
name string
loaded map[string]map[string][]string
want int
}{
{
name: "disagreement lasting several checks alerts once",
loaded: agree,
want: 1,
},
{
name: "disagreement in the loaded state does not alert",
loaded: disagree,
want: 0,
},
{
name: "nameserver new on the first check and disagreeing alerts once",
loaded: onlyA,
want: 1,
},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
t.Parallel()
// The hostname change detection uses only the notifier.
notifier := &mockNotifier{}
w := watcher.NewForTest(nil, nil, nil, nil, nil, notifier)
prev := hostnameState(tt.loaded)
for range 3 {
current := hostnameState(disagree)
w.DetectHostnameChanges(t.Context(), host, prev, current)
prev = current
}
got := 0
for _, n := range notifier.getNotifications() {
if n.Title == "Inconsistency: "+host {
got++
}
}
if got != tt.want {
t.Errorf("sent %d inconsistency alerts, want %d", got, tt.want)
}
})
}
}
// TestFirstCheckAfterRepeatedValuesLoaded saves a state file holding a
// hostname's CNAME once for every record type asked for, as checks did
// before each value was stored once, and two addresses each repeated,
// and loads it. A check that then finds each value once at each
// nameserver must notify nothing.
func TestFirstCheckAfterRepeatedValuesLoaded(t *testing.T) {
t.Parallel()
const (
cnameType = "CNAME"
cname = "c.example.net."
)
cfg := defaultTestConfig(t)
repeated := map[string][]string{
"A": {ip2, ip1, ip2, ip1},
cnameType: {cname, cname, cname, cname, cname, cname, cname, cname},
}
once := map[string][]string{"A": {ip1, ip2}, cnameType: {cname}}
saved := newTestDeps(t, cfg).state
saved.SetHostnameState(host, hostnameState(map[string]map[string][]string{
nsA: repeated, nsB: repeated,
}))
err := saved.Save()
if err != nil {
t.Fatalf("saving the state file: %v", err)
}
deps := newTestDeps(t, cfg)
err = deps.state.Load()
if err != nil {
t.Fatalf("loading the state file: %v", err)
}
prev, ok := deps.state.GetHostnameState(host)
if !ok {
t.Fatal("the state file has no state for " + host)
}
current := hostnameState(map[string]map[string][]string{
nsA: once, nsB: once,
})
// The hostname change detection uses only the notifier.
w := watcher.NewForTest(nil, nil, nil, nil, nil, deps.notifier)
w.DetectHostnameChanges(t.Context(), host, prev, current)
if got := deps.notifier.getNotifications(); len(got) != 0 {
t.Errorf("sent %v, want no notification", got)
}
}
+3 -2
View File
@@ -5,6 +5,7 @@ import (
"context"
"sneak.berlin/go/dnswatcher/internal/portcheck"
"sneak.berlin/go/dnswatcher/internal/resolver"
"sneak.berlin/go/dnswatcher/internal/tlscheck"
)
@@ -17,11 +18,11 @@ type DNSResolver interface {
) ([]string, error)
// LookupAllRecords queries all record types for a hostname,
// returning results keyed by nameserver then record type.
// returning each nameserver's response keyed by nameserver.
LookupAllRecords(
ctx context.Context,
hostname string,
) (map[string]map[string][]string, error)
) (map[string]*resolver.NameserverResponse, error)
// ResolveIPAddresses resolves a hostname to all IP addresses.
ResolveIPAddresses(
+215
View File
@@ -0,0 +1,215 @@
package watcher_test
import (
"maps"
"strings"
"testing"
"sneak.berlin/go/dnswatcher/internal/config"
"sneak.berlin/go/dnswatcher/internal/state"
"sneak.berlin/go/dnswatcher/internal/watcher"
)
// When one nameserver's A record changes and its TXT record does not,
// the record change and the inconsistency it starts name the A record
// alone, with its values written as plain text.
func TestChangeMessagesNameTheChangedType(t *testing.T) {
t.Parallel()
// A nameserver's records: this A address and the same TXT record.
records := func(address string) map[string][]string {
return map[string][]string{
"A": {address},
"TXT": {"v=spf1 -all"},
}
}
before := hostnameState(map[string]map[string][]string{
nsA: records(ip1),
nsB: records(ip1),
})
after := hostnameState(map[string]map[string][]string{
nsA: records(ip1),
nsB: records(ip2),
})
// The hostname change detection uses only the notifier.
notifier := &mockNotifier{}
w := watcher.NewForTest(nil, nil, nil, nil, nil, notifier)
w.DetectHostnameChanges(t.Context(), host, before, after)
want := map[string]string{
"Record Change: " + host: `Hostname: www.example.net
Nameserver: b.ns.example.net.
Type: A
Old: 192.0.2.1
New: 192.0.2.2`,
"Inconsistency: " + host: `Hostname: www.example.net
Type: A
a.ns.example.net.: 192.0.2.1
b.ns.example.net.: 192.0.2.2`,
}
notifications := notifier.getNotifications()
if len(notifications) != len(want) {
t.Fatalf(
"sent %d notifications, want %d: %v",
len(notifications), len(want), notifications,
)
}
for _, n := range notifications {
if n.Message != want[n.Title] {
t.Errorf(
"%s message:\n%s\nwant:\n%s",
n.Title, n.Message, want[n.Title],
)
}
}
}
// Every kind of notification about a configured apex domain's own
// records names it as a domain.
func TestDomainRecordNotificationsNameTheDomain(t *testing.T) {
t.Parallel()
notifier := &mockNotifier{}
w := watcher.NewForTest(
&config.Config{Domains: []string{domain}},
nil, nil, nil, nil, notifier,
)
// nsA's address changes, which also makes it differ from nsC; nsB
// fails; nsC answers again; nsD is gone.
nsD := "d.ns.example.net."
w.DetectHostnameChanges(t.Context(), domain,
saved(map[string]*state.NameserverRecordState{
nsA: answered(map[string][]string{"A": {ip1}}),
nsB: answered(map[string][]string{"A": {ip1}}),
nsC: failed(),
nsD: answered(map[string][]string{"A": {ip1}}),
}),
saved(map[string]*state.NameserverRecordState{
nsA: answered(map[string][]string{"A": {ip2}}),
nsB: failed(),
nsC: answered(map[string][]string{"A": {ip1}}),
}),
)
// The address at the end of its CNAME chain changes.
w.DetectHostnameChanges(
t.Context(), domain, cnameState(ip1), cnameState(ip2),
)
// NS Failure is sent for nsB failing and for nsD being gone.
want := map[string]int{
"Record Change": 1,
"Inconsistency": 1,
"NS Failure": 2,
"NS Recovery": 1,
"CNAME Address Change": 1,
}
sent := make(map[string]int)
for _, n := range notifier.getNotifications() {
kind, _, _ := strings.Cut(n.Title, ":")
sent[kind]++
if !strings.HasPrefix(n.Message, "Domain: "+domain+"\n") {
t.Errorf("%s message does not name the domain:\n%s",
n.Title, n.Message)
}
}
if !maps.Equal(sent, want) {
t.Errorf("sent %v, want %v", sent, want)
}
}
// The startup notification counts the configured domains and hostnames,
// although the state's hostnames also hold the apex domain's own
// records. Nothing is looked up: the watcher has no resolver.
func TestStartupNotificationCountsConfiguredNames(t *testing.T) {
t.Parallel()
cfg := defaultTestConfig(t)
cfg.SendTestNotification = true
cfg.Domains = []string{domain}
cfg.Hostnames = []string{host}
deps := newTestDeps(t, cfg)
w := watcher.NewForTest(cfg, deps.state, nil, nil, nil, deps.notifier)
// The state a check of both names saves.
deps.state.SetDomainState(domain, &state.DomainState{
Nameservers: []string{nsA},
})
for _, name := range []string{domain, host} {
deps.state.SetHostnameState(name, saved(
map[string]*state.NameserverRecordState{
nsA: answered(map[string][]string{"A": {ip1}}),
},
))
}
w.MaybeSendTestNotification(t.Context())
notifications := deps.notifier.getNotifications()
counts := "\nMonitoring 1 domain(s) and 1 hostname(s).\n"
if len(notifications) != 1 ||
!strings.Contains(notifications[0].Message, counts) {
t.Errorf("sent %v, want one message with %q", notifications, counts)
}
}
// A Port Change notification lists the configured apex domain and the
// hostname that resolve to the port's address on separate lines. The
// port checks read the saved hostname state and look nothing up, so the
// watcher has no resolver.
func TestPortChangeListsDomainsApartFromHostnames(t *testing.T) {
t.Parallel()
cfg := defaultTestConfig(t)
cfg.Domains = []string{domain}
cfg.Hostnames = []string{host}
deps := newTestDeps(t, cfg)
w := watcher.NewForTest(
cfg, deps.state, nil,
deps.portChecker, deps.tlsChecker, deps.notifier,
)
w.SetFirstRun(false)
// Both names resolve to ip1, whose port 443 the previous check
// found open. It is closed now.
for _, name := range []string{domain, host} {
deps.state.SetHostnameState(name, saved(
map[string]*state.NameserverRecordState{
nsA: answered(map[string][]string{"A": {ip1}}),
},
))
}
key := ip1 + ":443"
deps.state.SetPortState(key, &state.PortState{
Open: true, Hostnames: []string{domain, host},
})
deps.portChecker.closed = true
w.CheckAllPorts(t.Context())
title := "Port Change: " + key
want := `Domains: example.net
Hostnames: www.example.net
Address: 192.0.2.1:443
Port now closed`
got := deps.notifier.getNotifications()
if len(got) != 1 || got[0].Title != title || got[0].Message != want {
t.Errorf("sent %v, want one %q with message:\n%s", got, title, want)
}
}
+151
View File
@@ -0,0 +1,151 @@
package watcher_test
import (
"context"
"log/slog"
"reflect"
"testing"
"sneak.berlin/go/dnswatcher/internal/livednstest"
"sneak.berlin/go/dnswatcher/internal/resolver"
"sneak.berlin/go/dnswatcher/internal/watcher"
)
const domain = "example.net"
func TestNSAddressChangeAlerts(t *testing.T) {
t.Parallel()
// Each case is the nameserver addresses saved by the previous check
// and by the current one.
tests := []struct {
name string
prev, current map[string][]string
want int
}{
{
"same addresses",
map[string][]string{nsA: {ip1, ip2}},
map[string][]string{nsA: {ip1, ip2}},
0,
},
{
"same addresses in another order",
map[string][]string{nsA: {ip2, ip1}},
map[string][]string{nsA: {ip1, ip2}},
0,
},
{
"address replaced",
map[string][]string{nsA: {ip1}},
map[string][]string{nsA: {ip2}},
1,
},
{
"address added",
map[string][]string{nsA: {ip1}},
map[string][]string{nsA: {ip1, ip2}},
1,
},
{
"two nameservers changed",
map[string][]string{nsA: {ip1}, nsB: {ip2}},
map[string][]string{nsA: {ip3}, nsB: {ip3}},
2,
},
{
"nameserver added",
map[string][]string{nsA: {ip1}},
map[string][]string{nsA: {ip1}, nsB: {ip2}},
0,
},
{
"nameserver removed",
map[string][]string{nsA: {ip1}, nsB: {ip2}},
map[string][]string{nsA: {ip1}},
0,
},
{
"state file from before addresses were saved",
nil,
map[string][]string{nsA: {ip1}, nsB: {ip2}},
0,
},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
t.Parallel()
notifier := &mockNotifier{}
w := watcher.NewForTest(nil, nil, nil, nil, nil, notifier)
w.DetectNSAddressChanges(t.Context(), domain, tt.prev, tt.current)
got := len(notifier.getNotifications())
if got != tt.want {
t.Errorf("sent %d address changes, want %d", got, tt.want)
}
})
}
}
func TestNSAddressChangeAlertNamesDomainNameserverAndAddresses(
t *testing.T,
) {
t.Parallel()
notifier := &mockNotifier{}
w := watcher.NewForTest(nil, nil, nil, nil, nil, notifier)
w.DetectNSAddressChanges(
t.Context(), domain,
map[string][]string{nsA: {ip1}},
map[string][]string{nsA: {ip2, ip3}},
)
want := notification{
Title: "NS Address Change: " + domain,
Message: "Domain: " + domain + "\nNameserver: " + nsA +
"\nOld: " + ip1 + "\nNew: " + ip2 + ", " + ip3,
Priority: "warning",
}
got := notifier.getNotifications()
if len(got) != 1 || got[0] != want {
t.Errorf("sent %v, want %v", got, want)
}
}
// TestNameserverWithNoAddressKeepsPrevious looks up nameserver names
// with no address: two under .invalid, whose lookup fails with an
// error, and one that does not exist under a real zone, which live DNS
// answers with no address and no error. Each one with addresses saved
// by the previous check keeps them; the one without gets none.
func TestNameserverWithNoAddressKeepsPrevious(t *testing.T) {
t.Parallel()
w := watcher.NewForTest(
nil, nil, resolver.NewFromLogger(slog.Default()), nil, nil, nil,
)
nonexistentNS := "this-surely-does-not-exist-xyz." + testSmallDomain + "."
prev := map[string][]string{oldNS1: {oldIP}, nonexistentNS: {oldIP}}
var got map[string][]string
// The result is the same whether or not live DNS answers, so the
// lookup is not retried.
_ = livednstest.Run(func(ctx context.Context) error {
got = w.ResolveNameserverAddresses(
ctx, []string{oldNS1, oldNS2, nonexistentNS}, prev,
)
return nil
})
if !reflect.DeepEqual(got, prev) {
t.Errorf("saved %v, want %v", got, prev)
}
}
+448
View File
@@ -0,0 +1,448 @@
package watcher_test
import (
"context"
"fmt"
"log/slog"
"strings"
"testing"
"time"
"sneak.berlin/go/dnswatcher/internal/livednstest"
"sneak.berlin/go/dnswatcher/internal/resolver"
"sneak.berlin/go/dnswatcher/internal/state"
"sneak.berlin/go/dnswatcher/internal/watcher"
)
// answered is what a check saves for a nameserver that answered with
// these records.
func answered(records map[string][]string) *state.NameserverRecordState {
return &state.NameserverRecordState{Records: records, Status: "ok"}
}
// failed is what a check saves for a nameserver that did not answer.
func failed() *state.NameserverRecordState {
return &state.NameserverRecordState{
Records: map[string][]string{},
Status: "error",
Error: "all queries timed out",
}
}
// saved builds the hostname state a check saves.
func saved(
byNameserver map[string]*state.NameserverRecordState,
) *state.HostnameState {
return &state.HostnameState{RecordsByNameserver: byNameserver}
}
// alertCounts counts the hostname alerts sent, by kind.
type alertCounts struct {
failures, recoveries, recordChanges, inconsistencies int
}
// countAlerts runs the hostname change detection from the state loaded
// at startup through each check in turn, and counts the alerts sent.
func countAlerts(
t *testing.T,
loaded *state.HostnameState,
checks []*state.HostnameState,
) alertCounts {
t.Helper()
// The hostname change detection uses only the notifier.
notifier := &mockNotifier{}
w := watcher.NewForTest(nil, nil, nil, nil, nil, notifier)
prev := loaded
for _, current := range checks {
w.DetectHostnameChanges(t.Context(), host, prev, current)
prev = current
}
var got alertCounts
for _, n := range notifier.getNotifications() {
kind, _, _ := strings.Cut(n.Title, ":")
switch kind {
case "NS Failure":
got.failures++
case "NS Recovery":
got.recoveries++
case "Record Change":
got.recordChanges++
case "Inconsistency":
got.inconsistencies++
}
}
return got
}
func TestNSFailureAndRecoveryAlerts(t *testing.T) {
t.Parallel()
records := map[string][]string{"A": {ip1}}
bothAnswer := saved(map[string]*state.NameserverRecordState{
nsA: answered(records), nsB: answered(records),
})
bFails := saved(map[string]*state.NameserverRecordState{
nsA: answered(records), nsB: failed(),
})
onlyA := saved(map[string]*state.NameserverRecordState{
nsA: answered(records),
})
bAnswersNoRecords := saved(map[string]*state.NameserverRecordState{
nsA: answered(records), nsB: answered(map[string][]string{}),
})
bAnswersDifferently := saved(map[string]*state.NameserverRecordState{
nsA: answered(records), nsB: answered(map[string][]string{"A": {ip2}}),
})
// Each case starts from the state loaded at startup and runs the
// checks in order.
tests := []struct {
name string
loaded *state.HostnameState
checks []*state.HostnameState
want alertCounts
}{
{
"failure lasting several checks alerts once",
bothAnswer, []*state.HostnameState{bFails, bFails, bFails},
alertCounts{failures: 1},
},
{
"recovery alerts once",
bFails, []*state.HostnameState{bothAnswer, bothAnswer},
alertCounts{recoveries: 1},
},
{
"failing again after recovering alerts again",
bothAnswer, []*state.HostnameState{bFails, bothAnswer, bFails},
alertCounts{failures: 2, recoveries: 1},
},
{
"nameserver failing when first seen does not alert",
onlyA, []*state.HostnameState{bFails, bFails},
alertCounts{},
},
{
"answer with no records is a record change, not a failure",
bothAnswer, []*state.HostnameState{bAnswersNoRecords},
alertCounts{recordChanges: 1, inconsistencies: 1},
},
{
"recovered nameserver that answers differently disagrees",
bFails, []*state.HostnameState{bAnswersDifferently},
alertCounts{recoveries: 1, inconsistencies: 1},
},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
t.Parallel()
got := countAlerts(t, tt.loaded, tt.checks)
if got != tt.want {
t.Errorf("sent %+v, want %+v", got, tt.want)
}
})
}
}
func TestNSFailureAlertNamesHostnameNameserverAndReason(t *testing.T) {
t.Parallel()
records := map[string][]string{"A": {ip1}}
notifier := &mockNotifier{}
w := watcher.NewForTest(nil, nil, nil, nil, nil, notifier)
w.DetectHostnameChanges(
t.Context(), host,
saved(map[string]*state.NameserverRecordState{nsA: answered(records)}),
saved(map[string]*state.NameserverRecordState{nsA: failed()}),
)
notifications := notifier.getNotifications()
if len(notifications) != 1 {
t.Fatalf("sent %v, want one NS Failure", notifications)
}
msg := notifications[0].Message
if !strings.Contains(msg, host) || !strings.Contains(msg, nsA) ||
!strings.Contains(msg, failed().Error) {
t.Errorf(
"message %q does not name %s, %s and the reason",
msg, host, nsA,
)
}
}
// TestNameserverThatNeverAnswers asks a nameserver address where
// nothing answers, 192.0.2.1, and checks what the watcher saves for it.
// The deadline outlasts the resolver's first two-second try, as in the
// resolver's timeout test.
func TestNameserverThatNeverAnswers(t *testing.T) {
t.Parallel()
ctx, cancel := context.WithTimeout(t.Context(), 3*time.Second)
t.Cleanup(cancel)
res := resolver.NewFromLogger(slog.Default())
resp, err := res.QueryNameserverIP(ctx, nsA, "192.0.2.1", host)
if err != nil {
t.Fatal(err)
}
hs := watcher.BuildHostnameState(
map[string]*resolver.NameserverResponse{nsA: resp}, time.Now(),
)
got := hs.RecordsByNameserver[nsA]
if got.Status != "error" || got.Error == "" {
t.Errorf(
"saved status %q, error %q; want status error with a reason",
got.Status, got.Error,
)
}
}
// TestNameserverThatAnswersNXDOMAIN asks a real nameserver about a name
// that does not exist and checks what the watcher saves for it: NXDOMAIN
// is an answer, so the nameserver is saved as ok with no error.
func TestNameserverThatAnswersNXDOMAIN(t *testing.T) {
t.Parallel()
res := resolver.NewFromLogger(slog.Default())
name := "this-surely-does-not-exist-xyz." + testDomain
var (
ns string
resp *resolver.NameserverResponse
)
livednstest.Retry(t, "QueryNameserver("+name+")", func(ctx context.Context) error {
nameservers, err := res.LookupNS(ctx, testDomain)
if err != nil {
return err
}
ns = nameservers[0]
resp, err = res.QueryNameserver(ctx, ns, name)
if err != nil {
return err
}
// A timeout or a failure is no answer to check.
if resp.Status == resolver.StatusTimeout ||
resp.Status == resolver.StatusError {
return fmt.Errorf(
"%w: %s: %s", livednstest.ErrNoAnswer, ns, resp.Error,
)
}
return nil
})
if resp.Status != resolver.StatusNXDomain {
t.Fatalf("%s answered %q for %s, want NXDOMAIN", ns, resp.Status, name)
}
hs := watcher.BuildHostnameState(
map[string]*resolver.NameserverResponse{ns: resp}, time.Now(),
)
got := hs.RecordsByNameserver[ns]
if got.Status != "ok" || got.Error != "" {
t.Errorf(
"saved status %q, error %q; want status ok with no error",
got.Status, got.Error,
)
}
}
// TestNameserverThatRefuses asks a google.com nameserver about
// cloudflare.com, a zone it does not serve, which it refuses, and checks
// what the watcher saves for it: REFUSED is no answer, so the nameserver
// is saved as error with the reason.
func TestNameserverThatRefuses(t *testing.T) {
t.Parallel()
const reason = "server returned REFUSED"
res := resolver.NewFromLogger(slog.Default())
var (
ns string
resp *resolver.NameserverResponse
)
livednstest.Retry(
t,
"QueryNameserver(cloudflare.com)",
func(ctx context.Context) error {
nameservers, err := res.LookupNS(ctx, testDomain)
if err != nil {
return err
}
ns = nameservers[0]
resp, err = res.QueryNameserver(ctx, ns, "cloudflare.com")
if err != nil {
return err
}
// A timeout or a network error is no reply at all.
if resp.Status == resolver.StatusTimeout ||
strings.HasPrefix(resp.Error, "network error") {
return fmt.Errorf(
"%w: %s: %s", livednstest.ErrNoAnswer, ns, resp.Error,
)
}
return nil
},
)
if resp.Error != reason {
t.Fatalf(
"%s answered %q (%s) for cloudflare.com, want REFUSED",
ns, resp.Status, resp.Error,
)
}
hs := watcher.BuildHostnameState(
map[string]*resolver.NameserverResponse{ns: resp}, time.Now(),
)
got := hs.RecordsByNameserver[ns]
if got.Status != failed().Status || got.Error != reason {
t.Errorf(
"saved status %q, error %q; want status %q, error %q",
got.Status, got.Error, failed().Status, reason,
)
}
}
// TestPortStateWhenNoNameserverAnswered runs the port checks on
// hostname state built here, which gives the name no address. The port
// state saved for its old address is kept only when the name is a
// configured hostname or domain and none of its nameservers answered.
func TestPortStateWhenNoNameserverAnswered(t *testing.T) {
t.Parallel()
noneAnswered := saved(map[string]*state.NameserverRecordState{
nsA: failed(), nsB: failed(),
})
oneAnsweredNoAddress := saved(map[string]*state.NameserverRecordState{
nsA: answered(map[string][]string{}), nsB: failed(),
})
configured := []string{host}
tests := []struct {
name string
hostname *state.HostnameState
hostnames []string
domains []string
wantKept bool
}{
{"no nameserver answered", noneAnswered, configured, nil, true},
{
"no nameserver answered, configured as a domain",
noneAnswered, nil, configured, true,
},
{
"one answered with no address",
oneAnsweredNoAddress, configured, nil, false,
},
{"no nameserver answered, not configured", noneAnswered, nil, nil, false},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
t.Parallel()
cfg := defaultTestConfig(t)
cfg.Hostnames = tt.hostnames
cfg.Domains = tt.domains
// The port checks read the saved hostname state and look
// nothing up, so the watcher has no resolver.
deps := newTestDeps(t, cfg)
w := watcher.NewForTest(
cfg, deps.state, nil,
deps.portChecker, deps.tlsChecker, deps.notifier,
)
key := ip1 + ":443"
deps.state.SetHostnameState(host, tt.hostname)
deps.state.SetPortState(key, &state.PortState{
Open: true, Hostnames: []string{host},
})
w.CheckAllPorts(t.Context())
_, kept := deps.state.GetPortState(key)
if kept != tt.wantKept {
t.Errorf("port state %s kept: %v, want %v", key, kept, tt.wantKept)
}
})
}
}
// TestPortStateWhenNoNameserverAnsweredAndOtherNameMovesAway saves the
// port state of an address two configured hostnames resolve to. While
// none of the first one's nameservers answer, the port checks run with
// the other one still at that address, then after it moved away; the
// port state is kept both times.
func TestPortStateWhenNoNameserverAnsweredAndOtherNameMovesAway(
t *testing.T,
) {
t.Parallel()
const other = "mail.example.net"
cfg := defaultTestConfig(t)
cfg.Hostnames = []string{host, other}
// The port checks read the saved hostname state and look nothing
// up, so the watcher has no resolver.
deps := newTestDeps(t, cfg)
w := watcher.NewForTest(
cfg, deps.state, nil,
deps.portChecker, deps.tlsChecker, deps.notifier,
)
key := ip1 + ":443"
deps.state.SetPortState(key, &state.PortState{
Open: true, Hostnames: []string{host, other},
})
deps.state.SetHostnameState(host, saved(
map[string]*state.NameserverRecordState{nsA: failed(), nsB: failed()},
))
for _, otherIP := range []string{ip1, ip2} {
deps.state.SetHostnameState(other, saved(
map[string]*state.NameserverRecordState{
nsA: answered(map[string][]string{"A": {otherIP}}),
},
))
w.CheckAllPorts(t.Context())
if _, kept := deps.state.GetPortState(key); !kept {
t.Fatalf("port state %s removed with %s at %s", key, other, otherIP)
}
}
}
+162
View File
@@ -0,0 +1,162 @@
package watcher_test
import (
"maps"
"slices"
"testing"
"sneak.berlin/go/dnswatcher/internal/state"
"sneak.berlin/go/dnswatcher/internal/watcher"
)
// TestRemovedTargetsLeaveTheState loads a state saved while a domain
// and a hostname now removed from the configuration were still in it,
// and runs the removal that Run does before the first check. The
// removed names' domain, hostname and certificate entries are gone,
// the configured names' are kept, and nothing is notified. Nothing is
// looked up: the watcher has no resolver.
func TestRemovedTargetsLeaveTheState(t *testing.T) {
t.Parallel()
const (
removedDomain = "example.com"
removedHost = "www.example.com"
)
cfg := defaultTestConfig(t)
cfg.Domains = []string{domain}
cfg.Hostnames = []string{host}
deps := newTestDeps(t, cfg)
w := watcher.NewForTest(
cfg, deps.state, nil,
deps.portChecker, deps.tlsChecker, deps.notifier,
)
// The state a check of all four names saves, each name at ip1.
for _, name := range []string{domain, removedDomain} {
deps.state.SetDomainState(name, &state.DomainState{
Nameservers: []string{nsA},
})
}
for _, name := range []string{domain, host, removedDomain, removedHost} {
deps.state.SetHostnameState(name, saved(
map[string]*state.NameserverRecordState{
nsA: answered(map[string][]string{"A": {ip1}}),
},
))
deps.state.SetCertificateState(
ip1+":443:"+name, &state.CertificateState{Status: "ok"},
)
}
err := deps.state.Save()
if err != nil {
t.Fatalf("saving the state: %v", err)
}
err = deps.state.Load()
if err != nil {
t.Fatalf("loading the state: %v", err)
}
w.CleanupRemovedTargets()
snap := deps.state.GetSnapshot()
got := slices.Sorted(maps.Keys(snap.Domains))
if want := []string{domain}; !slices.Equal(got, want) {
t.Errorf("domain entries %v, want %v", got, want)
}
got = slices.Sorted(maps.Keys(snap.Hostnames))
if want := []string{domain, host}; !slices.Equal(got, want) {
t.Errorf("hostname entries %v, want %v", got, want)
}
got = slices.Sorted(maps.Keys(snap.Certificates))
if want := []string{
ip1 + ":443:" + domain, ip1 + ":443:" + host,
}; !slices.Equal(got, want) {
t.Errorf("certificate entries %v, want %v", got, want)
}
if sent := deps.notifier.getNotifications(); len(sent) != 0 {
t.Errorf("sent %v, want nothing", sent)
}
}
// TestCertificateStateForAnAddressGone runs the port checks on hostname
// state built here for a configured hostname, with certificate entries
// saved for it at ip1, ip2 and an IPv6 address. When its nameservers
// answered with ip1 and the IPv6 address, the entry for ip2 is removed.
// When none of them answered, its addresses are not known, and every
// entry is kept. Nothing is notified, and nothing is looked up.
func TestCertificateStateForAnAddressGone(t *testing.T) {
t.Parallel()
const ip6 = "2001:db8::1"
tests := []struct {
name string
hostname *state.HostnameState
want []string
}{
{
"answered without ip2",
saved(map[string]*state.NameserverRecordState{
nsA: answered(map[string][]string{
"A": {ip1}, "AAAA": {ip6},
}),
}),
[]string{ip1 + ":443:" + host, ip6 + ":443:" + host},
},
{
"no nameserver answered",
saved(map[string]*state.NameserverRecordState{
nsA: failed(), nsB: failed(),
}),
[]string{
ip1 + ":443:" + host,
ip2 + ":443:" + host,
ip6 + ":443:" + host,
},
},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
t.Parallel()
cfg := defaultTestConfig(t)
cfg.Hostnames = []string{host}
deps := newTestDeps(t, cfg)
w := watcher.NewForTest(
cfg, deps.state, nil,
deps.portChecker, deps.tlsChecker, deps.notifier,
)
w.SetFirstRun(false)
deps.state.SetHostnameState(host, tt.hostname)
for _, ip := range []string{ip1, ip2, ip6} {
deps.state.SetCertificateState(
ip+":443:"+host, &state.CertificateState{Status: "ok"},
)
}
w.CheckAllPorts(t.Context())
got := slices.Sorted(maps.Keys(deps.state.GetSnapshot().Certificates))
if !slices.Equal(got, tt.want) {
t.Errorf("certificate entries %v, want %v", got, tt.want)
}
if sent := deps.notifier.getNotifications(); len(sent) != 0 {
t.Errorf("sent %v, want nothing", sent)
}
})
}
}
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
+9
View File
@@ -0,0 +1,9 @@
{
"name": "dnswatcher-tooling",
"version": "0.0.0",
"private": true,
"description": "Pins the prettier that script/fmt and script/fmt-check run against this repo's markdown. Not a JavaScript project; nothing here is imported, published, or shipped.",
"devDependencies": {
"prettier": "3.9.6"
}
}
+13 -12
View File
@@ -3,18 +3,16 @@
# this repo. Idempotent: every install is guarded by a check so already
# installed tools are skipped. Base tooling comes from nix, apt, brew,
# or apk (detected in that order); assumes nothing is present.
# golangci-lint and goimports are installed via `go install` at the same
# pinned commits the Dockerfile uses (never "latest").
# goimports is not installed here: script/fmt and script/fmt-check-go
# run it with `go run` at a pinned commit.
# The linter is NOT installed here: golangci-lint runs via docker only
# (script/lint), pinned by image digest, so its only prerequisite is a
# working docker. Nor is prettier: script/fmt and
# script/fmt-check-markdown run it in a container from Dockerfile.fmt.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
# Pinned versions, 2026-08-07 (same pins as the Dockerfile)
# golangci-lint v2.12.2
GOLANGCI_LINT_REF="github.com/golangci/golangci-lint/v2/cmd/golangci-lint@c0d3ddc9cf3faa61a4e378e879ece580256d76e5"
# goimports v0.42.0
GOIMPORTS_REF="golang.org/x/tools/cmd/goimports@009367f5c17a8d4c45a961a3a509277190a9a6f0"
PKGMGR=""
SUDO=""
APT_UPDATED=""
@@ -69,10 +67,13 @@ main() {
if missing make; then pkg_install gnumake make make make; fi
if missing go; then pkg_install go golang go go; fi
# Lint/format tools, pinned via go install (installs into
# "$(go env GOPATH)/bin"; ensure that is on your PATH).
if missing golangci-lint; then go install "$GOLANGCI_LINT_REF"; fi
if missing goimports; then go install "$GOIMPORTS_REF"; fi
# Linting and the markdown formatter run via docker only. Warn,
# don't fail: building and testing work without it.
if missing docker; then
echo "bootstrap: WARNING: docker not found; install it to" \
"run make lint, make fmt, make fmt-check, make check" \
"and make docker." >&2
fi
go mod download
+14 -9
View File
@@ -1,18 +1,23 @@
#!/bin/sh
# script/cibuild: run the CI build. The Dockerfile runs make check, and
# the CHECK_EPOCH build argument below is fresh on every invocation, so
# the check layer is never served from the Docker layer cache: a
# successful build means the checks were executed and passed on this
# run, not on some earlier one. Only the check step and the steps after
# it are invalidated; the toolchain install and go mod download stay
# cached.
# script/cibuild: run the CI build. The Dockerfile's lint stage runs
# the Go half of make fmt-check and golangci-lint; its builder stage
# runs make test and make build. The markdown half of make fmt-check
# runs after that build, as its own build of Dockerfile.fmt, because
# there is no docker inside a docker build.
#
# --no-cache-filter=lint,builder runs both stages on every invocation;
# otherwise an unchanged tree is served from the layer cache and passes
# without linting or querying live DNS. script/fmt-check-markdown busts
# its own cache the same way.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd -P)"
ROOT="$(cd "$SCRIPT_DIR/.." && pwd -P)"
main() {
cd "$ROOT"
docker build --build-arg CHECK_EPOCH="$(date +%s%N)" .
docker build --no-cache-filter=lint,builder .
"$SCRIPT_DIR/fmt-check-markdown"
}
main "$@"
+14 -2
View File
@@ -1,6 +1,10 @@
#!/bin/sh
# script/docker: build the Docker image tagged with the project name.
# Identical in all repos; the tag comes from script/projectname.
# The tag comes from script/projectname.
#
# --no-cache-filter=lint,builder runs the lint stage and the builder
# stage (make test) on every invocation; otherwise an unchanged tree is
# served from the layer cache without linting or querying live DNS.
set -eu
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd -P)"
@@ -8,7 +12,15 @@ ROOT="$(cd "$SCRIPT_DIR/.." && pwd -P)"
main() {
cd "$ROOT"
docker build -t "$("$SCRIPT_DIR/projectname")" .
# Own line: a failing command substitution inside an argument does
# not trip `set -e`, so the inline form degrades silently to an
# empty constant. The VERSION build arg takes precedence over what
# the build would derive from the .git in its context.
version="$(git describe --tags --always --dirty 2>/dev/null || true)"
[ -n "$version" ] || version="unknown"
docker build --no-cache-filter=lint,builder \
--build-arg VERSION="$version" \
-t "$("$SCRIPT_DIR/projectname")" .
}
main "$@"
+54 -2
View File
@@ -1,13 +1,65 @@
#!/bin/sh
# script/fmt: format all files (writes).
# script/fmt: format all files (writes). Go with gofmt and goimports on
# the host, markdown with the prettier pinned by Dockerfile.fmt.
#
# goimports runs with `go run` at a pinned commit, never from PATH, so
# every machine formats with the same version and nothing installs it.
#
# The markdown pass is a `docker build --output type=local` rather than a
# `docker run -v`, so it needs no bind mount and behaves the same against
# a remote daemon; the formatted documents come back out of the build and
# are copied over the tree here.
#
# Unlike script/fmt-check-markdown this does not bust the cache: it is
# not a gate, and any edit to a document changes the COPY layer above the
# prettier step, so a cached result is a result over this exact tree.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
# goimports v0.42.0, 2026-08-07. Must match script/fmt-check-go.
GOIMPORTS_REF="golang.org/x/tools/cmd/goimports@009367f5c17a8d4c45a961a3a509277190a9a6f0"
# Must match the export stage name in Dockerfile.fmt.
stage=fmt-out
die() {
echo "script/fmt: $*" >&2
exit 1
}
main() {
cd "$ROOT"
gofmt -s -w .
goimports -w .
go run "$GOIMPORTS_REF" -w .
tmp="$(mktemp -d "${TMPDIR:-/tmp}/dnswatcher-fmt.XXXXXX")"
trap 'rm -rf "$tmp"' EXIT INT TERM
docker build \
--target "$stage" \
--output "type=local,dest=$tmp/out" \
-f Dockerfile.fmt .
# An empty export means prettier was handed nothing, which must not
# read as "already formatted".
(cd "$tmp/out" && find . -type f -name '*.md') |
sed 's|^\./||' | LC_ALL=C sort >"$tmp/files"
[ -s "$tmp/files" ] ||
die "the formatting build produced no markdown; the build" \
"context reached prettier empty"
# Copied only where the bytes differ, so an already-formatted tree
# keeps its timestamps and says nothing.
while IFS= read -r f; do
[ -n "$f" ] || continue
if [ -f "$f" ] && cmp -s "$tmp/out/$f" "$f"; then
continue
fi
cp "$tmp/out/$f" "$f"
echo "prettier: reformatted $f"
done <"$tmp/files"
}
main "$@"
+6 -10
View File
@@ -1,18 +1,14 @@
#!/bin/sh
# script/fmt-check: check formatting (read-only). Same scope as
# script/fmt, but fails instead of writing.
# script/fmt-check: check formatting (read-only). Same tools and scope
# as script/fmt, but fails instead of writing: the Go on the host, the
# markdown with prettier in a container.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd -P)"
main() {
cd "$ROOT"
files="$(gofmt -l .)"
if [ -n "$files" ]; then
echo "gofmt: files not formatted:" >&2
echo "$files" >&2
exit 1
fi
"$SCRIPT_DIR/fmt-check-go"
"$SCRIPT_DIR/fmt-check-markdown"
}
main "$@"
+31
View File
@@ -0,0 +1,31 @@
#!/bin/sh
# script/fmt-check-go: fail unless every Go source is formatted the way
# script/fmt would leave it, and name the files that are not. Read-only.
#
# Its own script because the Dockerfile's lint stage runs this half
# alone: there is no docker inside a docker build to run the markdown
# half in.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
# goimports v0.42.0, 2026-08-07. Must match script/fmt.
GOIMPORTS_REF="golang.org/x/tools/cmd/goimports@009367f5c17a8d4c45a961a3a509277190a9a6f0"
main() {
cd "$ROOT"
files="$(gofmt -s -l .)"
if [ -n "$files" ]; then
echo "gofmt: files not formatted:" >&2
echo "$files" >&2
exit 1
fi
files="$(go run "$GOIMPORTS_REF" -l .)"
if [ -n "$files" ]; then
echo "goimports: files not formatted:" >&2
echo "$files" >&2
exit 1
fi
}
main "$@"
+29
View File
@@ -0,0 +1,29 @@
#!/bin/sh
# script/fmt-check-markdown: fail unless every .md is formatted the way
# script/fmt would leave it. Read-only.
#
# prettier is never installed on the host: it runs in a container built
# from Dockerfile.fmt, pinned by package.json and yarn.lock.
# --no-cache-filter is here for the reason script/lint gives: a cached
# build checks nothing.
#
# Its own script because script/cibuild runs this half alone, after the
# Dockerfile's lint stage has checked the Go.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
# Must match the markdown check stage name in Dockerfile.fmt.
stage=fmt-check
main() {
cd "$ROOT"
docker build \
--progress=plain \
--no-cache-filter="$stage" \
--target "$stage" \
-f Dockerfile.fmt \
.
}
main "$@"
+14 -1
View File
@@ -7,7 +7,20 @@ ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
main() {
cd "$ROOT"
hook=".git/hooks/pre-commit"
# Stop if this directory is not the top of its own git checkout, for
# example a copy inside another repository, whose hook must not be
# replaced.
if [ "$(git rev-parse --show-toplevel)" != "$ROOT" ]; then
echo "install-precommit: $ROOT is not the top of a git checkout" >&2
exit 1
fi
# Ask git for the repository's own git directory: .git is a file, not
# a directory, in some checkouts (for example a clone made with
# --separate-git-dir). core.hooksPath is deliberately not followed, so
# the hook is never written outside this repository.
hooks="$(git rev-parse --git-common-dir)/hooks"
mkdir -p "$hooks"
hook="$hooks/pre-commit"
printf '#!/bin/sh\nset -e\nscript/precommit\n' > "$hook"
chmod +x "$hook"
echo "pre-commit hook installed: runs script/precommit"
+18 -2
View File
@@ -1,12 +1,28 @@
#!/bin/sh
# script/lint: run the linter.
# script/lint: run the linter. golangci-lint is never installed or run
# on the host: it runs via docker only, one way, everywhere. This
# builds Dockerfile.lint, which COPYs the repo into the digest-pinned
# golangci-lint image and lints as a build step, so a successful build
# means a clean lint.
#
# --no-cache-filter=lint forces the lint stage (source copy + linter
# run) to execute on every invocation. Without it an unchanged tree
# returns success in well under a second having linted nothing. The
# deps stage (base image + go mod download) stays cached, and no global
# cache invalidation is performed. --progress=plain keeps the linter's
# own output visible.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
main() {
cd "$ROOT"
golangci-lint run --config .golangci.yml ./...
docker build \
--progress=plain \
--no-cache-filter=lint \
--target lint \
-f Dockerfile.lint \
.
}
main "$@"
+24 -1
View File
@@ -1,12 +1,35 @@
#!/bin/sh
# script/test: run the test suite.
#
# -count=1 disables Go's test cache, and is load-bearing here. This
# suite queries live DNS on every run by policy (TESTING.md); a cached
# result is a replay of an earlier run's output with no query made at
# all. On an unchanged tree the whole suite would return success in
# under a second having resolved nothing, which makes the repeated-run
# green that is used as evidence for flakiness fixes worthless. Do not
# remove it.
#
# Conditional verbose rerun per REPO_POLICIES.md: run quiet first so
# CI and docker build logs stay readable, and rerun with -v only on
# failure. The rerun also carries -count=1 (a cached replay of the
# failure would show nothing new), and the exit status is forced to 1
# no matter how the rerun ends: the first failure already proved the
# suite broken, so a flaky test that passes the second time must not
# turn the build green.
#
# -timeout 90s is a deliberate backstop above the 60s hard cap on
# suite duration. Do not lower it.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
main() {
cd "$ROOT"
go test -v -race -timeout 30s -cover ./...
go test -count=1 -race -timeout 90s -cover ./... || {
echo "--- Rerunning with -v for details ---" >&2
go test -count=1 -race -timeout 90s -v ./... || true
exit 1
}
}
main "$@"
+8
View File
@@ -0,0 +1,8 @@
# THIS IS AN AUTOGENERATED FILE. DO NOT EDIT THIS FILE DIRECTLY.
# yarn lockfile v1
prettier@3.9.6:
version "3.9.6"
resolved "https://registry.yarnpkg.com/prettier/-/prettier-3.9.6.tgz#b3ea5146515d40fc53f18aa63f74dfab1e10dbf6"
integrity sha512-OpN0zzVdiaiAhxpuuj5efpIS4sY9j7bY6uR5mnj5yPzGkdkjNKSJeUThPb60Jw29QuAZgA4o+/iB49kFiaBX6g==