1 Commits
Author SHA1 Message Date
sneak 4ff5f85320 resolver: try servers in a random order on each resolution (closes #138)
check / check (push) Failing after 2m8s
Every resolution walked the root servers in a fixed order, so
a.root-servers.net got every first query and its timeouts were paid on
every lookup. Each list of servers the resolver walks is now walked in a
random order from rand.Shuffle, chosen anew each time; a server that does
not reply, refuses, or gives an error reply or a referral that leads no
closer is still passed over for the next. When a referral names a zone's
nameservers without their addresses, all of them are now looked up, not
only the first that resolves, so the zone is not given up because one
randomly picked nameserver failed. No test fails if the walk stops
shuffling: which server a live query reached is not observable.

Model: opus-5-5
2026-10-02 00:01:15 +00:00
5 changed files with 54 additions and 142 deletions
+46 -102
View File
@@ -12,7 +12,7 @@ dnswatcher watches configured DNS domains and hostnames for changes, monitors
TCP port availability, tracks TLS certificate expiry, and delivers real-time
notifications via Slack, Mattermost, and/or ntfy webhooks.
It resolves the names it watches itself via iterative (non-recursive) queries,
It performs all DNS resolution itself via iterative (non-recursive) queries,
tracing from root nameservers to authoritative servers directly—never relying on
upstream recursive resolvers.
@@ -69,14 +69,14 @@ notification endpoint set, changes show only on the dashboard; see
- Accepts a list of DNS domain names (apex domains, identified via the
[Public Suffix List](https://publicsuffix.org/)).
- Every **1 hour** by default, performs a full iterative trace from root servers
to discover all authoritative nameservers (NS records) for each domain.
- Every **1 hour**, performs a full iterative trace from root servers to
discover all authoritative nameservers (NS records) for each domain.
- Queries **every** discovered authoritative nameserver independently.
- Stores the domain's NS record set, as its parent zone's servers delegate it,
and the IPv4 and IPv6 addresses each nameserver's name resolves to.
- Stores the NS record set as observed by the delegation chain, and the IPv4 and
IPv6 addresses each nameserver's name resolves to.
- Any change triggers a notification:
- NS added to or removed from that set.
- NS address change: a nameserver that stays in the set resolves to
- NS added to or removed from the delegation.
- NS address change: a nameserver that stays in the delegation resolves to
different addresses than on the previous check. A nameserver added or
removed gets only the NS change notification. When the lookup of a
nameserver's addresses fails or finds none, its previous addresses are
@@ -86,7 +86,7 @@ notification endpoint set, changes show only on the dashboard; see
- Accepts a list of DNS hostnames (subdomains, distinguished from apex domains
via the Public Suffix List).
- Every **1 hour** by default, performs a full iterative trace to discover the
- Every **1 hour**, performs a full iterative trace to discover the
authoritative nameservers of the zone the hostname is in, which is not always
its last two labels (a name under `co.uk`, or in a delegated subdomain).
- Queries **each** authoritative nameserver independently for **all** record
@@ -125,16 +125,13 @@ notification endpoint set, changes show only on the dashboard; see
### TCP Port Monitoring
- For every configured domain and hostname, constructs a deduplicated list of
the IPv4 and IPv6 addresses in the A and AAAA records its authoritative
nameservers returned. A CNAME is not followed: a name whose CNAME points into
another zone usually has no addresses here, so its ports and certificate are
not checked.
all IPv4 and IPv6 addresses resolved via A, AAAA, and CNAME chain resolution
across all authoritative nameservers.
- Checks TCP connectivity on ports **80** and **443** for each IP address.
- Every **1 hour** by default, re-checks all ports.
- Every **1 hour**, re-checks all ports.
- Any change in port availability triggers a notification:
- Port transitioned from open to closed (or vice versa).
- New IP appeared (from DNS change): its port state is recorded without a
port notification; the DNS change notification shows the new address.
- New IP appeared (from DNS change) and its port state was recorded.
- IP disappeared (from DNS change) — noted in the DNS change notification;
port state for that IP is removed. When none of a name's nameservers
answered, its addresses are not known, so the port state saved for them is
@@ -142,14 +139,14 @@ notification endpoint set, changes show only on the dashboard; see
### TLS Certificate Monitoring
- Every **12 hours** by default, for each IP address listening on port 443,
connects via TLS using the correct SNI hostname.
- Every **12 hours**, for each IP address listening on port 443, connects via
TLS using the correct SNI hostname.
- Records the certificate's Subject CN, SANs, issuer, and expiry date.
- Any change triggers a notification:
- Certificate is expiring within **7 days** by default (warning, repeated
each check until renewed or expired).
- Certificate is expiring within **7 days** (warning, repeated each check
until renewed or expired).
- Certificate CN, issuer, or SANs changed (replacement detected, reports old
and new CN and issuer).
and new values).
- TLS connection failure to a previously-reachable IP:443 (handshake error,
timeout, connection refused after previously succeeding).
- TLS recovery: a previously-failing IP:443 now completes a handshake again.
@@ -181,25 +178,15 @@ includes:
- **NS recoveries**: Which nameserver recovered, which hostname/domain.
- **NS inconsistencies**: Which nameservers disagree, what each one returned,
which hostname affected.
- **Port changes**: Which IP:port, its new state, all associated hostnames.
- **TLS expiry warnings**: Expiry date and days remaining, CN, associated
hostname and IP.
- **TLS certificate changes**: Old and new CN and issuer, associated hostname
and IP. A change to the SANs alone is notified, but the SANs are not listed.
- **Port changes**: Which IP:port, old state, new state, all associated
hostnames.
- **TLS expiry warnings**: Which certificate, days remaining, CN, issuer,
associated hostname and IP.
- **TLS certificate changes**: Old and new CN/issuer/SANs, associated hostname
and IP.
- **TLS connection failures/recoveries**: Which IP:port, error details,
associated hostname.
Each endpoint is sent each notification on its own, in the background. A
delivery that fails (a network error, no reply within 10 seconds, or an HTTP
status of 400 or more) is retried up to 5 times: the first retry after about 1
second, each wait after that twice as long up to 60 seconds, every wait varied
at random by up to 25%. A delivery still failing after that is logged and
dropped.
The last 100 notifications, delivered or not, are kept in memory for the
dashboard's Recent alerts. They are not saved to the state file, so a restart
clears them.
### State Management
- All monitoring state is kept in memory and persisted to a JSON file on disk
@@ -243,16 +230,7 @@ dnswatcher exposes a lightweight HTTP API for operational visibility:
| `GET /.well-known/healthcheck` | Health check (JSON) |
| `GET /health` | Health check (JSON, legacy) |
| `GET /api/v1/status` | Current monitoring state |
| `GET /metrics` | Prometheus metrics, see below |
`/metrics` is served only when `DNSWATCHER_METRICS_USERNAME` is set, behind
Basic Auth. It has the Prometheus Go client's default metrics only (Go runtime,
process, and counts of `/metrics` requests); dnswatcher records no metrics of
its own.
Every route but `/metrics` may be read from a page on any origin: a cross-origin
`GET` gets `Access-Control-Allow-Origin: *`. Only `GET` is allowed cross-origin,
and without credentials. `/metrics` sends no CORS headers.
| `GET /metrics` | Prometheus metrics (optional) |
#### Server timeouts
@@ -343,8 +321,8 @@ following precedence (highest to lowest):
| `DNSWATCHER_TLS_INTERVAL` | TLS check interval, a positive duration such as `6h`; empty means the default, anything else stops startup | `12h` |
| `DNSWATCHER_TLS_EXPIRY_WARNING` | Days before expiry to warn | `7` |
| `DNSWATCHER_SENTRY_DSN` | Sentry DSN for error reporting | `""` |
| `DNSWATCHER_MAINTENANCE_MODE` | Only sets `maintenanceMode` in the health check response; changes nothing else | `false` |
| `DNSWATCHER_METRICS_USERNAME` | Basic auth username for /metrics, which is served only when this is set | `""` |
| `DNSWATCHER_MAINTENANCE_MODE` | Enable maintenance mode | `false` |
| `DNSWATCHER_METRICS_USERNAME` | Basic auth username for /metrics | `""` |
| `DNSWATCHER_METRICS_PASSWORD` | Basic auth password for /metrics | `""` |
| `DNSWATCHER_SEND_TEST_NOTIFICATION` | Send a test notification after first scan completes | `false` |
@@ -352,8 +330,7 @@ following precedence (highest to lowest):
monitoring targets are configured. A monitoring daemon with nothing to monitor
is a misconfiguration, so dnswatcher fails fast with a clear error message
rather than running silently. Set `DNSWATCHER_TARGETS` to a comma-separated list
of DNS names before starting. A name listed more than once, in any letter case
or with a trailing dot, is watched once.
of DNS names before starting.
**`/metrics` is rate limited.** Each client address may send it 30 requests a
minute, failed logins included; beyond that it answers `429 Too Many Requests`
@@ -395,15 +372,13 @@ DNSWATCHER_SEND_TEST_NOTIFICATION=true
## DNS Resolution Strategy
dnswatcher never uses the system's configured recursive resolver for the names
it watches. Instead, it performs full iterative resolution:
dnswatcher never uses the system's configured recursive resolver. Instead, it
performs full iterative resolution:
1. **Root servers**: Starts from the IPv4 addresses of the 13 root servers,
built into the binary; the list is not refreshed.
1. **Root servers**: Starts from the IANA root nameserver list (hardcoded, with
periodic refresh).
2. **TLD delegation**: Queries root servers for the TLD NS records.
3. **Domain delegation**: Queries TLD nameservers for the domain's NS records.
The delegation they give, from the domain's parent zone, is the domain's NS
record set.
4. **Authoritative query**: Queries all discovered authoritative nameservers
directly for the requested records.
@@ -419,13 +394,10 @@ This approach ensures:
- Independence from any upstream resolver's cache or filtering.
- Ability to detect split-horizon or inconsistent responses across authoritative
servers.
- Visibility into the full delegation chain.
CNAME chains are followed (with a depth limit to prevent loops) only to find the
addresses of nameservers. A watched name's records are stored as its nameservers
return them, CNAME included, without following it.
Sending a notification or a Sentry report is the one use of the system's
resolver: the HTTP client looks up the webhook's or Sentry's host name with it.
For hostname monitoring, the resolver follows CNAME chains (with a depth limit
to prevent loops) before collecting terminal A/AAAA records.
---
@@ -507,17 +479,12 @@ reachability:
A nameserver that answers NXDOMAIN or with no records has status `ok` and empty
`records`. A nameserver whose query failed, or that only referred it to other
nameservers, has status `error`, empty `records`, and the reason in `error`. A
certificate entry whose TLS connection or handshake failed likewise has status
`error`, the reason in `error`, and the certificate fields left empty or zero.
nameservers, has status `error`, empty `records`, and the reason in `error`.
`nameserverAddresses` lists, by nameserver, the sorted addresses its name
resolves to. A state file without it loads, and the next check fills it in
without a notification.
A port entry in the older format, with one `hostname` instead of the `hostnames`
list, loads as a list of that one name.
---
## Entrypoints
@@ -655,10 +622,9 @@ docker run -d \
- Port checks: every `DNSWATCHER_DNS_INTERVAL`, after DNS completes.
- TLS checks: every `DNSWATCHER_TLS_INTERVAL` (default 12h), after DNS
completes.
- Port and TLS checks use the IP addresses found by the DNS phase that
immediately precedes them. When that phase cannot find a name's
nameservers at all, the addresses an earlier check saved for the name are
used.
- Port and TLS checks always use freshly resolved IP addresses from the DNS
phase that immediately precedes them — never stale IPs from a previous
cycle.
4. **On change detection**: Send notifications to all configured endpoints,
update in-memory state, persist to disk.
5. **Shutdown**: The watcher stops checking and saves the final state to disk,
@@ -707,51 +673,29 @@ file, so it survives a restart without an external database.
cmd/dnswatcher/main.go Entry point (uber/fx bootstrap)
internal/
config/
config.go Viper-based configuration
classify.go Splits targets into domains and hostnames
(Public Suffix List)
config/config.go Viper-based configuration
globals/globals.go Build-time variables (version)
logger/logger.go slog structured logging (TTY detection)
healthcheck/healthcheck.go Health check service
middleware/middleware.go HTTP middleware (logging, CORS, security
middleware/middleware.go HTTP middleware (logging, CORS, security
headers, metrics auth and rate limit)
handlers/
handlers.go Shared handler setup and JSON responses
dashboard.go Web dashboard
templates/dashboard.html Dashboard template (embedded)
status.go /api/v1/status
healthcheck.go Health check handler
handlers/handlers.go HTTP request handlers
server/
server.go HTTP server lifecycle
routes.go Route definitions
state/state.go JSON file state persistence
resolver/
resolver.go Resolver setup and query status values
iterative.go Iterative DNS resolution engine
dns_client.go UDP and TCP DNS clients
errors.go Resolver errors
resolver/resolver.go Iterative DNS resolution engine
portcheck/portcheck.go TCP port connectivity checker
tlscheck/tlscheck.go TLS certificate inspector
notify/
notify.go Notification service (Slack, Mattermost, ntfy)
retry.go Delivery retries with backoff
history.go Last 100 notifications, for the dashboard
shutdown.go Waits for deliveries at shutdown
watcher/
watcher.go Main monitoring orchestrator and scheduler
interfaces.go The resolver, checkers and notifier it uses
tlscheck/tlscheck.go TLS certificate inspector
notify/notify.go Notification service (Slack, Mattermost, ntfy)
watcher/watcher.go Main monitoring orchestrator and scheduler
livednstest/livednstest.go Retry and concurrency limit for tests
against live DNS (imported only by tests)
static/
static.go Embeds the CSS served under /s/
css/tailwind.min.css Dashboard stylesheet
```
### Design Principles
- **No recursive resolvers**: The watched names are resolved iteratively,
- **No recursive resolvers**: All DNS resolution is performed iteratively,
tracing from root nameservers through the delegation chain to authoritative
servers.
- **No external database**: State is persisted as a single JSON file.
+1 -4
View File
@@ -21,10 +21,6 @@ trial run of the finished image: https://git.eeqj.de/sneak/dnswatcher/issues/149
- 2026-10-02: the resolver tries root servers, and every other server list it
walks, in a random order each time, not always from the top (closes #138).
- 2026-10-02: a name listed more than once in `DNSWATCHER_TARGETS`, in any
letter case or with a trailing dot, is watched once (closes #207).
- 2026-10-01: README checked against the code and corrected: metrics, CORS,
notification retries, CNAMEs, state file fields, Design tree (closes #108).
- 2026-10-01: a certificate within the expiry warning period is warned about on
every TLS check, where some checks used to skip it at random (closes #204).
- 2026-10-01: a domain's NS set is its delegation from the parent zone's
@@ -133,4 +129,5 @@ trial run of the finished image: https://git.eeqj.de/sneak/dnswatcher/issues/149
- 1.0 readiness: run it with a real config and read the logs:
https://git.eeqj.de/sneak/dnswatcher/issues/66
- README accuracy sweep: https://git.eeqj.de/sneak/dnswatcher/issues/108
- review toward 1.0: https://git.eeqj.de/sneak/dnswatcher/issues/144
+2 -7
View File
@@ -57,22 +57,17 @@ func ClassifyDNSName(name string) (DNSNameType, error) {
// ClassifyTargets splits a list of DNS names into apex domains and
// hostnames using the Public Suffix List. It returns an error if any
// name cannot be classified. A name given more than once, in any letter
// case or with a trailing dot, is kept once.
// name cannot be classified.
func ClassifyTargets(targets []string) ([]string, []string, error) {
var domains, hostnames []string
seen := make(map[string]bool)
for _, t := range targets {
normalized := strings.ToLower(strings.TrimSuffix(strings.TrimSpace(t), "."))
if normalized == "" || seen[normalized] {
if normalized == "" {
continue
}
seen[normalized] = true
typ, classErr := ClassifyDNSName(normalized)
if classErr != nil {
return nil, nil, classErr
-24
View File
@@ -1,7 +1,6 @@
package config_test
import (
"slices"
"testing"
"sneak.berlin/go/dnswatcher/internal/config"
@@ -94,29 +93,6 @@ func TestClassifyTargets(t *testing.T) {
}
}
func TestClassifyTargetsKeepsEachNameOnce(t *testing.T) {
t.Parallel()
domains, hostnames, err := config.ClassifyTargets([]string{
"example.org",
"Example.org.",
"www.example.org",
"EXAMPLE.ORG",
"WWW.Example.org.",
})
if err != nil {
t.Fatalf("unexpected error: %v", err)
}
if !slices.Equal(domains, []string{"example.org"}) {
t.Errorf("domains = %v, want [example.org]", domains)
}
if !slices.Equal(hostnames, []string{"www.example.org"}) {
t.Errorf("hostnames = %v, want [www.example.org]", hostnames)
}
}
func TestClassifyTargetsRejectsPublicSuffix(t *testing.T) {
t.Parallel()
+5 -5
View File
@@ -9,11 +9,11 @@
//
// 1. Bounded concurrency. Tests run in parallel and the build hosts
// have many cores, so without a limit every test starts its own
// iterative resolution at the same instant and they all send their
// first queries to the root servers within a few milliseconds of
// each other. Root servers rate-limit that, which shows up as a
// different arbitrary subset of tests failing on each run. Run caps
// how many live operations are in flight at once in one test binary.
// iterative resolution at the same instant and they all hit the
// first root server within a few milliseconds of each other. Root
// servers rate-limit that, which shows up as a different arbitrary
// subset of tests failing on each run. Run caps how many live
// operations are in flight at once in one test binary.
//
// 2. Retry with exponential backoff. Each live operation gets several
// attempts with its own timeout. An attempt is retried when it