At startup, before the first check, Run removes from the loaded state
the domain, hostname and certificate entries of names no longer in
DNSWATCHER_TARGETS, takes those names off each port entry's list of
names and removes a port entry left with none, so the dashboard,
/api/v1/status and the startup notification count only configured
names. A configured domain's own records, saved as a hostname entry
under its name, are kept. Nothing is notified. Each port check, next to
the removal of stale port entries, now also removes the certificate
entries for an address a name no longer resolves to, except while none
of its nameservers answered, as port entries already were.
Model: opus-5-5
The resolver lists in FailedTypes each record type whose query to a
nameserver got no usable reply (no reply, a code other than NOERROR or
NXDOMAIN, a referral, or a truncated reply whose TCP retry failed) and
logs it unless shutdown cut it short. A nameserver that answered no
type has failed.
The watcher saves such a type in failedTypes with the previous check's
records, leaves it out of the comparison with other nameservers on that
check, and compares it with the next answer. When the previous check
did not know its records either, it is also in unknownTypes and not
compared until it answers. A nameserver whose A, AAAA or CNAME query
failed is no answer when following a CNAME or resolving addresses.
Model: opus-5-5
A Port Change notification's `Hosts:` line listed a port's apex domains
and hostnames together. It now has a `Domains:` line and a `Hostnames:`
line, and leaves out one that would name nothing. A name is a domain
when it is a configured domain, the rule record notifications already
use to start `Domain:` or `Hostname:`; that rule is now one method,
isDomain, which both use. The port entries saved in the state are
unchanged. README describes the new lines.
Model: opus-5-5
A port entry in the state saves the apex domains that resolve to its
address with its hostnames. The dashboard's Ports table now has a
Domains column next to Hostnames, and a port entry in /api/v1/status
has a `domains` list, with `hostnames` no longer holding a domain. A
name is taken as a domain when it has a domain entry, as the dashboard
and API already tell a domain's own records from a hostname's. Both
read the split from one function, buildPorts. The state file is
unchanged; README says its port `hostnames` include domains.
Model: opus-5-5
Looking up a nameserver's own address followed only the addresses a
referral gave, so a nameserver whose zone is delegated without them,
such as a.ntpns.org of pool.ntp.org, never resolved. The walk to a
name's nameservers looked addresses up only when a referral gave none.
Both now go through queryZone, which asks the nameservers whose
addresses the referral gives first and, if none of them gives a usable
reply, looks up and asks the others. maxLookupDepth stops lookups three
deep, so delegations that point at each other still end; when the limit
is why no address was found, the error is ErrLookupDepthExceeded, not
"no address".
Model: opus-5-5
An apex domain's own records are still saved with the hostnames'
records, under the domain's name, so the port and TLS checks find its
addresses. Notifications about them now start `Domain:`, decided by the
configured domains. The dashboard and /api/v1/status, which read only
the saved state, take a hostname entry whose name also has a domain
entry as that domain's own records: the dashboard shows them in a second
table under Domains, the API in the domain's `recordsByNameserver`, and
neither lists or counts them as hostnames. The startup notification
counts domains and hostnames from the configuration. README says which
of a domain's own records are watched and how their changes are
notified.
Model: opus-5-5
/api/v1/status now gives `error` for each nameserver entry and
certificate entry whose status is `error`, copied from the state, which
already kept it. The dashboard shows that reason in place of the records
for a failed nameserver, which used to show the same `-` as one that
answered with no records, and across the CN, issuer and expiry cells
for a failed certificate, wrapped at a width of 20rem so the long TLS
error does not narrow the Endpoint column. The dashboard stylesheet is a
trimmed build, so the new markup uses only classes the page already had.
README Web Dashboard and HTTP API say so.
Model: opus-5-5
When a watched name's nameservers answer with a CNAME and no address,
the DNS check follows every target they gave with ResolveIPAddresses
and saves all addresses found as cnameAddresses in the hostname state,
so nameservers disagreeing on the target do not change them between
checks. Port and TLS checks use them. A change, also from or to none,
is notified as a CNAME address change; the first check from a state
file without them sends none. When a target cannot be followed, or
none of the name's nameservers answered, the last check's addresses
are kept. The domain check now runs the hostname check for the apex
instead of a copy of it.
Model: opus-5-5
A plain `docker build .`, which is how upaas builds, stamped `dev`:
`.dockerignore` left out `.git` and the builder declared
`ARG VERSION=dev`. `.dockerignore` now sends `.git` without
`.git/config`, which can hold a credential, and lists no tracked file,
which git would count as deleted. `ARG VERSION` has no default. The
Makefile takes a non-empty `VERSION` from the command line or the
environment, so a build arg still wins; otherwise `git describe` runs in
the builder, which trusts the checkout whoever owns it, as a context
sent as a tar archive keeps its owners. A new `make version` prints the
version; the build fails when the context carries `.git` and it comes
out empty, `dev` or `unknown`.
Model: opus-5-5
Every resolution walked the root servers in a fixed order, so
a.root-servers.net got every first query and its timeouts were paid on
every lookup. Each list of servers the resolver walks is now walked in a
random order from rand.Shuffle, chosen anew each time; a server that does
not reply, refuses, or gives an error reply or a referral that leads no
closer is still passed over for the next. When a referral names a zone's
nameservers without their addresses, all of them are now looked up, not
only the first that resolves, so the zone is not given up because the
first nameserver whose address was found gave no usable reply. No test
fails if the walk stops shuffling: which server a live query reached is
not observable.
Model: opus-5-5
Checked every README claim against the code on next and fixed the ones
that were wrong or missing: what /metrics serves and when, what
DNSWATCHER_MAINTENANCE_MODE does, CORS on the public routes,
notification retries and the in-memory alert history, the certificate
error field and old port entries in the state file, and the Design
tree's missing files. Also corrected: CNAMEs are not followed for
watched names, the root server list is never refreshed, the NS set is
the delegation from the domain's parent zone, notification contents,
and the system resolver being used for webhooks. Code problems found
are filed separately.
Model: opus-5-5
ClassifyTargets kept a name every time it appeared in DNSWATCHER_TARGETS,
so example.com,Example.com. put example.com in the domain list twice and
every check looked it up and checked its certificates twice, sending two
expiry warnings per TLS check (port checks were already grouped by address
and port). It now skips a name it has already kept, comparing after the
lower-casing and trailing-dot removal it already did; the list keeps the
order of first appearance.
Model: opus-5-5
REPO_POLICIES.md requires Getting Started, Rationale, Design and TODO
sections in the README, and it had none of them. Getting Started clones
the repository, builds the image and runs it watching example.com and
www.example.com; DNSWATCHER_TARGETS is the only setting it requires.
Rationale is drawn from what the README already says. The Architecture
section moves below Entrypoints and is renamed Design, its text
unchanged, so the required sections come in policy order. TODO points
to TODO.md and the 1.0 milestone instead of copying the list.
Model: opus-5-5
The port check removed the saved port state of every address no
configured name resolves to. A name whose nameservers all timed out
or failed is saved with no records, so its addresses looked gone and
lost their port state; when the nameservers answered again it was
recorded afresh, and a port that opened or closed meanwhile was not
notified.
An entry is now kept when one of the names saved on it is configured
and none of its nameservers answered on its last check, and such a
name stays on the entry when the port is checked again for another
name. A name whose nameservers answer with no addresses still loses
it, and so does a name no longer configured.
Model: opus-5-5
ResolveIPAddresses now returns an error, not no addresses, when no
nameserver of the name's zone answered. A nameserver with status
timeout or error is not an answer; one answer, even NXDOMAIN, is
enough for an empty result without an error.
When every server of a zone fails, FindAuthoritativeNameservers moves
on to the parent name, whose servers only refer the query onward. Such
a referral now has status error, so it is no answer either, and a
hostname's saved records show it as error. The only caller, the
nameserver address lookup, already keeps the previous addresses on an
error; its comment no longer says the resolver hides this case.
Model: opus-5-5
make fmt and make fmt-check now cover every Markdown file with prettier
(4-space tabs, proseWrap always), as template-app-go does: prettier is
pinned by package.json and yarn.lock and runs in a docker build on a
digest-pinned node image, never on the host. The check is forced to run
with --no-cache-filter, as script/lint is.
script/fmt-check is split into a Go half and a Markdown half because the
Dockerfile lint stage cannot run docker: that stage now runs the Go half
and script/cibuild runs the Markdown half after the build. *.md leaves
.dockerignore so documents reach the build context.
README.md, TESTING.md and TODO.md are reformatted by make fmt; apart from
the README Entrypoints entries, its make fmt line under Building and the
TODO entry, that diff is mechanical.
Model: opus-5-5
script/fmt-check now runs goimports in list mode and fails naming any
file it would change, and checks gofmt with -s, as script/fmt applies
it. Both scripts run goimports with `go run` at the commit that
script/bootstrap used to install, so a goimports on PATH is never used
and bootstrap no longer installs it. The pin is written in both
scripts; change them together. The first run on a machine, and every
Dockerfile lint stage run, downloads and builds goimports.
The Markdown half of the issue (prettier) is not done here: it needs
node in the lint image or a separate build, a decision for the owner.
Model: opus-5-5
A hostname's nameservers came from its last two labels, so a name under
co.uk was asked at the co.uk servers and a name in a delegated subdomain
at the parent's servers; both only refer onward. The hostname now goes
through FindAuthoritativeNameservers, which follows delegations for the
name and walks up its labels until it finds the zone it is in.
followDelegation now stops at an authoritative reply: that server holds
the zone, so its reply is not a referral. Without this, a CNAME answer
that also lists the zone's NS records in its authority section, as many
servers send, was followed as a referral until the delegation limit.
Model: opus-5-5
Each domain check now looks up the addresses every nameserver's name
resolves to, with the resolver's ResolveIPAddresses, and saves them
sorted in the domain's state. A nameserver that stays in the
delegation and resolves to different addresses sends one NS Address
Change notification naming the domain, the nameserver and the old and
new addresses. Added or removed nameservers get only the NS change
notification. A failed or empty lookup keeps the previous addresses,
because the resolver returns no address without an error when every
server it asks times out. State files without the field load, and the
next check fills it in silently. Watcher tests that run domain checks
use example.com, which has two nameservers, to stay within the
per-attempt limit.
Model: opus-5-5
The final save at shutdown came from the state's own stop hook, while
the watcher's stop hook only cancelled its run loop, so a check under
way could change state after that save or be cut off at exit. Run now
saves state as it returns, and the watcher's stop hook waits for Run,
bounded by the shutdown deadline. The state's own save stays; Save
holds the state lock for the whole write, so the two cannot overlap.
The start hook derives the watcher's context with WithoutCancel, so
the linter needs no exception. A new test stops a watcher built by New
and reads the change back from the state file, with no DNS.
Model: opus-5-5
DNSWATCHER_SENTRY_DSN was read but never used. This ports the Sentry
integration from gohttpserver with sentry-go v0.49.0. The server's start
hook calls sentry.Init when the DSN is set; a DSN Sentry cannot parse
fails the hook, so startup stops with the parse error. sentryhttp, with
Repanic, reports handler panics and passes them on to chi's Recoverer.
Shutdown sends queued reports once the HTTP server has stopped. The
client uses the older transport (DisableTelemetryBuffer): with the
default one, Flush can return before sending a report made just before
it. Client reports are off, so only panics are sent. The DSN is checked
at server start, not in config, so the config test is unchanged.
sentry-go raises several golang.org/x modules and moves go-spew and
go-difflib to untagged commits.
Model: opus-5-5
When shutdown cancels a check that is under way, the rest of the check
still runs with the cancelled context. The resolver already drops a
lookup the context cut short, but a cancelled connection attempt was
saved as a closed port or a failed certificate check and notified as
Port Change or TLS Failure. The watcher now drops a port or TLS check
result when its context was cancelled, the same way. The test runs a
check with the context already cancelled, using the real resolver and
the real port and TLS checkers; no query is sent and no connection is
made.
Model: opus-5-5
realIP took the first X-Forwarded-For entry, which the client itself
can write, so behind a proxy that appends to the header a client chose
the address dnswatcher logs and the /metrics rate limit counts. It now
walks the entries from the right past trusted proxies, using the
existing trusted-proxy check, and takes the first that is not one; the
leftmost when all are. All X-Forwarded-For header lines are read as one
list, since a proxy may add its own line instead of appending to the
client's. An empty entry where the client address belongs falls back
to the peer address, as an empty first entry did before. X-Real-IP is
unchanged.
Model: opus-5-5
LookupAllRecords now returns each nameserver's response, so the
watcher saves its status: ok when it answered, NXDOMAIN and no records
included, and error with the reason when it timed out, answered
SERVFAIL or REFUSED, or could not be reached. A nameserver that starts
failing sends NS Failure and one that answers again sends NS Recovery.
A failing nameserver is left out of the record change and
inconsistency comparisons. The resolver used to report REFUSED and
network errors as an answer with no records; they are now errors. A
lookup cut short by its context now returns an error instead of a
failure of the nameserver it was querying.
Model: opus-5-5
DNSWATCHER_DNS_INTERVAL and DNSWATCHER_TLS_INTERVAL were parsed with
time.ParseDuration and silently replaced by the default when that failed,
so a value like 5 or 1d gave hourly checks with no hint why, and zero or
negative values were accepted. Both now go through parseInterval, which
returns an error naming the variable and the value, and startup stops the
same way it does for invalid targets. An unset or empty variable still gets
its default from setupViper. The three tests that pinned the old fallback
are replaced. The README says what a valid value looks like.
Model: opus-5-5
/metrics is behind a password, and REPO_POLICIES.md requires rate
limiting on password logins. Each client address may now send it 30
requests a minute, counted by httprate before Basic Auth, so failed
logins use up the allowance and a request over it gets 429 without
the password being checked. The address is the one the existing
trusted-proxy logic in internal/middleware works out, with IPv6
addresses grouped by /64; an IPv4 address a proxy reports in
IPv6-mapped form counts as the plain IPv4 address. A Prometheus
server scraping every 15 seconds sends 4 requests a minute.
Model: opus-5-5
The Dockerfile builder stage takes ARG VERSION (default `dev`) and passes
it to `make build` on the command line, which overrides the Makefile's
`git describe` default. script/docker computes the version from
`git describe` on the host and passes it as --build-arg VERSION, because
.dockerignore leaves .git out of the build context and `git describe`
inside the build only ever produced `dev`. A build that passes no
argument, such as script/cibuild, still reports `dev`.
`logger.Identify`, which logs `starting` with the version, was never
called; `main` now calls it first, so the version is in the startup log.
Model: opus-5-5
The runtime image no longer sets USER. Its new entrypoint,
deploy/docker-entrypoint.sh, runs as root: it creates the data
directory if needed, gives it and everything in it to the dnswatcher
user (uid 10001) with mode 700 on the directory, then runs dnswatcher
as that user with su-exec. A bind-mounted host directory, whether
empty and root-owned or holding a state file left by another uid, no
longer has to be chowned first, and the README's upaas section now says
only which path to mount. The startup check that the data directory is
writable stays.
Model: opus-5-5
The live-DNS retry and concurrency limit is only for tests, but nothing
stopped program code from importing it and compiling it into the
binary. Its directory name now ends in test, and its import path is on
the test-support deny list in .golangci.yml, so make lint fails when
program code imports it. Every import and mention is updated to the new
name.
Model: opus-5-5
The watcher tests used a stand-in resolver and the resolver timeout test
a stand-in DNS client, against the rule that DNS is never mocked.
Watcher tests that look something up in DNS now run the real resolver
against live servers, each attempt on a new watcher. A DNS change is
tested by saving values live DNS never returns (names under .invalid,
192.0.2.1) in the state a check starts from, or by marking a real
nameserver failed. The timeout test queries 192.0.2.1, where nothing
answers. The live-DNS retry and concurrency limit moved from the
resolver tests to internal/livedns, so both packages share them.
NewFromLoggerWithClient had no other use and is gone. TESTING.md now
states the README's rule.
Model: opus-5-5
detectInconsistencies alerted for neighbouring pairs of nameservers whose
records differed, on every DNS check, for as long as they differed. It now
also takes the previous hostname state, compares every pair of nameservers,
and alerts for a pair that differs unless both were in that state and
already differed there, so a nameserver new on a check that answers
differently is reported once. The state loaded at startup is the previous
state for the first check, so a disagreement saved before a restart is not
reported again. The choice of pairs is tested on record data, and the alert
through the hostname change detection with the notifier stand-in and no
resolver. The README describes the new behaviour.
Model: opus-5-5
Nameservers may answer with DNS names in any letter case. For eeqj.de,
y.ns.joker.com answers in upper case while its peers answer in lower
case, so the inconsistency check fired on every cycle. extractRecordValue
now lower-cases CNAME, MX, SRV and NS targets, so the inconsistency check
and the record-change check both compare names regardless of case. A,
AAAA, TXT and CAA values are formatted as before. State saved before this
change can hold upper-case names, which report a one-time record change
on the first check after upgrading.
Model: opus-5-5
script/cibuild and script/docker were plain docker build. On an
unchanged tree the lint stage and the builder stage, which runs make
test, came from the layer cache, so the build passed without linting
or querying live DNS. Both scripts now pass
--no-cache-filter=lint,builder so those stages run on every build, as
script/lint already does for its own lint stage. Dependency downloads
inside those stages re-run each build. Each of the two stages in the
Dockerfile now notes that the scripts name it. README and TODO.md
updated to match.
Model: opus-4-8 (implementation); opus-5-5 (rework)
The runtime image runs as uid 10001, which owns /var/lib/dnswatcher. The
working directory is /, so config loading finds no .env or dnswatcher
config file there; the binary lives in /usr/local/bin. A Docker
HEALTHCHECK probes /.well-known/healthcheck every 10 seconds with busybox
wget, well inside the 60 seconds upaas waits.
Startup now fails with an error naming the data directory when it cannot
be written, instead of running with every save failing. The check creates
the directory if needed and writes and removes the temp file Save uses;
tests cover the create and the write failing.
README gains "Running under upaas": the prod branch, host directory
setup, network and port, environment and health check.
Model: opus-5-5
Adds a SecurityHeaders middleware and registers it globally, right after the request ID middleware, so every route gets the headers, including static files, /metrics and error responses.
It sets Strict-Transport-Security (one year, includeSubDomains), a Content-Security-Policy with default-src 'self', no scripts and frame-ancestors 'none', X-Frame-Options DENY, X-Content-Type-Options nosniff, Referrer-Policy no-referrer and a Permissions-Policy that turns every listed feature off.
HSTS is sent on every response, not only over TLS: the service runs behind a TLS-terminating proxy and REPO_POLICIES.md requires the application to send it. Referrer-Policy is stricter than the policy baseline because dashboard URLs can name internal hosts.
model: claude-opus-4-8 (implementation); claude-fable-5 (commit message)
The http.Server was built with only ReadHeaderTimeout set; the other three
timeouts were zero, which in net/http means no limit, so a peer could hold a
connection open past the header phase, responses had no write deadline, and
keep-alive connections were never reaped. ReadTimeout 15s, WriteTimeout 75s,
and IdleTimeout 120s now join ReadHeaderTimeout 10s as named constants.
WriteTimeout must stay above the 60s chimw.Timeout handler budget, because
net/http arms the write deadline once request headers are read; a test fails
the build if either number moves alone. The server literal moved into
newHTTPServer so the configuration can be asserted without binding a socket
(closes#99)
Model: opus-5
Shutdown waits for deliveries already in flight instead of dropping them.
Reviewed and green; squashed to next by the dispatcher, dnswatcher having no manager.
Model: opus-5
The repository had no licence file at all, which makes publicly readable
code all-rights-reserved by default: nobody may legally use it. That is a
1.0 blocker rather than a nicety, and `LICENSE` was the only file from
`REPO_POLICIES.md`'s required minimum still missing here.
The choice is standing org policy rather than a per-repo call: any public
repo lacking a licence gets MIT, while a private repo with no licence is
already all-rights-reserved and needs nothing. `sneak/dnswatcher` is
public, so MIT.
`LICENSE` carries the canonical MIT text byte-for-byte with only the
copyright line filled in; no clauses added, removed, reworded, or
reflowed. `README.md`'s first line now names the licence, which the
Description requirement in `REPO_POLICIES.md` calls for, and the License
section states MIT and points at the file instead of recording the
decision as pending.
`script/test` did not pass `-count=1`, so on an unchanged tree Go
served the whole suite from its test cache: exit 0 in ~0.2s with every
package marked `(cached)` and not one DNS query made. This repo's suite
exists to exercise live resolution on every run (`TESTING.md`), so that
green asserted nothing — and it is exactly the green used as evidence
that a flakiness fix works, since "run it a few times" stops being
runs after the first.
`-count=1` now disables caching on every invocation.
The conditional verbose rerun that `REPO_POLICIES.md` mandates was
missing at the same spot and is added here rather than left broken: the
primary run had been unconditionally `-v`, which is the failure mode
the policy exists to prevent (unreadable CI and `docker build` logs on
success). Tests now run quiet, and only a failure triggers the `-v`
rerun. The rerun carries `-count=1` too, so it cannot replay a cached
copy of the failure it is meant to diagnose, and its exit status is
discarded in favour of a forced 1: the first failure already proved the
suite broken, so a flake that passes the second time must not turn the
build green.
`-timeout 90s` is untouched. It is a deliberate backstop that must
strictly exceed the 60s hard cap on suite duration.
No special-casing for the Docker build, which also reaches this script
via `RUN make test`: a fresh container's test cache is empty, so
`-count=1` changes nothing there and carving out an exception would
only create a second code path that could drift.
Verified: three back-to-back `make test` runs on an unchanged tree,
zero `(cached)` markers, ~4.0-4.5s wall each (was ~0.2s cached),
comfortably inside the 20s target with `-race` and `-cover` both still
working and coverage percentages unchanged. The rerun-and-still-fail
path was exercised against a purpose-built flaky test that fails once
then passes: quiet failure, verbose rerun that genuinely re-executed,
exit 1 regardless. `make check` green.
golangci-lint is no longer installed or run on the host. script/lint is
now a thin wrapper that builds the new root Dockerfile.lint, which COPYs
the repo into the digest-pinned golangci/golangci-lint:v2.12.2 image and
lints as a build step, so a successful build is a clean lint. This works
even where the docker daemon is remote and bind mounts are impossible.
Dockerfile.lint is split into a deps stage (base image, go mod download)
and a lint stage (source copy, linter run). script/lint passes
--no-cache-filter=lint so the lint stage executes on every invocation:
caching is explicitly waived for linting, and a cached build lints
nothing. The deps stage stays cached and no global cache invalidation is
performed. --progress=plain keeps the linter's own output visible.
golangci-lint config verify is deliberately omitted: it fetches its JSON
schema over a live, unpinned HTTPS call, which would make linting
network-dependent and defeat hash-pinning.
script/bootstrap no longer installs golangci-lint and warns instead when
docker is absent. The goimports install stays, since script/fmt and
script/fmt-check still run it on the host.
The root Dockerfile ran make check in its builder stage, which would now
recurse into script/lint and shell out to docker build with no daemon
available. It gains its own lint stage on the same pinned image, invoked
directly, with the builder depending on it via COPY --from=lint and
running make fmt-check, make test and make build.
Adds a prominent "No DNS mocking. Ever." section near the top of `README.md`, per owner policy (sneak, 2026-08-07):
- DNS is never mocked in this project — no mock resolvers, fake DNS servers, or stubbed lookups, in tests or anywhere else.
- Tests exercise real iterative resolution against live nameservers by design.
- Flaky live tests are fixed with robustness (retries, multiple nameservers, timeouts) or explicit opt-in gating decided by the owner — never with mocks.
- Contributions introducing DNS mocks will be rejected.
Markdown-only change; matches the README's existing tone and hard-wrap style. `script/fmt` covers Go only, so no formatter output applies to this file. Verified via `script/cibuild` (docker build runs `make check` with the pinned toolchain) — green. A direct local `make check` shows 21 pre-existing `goconst` lint findings that come from a newer local `golangci-lint` (v2.12.2 vs the pinned v2.10.1) and are unrelated to this change.
Related: #93 is being reframed under this policy.
Co-authored-by: sneak <sneak@sneak.berlin>
Reviewed-on: #95
Co-authored-by: clawbot <clawbot@noreply.example.org>
Co-committed-by: clawbot <clawbot@noreply.example.org>
Closes [issue #88](#88).
Changes the default `DNSWATCHER_DATA_DIR` from the relative path `./data` to the absolute path `/var/lib/dnswatcher`, following the [Filesystem Hierarchy Standard](https://refspecs.linuxfoundation.org/FHS_3.0/fhs/ch05s08.html) convention for variable application state data.
## Changes
- **`internal/config/config.go`**: Changed the Viper default for `DATA_DIR` from `"./data"` to `"/var/lib/"+name`, where `name` is the application name ("dnswatcher"). This makes the default derived from the app name rather than hardcoded.
- **`internal/config/config_test.go`**: Updated `TestNew_DefaultValues` and `TestStatePath` to expect the new absolute default.
- **`README.md`**: Updated the environment variable table and `.env` example to show `/var/lib/dnswatcher` as the default.
The Dockerfile already set `ENV DNSWATCHER_DATA_DIR=/var/lib/dnswatcher` explicitly, so Docker deployments are unaffected. This change makes the code default consistent with the Docker configuration.
`docker build .` passes all checks (fmt, lint, tests, build).
Co-authored-by: user <user@Mac.lan guest wan>
Reviewed-on: #89
Co-authored-by: clawbot <clawbot@noreply.example.org>
Co-committed-by: clawbot <clawbot@noreply.example.org>
When set to a truthy value, sends a startup status notification to all configured notification channels after the first full scan completes on application startup. The notification is clearly an all-ok/success message showing the number of monitored domains, hostnames, ports, and certificates.
Changes:
- Added `SendTestNotification` config field reading `DNSWATCHER_SEND_TEST_NOTIFICATION`
- Added `maybeSendTestNotification()` in watcher, called after initial `RunOnce` in `Run`
- Added 3 watcher tests (enabled via Run, enabled via RunOnce alone, disabled)
- Added config tests for the new field
- Updated README: env var table, example .env, Docker example
Closes#84
Co-authored-by: user <user@Mac.lan guest wan>
Reviewed-on: #85
Co-authored-by: clawbot <clawbot@noreply.example.org>
Co-committed-by: clawbot <clawbot@noreply.example.org>
## Summary
Adds a read-only web dashboard at `GET /` that shows the current monitoring state and recent alerts. Unauthenticated, single-page, no navigation.
## What it shows
- **Summary bar**: counts of monitored domains, hostnames, ports, certificates
- **Domains**: nameservers with last-checked age
- **Hostnames**: per-nameserver DNS records, status badges, relative age
- **Ports**: open/closed state with associated hostnames and age
- **TLS Certificates**: CN, issuer, expiry (color-coded by urgency), status, age
- **Recent Alerts**: last 100 notifications in reverse chronological order with priority badges
Every data point displays its age (e.g. "5m ago") so freshness is visible at a glance. Auto-refreshes every 30 seconds.
## What it does NOT show
No secrets: webhook URLs, ntfy topics, Slack/Mattermost endpoints, API tokens, and configuration details are never exposed.
## Design
All assets (CSS) are embedded in the binary and served from `/s/`. Zero external HTTP requests at runtime — no CDN dependencies or third-party resources. Dark, technical aesthetic with saturated teals and blues on dark slate. Single page — everything on one screen.
## Implementation
- `internal/notify/history.go` — thread-safe ring buffer (`AlertHistory`) storing last 100 alerts
- `internal/notify/notify.go` — records each alert in history before dispatch; refactored `SendNotification` into smaller `dispatch*` helpers to satisfy funlen
- `internal/handlers/dashboard.go` — `HandleDashboard()` handler with embedded HTML template, helper functions (`relTime`, `formatRecords`, `expiryDays`, `joinStrings`)
- `internal/handlers/templates/dashboard.html` — Tailwind-styled single-page dashboard
- `internal/handlers/handlers.go` — added `State` and `Notify` dependencies via fx
- `internal/server/routes.go` — registered `GET /` route
- `static/` — embedded CSS assets served via `/s/` prefix
- `README.md` — documented the dashboard and new endpoint
## Tests
- `internal/notify/history_test.go` — empty, add+recent ordering, overflow beyond capacity
- `internal/handlers/dashboard_test.go` — `relTime`, `expiryDays`, `formatRecords`
- All existing tests pass unchanged
- `docker build .` passes
closes [#82](#82)
<!-- session: rework-pr-83 -->
Co-authored-by: user <user@Mac.lan guest wan>
Co-authored-by: clawbot <clawbot@noreply.git.eeqj.de>
Reviewed-on: #83
Co-authored-by: clawbot <clawbot@noreply.example.org>
Co-committed-by: clawbot <clawbot@noreply.example.org>
## Summary
Fixes documentation inaccuracies in README.md identified during QA audit.
### Changes
**API table (closes #67):**
- Removed `GET /api/v1/domains` and `GET /api/v1/hostnames` from the HTTP API table. These endpoints are not implemented — the only routes in `internal/server/routes.go` are `/health`, `/api/v1/status`, and `/metrics` (conditional).
**Feature claims (closes #68):**
- Removed "Inconsistency resolved" from hostname monitoring features. `detectInconsistencies()` detects current inconsistencies but has no state tracking to detect when they resolve.
- Removed `nxdomain` and `nodata` from the state status values table. While the resolver defines these constants, `buildHostnameState()` in the watcher only ever sets status to `"ok"`. Failed queries set `"error"` via the NS disappearance path. These values are never written to state.
- Removed "Empty response" (NODATA/NXDOMAIN) detection claim. Changes are caught generically by `detectRecordChanges()`, not with specific NODATA/NXDOMAIN labeling.
### What was NOT changed
- "Inconsistency detected" remains — this IS implemented in `detectInconsistencies()`.
- All other feature claims were verified against the code and are accurate.
- No Go source code was modified.
Co-authored-by: clawbot <clawbot@noreply.git.eeqj.de>
Co-authored-by: Jeffrey Paul <sneak@noreply.example.org>
Reviewed-on: #74
Co-authored-by: clawbot <clawbot@noreply.example.org>
Co-committed-by: clawbot <clawbot@noreply.example.org>
## Summary
When `DNSWATCHER_TARGETS` is empty (the default), dnswatcher previously started successfully and ran indefinitely monitoring nothing. This is a common misconfiguration — forgetting to set the variable or making a typo in its name — and gave no indication anything was wrong.
## Changes
- Added `ErrNoTargets` sentinel error in `internal/config/config.go`
- Extracted `parseAndValidateTargets()` helper to validate that at least one domain or hostname is configured after target classification
- If no targets are configured, dnswatcher now exits with a clear error: `"no monitoring targets configured: set DNSWATCHER_TARGETS environment variable"`
- Updated README.md to document that `DNSWATCHER_TARGETS` is required and dnswatcher will refuse to start without it
## How it works
The validation runs during config construction (via uber/fx), before the watcher or any other component starts. If `DNSWATCHER_TARGETS` is empty or contains only whitespace/empty entries, `buildConfig()` returns `ErrNoTargets`, which causes fx to fail startup with a clear error message.
This is fail-fast behavior: a monitoring daemon with nothing to monitor is a misconfiguration and should not silently run.
Closes #69
Co-authored-by: clawbot <clawbot@noreply.git.eeqj.de>
Reviewed-on: #75
Co-authored-by: clawbot <clawbot@noreply.example.org>
Co-committed-by: clawbot <clawbot@noreply.example.org>
## Summary
Port state keys are `ip:port` with a single `hostname` field. When multiple hostnames resolve to the same IP (shared hosting, CDN), only one hostname was associated. This caused orphaned port state when that hostname removed the IP from DNS while the IP remained valid for other hostnames.
## Changes
### State (`internal/state/state.go`)
- `PortState.Hostname` (string) → `PortState.Hostnames` ([]string)
- Custom `UnmarshalJSON` for backward compatibility: reads old single `hostname` field and migrates to a single-element `hostnames` slice
- Added `DeletePortState` and `GetAllPortKeys` methods for cleanup
### Watcher (`internal/watcher/watcher.go`)
- Refactored `checkAllPorts` into three phases:
1. Build IP:port → hostname associations from current DNS data
2. Check each unique IP:port once with all associated hostnames
3. Clean up stale port state entries with no hostname references
- Port change notifications now list all associated hostnames (`Hosts:` instead of `Host:`)
- Added `buildPortAssociations`, `parsePortKey`, and `cleanupStalePorts` helper functions
### README
- Updated state file format example: `hostname` → `hostnames` (array)
- Updated notification description to reflect multiple hostnames
## Backward Compatibility
Existing state files with the old single `hostname` string are handled gracefully via custom JSON unmarshaling — they are read as single-element `hostnames` slices.
Closes #55
Co-authored-by: clawbot <clawbot@noreply.eeqj.de>
Reviewed-on: #65
Co-authored-by: clawbot <clawbot@noreply.example.org>
Co-committed-by: clawbot <clawbot@noreply.example.org>
## Summary
DNS checks now always complete before port or TLS checks begin, ensuring those checks use freshly resolved IP addresses instead of potentially stale ones from a previous cycle.
## Problem
Port and TLS checks read IP addresses from state that was populated during the most recent DNS check. If DNS changes between cycles, port/TLS checks may target stale IPs. In particular, when the TLS ticker fired (every 12h), it ran `runTLSChecks` without refreshing DNS first — meaning TLS checks could use IPs that were up to 12 hours old.
## Changes
- **Extract `runDNSChecks()`** from the former `runDNSAndPortChecks()` so DNS resolution can be invoked independently as a prerequisite for any check type.
- **TLS ticker now runs DNS first**: When the TLS ticker fires, DNS checks run before TLS checks, ensuring fresh IPs.
- **`RunOnce` uses explicit 3-phase ordering**: DNS → ports → TLS. Port checks must complete before TLS because TLS checks only target IPs where port 443 is open.
- **New test `TestDNSRunsBeforePortAndTLSChecks`**: Verifies that when DNS IPs change between cycles, port and TLS checks pick up the new IPs.
- **README updated**: Monitoring lifecycle section now documents the DNS-first ordering guarantee.
## Check ordering
| Trigger | Phase 1 | Phase 2 | Phase 3 |
|---------|---------|---------|----------|
| Startup (`RunOnce`) | DNS | Ports | TLS |
| DNS ticker | DNS | Ports | — |
| TLS ticker | DNS | — | TLS |
closes #58
Co-authored-by: user <user@Mac.lan guest wan>
Reviewed-on: #64
Co-authored-by: clawbot <clawbot@noreply.example.org>
Co-committed-by: clawbot <clawbot@noreply.example.org>