The resolver lists in FailedTypes each record type whose query to a
nameserver got no usable reply (none after two tries, an error reply, a
referral, or a truncated reply whose TCP retry failed), and logs it with
the reason. A nameserver that answered no type has failed, as before.
The watcher saves such a type in failedTypes, keeping the previous
check's records, leaves it out of the comparison with other nameservers
on that check, and compares the kept records with the next answer. With
nothing to keep, it is also in unknownTypes and not compared until it
answers. Change messages leave out what was not compared. A nameserver
whose A, AAAA or CNAME query failed is no answer when following a CNAME
or resolving addresses.
Model: opus-5-5
A Record Change notification printed the nameserver's whole old and new
record sets in Go map syntax, and an Inconsistency notification the two
nameservers' whole sets, so a one-address change had to be found by eye
among kilobytes of unchanged TXT, CAA, MX and NS values. Both now list,
in sorted order of type, only the record types whose values differ: a
line naming the type, then each side's values separated by commas, or
none when that side has no records of that type. The dashboard's Recent
alerts shows the same text.
Model: opus-5-5
The startup notification ended "All notification channels are working.",
but it is written once and handed to every notification endpoint before
any delivery has succeeded or failed, so the claim was never checked and
was false whenever one endpoint refused it. It now says only that it is a
test sent to every configured notification endpoint. The startup
notification test checks the whole message.
Model: opus-5-5
The JSON log wrote a Go duration as a bare count of nanoseconds, so
the watcher starting line showed dnsInterval 120000000000 for 2m and a
delivery retry showed retryIn 1015437050. Each duration logged is now
passed through its String() form: dnsInterval and tlsInterval when the
watcher starts, retryIn on a delivery retry, and latency on a
succeeded port check. A test checks that retryIn is logged as the
text of the wait the retry actually took.
The request log's latency_ms is left as it is: its key names its
unit.
Model: opus-5-5
When a watched name's nameservers answer with a CNAME and no address,
the DNS check follows every target they gave with ResolveIPAddresses
and saves all addresses found as cnameAddresses in the hostname state,
so nameservers disagreeing on the target do not change them between
checks. Port and TLS checks use them. A change, also from or to none,
is notified as a CNAME address change; the first check from a state
file without them sends none. When a target cannot be followed, or
none of the name's nameservers answered, the last check's addresses
are kept. The domain check now runs the hostname check for the apex
instead of a copy of it.
Model: opus-5-5
An expiry warning was skipped when the last one for that hostname and
address was sent less than DNSWATCHER_TLS_INTERVAL ago. Each TLS check
runs after a DNS pass of varying length, so two checks can be less than
the interval apart, and a certificate about to expire was warned about on
every check or every other check, at random. TLS checks already start
once per interval, so the in-memory record of when each warning was sent
is removed and every check warns, as the README says.
The test that expected the second check to stay silent is replaced by one
that runs TLS checks on state built in the test, with no DNS.
Model: opus-5-5
The port check removed the saved port state of every address no
configured name resolves to. A name whose nameservers all timed out
or failed is saved with no records, so its addresses looked gone and
lost their port state; when the nameservers answered again it was
recorded afresh, and a port that opened or closed meanwhile was not
notified.
An entry is now kept when one of the names saved on it is configured
and none of its nameservers answered on its last check, and such a
name stays on the entry when the port is checked again for another
name. A name whose nameservers answer with no addresses still loses
it, and so does a name no longer configured.
Model: opus-5-5
ResolveIPAddresses now returns an error, not no addresses, when no
nameserver of the name's zone answered. A nameserver with status
timeout or error is not an answer; one answer, even NXDOMAIN, is
enough for an empty result without an error.
When every server of a zone fails, FindAuthoritativeNameservers moves
on to the parent name, whose servers only refer the query onward. Such
a referral now has status error, so it is no answer either, and a
hostname's saved records show it as error. The only caller, the
nameserver address lookup, already keeps the previous addresses on an
error; its comment no longer says the resolver hides this case.
Model: opus-5-5
Each domain check now looks up the addresses every nameserver's name
resolves to, with the resolver's ResolveIPAddresses, and saves them
sorted in the domain's state. A nameserver that stays in the
delegation and resolves to different addresses sends one NS Address
Change notification naming the domain, the nameserver and the old and
new addresses. Added or removed nameservers get only the NS change
notification. A failed or empty lookup keeps the previous addresses,
because the resolver returns no address without an error when every
server it asks times out. State files without the field load, and the
next check fills it in silently. Watcher tests that run domain checks
use example.com, which has two nameservers, to stay within the
per-attempt limit.
Model: opus-5-5
The final save at shutdown came from the state's own stop hook, while
the watcher's stop hook only cancelled its run loop, so a check under
way could change state after that save or be cut off at exit. Run now
saves state as it returns, and the watcher's stop hook waits for Run,
bounded by the shutdown deadline. The state's own save stays; Save
holds the state lock for the whole write, so the two cannot overlap.
The start hook derives the watcher's context with WithoutCancel, so
the linter needs no exception. A new test stops a watcher built by New
and reads the change back from the state file, with no DNS.
Model: opus-5-5
When shutdown cancels a check that is under way, the rest of the check
still runs with the cancelled context. The resolver already drops a
lookup the context cut short, but a cancelled connection attempt was
saved as a closed port or a failed certificate check and notified as
Port Change or TLS Failure. The watcher now drops a port or TLS check
result when its context was cancelled, the same way. The test runs a
check with the context already cancelled, using the real resolver and
the real port and TLS checkers; no query is sent and no connection is
made.
Model: opus-5-5
LookupAllRecords now returns each nameserver's response, so the
watcher saves its status: ok when it answered, NXDOMAIN and no records
included, and error with the reason when it timed out, answered
SERVFAIL or REFUSED, or could not be reached. A nameserver that starts
failing sends NS Failure and one that answers again sends NS Recovery.
A failing nameserver is left out of the record change and
inconsistency comparisons. The resolver used to report REFUSED and
network errors as an answer with no records; they are now errors. A
lookup cut short by its context now returns an error instead of a
failure of the nameserver it was querying.
Model: opus-5-5
state.NewForTest, state.NewForTestWithDataDir and watcher.NewForTest
were compiled into and exported by the production packages
internal/state and internal/watcher. NewForTestWithDataDir and
watcher.NewForTest now live in their package's export_test.go. Tests in
other packages cannot see those files, so the watcher and middleware
tests build their State with state.New and a temporary data directory;
the watcher tests no longer try to save to /state.json.
state.NewForTest, whose State saved to /, is deleted: the state tests
that used it pass t.TempDir() to NewForTestWithDataDir, and the test of
the helper itself is gone.
Model: opus-5-5
detectInconsistencies alerted for neighbouring pairs of nameservers whose
records differed, on every DNS check, for as long as they differed. It now
also takes the previous hostname state, compares every pair of nameservers,
and alerts for a pair that differs unless both were in that state and
already differed there, so a nameserver new on a check that answers
differently is reported once. The state loaded at startup is the previous
state for the first check, so a disagreement saved before a restart is not
reported again. The choice of pairs is tested on record data, and the alert
through the hostname change detection with the notifier stand-in and no
resolver. The README describes the new behaviour.
Model: opus-5-5
Updates golangci-lint to v2.12.2 and sets `.golangci.yml` to the org-standard v2-schema config already deployed across the org's repos. The config change is owner-authorized (see #96 (comment) and #96 (comment)); the same file is being landed as canonical via prompts PR #24 (sneak/prompts#24).
## Changes
- **Commit-pinned installs**: golangci-lint pinned to commit `c0d3ddc9cf3faa61a4e378e879ece580256d76e5` (v2.12.2, released 2026-05-06) in `Dockerfile` and `script/bootstrap`.
- **`.golangci.yml` set to the org-standard v2 config** (sha256 `021cc83f4e6fc7c31b95b34b846723dfcf20b66b7baeea1dc40406e643346bcb`), byte-identical to the file used across the org's other repos. Settings live under `linters.settings`, so the `lll`/`funlen`/`cyclop`/`dupl` thresholds are actually applied (under the old hybrid file, v2 silently ignored the top-level `linters-settings` block).
- **Lint fixes** required by the now-active thresholds:
- `goconst`: shared constants for repeated status/priority/DNS-fixture strings in `internal/watcher/watcher.go` and the notify, state, and watcher tests
- `dupl`: consolidated duplicated ntfy/slack HTTP-error tests and SendNotification endpoint-error tests behind shared helpers in `internal/notify/delivery_test.go`
- `lll`: wrapped long test table entries and comments in `internal/config/classify_test.go`, `internal/notify/history_test.go`, `internal/state/state_test.go`, `internal/watcher/watcher_test.go`; shortened one inline nolint justification in `internal/notify/retry.go`
- **`TODO.md`**: Completed Steps entry updated in the same commit.
- Rebased onto current `main` (`f79cd98`); the branch is one clean commit.
## Notes
- v2.12 deprecates the `gomodguard` linter in favor of `gomodguard_v2`. The org-standard config does not disable the deprecated linter, so golangci-lint may emit an informational deprecation warning; this is accepted by the owner and does not affect the exit status (this exact config+code combination was CI-green at `dea7e44`).
## Verification
- `make check` exits 0 (fmt-check, tests, lint)
- `make lint`: 0 issues; no deprecation warning surfaced in the runs performed
- sha256 of `.golangci.yml` at HEAD verified equal to `021cc83f4e6fc7c31b95b34b846723dfcf20b66b7baeea1dc40406e643346bcb`
Co-authored-by: sneak <sneak@sneak.berlin>
Reviewed-on: #96
Co-authored-by: clawbot <clawbot@noreply.example.org>
Co-committed-by: clawbot <clawbot@noreply.example.org>
When set to a truthy value, sends a startup status notification to all configured notification channels after the first full scan completes on application startup. The notification is clearly an all-ok/success message showing the number of monitored domains, hostnames, ports, and certificates.
Changes:
- Added `SendTestNotification` config field reading `DNSWATCHER_SEND_TEST_NOTIFICATION`
- Added `maybeSendTestNotification()` in watcher, called after initial `RunOnce` in `Run`
- Added 3 watcher tests (enabled via Run, enabled via RunOnce alone, disabled)
- Added config tests for the new field
- Updated README: env var table, example .env, Docker example
Closes#84
Co-authored-by: user <user@Mac.lan guest wan>
Reviewed-on: #85
Co-authored-by: clawbot <clawbot@noreply.example.org>
Co-committed-by: clawbot <clawbot@noreply.example.org>
## Summary
The `OnStart` hook previously derived the watcher's context from the fx startup context (`startCtx`) via `context.WithoutCancel()`. While `WithoutCancel` strips cancellation and deadline, using `context.Background()` makes the intent explicit: the watcher's monitoring loop must outlive the fx startup phase and is controlled solely by the `cancel` func called in `OnStop`.
## Changes
- Replace `context.WithCancel(context.WithoutCancel(startCtx))` with `context.WithCancel(context.Background())`
- Add explanatory comment documenting why the watcher context is not derived from the startup context
- Unused `startCtx` parameter changed to `_`
Closes#53
Co-authored-by: clawbot <clawbot@noreply.git.eeqj.de>
Co-authored-by: Jeffrey Paul <sneak@noreply.example.org>
Reviewed-on: #63
Co-authored-by: clawbot <clawbot@noreply.example.org>
Co-committed-by: clawbot <clawbot@noreply.example.org>
## Summary
Port state keys are `ip:port` with a single `hostname` field. When multiple hostnames resolve to the same IP (shared hosting, CDN), only one hostname was associated. This caused orphaned port state when that hostname removed the IP from DNS while the IP remained valid for other hostnames.
## Changes
### State (`internal/state/state.go`)
- `PortState.Hostname` (string) → `PortState.Hostnames` ([]string)
- Custom `UnmarshalJSON` for backward compatibility: reads old single `hostname` field and migrates to a single-element `hostnames` slice
- Added `DeletePortState` and `GetAllPortKeys` methods for cleanup
### Watcher (`internal/watcher/watcher.go`)
- Refactored `checkAllPorts` into three phases:
1. Build IP:port → hostname associations from current DNS data
2. Check each unique IP:port once with all associated hostnames
3. Clean up stale port state entries with no hostname references
- Port change notifications now list all associated hostnames (`Hosts:` instead of `Host:`)
- Added `buildPortAssociations`, `parsePortKey`, and `cleanupStalePorts` helper functions
### README
- Updated state file format example: `hostname` → `hostnames` (array)
- Updated notification description to reflect multiple hostnames
## Backward Compatibility
Existing state files with the old single `hostname` string are handled gracefully via custom JSON unmarshaling — they are read as single-element `hostnames` slices.
Closes #55
Co-authored-by: clawbot <clawbot@noreply.eeqj.de>
Reviewed-on: #65
Co-authored-by: clawbot <clawbot@noreply.example.org>
Co-committed-by: clawbot <clawbot@noreply.example.org>
## Summary
DNS checks now always complete before port or TLS checks begin, ensuring those checks use freshly resolved IP addresses instead of potentially stale ones from a previous cycle.
## Problem
Port and TLS checks read IP addresses from state that was populated during the most recent DNS check. If DNS changes between cycles, port/TLS checks may target stale IPs. In particular, when the TLS ticker fired (every 12h), it ran `runTLSChecks` without refreshing DNS first — meaning TLS checks could use IPs that were up to 12 hours old.
## Changes
- **Extract `runDNSChecks()`** from the former `runDNSAndPortChecks()` so DNS resolution can be invoked independently as a prerequisite for any check type.
- **TLS ticker now runs DNS first**: When the TLS ticker fires, DNS checks run before TLS checks, ensuring fresh IPs.
- **`RunOnce` uses explicit 3-phase ordering**: DNS → ports → TLS. Port checks must complete before TLS because TLS checks only target IPs where port 443 is open.
- **New test `TestDNSRunsBeforePortAndTLSChecks`**: Verifies that when DNS IPs change between cycles, port and TLS checks pick up the new IPs.
- **README updated**: Monitoring lifecycle section now documents the DNS-first ordering guarantee.
## Check ordering
| Trigger | Phase 1 | Phase 2 | Phase 3 |
|---------|---------|---------|----------|
| Startup (`RunOnce`) | DNS | Ports | TLS |
| DNS ticker | DNS | Ports | — |
| TLS ticker | DNS | — | TLS |
closes #58
Co-authored-by: user <user@Mac.lan guest wan>
Reviewed-on: #64
Co-authored-by: clawbot <clawbot@noreply.example.org>
Co-committed-by: clawbot <clawbot@noreply.example.org>
checkTLSExpiry fired every monitoring cycle with no deduplication,
causing notification spam for expiring certificates. Added an
in-memory map tracking the last notification time per domain/IP
pair, suppressing re-notification within the TLS check interval.
Added TestTLSExpiryWarningDedup to verify deduplication works.
collectIPs only reads HostnameState, but checkDomain only stored
DomainState (nameservers). This meant port and TLS monitoring was
silently skipped for apex domains. Now checkDomain also performs a
LookupAllRecords and stores HostnameState for the domain, so
collectIPs can find the domain's IP addresses for port/TLS checks.
Added TestDomainPortAndTLSChecks to verify the fix.
- CheckPorts now runs all port checks concurrently using errgroup
- Added port number validation (1-65535) with ErrInvalidPort sentinel error
- Updated PortChecker interface to use *PortResult return type
- Added tests for invalid port numbers (0, negative, >65535)
- All checks pass (make check clean)
Implements the full monitoring loop:
- Immediate checks on startup, then periodic DNS+port and TLS cycles
- Domain NS change detection with notifications
- Per-nameserver hostname record tracking with change/failure/recovery
and inconsistency detection
- TCP port 80/443 monitoring with state change notifications
- TLS certificate monitoring with change, expiry, and failure detection
- State persistence after each cycle
- First run establishes baseline without notifications
- Graceful shutdown via context cancellation
Defines DNSResolver, PortChecker, TLSChecker, and Notifier interfaces
for dependency injection. Updates main.go fx wiring and resolver stub
signature to match per-NS record format.
Closes#2
Full project structure following upaas conventions: uber/fx DI, go-chi
routing, slog logging, Viper config. State persisted as JSON file with
per-nameserver record tracking for inconsistency detection. Stub
implementations for resolver, portcheck, tlscheck, and watcher.