check / check (push) Failing after 37s
When a watched name's nameservers answer with a CNAME and no address, the DNS check follows every CNAME target they gave with ResolveIPAddresses and saves the addresses found for all of them in the hostname state as cnameAddresses, so nameservers that disagree on the target do not change them from check to check. The port and TLS checks use them. A change in them is notified as a CNAME address change, also from or to none. A state file without the field loads them as not known (nil), so its first check sends nothing for them. When a target cannot be followed, or none of the name's nameservers answered, the last check's addresses are kept. The domain check now runs the hostname check for the apex instead of a copy of it. Model: opus-5-5
812 lines
39 KiB
Markdown
812 lines
39 KiB
Markdown
# dnswatcher
|
|
|
|
dnswatcher is an MIT-licensed, pre-1.0 Go daemon by
|
|
[@sneak](https://sneak.berlin) that monitors DNS records, TCP port availability,
|
|
and TLS certificates, delivering real-time change notifications via Slack,
|
|
Mattermost, and ntfy webhooks.
|
|
|
|
> ⚠️ Pre-1.0 software. APIs, configuration, and behavior may change without
|
|
> notice.
|
|
|
|
dnswatcher watches configured DNS domains and hostnames for changes, monitors
|
|
TCP port availability, tracks TLS certificate expiry, and delivers real-time
|
|
notifications via Slack, Mattermost, and/or ntfy webhooks.
|
|
|
|
It resolves the names it watches itself via iterative (non-recursive) queries,
|
|
tracing from root nameservers to authoritative servers directly—never relying on
|
|
upstream recursive resolvers.
|
|
|
|
State is persisted to a local JSON file so that monitoring survives restarts
|
|
without requiring an external database.
|
|
|
|
---
|
|
|
|
## No DNS mocking. Ever.
|
|
|
|
**DNS is never mocked in this project — not in tests, not anywhere else.** No
|
|
mock resolvers, no fake DNS servers, no stubbed lookups.
|
|
|
|
dnswatcher's entire purpose is correct behavior against the real DNS. Tests
|
|
exercise real iterative resolution against live nameservers by design; a test
|
|
suite that passes against a mock proves nothing about the one thing this program
|
|
exists to do.
|
|
|
|
When live tests are flaky, that is a robustness problem, and it gets fixed with
|
|
robustness: retries with backoff, querying multiple independent nameservers,
|
|
longer timeouts — or explicit opt-in gating decided by the project owner. Never
|
|
with mocks.
|
|
|
|
Contributions that introduce mocked, faked, or stubbed DNS will be rejected.
|
|
|
|
---
|
|
|
|
## Getting Started
|
|
|
|
You need git and Docker. This builds the image and runs dnswatcher watching
|
|
`example.com` and `www.example.com`:
|
|
|
|
```sh
|
|
git clone https://git.eeqj.de/sneak/dnswatcher.git
|
|
cd dnswatcher
|
|
docker build -t dnswatcher .
|
|
docker run -d --name dnswatcher \
|
|
-p 8080:8080 \
|
|
-v dnswatcher-data:/var/lib/dnswatcher \
|
|
-e DNSWATCHER_TARGETS=example.com,www.example.com \
|
|
dnswatcher
|
|
```
|
|
|
|
The build also runs the linter and the test suite, which queries live DNS. Once
|
|
the container is running, the dashboard is at <http://localhost:8080/>. With no
|
|
notification endpoint set, changes show only on the dashboard; see
|
|
[Configuration](#configuration) to add one.
|
|
|
|
---
|
|
|
|
## Features
|
|
|
|
### DNS Domain Monitoring (Apex Domains)
|
|
|
|
- Accepts a list of DNS domain names (apex domains, identified via the
|
|
[Public Suffix List](https://publicsuffix.org/)).
|
|
- Every **1 hour** by default, performs a full iterative trace from root servers
|
|
to discover all authoritative nameservers (NS records) for each domain.
|
|
- Queries **every** discovered authoritative nameserver independently.
|
|
- Stores the domain's NS record set, as its parent zone's servers delegate it,
|
|
and the IPv4 and IPv6 addresses each nameserver's name resolves to.
|
|
- Any change triggers a notification:
|
|
- NS added to or removed from that set.
|
|
- NS address change: a nameserver that stays in the set resolves to
|
|
different addresses than on the previous check. A nameserver added or
|
|
removed gets only the NS change notification. When the lookup of a
|
|
nameserver's addresses fails or finds none, its previous addresses are
|
|
kept and nothing is sent.
|
|
|
|
### DNS Hostname Monitoring (Subdomains)
|
|
|
|
- Accepts a list of DNS hostnames (subdomains, distinguished from apex domains
|
|
via the Public Suffix List).
|
|
- Every **1 hour** by default, performs a full iterative trace to discover the
|
|
authoritative nameservers of the zone the hostname is in, which is not always
|
|
its last two labels (a name under `co.uk`, or in a delegated subdomain).
|
|
- Queries **each** authoritative nameserver independently for **all** record
|
|
types: A, AAAA, CNAME, MX, TXT, SRV, CAA, NS.
|
|
- Stores results **per nameserver**. The state for a hostname is not a merged
|
|
view — it is a map from nameserver to record set.
|
|
- DNS names inside record values (CNAME, MX, SRV and NS targets) are stored in
|
|
lower case, because names are case-insensitive and nameservers may answer in
|
|
any letter case. TXT and CAA values keep their letter case; they are not
|
|
lower-cased.
|
|
- Any observable change in any nameserver's response triggers a notification.
|
|
This includes:
|
|
- **Record change**: A nameserver returns different records than it did on
|
|
the previous check (additions, removals, value changes).
|
|
- **NS query failure**: A nameserver that previously responded becomes
|
|
unreachable (timeout, SERVFAIL, REFUSED, network error). This is distinct
|
|
from "responded with no records": a nameserver that answers NXDOMAIN or
|
|
with no records has responded. The alert is sent once, on the check where
|
|
it starts failing. A failing nameserver gives no records, so it is not
|
|
reported as a record change or compared for inconsistency. A nameserver
|
|
that is already failing on the first check that sees it is recorded
|
|
silently.
|
|
- **NS recovery**: A previously-unreachable nameserver starts responding
|
|
again. Its records are not compared with those from before it failed, so a
|
|
change made while it was failing is not reported as a record change.
|
|
- **Inconsistency detected**: Two nameservers return different record sets
|
|
for the same hostname and did not already differ on the previous check.
|
|
Every pair of nameservers is compared. The alert is sent once for each
|
|
such pair, on the check where they start to disagree, and not again while
|
|
they keep disagreeing, including after a restart. A nameserver that was
|
|
not in the previous check (newly added, or back after dropping out), or
|
|
failed on it, and answers differently is reported on the check where it
|
|
answers. If a pair agrees again and later disagrees, the alert is sent
|
|
again.
|
|
- **CNAME address change**: The addresses at the end of a name's CNAME chain
|
|
differ from those of the previous check. They are found when its
|
|
nameservers answer with a CNAME and no address; a name that answers with
|
|
an address has none. A change from or to no addresses is sent too, as when
|
|
a name moves between A records and a CNAME. Nothing is sent when the
|
|
previous addresses were kept because a chain could not be followed or none
|
|
of the name's nameservers answered. The first check after loading a state
|
|
file without `cnameAddresses` sends nothing: it saves the addresses it
|
|
finds for the next check to compare.
|
|
|
|
### TCP Port Monitoring
|
|
|
|
- For every configured domain and hostname, constructs a deduplicated list of
|
|
the IPv4 and IPv6 addresses in the A and AAAA records its authoritative
|
|
nameservers returned. When they returned a CNAME and no address, the CNAME
|
|
chain is followed and the addresses at its end are used, and a change in those
|
|
is notified as a CNAME address change. When the nameservers gave different
|
|
CNAME targets, each is followed and the addresses of all are used. When a
|
|
chain cannot be followed, or none of the name's nameservers answered, the
|
|
addresses the last check found at its end are used.
|
|
- Checks TCP connectivity on ports **80** and **443** for each IP address.
|
|
- Every **1 hour** by default, re-checks all ports.
|
|
- Any change in port availability triggers a notification:
|
|
- Port transitioned from open to closed (or vice versa).
|
|
- New IP appeared (from DNS change): its port state is recorded without a
|
|
port notification; the DNS change notification shows the new address.
|
|
- IP disappeared (from DNS change) — noted in the DNS change notification;
|
|
port state for that IP is removed. When none of a name's nameservers
|
|
answered, its addresses are not known, so the port state saved for them is
|
|
kept.
|
|
|
|
### TLS Certificate Monitoring
|
|
|
|
- Every **12 hours** by default, for each IP address listening on port 443,
|
|
connects via TLS using the correct SNI hostname.
|
|
- Records the certificate's Subject CN, SANs, issuer, and expiry date.
|
|
- Any change triggers a notification:
|
|
- Certificate is expiring within **7 days** by default (warning, repeated
|
|
each check until renewed or expired).
|
|
- Certificate CN, issuer, or SANs changed (replacement detected, reports old
|
|
and new CN and issuer).
|
|
- TLS connection failure to a previously-reachable IP:443 (handshake error,
|
|
timeout, connection refused after previously succeeding).
|
|
- TLS recovery: a previously-failing IP:443 now completes a handshake again.
|
|
|
|
### Notifications
|
|
|
|
**Every observable state change produces a notification.** dnswatcher is
|
|
designed as a real-time change feed — degradations, failures, recoveries, and
|
|
routine changes are all reported equally.
|
|
|
|
Supported notification backends:
|
|
|
|
| Backend | Configuration | Payload Format |
|
|
| -------------- | ------------------------------------------ | ---------------------------- |
|
|
| **Slack** | Incoming Webhook URL | Attachments with color |
|
|
| **Mattermost** | Incoming Webhook URL | Slack-compatible attachments |
|
|
| **ntfy** | Topic URL (e.g. `https://ntfy.sh/mytopic`) | Title + body + priority |
|
|
|
|
All configured endpoints receive every notification. Notification content
|
|
includes:
|
|
|
|
- **DNS record changes**: Which hostname, which nameserver, what record type,
|
|
old values, new values.
|
|
- **DNS NS changes**: Which domain, which nameservers were added/removed.
|
|
- **NS address changes**: Which domain, which nameserver, its old and new
|
|
addresses.
|
|
- **CNAME address changes**: Which hostname, the old and new addresses at the
|
|
end of its CNAME chain.
|
|
- **NS query failures**: Which nameserver failed, error type (timeout, SERVFAIL,
|
|
REFUSED, network error), which hostname/domain affected.
|
|
- **NS recoveries**: Which nameserver recovered, which hostname/domain.
|
|
- **NS inconsistencies**: Which nameservers disagree, what each one returned,
|
|
which hostname affected.
|
|
- **Port changes**: Which IP:port, its new state, all associated hostnames.
|
|
- **TLS expiry warnings**: Expiry date and days remaining, CN, associated
|
|
hostname and IP.
|
|
- **TLS certificate changes**: Old and new CN and issuer, associated hostname
|
|
and IP. A change to the SANs alone is notified, but the SANs are not listed.
|
|
- **TLS connection failures/recoveries**: Which IP:port, error details,
|
|
associated hostname.
|
|
|
|
Each endpoint is sent each notification on its own, in the background. A
|
|
delivery that fails (a network error, no reply within 10 seconds, or an HTTP
|
|
status of 400 or more) is retried up to 5 times: the first retry after about 1
|
|
second, each wait after that twice as long up to 60 seconds, every wait varied
|
|
at random by up to 25%. A delivery still failing after that is logged and
|
|
dropped.
|
|
|
|
The last 100 notifications, delivered or not, are kept in memory for the
|
|
dashboard's Recent alerts. They are not saved to the state file, so a restart
|
|
clears them.
|
|
|
|
### State Management
|
|
|
|
- All monitoring state is kept in memory and persisted to a JSON file on disk
|
|
(`DATA_DIR/state.json`).
|
|
- State is loaded on startup to resume monitoring without triggering
|
|
false-positive change notifications.
|
|
- State is written atomically (write to temp file, then rename) to prevent
|
|
corruption.
|
|
|
|
### Web Dashboard
|
|
|
|
dnswatcher includes an unauthenticated, read-only web dashboard at the root URL
|
|
(`/`). It displays:
|
|
|
|
- **Summary counts** for monitored domains, hostnames, ports, and certificates.
|
|
- **Domains** with their discovered nameservers.
|
|
- **Hostnames** with per-nameserver DNS records and status.
|
|
- **Ports** with open/closed state and associated hostnames.
|
|
- **TLS certificates** with CN, issuer, expiry, and status.
|
|
- **Recent alerts** (last 100 notifications sent since the process started),
|
|
displayed in reverse chronological order.
|
|
|
|
Every data point shows its age (e.g. "5m ago") so you can tell at a glance how
|
|
fresh the information is. The page auto-refreshes every 30 seconds.
|
|
|
|
The dashboard intentionally does not expose any configuration details such as
|
|
webhook URLs, notification endpoints, or API tokens.
|
|
|
|
All assets (CSS) are embedded in the binary and served from the application
|
|
itself. The dashboard makes zero external HTTP requests — no CDN dependencies or
|
|
third-party resources are loaded at runtime.
|
|
|
|
### HTTP API
|
|
|
|
dnswatcher exposes a lightweight HTTP API for operational visibility:
|
|
|
|
| Endpoint | Description |
|
|
| ------------------------------ | ----------------------------- |
|
|
| `GET /` | Web dashboard (HTML) |
|
|
| `GET /s/...` | Static assets (embedded CSS) |
|
|
| `GET /.well-known/healthcheck` | Health check (JSON) |
|
|
| `GET /health` | Health check (JSON, legacy) |
|
|
| `GET /api/v1/status` | Current monitoring state |
|
|
| `GET /metrics` | Prometheus metrics, see below |
|
|
|
|
`/metrics` is served only when `DNSWATCHER_METRICS_USERNAME` is set, behind
|
|
Basic Auth. It has the Prometheus Go client's default metrics only (Go runtime,
|
|
process, and counts of `/metrics` requests); dnswatcher records no metrics of
|
|
its own.
|
|
|
|
Every route but `/metrics` may be read from a page on any origin: a cross-origin
|
|
`GET` gets `Access-Control-Allow-Origin: *`. Only `GET` is allowed cross-origin,
|
|
and without credentials. `/metrics` sends no CORS headers.
|
|
|
|
#### Server timeouts
|
|
|
|
The HTTP server sets all four socket-level timeouts. These are compile-time
|
|
constants in `internal/server/server.go`, not configurable via environment
|
|
variables.
|
|
|
|
| Timeout | Value | Purpose |
|
|
| ------------------- | ----- | --------------------------------------------- |
|
|
| `ReadHeaderTimeout` | 10s | Bounds the request header read (slowloris) |
|
|
| `ReadTimeout` | 15s | Bounds the whole request read, headers + body |
|
|
| `WriteTimeout` | 75s | Bounds handler execution plus response flush |
|
|
| `IdleTimeout` | 120s | Reaps idle keep-alive connections |
|
|
|
|
These are distinct from the 60s per-request handler budget applied by
|
|
`chimw.Timeout` in `internal/server/routes.go`, which cancels the request
|
|
context but does not touch the socket. `WriteTimeout` is deliberately larger
|
|
than that budget: the write deadline is armed once request headers are read, so
|
|
a smaller value would sever the connection before a handler using its full
|
|
budget could respond. `IdleTimeout` exceeds common Prometheus scrape intervals
|
|
so the scraper reuses its connection.
|
|
|
|
### Security Headers
|
|
|
|
Every response — the dashboard, the static assets under `/s/...`, the
|
|
healthchecks, the JSON API, and `/metrics` — carries the following headers, set
|
|
by a global middleware:
|
|
|
|
| Header | Value |
|
|
| --------------------------- | ------------------------------------- |
|
|
| `Strict-Transport-Security` | `max-age=31536000; includeSubDomains` |
|
|
| `Content-Security-Policy` | see below |
|
|
| `X-Frame-Options` | `DENY` |
|
|
| `X-Content-Type-Options` | `nosniff` |
|
|
| `Referrer-Policy` | `no-referrer` |
|
|
| `Permissions-Policy` | all unused browser features denied |
|
|
|
|
The content security policy is:
|
|
|
|
```
|
|
default-src 'self'; script-src 'none'; style-src 'self'; img-src 'self';
|
|
font-src 'none'; connect-src 'none'; object-src 'none'; base-uri 'none';
|
|
form-action 'none'; frame-ancestors 'none'
|
|
```
|
|
|
|
The dashboard ships no JavaScript (the 30-second refresh is a
|
|
`<meta http-equiv="refresh">`), no inline styles, no inline event handlers, and
|
|
no images; its only subresource is the embedded stylesheet at
|
|
`/s/css/tailwind.min.css`, which `style-src 'self'` permits. The policy
|
|
therefore needs neither `unsafe-inline` nor `unsafe-eval`.
|
|
`frame-ancestors 'none'` is the primary anti-framing control, with
|
|
`X-Frame-Options: DENY` retained as the legacy fallback.
|
|
|
|
HSTS is emitted unconditionally, including over plain HTTP. dnswatcher is
|
|
expected to run behind a TLS-terminating reverse proxy, and the browser must
|
|
still be told to enforce HTTPS end to end, so the header is never gated on
|
|
whether the request itself arrived over TLS.
|
|
|
|
`Referrer-Policy: no-referrer` is stricter than the
|
|
`strict-origin-when-cross-origin` baseline: the dashboard has no cross-origin
|
|
navigation needs, and its URL may name internal hosts.
|
|
|
|
---
|
|
|
|
## Configuration
|
|
|
|
Configuration is loaded via [Viper](https://github.com/spf13/viper) with the
|
|
following precedence (highest to lowest):
|
|
|
|
1. Environment variables (prefixed with `DNSWATCHER_`)
|
|
2. `.env` file (loaded via godotenv)
|
|
3. Config file: `/etc/dnswatcher/dnswatcher.yaml`,
|
|
`~/.config/dnswatcher/dnswatcher.yaml`, or `./dnswatcher.yaml`
|
|
4. Defaults
|
|
|
|
### Environment Variables
|
|
|
|
| Variable | Description | Default |
|
|
| ----------------------------------- | ----------------------------------------------------------------------------------------------------------- | --------------------- |
|
|
| `PORT` | HTTP listen port | `8080` |
|
|
| `DNSWATCHER_DEBUG` | Enable debug logging | `false` |
|
|
| `DNSWATCHER_DATA_DIR` | Directory for state file | `/var/lib/dnswatcher` |
|
|
| `DNSWATCHER_TARGETS` | Comma-separated DNS names (auto-classified via PSL) | `""` |
|
|
| `DNSWATCHER_SLACK_WEBHOOK` | Slack incoming webhook URL | `""` |
|
|
| `DNSWATCHER_MATTERMOST_WEBHOOK` | Mattermost incoming webhook URL | `""` |
|
|
| `DNSWATCHER_NTFY_TOPIC` | ntfy topic URL | `""` |
|
|
| `DNSWATCHER_DNS_INTERVAL` | DNS check interval, a positive duration such as `30m`; empty means the default, anything else stops startup | `1h` |
|
|
| `DNSWATCHER_TLS_INTERVAL` | TLS check interval, a positive duration such as `6h`; empty means the default, anything else stops startup | `12h` |
|
|
| `DNSWATCHER_TLS_EXPIRY_WARNING` | Days before expiry to warn | `7` |
|
|
| `DNSWATCHER_SENTRY_DSN` | Sentry DSN for error reporting | `""` |
|
|
| `DNSWATCHER_MAINTENANCE_MODE` | Only sets `maintenanceMode` in the health check response; changes nothing else | `false` |
|
|
| `DNSWATCHER_METRICS_USERNAME` | Basic auth username for /metrics, which is served only when this is set | `""` |
|
|
| `DNSWATCHER_METRICS_PASSWORD` | Basic auth password for /metrics | `""` |
|
|
| `DNSWATCHER_SEND_TEST_NOTIFICATION` | Send a test notification after first scan completes | `false` |
|
|
|
|
**`DNSWATCHER_TARGETS` is required.** dnswatcher will refuse to start if no
|
|
monitoring targets are configured. A monitoring daemon with nothing to monitor
|
|
is a misconfiguration, so dnswatcher fails fast with a clear error message
|
|
rather than running silently. Set `DNSWATCHER_TARGETS` to a comma-separated list
|
|
of DNS names before starting. A name listed more than once, in any letter case
|
|
or with a trailing dot, is watched once.
|
|
|
|
**`/metrics` is rate limited.** Each client address may send it 30 requests a
|
|
minute, failed logins included; beyond that it answers `429 Too Many Requests`
|
|
without checking the password. A Prometheus server scraping every 15 seconds
|
|
sends 4 a minute. IPv6 addresses in one /64 count as one client. When the
|
|
request comes from a private or loopback address, such as a reverse proxy's, the
|
|
client address is taken from the `X-Real-IP` header the proxy sets, or else from
|
|
`X-Forwarded-For`, as the last address in it that is not private or loopback. A
|
|
proxy that sets neither makes all its clients share one allowance.
|
|
|
|
**`DNSWATCHER_DNS_INTERVAL` and `DNSWATCHER_TLS_INTERVAL`** take a positive
|
|
duration: a number followed by a unit such as `s`, `m` or `h`, for example
|
|
`90s`, `30m`, `1h` or `1h30m`. There is no unit for days; write `24h`. An unset
|
|
or empty variable (`DNSWATCHER_DNS_INTERVAL=`) means the default. If either is
|
|
set to anything else, including a bare number or a zero or negative duration,
|
|
dnswatcher refuses to start with an error naming the variable and the value.
|
|
|
|
**`DNSWATCHER_SENTRY_DSN` reports crashes in HTTP requests to Sentry.** When it
|
|
is set, a panic in an HTTP request handler is sent to Sentry, and the request
|
|
still gets a `500 Internal Server Error` answer. Nothing else is sent to Sentry:
|
|
DNS, port and TLS problems are reported as notifications. A value Sentry cannot
|
|
parse stops dnswatcher at startup. At shutdown, reports not yet sent are sent,
|
|
waiting at most 2 seconds.
|
|
|
|
### Example `.env`
|
|
|
|
```sh
|
|
PORT=8080
|
|
DNSWATCHER_DEBUG=false
|
|
DNSWATCHER_DATA_DIR=/var/lib/dnswatcher
|
|
DNSWATCHER_TARGETS=example.com,example.org,www.example.com,api.example.com,mail.example.org
|
|
DNSWATCHER_SLACK_WEBHOOK=https://hooks.slack.com/services/T.../B.../xxx
|
|
DNSWATCHER_MATTERMOST_WEBHOOK=https://mattermost.example.com/hooks/xxx
|
|
DNSWATCHER_NTFY_TOPIC=https://ntfy.sh/my-dns-alerts
|
|
DNSWATCHER_SEND_TEST_NOTIFICATION=true
|
|
```
|
|
|
|
---
|
|
|
|
## DNS Resolution Strategy
|
|
|
|
dnswatcher never uses the system's configured recursive resolver for the names
|
|
it watches. Instead, it performs full iterative resolution:
|
|
|
|
1. **Root servers**: Starts from the IPv4 addresses of the 13 root servers,
|
|
built into the binary; the list is not refreshed.
|
|
2. **TLD delegation**: Queries root servers for the TLD NS records.
|
|
3. **Domain delegation**: Queries TLD nameservers for the domain's NS records.
|
|
The delegation they give, from the domain's parent zone, is the domain's NS
|
|
record set.
|
|
4. **Authoritative query**: Queries all discovered authoritative nameservers
|
|
directly for the requested records.
|
|
|
|
In steps 2 and 3 the servers are asked one at a time in a random order, chosen
|
|
anew each time, so no one root server gets every first query. A server that does
|
|
not reply, refuses the query, or gives an error reply such as SERVFAIL or a
|
|
referral that leads no closer to the name is passed over for the next one. When
|
|
a referral names a zone's nameservers without their addresses, the addresses of
|
|
all of them are looked up, so that each can be asked.
|
|
|
|
This approach ensures:
|
|
|
|
- Independence from any upstream resolver's cache or filtering.
|
|
- Ability to detect split-horizon or inconsistent responses across authoritative
|
|
servers.
|
|
|
|
A watched name's records are stored as its nameservers return them, CNAME
|
|
included. When they return a CNAME and no address, the chain of every CNAME
|
|
target they gave is followed (with a depth limit to prevent loops) to the A and
|
|
AAAA records at its end, and the port and TLS checks use those addresses.
|
|
Nameservers' addresses are also found by following CNAME chains.
|
|
|
|
Sending a notification or a Sentry report is the one use of the system's
|
|
resolver: the HTTP client looks up the webhook's or Sentry's host name with it.
|
|
|
|
---
|
|
|
|
## State File Format
|
|
|
|
The state file (`DATA_DIR/state.json`) contains the complete monitoring
|
|
snapshot. Hostname records are stored **per authoritative nameserver**, not as a
|
|
merged view, to enable inconsistency detection.
|
|
|
|
```json
|
|
{
|
|
"version": 1,
|
|
"lastUpdated": "2026-02-19T12:00:00Z",
|
|
"domains": {
|
|
"example.com": {
|
|
"nameservers": ["ns1.example.com.", "ns2.example.com."],
|
|
"nameserverAddresses": {
|
|
"ns1.example.com.": ["192.0.2.53", "2001:db8::53"],
|
|
"ns2.example.com.": ["198.51.100.53"]
|
|
},
|
|
"lastChecked": "2026-02-19T12:00:00Z"
|
|
}
|
|
},
|
|
"hostnames": {
|
|
"www.example.com": {
|
|
"recordsByNameserver": {
|
|
"ns1.example.com.": {
|
|
"records": {
|
|
"A": ["93.184.216.34"],
|
|
"AAAA": ["2606:2800:220:1:248:1893:25c8:1946"]
|
|
},
|
|
"status": "ok",
|
|
"lastChecked": "2026-02-19T12:00:00Z"
|
|
},
|
|
"ns2.example.com.": {
|
|
"records": {
|
|
"A": ["93.184.216.34"],
|
|
"AAAA": ["2606:2800:220:1:248:1893:25c8:1946"]
|
|
},
|
|
"status": "ok",
|
|
"lastChecked": "2026-02-19T12:00:00Z"
|
|
}
|
|
},
|
|
"cnameAddresses": [],
|
|
"lastChecked": "2026-02-19T12:00:00Z"
|
|
}
|
|
},
|
|
"ports": {
|
|
"93.184.216.34:80": {
|
|
"open": true,
|
|
"hostnames": ["www.example.com"],
|
|
"lastChecked": "2026-02-19T12:00:00Z"
|
|
},
|
|
"93.184.216.34:443": {
|
|
"open": true,
|
|
"hostnames": ["www.example.com"],
|
|
"lastChecked": "2026-02-19T12:00:00Z"
|
|
}
|
|
},
|
|
"certificates": {
|
|
"93.184.216.34:443:www.example.com": {
|
|
"commonName": "www.example.com",
|
|
"issuer": "DigiCert TLS RSA SHA256 2020 CA1",
|
|
"notAfter": "2027-01-15T23:59:59Z",
|
|
"subjectAlternativeNames": ["www.example.com"],
|
|
"status": "ok",
|
|
"lastChecked": "2026-02-19T06:00:00Z"
|
|
}
|
|
}
|
|
}
|
|
```
|
|
|
|
The `status` field for each per-nameserver entry and certificate entry tracks
|
|
reachability:
|
|
|
|
| Status | Meaning |
|
|
| ------- | -------------------------------------------------------- |
|
|
| `ok` | Query succeeded, records are current |
|
|
| `error` | Query failed (timeout, SERVFAIL, REFUSED, network error) |
|
|
|
|
A nameserver that answers NXDOMAIN or with no records has status `ok` and empty
|
|
`records`. A nameserver whose query failed, or that only referred it to other
|
|
nameservers, has status `error`, empty `records`, and the reason in `error`. A
|
|
certificate entry whose TLS connection or handshake failed likewise has status
|
|
`error`, the reason in `error`, and the certificate fields left empty or zero.
|
|
|
|
`nameserverAddresses` lists, by nameserver, the sorted addresses its name
|
|
resolves to. A state file without it loads, and the next check fills it in
|
|
without a notification.
|
|
|
|
`cnameAddresses` lists the sorted addresses at the end of the chain of every
|
|
CNAME target a hostname's nameservers gave, found when they answered with a
|
|
CNAME and no address; it is empty when they answered with an address. When a
|
|
chain cannot be followed, or none of the name's nameservers answered, the
|
|
previous check's list is kept, or `null` when no earlier check saved one. A
|
|
state file without it loads, and the first check after that saves it without a
|
|
notification.
|
|
|
|
A port entry in the older format, with one `hostname` instead of the `hostnames`
|
|
list, loads as a list of that one name.
|
|
|
|
---
|
|
|
|
## Entrypoints
|
|
|
|
This repository adheres to the
|
|
[Scripts to Rule Them All](https://github.com/github/scripts-to-rule-them-all)
|
|
standard: normalized scripts in `script/` are the entrypoints for the
|
|
development workflow, and the Makefile targets are thin shims that call them. We
|
|
provide:
|
|
|
|
- `script/bootstrap` — install all dependencies (go, `go mod download`). It does
|
|
not install golangci-lint or prettier: both run in Docker, see `script/lint`
|
|
and `script/fmt` below.
|
|
- `script/setup` — make a fresh clone ready for development: bootstrap plus the
|
|
git pre-commit hook
|
|
- `script/projectname` — print the project name (used for the Docker image tag)
|
|
- `script/test` — run the test suite (race detector, coverage). Caching is
|
|
waived for testing, exactly as it is for linting: `-count=1` forces every
|
|
invocation to execute, because the suite queries live DNS and a cached pass
|
|
queries nothing. Failures are rerun with `-v` automatically, and the build
|
|
fails even if that rerun passes.
|
|
- `script/lint` — run golangci-lint, always inside Docker: it builds
|
|
`Dockerfile.lint`, which COPYs the repo into the digest-pinned `golangci-lint`
|
|
image and lints as a build step, so a successful build is a clean lint. The
|
|
linter is never installed or run on the host, and Docker is the only
|
|
prerequisite. Caching is waived for linting: the lint stage is forced to
|
|
execute on every run with `--no-cache-filter`, because a cached build lints
|
|
nothing.
|
|
- `script/fmt` — format all code (gofmt -s, goimports) and all Markdown
|
|
(prettier). goimports runs with `go run` at a pinned commit, never from your
|
|
`PATH`. prettier runs inside Docker, built from `Dockerfile.fmt` on a
|
|
digest-pinned node image, at the version pinned by `package.json` and
|
|
`yarn.lock`; it is never installed on the host.
|
|
- `script/fmt-check` — check formatting (read-only) with the same tools, failing
|
|
on any file `script/fmt` would change. It runs the two scripts below.
|
|
- `script/fmt-check-go` — the gofmt and goimports half, on the host. The
|
|
`Dockerfile` lint stage runs it.
|
|
- `script/fmt-check-markdown` — the prettier half, inside Docker, forced to
|
|
execute on every run with `--no-cache-filter`
|
|
- `script/check` — run test, lint, and fmt-check
|
|
- `script/docker` — build the Docker image tagged via `script/projectname`, with
|
|
`--no-cache-filter=lint,builder` so the lint stage and the builder stage,
|
|
which runs the tests, run on every invocation, and with the version from
|
|
`git describe` passed as `--build-arg VERSION`
|
|
- `script/cibuild` — CI entrypoint: `docker build` with
|
|
`--no-cache-filter=lint,builder`, so the lint stage and the builder stage,
|
|
which runs the tests, run on every invocation, because a cached build lints
|
|
nothing and queries no DNS; then `script/fmt-check-markdown`
|
|
- `script/precommit` — run by the git pre-commit hook; `go mod tidy` guard, then
|
|
`script/check`
|
|
- `script/install-precommit` — install the git pre-commit hook
|
|
|
|
## Building
|
|
|
|
```sh
|
|
make build # Build binary to bin/dnswatcher
|
|
make test # Run tests with race detector
|
|
make lint # Run golangci-lint in Docker (requires docker)
|
|
make fmt # Format code and Markdown (requires docker)
|
|
make check # Run all checks (test, lint, fmt-check)
|
|
make clean # Remove build artifacts
|
|
```
|
|
|
|
### Build-Time Variables
|
|
|
|
`make build` sets the version with `-ldflags "-X main.Version=..."`, taking it
|
|
from `git describe --tags --always --dirty`, or from `VERSION` when given on the
|
|
command line (`make build VERSION=1.2.3`). The version appears in the startup
|
|
log and in the health check response.
|
|
|
|
The Docker image has no `.git`, so the `Dockerfile` takes the version as
|
|
`--build-arg VERSION`. `make docker` passes it; a plain `docker build` passes
|
|
none, and that image reports `dev`.
|
|
|
|
---
|
|
|
|
## Docker
|
|
|
|
```sh
|
|
docker build -t dnswatcher .
|
|
docker run -d \
|
|
-p 8080:8080 \
|
|
-v dnswatcher-data:/var/lib/dnswatcher \
|
|
-e DNSWATCHER_TARGETS=example.com,www.example.com \
|
|
-e DNSWATCHER_NTFY_TOPIC=https://ntfy.sh/my-alerts \
|
|
-e DNSWATCHER_SEND_TEST_NOTIFICATION=true \
|
|
dnswatcher
|
|
```
|
|
|
|
---
|
|
|
|
## Running under upaas
|
|
|
|
[upaas](https://git.eeqj.de/sneak/upaas) builds the image from this repository's
|
|
`Dockerfile` and runs it. The app needs:
|
|
|
|
- **Branch:** `prod`. `prod` is cut from `main`, and merging a `main` to `prod`
|
|
pull request is a deploy.
|
|
- **Volume:** one host directory mounted at `/var/lib/dnswatcher`, where the
|
|
state file lives.
|
|
- **Network and port:** the dashboard is unauthenticated and shows every watched
|
|
name and recent alert, and upaas publishes every mapped port on all interfaces
|
|
of the host ([upaas issue 113](https://git.eeqj.de/sneak/upaas/issues/113)).
|
|
Add a port mapping to container port `8080` only if the dashboard should be
|
|
public. Otherwise add none: set the app's Docker network in upaas to your
|
|
reverse proxy's Docker network, and the proxy reaches the app at `upaas-`
|
|
followed by the app name, port `8080`.
|
|
- **Required environment:** `DNSWATCHER_TARGETS`, a comma-separated list of the
|
|
domains and hostnames to watch. dnswatcher refuses to start without it.
|
|
- **Recommended environment:** at least one notification endpoint
|
|
(`DNSWATCHER_SLACK_WEBHOOK`, `DNSWATCHER_MATTERMOST_WEBHOOK`,
|
|
`DNSWATCHER_NTFY_TOPIC`); without one, changes show only on the dashboard.
|
|
`DNSWATCHER_METRICS_USERNAME` and `DNSWATCHER_METRICS_PASSWORD` serve
|
|
`/metrics` behind basic auth.
|
|
- **Leave unset:** `DNSWATCHER_DATA_DIR`, which the image sets to
|
|
`/var/lib/dnswatcher`, and `PORT`, which defaults to `8080`. Every setting
|
|
comes from the environment; the image holds no config file.
|
|
- **Health check:** the image's own, which requests `/.well-known/healthcheck`
|
|
every 10 seconds. upaas reads the container's health 60 seconds after a deploy
|
|
and marks the deploy failed unless it is `healthy`.
|
|
|
|
---
|
|
|
|
## Monitoring Lifecycle
|
|
|
|
1. **Startup**: Check that the data directory can be written, and exit with an
|
|
error naming it if not. Load state from disk. If no state file exists, start
|
|
with empty state (first check will establish baseline without triggering
|
|
change notifications).
|
|
2. **Initial check**: Immediately perform all DNS, port, and TLS checks on
|
|
startup.
|
|
3. **Periodic checks** (DNS always runs first):
|
|
- DNS checks: every `DNSWATCHER_DNS_INTERVAL` (default 1h). Also re-run
|
|
before every TLS check cycle to ensure fresh IPs.
|
|
- Port checks: every `DNSWATCHER_DNS_INTERVAL`, after DNS completes.
|
|
- TLS checks: every `DNSWATCHER_TLS_INTERVAL` (default 12h), after DNS
|
|
completes.
|
|
- Port and TLS checks use the IP addresses found by the DNS phase that
|
|
immediately precedes them. When that phase cannot find a name's
|
|
nameservers at all, the addresses an earlier check saved for the name are
|
|
used. When it cannot follow a name's CNAME chain, or none of the name's
|
|
nameservers answered, the addresses an earlier check found at the end of
|
|
the chain are used.
|
|
4. **On change detection**: Send notifications to all configured endpoints,
|
|
update in-memory state, persist to disk.
|
|
5. **Shutdown**: The watcher stops checking and saves the final state to disk,
|
|
and shutdown waits for that save before it goes on. Then it waits for
|
|
in-flight notification deliveries to complete. Both waits share the fx
|
|
shutdown timeout (15s by default): deliveries still retrying against an
|
|
unreachable endpoint when that expires are abandoned, and the number
|
|
abandoned is logged at warn level rather than dropped silently. Notifications
|
|
generated after shutdown has begun are refused and logged, so a late burst
|
|
cannot extend the shutdown. A DNS lookup, port check or TLS check that
|
|
shutdown cuts short saves nothing and sends no notification.
|
|
|
|
---
|
|
|
|
## Planned Future Features (Post-1.0)
|
|
|
|
- **DNSSEC validation**: Validate the DNSSEC chain of trust during iterative
|
|
resolution and report DNSSEC failures as notifications.
|
|
|
|
---
|
|
|
|
## Project Structure
|
|
|
|
Follows the conventions defined in `REPO_POLICIES.md`, adapted from the
|
|
[upaas](https://git.eeqj.de/sneak/upaas) project template. Uses uber/fx for
|
|
dependency injection, go-chi for HTTP routing, slog for logging, and Viper for
|
|
configuration.
|
|
|
|
---
|
|
|
|
## Rationale
|
|
|
|
dnswatcher exists to report changes to the DNS records, TCP port availability
|
|
and TLS certificates of its configured domains and hostnames, failures and
|
|
recoveries included: it is designed as a real-time change feed. It queries the
|
|
authoritative nameservers directly, tracing from the root, instead of a
|
|
recursive resolver, so no resolver's cache or filtering hides a change and
|
|
nameservers that disagree with each other are seen. Its state is a single JSON
|
|
file, so it survives a restart without an external database.
|
|
|
|
---
|
|
|
|
## Design
|
|
|
|
```
|
|
cmd/dnswatcher/main.go Entry point (uber/fx bootstrap)
|
|
|
|
internal/
|
|
config/
|
|
config.go Viper-based configuration
|
|
classify.go Splits targets into domains and hostnames
|
|
(Public Suffix List)
|
|
globals/globals.go Build-time variables (version)
|
|
logger/logger.go slog structured logging (TTY detection)
|
|
healthcheck/healthcheck.go Health check service
|
|
middleware/middleware.go HTTP middleware (logging, CORS, security
|
|
headers, metrics auth and rate limit)
|
|
handlers/
|
|
handlers.go Shared handler setup and JSON responses
|
|
dashboard.go Web dashboard
|
|
templates/dashboard.html Dashboard template (embedded)
|
|
status.go /api/v1/status
|
|
healthcheck.go Health check handler
|
|
server/
|
|
server.go HTTP server lifecycle
|
|
routes.go Route definitions
|
|
state/state.go JSON file state persistence
|
|
resolver/
|
|
resolver.go Resolver setup and query status values
|
|
iterative.go Iterative DNS resolution engine
|
|
dns_client.go UDP and TCP DNS clients
|
|
errors.go Resolver errors
|
|
portcheck/portcheck.go TCP port connectivity checker
|
|
tlscheck/tlscheck.go TLS certificate inspector
|
|
notify/
|
|
notify.go Notification service (Slack, Mattermost, ntfy)
|
|
retry.go Delivery retries with backoff
|
|
history.go Last 100 notifications, for the dashboard
|
|
shutdown.go Waits for deliveries at shutdown
|
|
watcher/
|
|
watcher.go Main monitoring orchestrator and scheduler
|
|
interfaces.go The resolver, checkers and notifier it uses
|
|
livednstest/livednstest.go Retry and concurrency limit for tests
|
|
against live DNS (imported only by tests)
|
|
|
|
static/
|
|
static.go Embeds the CSS served under /s/
|
|
css/tailwind.min.css Dashboard stylesheet
|
|
```
|
|
|
|
### Design Principles
|
|
|
|
- **No recursive resolvers**: The watched names are resolved iteratively,
|
|
tracing from root nameservers through the delegation chain to authoritative
|
|
servers.
|
|
- **No external database**: State is persisted as a single JSON file.
|
|
- **Dependency injection**: All components are wired via
|
|
[uber/fx](https://github.com/uber-go/fx).
|
|
- **Structured logging**: All logs use `log/slog` with JSON output in production
|
|
(TTY detection for development).
|
|
- **Graceful shutdown**: All background goroutines respect context cancellation
|
|
and the fx lifecycle. In-flight notification deliveries are drained on
|
|
shutdown, bounded by the shutdown timeout.
|
|
|
|
---
|
|
|
|
## TODO
|
|
|
|
[`TODO.md`](./TODO.md) names the next step and the steps planned after it. The
|
|
work for 1.0 is tracked as issues on the
|
|
[1.0 milestone](https://git.eeqj.de/sneak/dnswatcher/milestone/7).
|
|
|
|
---
|
|
|
|
## License
|
|
|
|
dnswatcher is released under the MIT License, Copyright (c) 2026
|
|
[@sneak](https://sneak.berlin). See the [`LICENSE`](./LICENSE) file in the
|
|
repository root for the full text.
|
|
|
|
## Author
|
|
|
|
[@sneak](https://sneak.berlin)
|