check / check (push) Failing after 2m7s
ResolveIPAddresses now returns an error, not no addresses, when no nameserver of the name's zone answered. A nameserver with status timeout or error is not an answer; one answer, even NXDOMAIN, is enough for an empty result without an error. When every server of a zone fails, FindAuthoritativeNameservers moves on to the parent name, whose servers only refer the query onward. Such a referral now has status error, so it is no answer either, and a hostname's saved records show it as error. The only caller, the nameserver address lookup, already keeps the previous addresses on an error; its comment no longer says the resolver hides this case. Model: opus-5-5
677 lines
32 KiB
Markdown
677 lines
32 KiB
Markdown
# dnswatcher
|
|
|
|
dnswatcher is an MIT-licensed, pre-1.0 Go daemon by
|
|
[@sneak](https://sneak.berlin) that monitors DNS records, TCP port availability,
|
|
and TLS certificates, delivering real-time change notifications via Slack,
|
|
Mattermost, and ntfy webhooks.
|
|
|
|
> ⚠️ Pre-1.0 software. APIs, configuration, and behavior may change without
|
|
> notice.
|
|
|
|
dnswatcher watches configured DNS domains and hostnames for changes, monitors
|
|
TCP port availability, tracks TLS certificate expiry, and delivers real-time
|
|
notifications via Slack, Mattermost, and/or ntfy webhooks.
|
|
|
|
It performs all DNS resolution itself via iterative (non-recursive) queries,
|
|
tracing from root nameservers to authoritative servers directly—never relying on
|
|
upstream recursive resolvers.
|
|
|
|
State is persisted to a local JSON file so that monitoring survives restarts
|
|
without requiring an external database.
|
|
|
|
---
|
|
|
|
## No DNS mocking. Ever.
|
|
|
|
**DNS is never mocked in this project — not in tests, not anywhere else.** No
|
|
mock resolvers, no fake DNS servers, no stubbed lookups.
|
|
|
|
dnswatcher's entire purpose is correct behavior against the real DNS. Tests
|
|
exercise real iterative resolution against live nameservers by design; a test
|
|
suite that passes against a mock proves nothing about the one thing this program
|
|
exists to do.
|
|
|
|
When live tests are flaky, that is a robustness problem, and it gets fixed with
|
|
robustness: retries with backoff, querying multiple independent nameservers,
|
|
longer timeouts — or explicit opt-in gating decided by the project owner. Never
|
|
with mocks.
|
|
|
|
Contributions that introduce mocked, faked, or stubbed DNS will be rejected.
|
|
|
|
---
|
|
|
|
## Features
|
|
|
|
### DNS Domain Monitoring (Apex Domains)
|
|
|
|
- Accepts a list of DNS domain names (apex domains, identified via the
|
|
[Public Suffix List](https://publicsuffix.org/)).
|
|
- Every **1 hour**, performs a full iterative trace from root servers to
|
|
discover all authoritative nameservers (NS records) for each domain.
|
|
- Queries **every** discovered authoritative nameserver independently.
|
|
- Stores the NS record set as observed by the delegation chain, and the IPv4 and
|
|
IPv6 addresses each nameserver's name resolves to.
|
|
- Any change triggers a notification:
|
|
- NS added to or removed from the delegation.
|
|
- NS address change: a nameserver that stays in the delegation resolves to
|
|
different addresses than on the previous check. A nameserver added or
|
|
removed gets only the NS change notification. When the lookup of a
|
|
nameserver's addresses fails or finds none, its previous addresses are
|
|
kept and nothing is sent.
|
|
|
|
### DNS Hostname Monitoring (Subdomains)
|
|
|
|
- Accepts a list of DNS hostnames (subdomains, distinguished from apex domains
|
|
via the Public Suffix List).
|
|
- Every **1 hour**, performs a full iterative trace to discover the
|
|
authoritative nameservers of the zone the hostname is in, which is not always
|
|
its last two labels (a name under `co.uk`, or in a delegated subdomain).
|
|
- Queries **each** authoritative nameserver independently for **all** record
|
|
types: A, AAAA, CNAME, MX, TXT, SRV, CAA, NS.
|
|
- Stores results **per nameserver**. The state for a hostname is not a merged
|
|
view — it is a map from nameserver to record set.
|
|
- DNS names inside record values (CNAME, MX, SRV and NS targets) are stored in
|
|
lower case, because names are case-insensitive and nameservers may answer in
|
|
any letter case. TXT and CAA values keep their letter case; they are not
|
|
lower-cased.
|
|
- Any observable change in any nameserver's response triggers a notification.
|
|
This includes:
|
|
- **Record change**: A nameserver returns different records than it did on
|
|
the previous check (additions, removals, value changes).
|
|
- **NS query failure**: A nameserver that previously responded becomes
|
|
unreachable (timeout, SERVFAIL, REFUSED, network error). This is distinct
|
|
from "responded with no records": a nameserver that answers NXDOMAIN or
|
|
with no records has responded. The alert is sent once, on the check where
|
|
it starts failing. A failing nameserver gives no records, so it is not
|
|
reported as a record change or compared for inconsistency. A nameserver
|
|
that is already failing on the first check that sees it is recorded
|
|
silently.
|
|
- **NS recovery**: A previously-unreachable nameserver starts responding
|
|
again. Its records are not compared with those from before it failed, so a
|
|
change made while it was failing is not reported as a record change.
|
|
- **Inconsistency detected**: Two nameservers return different record sets
|
|
for the same hostname and did not already differ on the previous check.
|
|
Every pair of nameservers is compared. The alert is sent once for each
|
|
such pair, on the check where they start to disagree, and not again while
|
|
they keep disagreeing, including after a restart. A nameserver that was
|
|
not in the previous check (newly added, or back after dropping out), or
|
|
failed on it, and answers differently is reported on the check where it
|
|
answers. If a pair agrees again and later disagrees, the alert is sent
|
|
again.
|
|
|
|
### TCP Port Monitoring
|
|
|
|
- For every configured domain and hostname, constructs a deduplicated list of
|
|
all IPv4 and IPv6 addresses resolved via A, AAAA, and CNAME chain resolution
|
|
across all authoritative nameservers.
|
|
- Checks TCP connectivity on ports **80** and **443** for each IP address.
|
|
- Every **1 hour**, re-checks all ports.
|
|
- Any change in port availability triggers a notification:
|
|
- Port transitioned from open to closed (or vice versa).
|
|
- New IP appeared (from DNS change) and its port state was recorded.
|
|
- IP disappeared (from DNS change) — noted in the DNS change notification;
|
|
port state for that IP is removed.
|
|
|
|
### TLS Certificate Monitoring
|
|
|
|
- Every **12 hours**, for each IP address listening on port 443, connects via
|
|
TLS using the correct SNI hostname.
|
|
- Records the certificate's Subject CN, SANs, issuer, and expiry date.
|
|
- Any change triggers a notification:
|
|
- Certificate is expiring within **7 days** (warning, repeated each check
|
|
until renewed or expired).
|
|
- Certificate CN, issuer, or SANs changed (replacement detected, reports old
|
|
and new values).
|
|
- TLS connection failure to a previously-reachable IP:443 (handshake error,
|
|
timeout, connection refused after previously succeeding).
|
|
- TLS recovery: a previously-failing IP:443 now completes a handshake again.
|
|
|
|
### Notifications
|
|
|
|
**Every observable state change produces a notification.** dnswatcher is
|
|
designed as a real-time change feed — degradations, failures, recoveries, and
|
|
routine changes are all reported equally.
|
|
|
|
Supported notification backends:
|
|
|
|
| Backend | Configuration | Payload Format |
|
|
| -------------- | ------------------------------------------ | ---------------------------- |
|
|
| **Slack** | Incoming Webhook URL | Attachments with color |
|
|
| **Mattermost** | Incoming Webhook URL | Slack-compatible attachments |
|
|
| **ntfy** | Topic URL (e.g. `https://ntfy.sh/mytopic`) | Title + body + priority |
|
|
|
|
All configured endpoints receive every notification. Notification content
|
|
includes:
|
|
|
|
- **DNS record changes**: Which hostname, which nameserver, what record type,
|
|
old values, new values.
|
|
- **DNS NS changes**: Which domain, which nameservers were added/removed.
|
|
- **NS address changes**: Which domain, which nameserver, its old and new
|
|
addresses.
|
|
- **NS query failures**: Which nameserver failed, error type (timeout, SERVFAIL,
|
|
REFUSED, network error), which hostname/domain affected.
|
|
- **NS recoveries**: Which nameserver recovered, which hostname/domain.
|
|
- **NS inconsistencies**: Which nameservers disagree, what each one returned,
|
|
which hostname affected.
|
|
- **Port changes**: Which IP:port, old state, new state, all associated
|
|
hostnames.
|
|
- **TLS expiry warnings**: Which certificate, days remaining, CN, issuer,
|
|
associated hostname and IP.
|
|
- **TLS certificate changes**: Old and new CN/issuer/SANs, associated hostname
|
|
and IP.
|
|
- **TLS connection failures/recoveries**: Which IP:port, error details,
|
|
associated hostname.
|
|
|
|
### State Management
|
|
|
|
- All monitoring state is kept in memory and persisted to a JSON file on disk
|
|
(`DATA_DIR/state.json`).
|
|
- State is loaded on startup to resume monitoring without triggering
|
|
false-positive change notifications.
|
|
- State is written atomically (write to temp file, then rename) to prevent
|
|
corruption.
|
|
|
|
### Web Dashboard
|
|
|
|
dnswatcher includes an unauthenticated, read-only web dashboard at the root URL
|
|
(`/`). It displays:
|
|
|
|
- **Summary counts** for monitored domains, hostnames, ports, and certificates.
|
|
- **Domains** with their discovered nameservers.
|
|
- **Hostnames** with per-nameserver DNS records and status.
|
|
- **Ports** with open/closed state and associated hostnames.
|
|
- **TLS certificates** with CN, issuer, expiry, and status.
|
|
- **Recent alerts** (last 100 notifications sent since the process started),
|
|
displayed in reverse chronological order.
|
|
|
|
Every data point shows its age (e.g. "5m ago") so you can tell at a glance how
|
|
fresh the information is. The page auto-refreshes every 30 seconds.
|
|
|
|
The dashboard intentionally does not expose any configuration details such as
|
|
webhook URLs, notification endpoints, or API tokens.
|
|
|
|
All assets (CSS) are embedded in the binary and served from the application
|
|
itself. The dashboard makes zero external HTTP requests — no CDN dependencies or
|
|
third-party resources are loaded at runtime.
|
|
|
|
### HTTP API
|
|
|
|
dnswatcher exposes a lightweight HTTP API for operational visibility:
|
|
|
|
| Endpoint | Description |
|
|
| ------------------------------ | ----------------------------- |
|
|
| `GET /` | Web dashboard (HTML) |
|
|
| `GET /s/...` | Static assets (embedded CSS) |
|
|
| `GET /.well-known/healthcheck` | Health check (JSON) |
|
|
| `GET /health` | Health check (JSON, legacy) |
|
|
| `GET /api/v1/status` | Current monitoring state |
|
|
| `GET /metrics` | Prometheus metrics (optional) |
|
|
|
|
#### Server timeouts
|
|
|
|
The HTTP server sets all four socket-level timeouts. These are compile-time
|
|
constants in `internal/server/server.go`, not configurable via environment
|
|
variables.
|
|
|
|
| Timeout | Value | Purpose |
|
|
| ------------------- | ----- | --------------------------------------------- |
|
|
| `ReadHeaderTimeout` | 10s | Bounds the request header read (slowloris) |
|
|
| `ReadTimeout` | 15s | Bounds the whole request read, headers + body |
|
|
| `WriteTimeout` | 75s | Bounds handler execution plus response flush |
|
|
| `IdleTimeout` | 120s | Reaps idle keep-alive connections |
|
|
|
|
These are distinct from the 60s per-request handler budget applied by
|
|
`chimw.Timeout` in `internal/server/routes.go`, which cancels the request
|
|
context but does not touch the socket. `WriteTimeout` is deliberately larger
|
|
than that budget: the write deadline is armed once request headers are read, so
|
|
a smaller value would sever the connection before a handler using its full
|
|
budget could respond. `IdleTimeout` exceeds common Prometheus scrape intervals
|
|
so the scraper reuses its connection.
|
|
|
|
### Security Headers
|
|
|
|
Every response — the dashboard, the static assets under `/s/...`, the
|
|
healthchecks, the JSON API, and `/metrics` — carries the following headers, set
|
|
by a global middleware:
|
|
|
|
| Header | Value |
|
|
| --------------------------- | ------------------------------------- |
|
|
| `Strict-Transport-Security` | `max-age=31536000; includeSubDomains` |
|
|
| `Content-Security-Policy` | see below |
|
|
| `X-Frame-Options` | `DENY` |
|
|
| `X-Content-Type-Options` | `nosniff` |
|
|
| `Referrer-Policy` | `no-referrer` |
|
|
| `Permissions-Policy` | all unused browser features denied |
|
|
|
|
The content security policy is:
|
|
|
|
```
|
|
default-src 'self'; script-src 'none'; style-src 'self'; img-src 'self';
|
|
font-src 'none'; connect-src 'none'; object-src 'none'; base-uri 'none';
|
|
form-action 'none'; frame-ancestors 'none'
|
|
```
|
|
|
|
The dashboard ships no JavaScript (the 30-second refresh is a
|
|
`<meta http-equiv="refresh">`), no inline styles, no inline event handlers, and
|
|
no images; its only subresource is the embedded stylesheet at
|
|
`/s/css/tailwind.min.css`, which `style-src 'self'` permits. The policy
|
|
therefore needs neither `unsafe-inline` nor `unsafe-eval`.
|
|
`frame-ancestors 'none'` is the primary anti-framing control, with
|
|
`X-Frame-Options: DENY` retained as the legacy fallback.
|
|
|
|
HSTS is emitted unconditionally, including over plain HTTP. dnswatcher is
|
|
expected to run behind a TLS-terminating reverse proxy, and the browser must
|
|
still be told to enforce HTTPS end to end, so the header is never gated on
|
|
whether the request itself arrived over TLS.
|
|
|
|
`Referrer-Policy: no-referrer` is stricter than the
|
|
`strict-origin-when-cross-origin` baseline: the dashboard has no cross-origin
|
|
navigation needs, and its URL may name internal hosts.
|
|
|
|
---
|
|
|
|
## Architecture
|
|
|
|
```
|
|
cmd/dnswatcher/main.go Entry point (uber/fx bootstrap)
|
|
|
|
internal/
|
|
config/config.go Viper-based configuration
|
|
globals/globals.go Build-time variables (version)
|
|
logger/logger.go slog structured logging (TTY detection)
|
|
healthcheck/healthcheck.go Health check service
|
|
middleware/middleware.go HTTP middleware (logging, CORS, security
|
|
headers, metrics auth and rate limit)
|
|
handlers/handlers.go HTTP request handlers
|
|
server/
|
|
server.go HTTP server lifecycle
|
|
routes.go Route definitions
|
|
state/state.go JSON file state persistence
|
|
resolver/resolver.go Iterative DNS resolution engine
|
|
portcheck/portcheck.go TCP port connectivity checker
|
|
tlscheck/tlscheck.go TLS certificate inspector
|
|
notify/notify.go Notification service (Slack, Mattermost, ntfy)
|
|
watcher/watcher.go Main monitoring orchestrator and scheduler
|
|
livednstest/livednstest.go Retry and concurrency limit for tests
|
|
against live DNS (imported only by tests)
|
|
```
|
|
|
|
### Design Principles
|
|
|
|
- **No recursive resolvers**: All DNS resolution is performed iteratively,
|
|
tracing from root nameservers through the delegation chain to authoritative
|
|
servers.
|
|
- **No external database**: State is persisted as a single JSON file.
|
|
- **Dependency injection**: All components are wired via
|
|
[uber/fx](https://github.com/uber-go/fx).
|
|
- **Structured logging**: All logs use `log/slog` with JSON output in production
|
|
(TTY detection for development).
|
|
- **Graceful shutdown**: All background goroutines respect context cancellation
|
|
and the fx lifecycle. In-flight notification deliveries are drained on
|
|
shutdown, bounded by the shutdown timeout.
|
|
|
|
---
|
|
|
|
## Configuration
|
|
|
|
Configuration is loaded via [Viper](https://github.com/spf13/viper) with the
|
|
following precedence (highest to lowest):
|
|
|
|
1. Environment variables (prefixed with `DNSWATCHER_`)
|
|
2. `.env` file (loaded via godotenv)
|
|
3. Config file: `/etc/dnswatcher/dnswatcher.yaml`,
|
|
`~/.config/dnswatcher/dnswatcher.yaml`, or `./dnswatcher.yaml`
|
|
4. Defaults
|
|
|
|
### Environment Variables
|
|
|
|
| Variable | Description | Default |
|
|
| ----------------------------------- | ----------------------------------------------------------------------------------------------------------- | --------------------- |
|
|
| `PORT` | HTTP listen port | `8080` |
|
|
| `DNSWATCHER_DEBUG` | Enable debug logging | `false` |
|
|
| `DNSWATCHER_DATA_DIR` | Directory for state file | `/var/lib/dnswatcher` |
|
|
| `DNSWATCHER_TARGETS` | Comma-separated DNS names (auto-classified via PSL) | `""` |
|
|
| `DNSWATCHER_SLACK_WEBHOOK` | Slack incoming webhook URL | `""` |
|
|
| `DNSWATCHER_MATTERMOST_WEBHOOK` | Mattermost incoming webhook URL | `""` |
|
|
| `DNSWATCHER_NTFY_TOPIC` | ntfy topic URL | `""` |
|
|
| `DNSWATCHER_DNS_INTERVAL` | DNS check interval, a positive duration such as `30m`; empty means the default, anything else stops startup | `1h` |
|
|
| `DNSWATCHER_TLS_INTERVAL` | TLS check interval, a positive duration such as `6h`; empty means the default, anything else stops startup | `12h` |
|
|
| `DNSWATCHER_TLS_EXPIRY_WARNING` | Days before expiry to warn | `7` |
|
|
| `DNSWATCHER_SENTRY_DSN` | Sentry DSN for error reporting | `""` |
|
|
| `DNSWATCHER_MAINTENANCE_MODE` | Enable maintenance mode | `false` |
|
|
| `DNSWATCHER_METRICS_USERNAME` | Basic auth username for /metrics | `""` |
|
|
| `DNSWATCHER_METRICS_PASSWORD` | Basic auth password for /metrics | `""` |
|
|
| `DNSWATCHER_SEND_TEST_NOTIFICATION` | Send a test notification after first scan completes | `false` |
|
|
|
|
**`DNSWATCHER_TARGETS` is required.** dnswatcher will refuse to start if no
|
|
monitoring targets are configured. A monitoring daemon with nothing to monitor
|
|
is a misconfiguration, so dnswatcher fails fast with a clear error message
|
|
rather than running silently. Set `DNSWATCHER_TARGETS` to a comma-separated list
|
|
of DNS names before starting.
|
|
|
|
**`/metrics` is rate limited.** Each client address may send it 30 requests a
|
|
minute, failed logins included; beyond that it answers `429 Too Many Requests`
|
|
without checking the password. A Prometheus server scraping every 15 seconds
|
|
sends 4 a minute. IPv6 addresses in one /64 count as one client. When the
|
|
request comes from a private or loopback address, such as a reverse proxy's, the
|
|
client address is taken from the `X-Real-IP` header the proxy sets, or else from
|
|
`X-Forwarded-For`, as the last address in it that is not private or loopback. A
|
|
proxy that sets neither makes all its clients share one allowance.
|
|
|
|
**`DNSWATCHER_DNS_INTERVAL` and `DNSWATCHER_TLS_INTERVAL`** take a positive
|
|
duration: a number followed by a unit such as `s`, `m` or `h`, for example
|
|
`90s`, `30m`, `1h` or `1h30m`. There is no unit for days; write `24h`. An unset
|
|
or empty variable (`DNSWATCHER_DNS_INTERVAL=`) means the default. If either is
|
|
set to anything else, including a bare number or a zero or negative duration,
|
|
dnswatcher refuses to start with an error naming the variable and the value.
|
|
|
|
**`DNSWATCHER_SENTRY_DSN` reports crashes in HTTP requests to Sentry.** When it
|
|
is set, a panic in an HTTP request handler is sent to Sentry, and the request
|
|
still gets a `500 Internal Server Error` answer. Nothing else is sent to Sentry:
|
|
DNS, port and TLS problems are reported as notifications. A value Sentry cannot
|
|
parse stops dnswatcher at startup. At shutdown, reports not yet sent are sent,
|
|
waiting at most 2 seconds.
|
|
|
|
### Example `.env`
|
|
|
|
```sh
|
|
PORT=8080
|
|
DNSWATCHER_DEBUG=false
|
|
DNSWATCHER_DATA_DIR=/var/lib/dnswatcher
|
|
DNSWATCHER_TARGETS=example.com,example.org,www.example.com,api.example.com,mail.example.org
|
|
DNSWATCHER_SLACK_WEBHOOK=https://hooks.slack.com/services/T.../B.../xxx
|
|
DNSWATCHER_MATTERMOST_WEBHOOK=https://mattermost.example.com/hooks/xxx
|
|
DNSWATCHER_NTFY_TOPIC=https://ntfy.sh/my-dns-alerts
|
|
DNSWATCHER_SEND_TEST_NOTIFICATION=true
|
|
```
|
|
|
|
---
|
|
|
|
## DNS Resolution Strategy
|
|
|
|
dnswatcher never uses the system's configured recursive resolver. Instead, it
|
|
performs full iterative resolution:
|
|
|
|
1. **Root servers**: Starts from the IANA root nameserver list (hardcoded, with
|
|
periodic refresh).
|
|
2. **TLD delegation**: Queries root servers for the TLD NS records.
|
|
3. **Domain delegation**: Queries TLD nameservers for the domain's NS records.
|
|
4. **Authoritative query**: Queries all discovered authoritative nameservers
|
|
directly for the requested records.
|
|
|
|
This approach ensures:
|
|
|
|
- Independence from any upstream resolver's cache or filtering.
|
|
- Ability to detect split-horizon or inconsistent responses across authoritative
|
|
servers.
|
|
- Visibility into the full delegation chain.
|
|
|
|
For hostname monitoring, the resolver follows CNAME chains (with a depth limit
|
|
to prevent loops) before collecting terminal A/AAAA records.
|
|
|
|
---
|
|
|
|
## State File Format
|
|
|
|
The state file (`DATA_DIR/state.json`) contains the complete monitoring
|
|
snapshot. Hostname records are stored **per authoritative nameserver**, not as a
|
|
merged view, to enable inconsistency detection.
|
|
|
|
```json
|
|
{
|
|
"version": 1,
|
|
"lastUpdated": "2026-02-19T12:00:00Z",
|
|
"domains": {
|
|
"example.com": {
|
|
"nameservers": ["ns1.example.com.", "ns2.example.com."],
|
|
"nameserverAddresses": {
|
|
"ns1.example.com.": ["192.0.2.53", "2001:db8::53"],
|
|
"ns2.example.com.": ["198.51.100.53"]
|
|
},
|
|
"lastChecked": "2026-02-19T12:00:00Z"
|
|
}
|
|
},
|
|
"hostnames": {
|
|
"www.example.com": {
|
|
"recordsByNameserver": {
|
|
"ns1.example.com.": {
|
|
"records": {
|
|
"A": ["93.184.216.34"],
|
|
"AAAA": ["2606:2800:220:1:248:1893:25c8:1946"]
|
|
},
|
|
"status": "ok",
|
|
"lastChecked": "2026-02-19T12:00:00Z"
|
|
},
|
|
"ns2.example.com.": {
|
|
"records": {
|
|
"A": ["93.184.216.34"],
|
|
"AAAA": ["2606:2800:220:1:248:1893:25c8:1946"]
|
|
},
|
|
"status": "ok",
|
|
"lastChecked": "2026-02-19T12:00:00Z"
|
|
}
|
|
},
|
|
"lastChecked": "2026-02-19T12:00:00Z"
|
|
}
|
|
},
|
|
"ports": {
|
|
"93.184.216.34:80": {
|
|
"open": true,
|
|
"hostnames": ["www.example.com"],
|
|
"lastChecked": "2026-02-19T12:00:00Z"
|
|
},
|
|
"93.184.216.34:443": {
|
|
"open": true,
|
|
"hostnames": ["www.example.com"],
|
|
"lastChecked": "2026-02-19T12:00:00Z"
|
|
}
|
|
},
|
|
"certificates": {
|
|
"93.184.216.34:443:www.example.com": {
|
|
"commonName": "www.example.com",
|
|
"issuer": "DigiCert TLS RSA SHA256 2020 CA1",
|
|
"notAfter": "2027-01-15T23:59:59Z",
|
|
"subjectAlternativeNames": ["www.example.com"],
|
|
"status": "ok",
|
|
"lastChecked": "2026-02-19T06:00:00Z"
|
|
}
|
|
}
|
|
}
|
|
```
|
|
|
|
The `status` field for each per-nameserver entry and certificate entry tracks
|
|
reachability:
|
|
|
|
| Status | Meaning |
|
|
| ------- | -------------------------------------------------------- |
|
|
| `ok` | Query succeeded, records are current |
|
|
| `error` | Query failed (timeout, SERVFAIL, REFUSED, network error) |
|
|
|
|
A nameserver that answers NXDOMAIN or with no records has status `ok` and empty
|
|
`records`. A nameserver whose query failed, or that only referred it to other
|
|
nameservers, has status `error`, empty `records`, and the reason in `error`.
|
|
|
|
`nameserverAddresses` lists, by nameserver, the sorted addresses its name
|
|
resolves to. A state file without it loads, and the next check fills it in
|
|
without a notification.
|
|
|
|
---
|
|
|
|
## Entrypoints
|
|
|
|
This repository adheres to the
|
|
[Scripts to Rule Them All](https://github.com/github/scripts-to-rule-them-all)
|
|
standard: normalized scripts in `script/` are the entrypoints for the
|
|
development workflow, and the Makefile targets are thin shims that call them. We
|
|
provide:
|
|
|
|
- `script/bootstrap` — install all dependencies (go, `go mod download`). It does
|
|
not install golangci-lint or prettier: both run in Docker, see `script/lint`
|
|
and `script/fmt` below.
|
|
- `script/setup` — make a fresh clone ready for development: bootstrap plus the
|
|
git pre-commit hook
|
|
- `script/projectname` — print the project name (used for the Docker image tag)
|
|
- `script/test` — run the test suite (race detector, coverage). Caching is
|
|
waived for testing, exactly as it is for linting: `-count=1` forces every
|
|
invocation to execute, because the suite queries live DNS and a cached pass
|
|
queries nothing. Failures are rerun with `-v` automatically, and the build
|
|
fails even if that rerun passes.
|
|
- `script/lint` — run golangci-lint, always inside Docker: it builds
|
|
`Dockerfile.lint`, which COPYs the repo into the digest-pinned `golangci-lint`
|
|
image and lints as a build step, so a successful build is a clean lint. The
|
|
linter is never installed or run on the host, and Docker is the only
|
|
prerequisite. Caching is waived for linting: the lint stage is forced to
|
|
execute on every run with `--no-cache-filter`, because a cached build lints
|
|
nothing.
|
|
- `script/fmt` — format all code (gofmt -s, goimports) and all Markdown
|
|
(prettier). goimports runs with `go run` at a pinned commit, never from your
|
|
`PATH`. prettier runs inside Docker, built from `Dockerfile.fmt` on a
|
|
digest-pinned node image, at the version pinned by `package.json` and
|
|
`yarn.lock`; it is never installed on the host.
|
|
- `script/fmt-check` — check formatting (read-only) with the same tools, failing
|
|
on any file `script/fmt` would change. It runs the two scripts below.
|
|
- `script/fmt-check-go` — the gofmt and goimports half, on the host. The
|
|
`Dockerfile` lint stage runs it.
|
|
- `script/fmt-check-markdown` — the prettier half, inside Docker, forced to
|
|
execute on every run with `--no-cache-filter`
|
|
- `script/check` — run test, lint, and fmt-check
|
|
- `script/docker` — build the Docker image tagged via `script/projectname`, with
|
|
`--no-cache-filter=lint,builder` so the lint stage and the builder stage,
|
|
which runs the tests, run on every invocation, and with the version from
|
|
`git describe` passed as `--build-arg VERSION`
|
|
- `script/cibuild` — CI entrypoint: `docker build` with
|
|
`--no-cache-filter=lint,builder`, so the lint stage and the builder stage,
|
|
which runs the tests, run on every invocation, because a cached build lints
|
|
nothing and queries no DNS; then `script/fmt-check-markdown`
|
|
- `script/precommit` — run by the git pre-commit hook; `go mod tidy` guard, then
|
|
`script/check`
|
|
- `script/install-precommit` — install the git pre-commit hook
|
|
|
|
## Building
|
|
|
|
```sh
|
|
make build # Build binary to bin/dnswatcher
|
|
make test # Run tests with race detector
|
|
make lint # Run golangci-lint in Docker (requires docker)
|
|
make fmt # Format code and Markdown (requires docker)
|
|
make check # Run all checks (test, lint, fmt-check)
|
|
make clean # Remove build artifacts
|
|
```
|
|
|
|
### Build-Time Variables
|
|
|
|
`make build` sets the version with `-ldflags "-X main.Version=..."`, taking it
|
|
from `git describe --tags --always --dirty`, or from `VERSION` when given on the
|
|
command line (`make build VERSION=1.2.3`). The version appears in the startup
|
|
log and in the health check response.
|
|
|
|
The Docker image has no `.git`, so the `Dockerfile` takes the version as
|
|
`--build-arg VERSION`. `make docker` passes it; a plain `docker build` passes
|
|
none, and that image reports `dev`.
|
|
|
|
---
|
|
|
|
## Docker
|
|
|
|
```sh
|
|
docker build -t dnswatcher .
|
|
docker run -d \
|
|
-p 8080:8080 \
|
|
-v dnswatcher-data:/var/lib/dnswatcher \
|
|
-e DNSWATCHER_TARGETS=example.com,www.example.com \
|
|
-e DNSWATCHER_NTFY_TOPIC=https://ntfy.sh/my-alerts \
|
|
-e DNSWATCHER_SEND_TEST_NOTIFICATION=true \
|
|
dnswatcher
|
|
```
|
|
|
|
---
|
|
|
|
## Running under upaas
|
|
|
|
[upaas](https://git.eeqj.de/sneak/upaas) builds the image from this repository's
|
|
`Dockerfile` and runs it. The app needs:
|
|
|
|
- **Branch:** `prod`. `prod` is cut from `main`, and merging a `main` to `prod`
|
|
pull request is a deploy.
|
|
- **Volume:** one host directory mounted at `/var/lib/dnswatcher`, where the
|
|
state file lives.
|
|
- **Network and port:** the dashboard is unauthenticated and shows every watched
|
|
name and recent alert, and upaas publishes every mapped port on all interfaces
|
|
of the host ([upaas issue 113](https://git.eeqj.de/sneak/upaas/issues/113)).
|
|
Add a port mapping to container port `8080` only if the dashboard should be
|
|
public. Otherwise add none: set the app's Docker network in upaas to your
|
|
reverse proxy's Docker network, and the proxy reaches the app at `upaas-`
|
|
followed by the app name, port `8080`.
|
|
- **Required environment:** `DNSWATCHER_TARGETS`, a comma-separated list of the
|
|
domains and hostnames to watch. dnswatcher refuses to start without it.
|
|
- **Recommended environment:** at least one notification endpoint
|
|
(`DNSWATCHER_SLACK_WEBHOOK`, `DNSWATCHER_MATTERMOST_WEBHOOK`,
|
|
`DNSWATCHER_NTFY_TOPIC`); without one, changes show only on the dashboard.
|
|
`DNSWATCHER_METRICS_USERNAME` and `DNSWATCHER_METRICS_PASSWORD` serve
|
|
`/metrics` behind basic auth.
|
|
- **Leave unset:** `DNSWATCHER_DATA_DIR`, which the image sets to
|
|
`/var/lib/dnswatcher`, and `PORT`, which defaults to `8080`. Every setting
|
|
comes from the environment; the image holds no config file.
|
|
- **Health check:** the image's own, which requests `/.well-known/healthcheck`
|
|
every 10 seconds. upaas reads the container's health 60 seconds after a deploy
|
|
and marks the deploy failed unless it is `healthy`.
|
|
|
|
---
|
|
|
|
## Monitoring Lifecycle
|
|
|
|
1. **Startup**: Check that the data directory can be written, and exit with an
|
|
error naming it if not. Load state from disk. If no state file exists, start
|
|
with empty state (first check will establish baseline without triggering
|
|
change notifications).
|
|
2. **Initial check**: Immediately perform all DNS, port, and TLS checks on
|
|
startup.
|
|
3. **Periodic checks** (DNS always runs first):
|
|
- DNS checks: every `DNSWATCHER_DNS_INTERVAL` (default 1h). Also re-run
|
|
before every TLS check cycle to ensure fresh IPs.
|
|
- Port checks: every `DNSWATCHER_DNS_INTERVAL`, after DNS completes.
|
|
- TLS checks: every `DNSWATCHER_TLS_INTERVAL` (default 12h), after DNS
|
|
completes.
|
|
- Port and TLS checks always use freshly resolved IP addresses from the DNS
|
|
phase that immediately precedes them — never stale IPs from a previous
|
|
cycle.
|
|
4. **On change detection**: Send notifications to all configured endpoints,
|
|
update in-memory state, persist to disk.
|
|
5. **Shutdown**: The watcher stops checking and saves the final state to disk,
|
|
and shutdown waits for that save before it goes on. Then it waits for
|
|
in-flight notification deliveries to complete. Both waits share the fx
|
|
shutdown timeout (15s by default): deliveries still retrying against an
|
|
unreachable endpoint when that expires are abandoned, and the number
|
|
abandoned is logged at warn level rather than dropped silently. Notifications
|
|
generated after shutdown has begun are refused and logged, so a late burst
|
|
cannot extend the shutdown. A DNS lookup, port check or TLS check that
|
|
shutdown cuts short saves nothing and sends no notification.
|
|
|
|
---
|
|
|
|
## Planned Future Features (Post-1.0)
|
|
|
|
- **DNSSEC validation**: Validate the DNSSEC chain of trust during iterative
|
|
resolution and report DNSSEC failures as notifications.
|
|
|
|
---
|
|
|
|
## Project Structure
|
|
|
|
Follows the conventions defined in `REPO_POLICIES.md`, adapted from the
|
|
[upaas](https://git.eeqj.de/sneak/upaas) project template. Uses uber/fx for
|
|
dependency injection, go-chi for HTTP routing, slog for logging, and Viper for
|
|
configuration.
|
|
|
|
---
|
|
|
|
## License
|
|
|
|
dnswatcher is released under the MIT License, Copyright (c) 2026
|
|
[@sneak](https://sneak.berlin). See the [`LICENSE`](./LICENSE) file in the
|
|
repository root for the full text.
|
|
|
|
## Author
|
|
|
|
[@sneak](https://sneak.berlin)
|