Files
dnswatcher/README.md
T
clawbot 3182fc99a6
check / check (push) Canceled after 0s
docker: a plain docker build . stamps the git version (closes #210)
A plain `docker build .`, which is how upaas builds, stamped `dev`:
`.dockerignore` left out `.git` and the builder declared
`ARG VERSION=dev`. `.dockerignore` now sends `.git` without
`.git/config`, which can hold a credential, and lists no tracked file,
which git would count as deleted. `ARG VERSION` has no default. The
Makefile takes a non-empty `VERSION` from the command line or the
environment, so a build arg still wins; otherwise `git describe` runs in
the builder, which trusts the checkout whoever owns it, as a context
sent as a tar archive keeps its owners. A new `make version` prints the
version; the build fails when the context carries `.git` and it comes
out empty, `dev` or `unknown`.

Model: opus-5-5
2026-10-02 06:27:47 +02:00

799 lines
38 KiB
Markdown

# dnswatcher
dnswatcher is an MIT-licensed, pre-1.0 Go daemon by
[@sneak](https://sneak.berlin) that monitors DNS records, TCP port availability,
and TLS certificates, delivering real-time change notifications via Slack,
Mattermost, and ntfy webhooks.
> ⚠️ Pre-1.0 software. APIs, configuration, and behavior may change without
> notice.
dnswatcher watches configured DNS domains and hostnames for changes, monitors
TCP port availability, tracks TLS certificate expiry, and delivers real-time
notifications via Slack, Mattermost, and/or ntfy webhooks.
It resolves the names it watches itself via iterative (non-recursive) queries,
tracing from root nameservers to authoritative servers directly—never relying on
upstream recursive resolvers.
State is persisted to a local JSON file so that monitoring survives restarts
without requiring an external database.
---
## No DNS mocking. Ever.
**DNS is never mocked in this project — not in tests, not anywhere else.** No
mock resolvers, no fake DNS servers, no stubbed lookups.
dnswatcher's entire purpose is correct behavior against the real DNS. Tests
exercise real iterative resolution against live nameservers by design; a test
suite that passes against a mock proves nothing about the one thing this program
exists to do.
When live tests are flaky, that is a robustness problem, and it gets fixed with
robustness: retries with backoff, querying multiple independent nameservers,
longer timeouts — or explicit opt-in gating decided by the project owner. Never
with mocks.
Contributions that introduce mocked, faked, or stubbed DNS will be rejected.
---
## Getting Started
You need git and Docker. This builds the image and runs dnswatcher watching
`example.com` and `www.example.com`:
```sh
git clone https://git.eeqj.de/sneak/dnswatcher.git
cd dnswatcher
docker build -t dnswatcher .
docker run -d --name dnswatcher \
-p 8080:8080 \
-v dnswatcher-data:/var/lib/dnswatcher \
-e DNSWATCHER_TARGETS=example.com,www.example.com \
dnswatcher
```
The build also runs the linter and the test suite, which queries live DNS. Once
the container is running, the dashboard is at <http://localhost:8080/>. With no
notification endpoint set, changes show only on the dashboard; see
[Configuration](#configuration) to add one.
---
## Features
### DNS Domain Monitoring (Apex Domains)
- Accepts a list of DNS domain names (apex domains, identified via the
[Public Suffix List](https://publicsuffix.org/)).
- Every **1 hour** by default, performs a full iterative trace from root servers
to discover all authoritative nameservers (NS records) for each domain.
- Queries **every** discovered authoritative nameserver independently.
- Stores the domain's NS record set, as its parent zone's servers delegate it,
and the IPv4 and IPv6 addresses each nameserver's name resolves to.
- Any change triggers a notification:
- NS added to or removed from that set.
- NS address change: a nameserver that stays in the set resolves to
different addresses than on the previous check. A nameserver added or
removed gets only the NS change notification. When the lookup of a
nameserver's addresses fails or finds none, its previous addresses are
kept and nothing is sent.
### DNS Hostname Monitoring (Subdomains)
- Accepts a list of DNS hostnames (subdomains, distinguished from apex domains
via the Public Suffix List).
- Every **1 hour** by default, performs a full iterative trace to discover the
authoritative nameservers of the zone the hostname is in, which is not always
its last two labels (a name under `co.uk`, or in a delegated subdomain).
- Queries **each** authoritative nameserver independently for **all** record
types: A, AAAA, CNAME, MX, TXT, SRV, CAA, NS.
- Stores results **per nameserver**. The state for a hostname is not a merged
view — it is a map from nameserver to record set.
- DNS names inside record values (CNAME, MX, SRV and NS targets) are stored in
lower case, because names are case-insensitive and nameservers may answer in
any letter case. TXT and CAA values keep their letter case; they are not
lower-cased.
- Any observable change in any nameserver's response triggers a notification.
This includes:
- **Record change**: A nameserver returns different records than it did on
the previous check (additions, removals, value changes).
- **NS query failure**: A nameserver that previously responded becomes
unreachable (timeout, SERVFAIL, REFUSED, network error). This is distinct
from "responded with no records": a nameserver that answers NXDOMAIN or
with no records has responded. The alert is sent once, on the check where
it starts failing. A failing nameserver gives no records, so it is not
reported as a record change or compared for inconsistency. A nameserver
that is already failing on the first check that sees it is recorded
silently.
- **NS recovery**: A previously-unreachable nameserver starts responding
again. Its records are not compared with those from before it failed, so a
change made while it was failing is not reported as a record change.
- **Inconsistency detected**: Two nameservers return different record sets
for the same hostname and did not already differ on the previous check.
Every pair of nameservers is compared. The alert is sent once for each
such pair, on the check where they start to disagree, and not again while
they keep disagreeing, including after a restart. A nameserver that was
not in the previous check (newly added, or back after dropping out), or
failed on it, and answers differently is reported on the check where it
answers. If a pair agrees again and later disagrees, the alert is sent
again.
### TCP Port Monitoring
- For every configured domain and hostname, constructs a deduplicated list of
the IPv4 and IPv6 addresses in the A and AAAA records its authoritative
nameservers returned. A CNAME is not followed: a name whose CNAME points into
another zone usually has no addresses here, so its ports and certificate are
not checked.
- Checks TCP connectivity on ports **80** and **443** for each IP address.
- Every **1 hour** by default, re-checks all ports.
- Any change in port availability triggers a notification:
- Port transitioned from open to closed (or vice versa).
- New IP appeared (from DNS change): its port state is recorded without a
port notification; the DNS change notification shows the new address.
- IP disappeared (from DNS change) — noted in the DNS change notification;
port state for that IP is removed. When none of a name's nameservers
answered, its addresses are not known, so the port state saved for them is
kept.
### TLS Certificate Monitoring
- Every **12 hours** by default, for each IP address listening on port 443,
connects via TLS using the correct SNI hostname.
- Records the certificate's Subject CN, SANs, issuer, and expiry date.
- Any change triggers a notification:
- Certificate is expiring within **7 days** by default (warning, repeated
each check until renewed or expired).
- Certificate CN, issuer, or SANs changed (replacement detected, reports old
and new CN and issuer).
- TLS connection failure to a previously-reachable IP:443 (handshake error,
timeout, connection refused after previously succeeding).
- TLS recovery: a previously-failing IP:443 now completes a handshake again.
### Notifications
**Every observable state change produces a notification.** dnswatcher is
designed as a real-time change feed — degradations, failures, recoveries, and
routine changes are all reported equally.
Supported notification backends:
| Backend | Configuration | Payload Format |
| -------------- | ------------------------------------------ | ---------------------------- |
| **Slack** | Incoming Webhook URL | Attachments with color |
| **Mattermost** | Incoming Webhook URL | Slack-compatible attachments |
| **ntfy** | Topic URL (e.g. `https://ntfy.sh/mytopic`) | Title + body + priority |
All configured endpoints receive every notification. Notification content
includes:
- **DNS record changes**: Which hostname, which nameserver, what record type,
old values, new values.
- **DNS NS changes**: Which domain, which nameservers were added/removed.
- **NS address changes**: Which domain, which nameserver, its old and new
addresses.
- **NS query failures**: Which nameserver failed, error type (timeout, SERVFAIL,
REFUSED, network error), which hostname/domain affected.
- **NS recoveries**: Which nameserver recovered, which hostname/domain.
- **NS inconsistencies**: Which nameservers disagree, what each one returned,
which hostname affected.
- **Port changes**: Which IP:port, its new state, all associated hostnames.
- **TLS expiry warnings**: Expiry date and days remaining, CN, associated
hostname and IP.
- **TLS certificate changes**: Old and new CN and issuer, associated hostname
and IP. A change to the SANs alone is notified, but the SANs are not listed.
- **TLS connection failures/recoveries**: Which IP:port, error details,
associated hostname.
Each endpoint is sent each notification on its own, in the background. A
delivery that fails (a network error, no reply within 10 seconds, or an HTTP
status of 400 or more) is retried up to 5 times: the first retry after about 1
second, each wait after that twice as long up to 60 seconds, every wait varied
at random by up to 25%. A delivery still failing after that is logged and
dropped.
The last 100 notifications, delivered or not, are kept in memory for the
dashboard's Recent alerts. They are not saved to the state file, so a restart
clears them.
### State Management
- All monitoring state is kept in memory and persisted to a JSON file on disk
(`DATA_DIR/state.json`).
- State is loaded on startup to resume monitoring without triggering
false-positive change notifications.
- State is written atomically (write to temp file, then rename) to prevent
corruption.
### Web Dashboard
dnswatcher includes an unauthenticated, read-only web dashboard at the root URL
(`/`). It displays:
- **Summary counts** for monitored domains, hostnames, ports, and certificates.
- **Domains** with their discovered nameservers.
- **Hostnames** with per-nameserver DNS records and status.
- **Ports** with open/closed state and associated hostnames.
- **TLS certificates** with CN, issuer, expiry, and status.
- **Recent alerts** (last 100 notifications sent since the process started),
displayed in reverse chronological order.
Every data point shows its age (e.g. "5m ago") so you can tell at a glance how
fresh the information is. The page auto-refreshes every 30 seconds.
The dashboard intentionally does not expose any configuration details such as
webhook URLs, notification endpoints, or API tokens.
All assets (CSS) are embedded in the binary and served from the application
itself. The dashboard makes zero external HTTP requests — no CDN dependencies or
third-party resources are loaded at runtime.
### HTTP API
dnswatcher exposes a lightweight HTTP API for operational visibility:
| Endpoint | Description |
| ------------------------------ | ----------------------------- |
| `GET /` | Web dashboard (HTML) |
| `GET /s/...` | Static assets (embedded CSS) |
| `GET /.well-known/healthcheck` | Health check (JSON) |
| `GET /health` | Health check (JSON, legacy) |
| `GET /api/v1/status` | Current monitoring state |
| `GET /metrics` | Prometheus metrics, see below |
`/metrics` is served only when `DNSWATCHER_METRICS_USERNAME` is set, behind
Basic Auth. It has the Prometheus Go client's default metrics only (Go runtime,
process, and counts of `/metrics` requests); dnswatcher records no metrics of
its own.
Every route but `/metrics` may be read from a page on any origin: a cross-origin
`GET` gets `Access-Control-Allow-Origin: *`. Only `GET` is allowed cross-origin,
and without credentials. `/metrics` sends no CORS headers.
#### Server timeouts
The HTTP server sets all four socket-level timeouts. These are compile-time
constants in `internal/server/server.go`, not configurable via environment
variables.
| Timeout | Value | Purpose |
| ------------------- | ----- | --------------------------------------------- |
| `ReadHeaderTimeout` | 10s | Bounds the request header read (slowloris) |
| `ReadTimeout` | 15s | Bounds the whole request read, headers + body |
| `WriteTimeout` | 75s | Bounds handler execution plus response flush |
| `IdleTimeout` | 120s | Reaps idle keep-alive connections |
These are distinct from the 60s per-request handler budget applied by
`chimw.Timeout` in `internal/server/routes.go`, which cancels the request
context but does not touch the socket. `WriteTimeout` is deliberately larger
than that budget: the write deadline is armed once request headers are read, so
a smaller value would sever the connection before a handler using its full
budget could respond. `IdleTimeout` exceeds common Prometheus scrape intervals
so the scraper reuses its connection.
### Security Headers
Every response — the dashboard, the static assets under `/s/...`, the
healthchecks, the JSON API, and `/metrics` — carries the following headers, set
by a global middleware:
| Header | Value |
| --------------------------- | ------------------------------------- |
| `Strict-Transport-Security` | `max-age=31536000; includeSubDomains` |
| `Content-Security-Policy` | see below |
| `X-Frame-Options` | `DENY` |
| `X-Content-Type-Options` | `nosniff` |
| `Referrer-Policy` | `no-referrer` |
| `Permissions-Policy` | all unused browser features denied |
The content security policy is:
```
default-src 'self'; script-src 'none'; style-src 'self'; img-src 'self';
font-src 'none'; connect-src 'none'; object-src 'none'; base-uri 'none';
form-action 'none'; frame-ancestors 'none'
```
The dashboard ships no JavaScript (the 30-second refresh is a
`<meta http-equiv="refresh">`), no inline styles, no inline event handlers, and
no images; its only subresource is the embedded stylesheet at
`/s/css/tailwind.min.css`, which `style-src 'self'` permits. The policy
therefore needs neither `unsafe-inline` nor `unsafe-eval`.
`frame-ancestors 'none'` is the primary anti-framing control, with
`X-Frame-Options: DENY` retained as the legacy fallback.
HSTS is emitted unconditionally, including over plain HTTP. dnswatcher is
expected to run behind a TLS-terminating reverse proxy, and the browser must
still be told to enforce HTTPS end to end, so the header is never gated on
whether the request itself arrived over TLS.
`Referrer-Policy: no-referrer` is stricter than the
`strict-origin-when-cross-origin` baseline: the dashboard has no cross-origin
navigation needs, and its URL may name internal hosts.
---
## Configuration
Configuration is loaded via [Viper](https://github.com/spf13/viper) with the
following precedence (highest to lowest):
1. Environment variables (prefixed with `DNSWATCHER_`)
2. `.env` file (loaded via godotenv)
3. Config file: `/etc/dnswatcher/dnswatcher.yaml`,
`~/.config/dnswatcher/dnswatcher.yaml`, or `./dnswatcher.yaml`
4. Defaults
### Environment Variables
| Variable | Description | Default |
| ----------------------------------- | ----------------------------------------------------------------------------------------------------------- | --------------------- |
| `PORT` | HTTP listen port | `8080` |
| `DNSWATCHER_DEBUG` | Enable debug logging | `false` |
| `DNSWATCHER_DATA_DIR` | Directory for state file | `/var/lib/dnswatcher` |
| `DNSWATCHER_TARGETS` | Comma-separated DNS names (auto-classified via PSL) | `""` |
| `DNSWATCHER_SLACK_WEBHOOK` | Slack incoming webhook URL | `""` |
| `DNSWATCHER_MATTERMOST_WEBHOOK` | Mattermost incoming webhook URL | `""` |
| `DNSWATCHER_NTFY_TOPIC` | ntfy topic URL | `""` |
| `DNSWATCHER_DNS_INTERVAL` | DNS check interval, a positive duration such as `30m`; empty means the default, anything else stops startup | `1h` |
| `DNSWATCHER_TLS_INTERVAL` | TLS check interval, a positive duration such as `6h`; empty means the default, anything else stops startup | `12h` |
| `DNSWATCHER_TLS_EXPIRY_WARNING` | Days before expiry to warn | `7` |
| `DNSWATCHER_SENTRY_DSN` | Sentry DSN for error reporting | `""` |
| `DNSWATCHER_MAINTENANCE_MODE` | Only sets `maintenanceMode` in the health check response; changes nothing else | `false` |
| `DNSWATCHER_METRICS_USERNAME` | Basic auth username for /metrics, which is served only when this is set | `""` |
| `DNSWATCHER_METRICS_PASSWORD` | Basic auth password for /metrics | `""` |
| `DNSWATCHER_SEND_TEST_NOTIFICATION` | Send a test notification after first scan completes | `false` |
**`DNSWATCHER_TARGETS` is required.** dnswatcher will refuse to start if no
monitoring targets are configured. A monitoring daemon with nothing to monitor
is a misconfiguration, so dnswatcher fails fast with a clear error message
rather than running silently. Set `DNSWATCHER_TARGETS` to a comma-separated list
of DNS names before starting. A name listed more than once, in any letter case
or with a trailing dot, is watched once.
**`/metrics` is rate limited.** Each client address may send it 30 requests a
minute, failed logins included; beyond that it answers `429 Too Many Requests`
without checking the password. A Prometheus server scraping every 15 seconds
sends 4 a minute. IPv6 addresses in one /64 count as one client. When the
request comes from a private or loopback address, such as a reverse proxy's, the
client address is taken from the `X-Real-IP` header the proxy sets, or else from
`X-Forwarded-For`, as the last address in it that is not private or loopback. A
proxy that sets neither makes all its clients share one allowance.
**`DNSWATCHER_DNS_INTERVAL` and `DNSWATCHER_TLS_INTERVAL`** take a positive
duration: a number followed by a unit such as `s`, `m` or `h`, for example
`90s`, `30m`, `1h` or `1h30m`. There is no unit for days; write `24h`. An unset
or empty variable (`DNSWATCHER_DNS_INTERVAL=`) means the default. If either is
set to anything else, including a bare number or a zero or negative duration,
dnswatcher refuses to start with an error naming the variable and the value.
**`DNSWATCHER_SENTRY_DSN` reports crashes in HTTP requests to Sentry.** When it
is set, a panic in an HTTP request handler is sent to Sentry, and the request
still gets a `500 Internal Server Error` answer. Nothing else is sent to Sentry:
DNS, port and TLS problems are reported as notifications. A value Sentry cannot
parse stops dnswatcher at startup. At shutdown, reports not yet sent are sent,
waiting at most 2 seconds.
### Example `.env`
```sh
PORT=8080
DNSWATCHER_DEBUG=false
DNSWATCHER_DATA_DIR=/var/lib/dnswatcher
DNSWATCHER_TARGETS=example.com,example.org,www.example.com,api.example.com,mail.example.org
DNSWATCHER_SLACK_WEBHOOK=https://hooks.slack.com/services/T.../B.../xxx
DNSWATCHER_MATTERMOST_WEBHOOK=https://mattermost.example.com/hooks/xxx
DNSWATCHER_NTFY_TOPIC=https://ntfy.sh/my-dns-alerts
DNSWATCHER_SEND_TEST_NOTIFICATION=true
```
---
## DNS Resolution Strategy
dnswatcher never uses the system's configured recursive resolver for the names
it watches. Instead, it performs full iterative resolution:
1. **Root servers**: Starts from the IPv4 addresses of the 13 root servers,
built into the binary; the list is not refreshed.
2. **TLD delegation**: Queries root servers for the TLD NS records.
3. **Domain delegation**: Queries TLD nameservers for the domain's NS records.
The delegation they give, from the domain's parent zone, is the domain's NS
record set.
4. **Authoritative query**: Queries all discovered authoritative nameservers
directly for the requested records.
In steps 2 and 3 the servers are asked one at a time in a random order, chosen
anew each time, so no one root server gets every first query. A server that does
not reply, refuses the query, or gives an error reply such as SERVFAIL or a
referral that leads no closer to the name is passed over for the next one. When
a referral names a zone's nameservers without their addresses, the addresses of
all of them are looked up, so that each can be asked.
This approach ensures:
- Independence from any upstream resolver's cache or filtering.
- Ability to detect split-horizon or inconsistent responses across authoritative
servers.
CNAME chains are followed (with a depth limit to prevent loops) only to find the
addresses of nameservers. A watched name's records are stored as its nameservers
return them, CNAME included, without following it.
Sending a notification or a Sentry report is the one use of the system's
resolver: the HTTP client looks up the webhook's or Sentry's host name with it.
---
## State File Format
The state file (`DATA_DIR/state.json`) contains the complete monitoring
snapshot. Hostname records are stored **per authoritative nameserver**, not as a
merged view, to enable inconsistency detection.
```json
{
"version": 1,
"lastUpdated": "2026-02-19T12:00:00Z",
"domains": {
"example.com": {
"nameservers": ["ns1.example.com.", "ns2.example.com."],
"nameserverAddresses": {
"ns1.example.com.": ["192.0.2.53", "2001:db8::53"],
"ns2.example.com.": ["198.51.100.53"]
},
"lastChecked": "2026-02-19T12:00:00Z"
}
},
"hostnames": {
"www.example.com": {
"recordsByNameserver": {
"ns1.example.com.": {
"records": {
"A": ["93.184.216.34"],
"AAAA": ["2606:2800:220:1:248:1893:25c8:1946"]
},
"status": "ok",
"lastChecked": "2026-02-19T12:00:00Z"
},
"ns2.example.com.": {
"records": {
"A": ["93.184.216.34"],
"AAAA": ["2606:2800:220:1:248:1893:25c8:1946"]
},
"status": "ok",
"lastChecked": "2026-02-19T12:00:00Z"
}
},
"lastChecked": "2026-02-19T12:00:00Z"
}
},
"ports": {
"93.184.216.34:80": {
"open": true,
"hostnames": ["www.example.com"],
"lastChecked": "2026-02-19T12:00:00Z"
},
"93.184.216.34:443": {
"open": true,
"hostnames": ["www.example.com"],
"lastChecked": "2026-02-19T12:00:00Z"
}
},
"certificates": {
"93.184.216.34:443:www.example.com": {
"commonName": "www.example.com",
"issuer": "DigiCert TLS RSA SHA256 2020 CA1",
"notAfter": "2027-01-15T23:59:59Z",
"subjectAlternativeNames": ["www.example.com"],
"status": "ok",
"lastChecked": "2026-02-19T06:00:00Z"
}
}
}
```
The `status` field for each per-nameserver entry and certificate entry tracks
reachability:
| Status | Meaning |
| ------- | -------------------------------------------------------- |
| `ok` | Query succeeded, records are current |
| `error` | Query failed (timeout, SERVFAIL, REFUSED, network error) |
A nameserver that answers NXDOMAIN or with no records has status `ok` and empty
`records`. A nameserver whose query failed, or that only referred it to other
nameservers, has status `error`, empty `records`, and the reason in `error`. A
certificate entry whose TLS connection or handshake failed likewise has status
`error`, the reason in `error`, and the certificate fields left empty or zero.
`nameserverAddresses` lists, by nameserver, the sorted addresses its name
resolves to. A state file without it loads, and the next check fills it in
without a notification.
A port entry in the older format, with one `hostname` instead of the `hostnames`
list, loads as a list of that one name.
---
## Entrypoints
This repository adheres to the
[Scripts to Rule Them All](https://github.com/github/scripts-to-rule-them-all)
standard: normalized scripts in `script/` are the entrypoints for the
development workflow, and the Makefile targets are thin shims that call them. We
provide:
- `script/bootstrap` — install all dependencies (go, `go mod download`). It does
not install golangci-lint or prettier: both run in Docker, see `script/lint`
and `script/fmt` below.
- `script/setup` — make a fresh clone ready for development: bootstrap plus the
git pre-commit hook
- `script/projectname` — print the project name (used for the Docker image tag)
- `script/test` — run the test suite (race detector, coverage). Caching is
waived for testing, exactly as it is for linting: `-count=1` forces every
invocation to execute, because the suite queries live DNS and a cached pass
queries nothing. Failures are rerun with `-v` automatically, and the build
fails even if that rerun passes.
- `script/lint` — run golangci-lint, always inside Docker: it builds
`Dockerfile.lint`, which COPYs the repo into the digest-pinned `golangci-lint`
image and lints as a build step, so a successful build is a clean lint. The
linter is never installed or run on the host, and Docker is the only
prerequisite. Caching is waived for linting: the lint stage is forced to
execute on every run with `--no-cache-filter`, because a cached build lints
nothing.
- `script/fmt` — format all code (gofmt -s, goimports) and all Markdown
(prettier). goimports runs with `go run` at a pinned commit, never from your
`PATH`. prettier runs inside Docker, built from `Dockerfile.fmt` on a
digest-pinned node image, at the version pinned by `package.json` and
`yarn.lock`; it is never installed on the host.
- `script/fmt-check` — check formatting (read-only) with the same tools, failing
on any file `script/fmt` would change. It runs the two scripts below.
- `script/fmt-check-go` — the gofmt and goimports half, on the host. The
`Dockerfile` lint stage runs it.
- `script/fmt-check-markdown` — the prettier half, inside Docker, forced to
execute on every run with `--no-cache-filter`
- `script/check` — run test, lint, and fmt-check
- `script/docker` — build the Docker image tagged via `script/projectname`, with
`--no-cache-filter=lint,builder` so the lint stage and the builder stage,
which runs the tests, run on every invocation, and with the version from
`git describe` passed as `--build-arg VERSION`
- `script/cibuild` — CI entrypoint: `docker build` with
`--no-cache-filter=lint,builder`, so the lint stage and the builder stage,
which runs the tests, run on every invocation, because a cached build lints
nothing and queries no DNS; then `script/fmt-check-markdown`
- `script/precommit` — run by the git pre-commit hook; `go mod tidy` guard, then
`script/check`
- `script/install-precommit` — install the git pre-commit hook
## Building
```sh
make build # Build binary to bin/dnswatcher
make version # Print the version make build stamps
make test # Run tests with race detector
make lint # Run golangci-lint in Docker (requires docker)
make fmt # Format code and Markdown (requires docker)
make check # Run all checks (test, lint, fmt-check)
make clean # Remove build artifacts
```
### Build-Time Variables
`make build` sets the version with `-ldflags "-X main.Version=..."`, taking it
from `VERSION` when given on the command line (`make build VERSION=1.2.3`) or in
the environment, otherwise from `git describe --tags --always --dirty`, and
`dev` without git metadata. An empty `VERSION` counts as not given. The version
appears in the startup log and in the health check response.
The image takes it the same way, from the `.git` the build context carries, so a
plain `docker build .` of a clone stamps the commit it was built from; a clone
without tags stamps the short commit. A clone made with `--depth 1` carries at
most a tag on its own commit, so such a clone of an untagged commit stamps the
short commit. In a build from a directory, `.dockerignore` keeps out
`.git/config`, which `git describe` does not need and which can hold a
credential. Docker does not apply `.dockerignore` to a context sent as a tar
archive, as upaas sends it, so that context carries `.git/config` into the
build. It also keeps its files' owners, so git in the build trusts the checkout
whoever owns it. A non-empty `--build-arg VERSION=...` takes precedence;
`make docker` passes the version `git describe` gives on the host. The build
fails when the context carries `.git`, as a directory or as a file, and the
version comes out empty, `dev` or `unknown`. `.dockerignore` must list no
tracked file: git in the build would see it as deleted and mark the version
`-dirty`.
---
## Docker
```sh
docker build -t dnswatcher .
docker run -d \
-p 8080:8080 \
-v dnswatcher-data:/var/lib/dnswatcher \
-e DNSWATCHER_TARGETS=example.com,www.example.com \
-e DNSWATCHER_NTFY_TOPIC=https://ntfy.sh/my-alerts \
-e DNSWATCHER_SEND_TEST_NOTIFICATION=true \
dnswatcher
```
---
## Running under upaas
[upaas](https://git.eeqj.de/sneak/upaas) builds the image from this repository's
`Dockerfile` and runs it. The app needs:
- **Branch:** `prod`. `prod` is cut from `main`, and merging a `main` to `prod`
pull request is a deploy.
- **Volume:** one host directory mounted at `/var/lib/dnswatcher`, where the
state file lives.
- **Network and port:** the dashboard is unauthenticated and shows every watched
name and recent alert, and upaas publishes every mapped port on all interfaces
of the host ([upaas issue 113](https://git.eeqj.de/sneak/upaas/issues/113)).
Add a port mapping to container port `8080` only if the dashboard should be
public. Otherwise add none: set the app's Docker network in upaas to your
reverse proxy's Docker network, and the proxy reaches the app at `upaas-`
followed by the app name, port `8080`.
- **Required environment:** `DNSWATCHER_TARGETS`, a comma-separated list of the
domains and hostnames to watch. dnswatcher refuses to start without it.
- **Recommended environment:** at least one notification endpoint
(`DNSWATCHER_SLACK_WEBHOOK`, `DNSWATCHER_MATTERMOST_WEBHOOK`,
`DNSWATCHER_NTFY_TOPIC`); without one, changes show only on the dashboard.
`DNSWATCHER_METRICS_USERNAME` and `DNSWATCHER_METRICS_PASSWORD` serve
`/metrics` behind basic auth.
- **Leave unset:** `DNSWATCHER_DATA_DIR`, which the image sets to
`/var/lib/dnswatcher`, and `PORT`, which defaults to `8080`. Every setting
comes from the environment; the image holds no config file.
- **Health check:** the image's own, which requests `/.well-known/healthcheck`
every 10 seconds. upaas reads the container's health 60 seconds after a deploy
and marks the deploy failed unless it is `healthy`.
---
## Monitoring Lifecycle
1. **Startup**: Check that the data directory can be written, and exit with an
error naming it if not. Load state from disk. If no state file exists, start
with empty state (first check will establish baseline without triggering
change notifications).
2. **Initial check**: Immediately perform all DNS, port, and TLS checks on
startup.
3. **Periodic checks** (DNS always runs first):
- DNS checks: every `DNSWATCHER_DNS_INTERVAL` (default 1h). Also re-run
before every TLS check cycle to ensure fresh IPs.
- Port checks: every `DNSWATCHER_DNS_INTERVAL`, after DNS completes.
- TLS checks: every `DNSWATCHER_TLS_INTERVAL` (default 12h), after DNS
completes.
- Port and TLS checks use the IP addresses found by the DNS phase that
immediately precedes them. When that phase cannot find a name's
nameservers at all, the addresses an earlier check saved for the name are
used.
4. **On change detection**: Send notifications to all configured endpoints,
update in-memory state, persist to disk.
5. **Shutdown**: The watcher stops checking and saves the final state to disk,
and shutdown waits for that save before it goes on. Then it waits for
in-flight notification deliveries to complete. Both waits share the fx
shutdown timeout (15s by default): deliveries still retrying against an
unreachable endpoint when that expires are abandoned, and the number
abandoned is logged at warn level rather than dropped silently. Notifications
generated after shutdown has begun are refused and logged, so a late burst
cannot extend the shutdown. A DNS lookup, port check or TLS check that
shutdown cuts short saves nothing and sends no notification.
---
## Planned Future Features (Post-1.0)
- **DNSSEC validation**: Validate the DNSSEC chain of trust during iterative
resolution and report DNSSEC failures as notifications.
---
## Project Structure
Follows the conventions defined in `REPO_POLICIES.md`, adapted from the
[upaas](https://git.eeqj.de/sneak/upaas) project template. Uses uber/fx for
dependency injection, go-chi for HTTP routing, slog for logging, and Viper for
configuration.
---
## Rationale
dnswatcher exists to report changes to the DNS records, TCP port availability
and TLS certificates of its configured domains and hostnames, failures and
recoveries included: it is designed as a real-time change feed. It queries the
authoritative nameservers directly, tracing from the root, instead of a
recursive resolver, so no resolver's cache or filtering hides a change and
nameservers that disagree with each other are seen. Its state is a single JSON
file, so it survives a restart without an external database.
---
## Design
```
cmd/dnswatcher/main.go Entry point (uber/fx bootstrap)
internal/
config/
config.go Viper-based configuration
classify.go Splits targets into domains and hostnames
(Public Suffix List)
globals/globals.go Build-time variables (version)
logger/logger.go slog structured logging (TTY detection)
healthcheck/healthcheck.go Health check service
middleware/middleware.go HTTP middleware (logging, CORS, security
headers, metrics auth and rate limit)
handlers/
handlers.go Shared handler setup and JSON responses
dashboard.go Web dashboard
templates/dashboard.html Dashboard template (embedded)
status.go /api/v1/status
healthcheck.go Health check handler
server/
server.go HTTP server lifecycle
routes.go Route definitions
state/state.go JSON file state persistence
resolver/
resolver.go Resolver setup and query status values
iterative.go Iterative DNS resolution engine
dns_client.go UDP and TCP DNS clients
errors.go Resolver errors
portcheck/portcheck.go TCP port connectivity checker
tlscheck/tlscheck.go TLS certificate inspector
notify/
notify.go Notification service (Slack, Mattermost, ntfy)
retry.go Delivery retries with backoff
history.go Last 100 notifications, for the dashboard
shutdown.go Waits for deliveries at shutdown
watcher/
watcher.go Main monitoring orchestrator and scheduler
interfaces.go The resolver, checkers and notifier it uses
livednstest/livednstest.go Retry and concurrency limit for tests
against live DNS (imported only by tests)
static/
static.go Embeds the CSS served under /s/
css/tailwind.min.css Dashboard stylesheet
```
### Design Principles
- **No recursive resolvers**: The watched names are resolved iteratively,
tracing from root nameservers through the delegation chain to authoritative
servers.
- **No external database**: State is persisted as a single JSON file.
- **Dependency injection**: All components are wired via
[uber/fx](https://github.com/uber-go/fx).
- **Structured logging**: All logs use `log/slog` with JSON output in production
(TTY detection for development).
- **Graceful shutdown**: All background goroutines respect context cancellation
and the fx lifecycle. In-flight notification deliveries are drained on
shutdown, bounded by the shutdown timeout.
---
## TODO
[`TODO.md`](./TODO.md) names the next step and the steps planned after it. The
work for 1.0 is tracked as issues on the
[1.0 milestone](https://git.eeqj.de/sneak/dnswatcher/milestone/7).
---
## License
dnswatcher is released under the MIT License, Copyright (c) 2026
[@sneak](https://sneak.berlin). See the [`LICENSE`](./LICENSE) file in the
repository root for the full text.
## Author
[@sneak](https://sneak.berlin)