Compare commits
12 Commits
67903257e8
...
clawbot/do
| Author | SHA1 | Date | |
|---|---|---|---|
| def52ae092 | |||
| d61d9dc1c1 | |||
| b0a011f6b4 | |||
| 5976a4a98f | |||
| b2c9acdaa6 | |||
| af3703d748 | |||
| 322d9a6d6b | |||
| b9f7db6901 | |||
| 48cf93ec7e | |||
| 8d64259283 | |||
| bde32d3ee6 | |||
| 62576f6fc6 |
12
Dockerfile
12
Dockerfile
@@ -109,6 +109,18 @@ USER webhooker
|
||||
|
||||
EXPOSE 8080
|
||||
|
||||
# The binary defaults BIND_ADDRESS to 127.0.0.1, which is right for a
|
||||
# bare host: the cleartext listener serves the admin UI and the
|
||||
# unauthenticated receiver, so it must not appear on every interface
|
||||
# of a machine that configured nothing. A container is the other case.
|
||||
# Its network namespace is already the isolation boundary, so binding
|
||||
# every address inside it exposes nothing; what decides exposure is
|
||||
# the publish flag, and `-p 127.0.0.1:8080:8080` is the operator's
|
||||
# control there. Shipping the image on loopback would buy no security
|
||||
# and would make the process unreachable through its own published
|
||||
# port.
|
||||
ENV BIND_ADDRESS=0.0.0.0
|
||||
|
||||
HEALTHCHECK --interval=30s --timeout=3s --start-period=5s --retries=3 \
|
||||
CMD wget --no-verbose --tries=1 --spider http://localhost:8080/.well-known/healthcheck || exit 1
|
||||
|
||||
|
||||
522
README.md
522
README.md
@@ -7,6 +7,13 @@ services, durably stores them, and delivers them to configured targets
|
||||
with retry support, logging, and observability. Category: infrastructure
|
||||
/ web service. License: MIT.
|
||||
|
||||
Each entrypoint is a version 4 UUID served at `/webhook/{uuid}`, and
|
||||
that UUID is the entrypoint's only credential. webhooker does not use
|
||||
shared secrets, HMAC signatures or token headers on the receiver, and
|
||||
will not add them — read
|
||||
[The entrypoint URL is the authentication secret](#the-entrypoint-url-is-the-authentication-secret)
|
||||
before deploying one.
|
||||
|
||||
## Getting Started
|
||||
|
||||
### Prerequisites
|
||||
@@ -71,8 +78,18 @@ make clean # Remove bin/
|
||||
### Configuration
|
||||
|
||||
All configuration is via environment variables. For local development,
|
||||
you can place variables in a `.env` file in the project root (loaded
|
||||
automatically via `godotenv/autoload`).
|
||||
you can place variables in a `.env` file in the process working
|
||||
directory, read once at startup before anything else looks at the
|
||||
environment.
|
||||
|
||||
The file is optional and having none is the normal case for a
|
||||
deployment. A file that is there but cannot be parsed aborts startup
|
||||
with a message naming it, because a single malformed line makes none
|
||||
of the file apply: every variable in it silently reverts to its
|
||||
default, which is exactly the failure [Invalid values abort
|
||||
startup](#invalid-values-abort-startup) exists to prevent, for all of
|
||||
them at once. A variable already present in the real environment wins
|
||||
over the file's value for the same name.
|
||||
|
||||
The environment is selected by setting `WEBHOOKER_ENVIRONMENT` to `dev`
|
||||
or `prod` (default: `dev`). The setting controls exactly one behavior:
|
||||
@@ -119,12 +136,13 @@ TTY detection, and security headers are always applied.
|
||||
| ----------------------- | ----------------------------------- | -------- |
|
||||
| `WEBHOOKER_ENVIRONMENT` | `dev` or `prod` | `dev` |
|
||||
| `PORT` | HTTP listen port | `8080` |
|
||||
| `BIND_ADDRESS` | IP address the HTTP listener binds. Loopback by default, so the cleartext listener is not published on every interface. The Docker image ships `0.0.0.0` instead. See [Bind address](#bind-address) | `127.0.0.1` (image: `0.0.0.0`) |
|
||||
| `DATA_DIR` | Directory for all SQLite databases | `/var/lib/webhooker` |
|
||||
| `DEBUG` | Enable debug logging | `false` |
|
||||
| `MAINTENANCE_MODE` | Report `maintenanceMode: true` in the healthcheck JSON. It does not change how any request is served — no maintenance page exists | `false` |
|
||||
| `METRICS_USERNAME` | Basic auth username for `/metrics`. Must be set together with `METRICS_PASSWORD`; one without the other fails startup | `""` |
|
||||
| `METRICS_PASSWORD` | Basic auth password for `/metrics`. Must be set together with `METRICS_USERNAME`; one without the other fails startup | `""` |
|
||||
| `SENTRY_DSN` | Sentry error reporting DSN | `""` |
|
||||
| `SENTRY_DSN` | Sentry error reporting DSN. Unset leaves error reporting off; a value the Sentry SDK cannot parse fails startup rather than serving with reporting silently off | `""` |
|
||||
| `RETENTION_SWEEP_INTERVAL` | How often the retention reaper and archive sweeper run (Go duration, must be positive) | `1h` |
|
||||
| `SESSION_IDLE_TIMEOUT` | Idle session timeout (Go duration) | `24h` |
|
||||
| `RECEIVER_RATE_LIMIT` | Receiver requests/minute per IP per entrypoint (10x that per IP across the route) | `120` |
|
||||
@@ -240,6 +258,68 @@ A set but unparseable value aborts startup. When the list is non-empty
|
||||
webhooker logs it at startup, blocks and all, so the hole is visible in
|
||||
the log of any deployment that has one.
|
||||
|
||||
#### Bind address
|
||||
|
||||
`BIND_ADDRESS` is the IP address the HTTP listener binds. The binary
|
||||
defaults to `127.0.0.1`, so a bare webhooker is reachable only from the
|
||||
host it runs on. The Docker image ships `ENV BIND_ADDRESS=0.0.0.0`
|
||||
instead — see below for why the two differ.
|
||||
|
||||
That listener speaks **cleartext**, and it serves both the admin UI and
|
||||
the unauthenticated webhook receiver. webhooker terminates no TLS
|
||||
itself; a production deployment puts a reverse proxy in front of it
|
||||
(see
|
||||
[Deployment behind a reverse proxy](#deployment-behind-a-reverse-proxy)),
|
||||
and the proxy reaches it over loopback. A default that bound every
|
||||
interface would leave that cleartext port answering the internet
|
||||
alongside the proxy — the admin login form and the receiver, in the
|
||||
clear, on a port nobody chose to publish. Reaching webhooker from
|
||||
another host is therefore something you configure, not something you
|
||||
get by default.
|
||||
|
||||
**In a container the answer is `0.0.0.0`, which is why the image ships
|
||||
that.** A container's network namespace is already the boundary the
|
||||
loopback default is reaching for: nothing outside the container gets to
|
||||
`0.0.0.0:8080` because of the namespace, whatever the process bound.
|
||||
Exposure is decided at the publish flag instead — `-p
|
||||
127.0.0.1:8080:8080` rather than `-p 8080:8080` — which is the
|
||||
operator's to choose and is what
|
||||
[Running with Docker](#running-with-docker) shows. A loopback bind
|
||||
inside a container buys nothing and makes the process unreachable
|
||||
through its own published port.
|
||||
|
||||
The value must be an IP address literal:
|
||||
|
||||
- `127.0.0.1` — loopback only (the binary's default). Use this with a
|
||||
reverse proxy on the same host.
|
||||
- `0.0.0.0` — every IPv4 address. The image's default; on a bare host,
|
||||
only behind a firewall on the port.
|
||||
- `::` — every address, IPv6 and (on Linux, with the default
|
||||
`net.ipv6.bindv6only=0`) IPv4 as well.
|
||||
- A specific address such as `10.0.0.5` — that interface only.
|
||||
|
||||
An **empty** value is treated as unset, as everywhere else here, and
|
||||
takes the default. In a container that matters: `BIND_ADDRESS=` throws
|
||||
away the image's `0.0.0.0` and falls back to the binary's
|
||||
`127.0.0.1`, which is the one quiet failure this setting has — see
|
||||
[Running with Docker](#running-with-docker).
|
||||
|
||||
Hostnames are **not** accepted. `localhost` aborts startup rather than
|
||||
being resolved: which of `127.0.0.1` and `::1` it means differs by
|
||||
host, a name can resolve to several addresses of which only one could
|
||||
be bound, and the answer can change under a running process. A value
|
||||
carrying a port (`127.0.0.1:8080`) is likewise rejected — the port is
|
||||
`PORT`'s business. Any unparseable value aborts startup; see
|
||||
[Invalid values abort startup](#invalid-values-abort-startup).
|
||||
|
||||
An address that parses but is not assigned to this host — say
|
||||
`10.0.0.5` on a machine that has no such interface — is a valid
|
||||
literal, so it reaches the listener and fails there. The process logs
|
||||
the bind error and exits non-zero rather than staying up with nothing
|
||||
listening. The effective value is in the `bindAddress` field of the
|
||||
startup log line, which is the way to check what a running deployment
|
||||
actually bound.
|
||||
|
||||
#### Metrics credentials
|
||||
|
||||
`METRICS_USERNAME` and `METRICS_PASSWORD` are set together or not at
|
||||
@@ -397,11 +477,24 @@ additionally be a number in the range 1–65535,
|
||||
`RECEIVER_RATE_LIMIT` must be at least 1,
|
||||
`RETENTION_SWEEP_INTERVAL` must be greater than zero (it is a ticker
|
||||
period, so `0s` or a negative value would crash the reaper after
|
||||
startup), and every entry in `TRUSTED_PROXIES` and
|
||||
`ALLOWED_EGRESS_CIDRS` must be a CIDR block or a bare IP address.
|
||||
startup), every entry in `TRUSTED_PROXIES` and
|
||||
`ALLOWED_EGRESS_CIDRS` must be a CIDR block or a bare IP address, and
|
||||
`BIND_ADDRESS` must be an IP address literal — `localhost`,
|
||||
`127.0.0.1:8080` and `10.0.0.0/8` are each rejected rather than
|
||||
resolved, split, or narrowed to something they do not say — and
|
||||
`SENTRY_DSN` must parse as a Sentry DSN.
|
||||
`SESSION_IDLE_TIMEOUT` is the exception: a
|
||||
non-positive value there means idle expiry is disabled, not invalid.
|
||||
|
||||
`SENTRY_DSN` is checked with the Sentry SDK's own parser, the same call
|
||||
the SDK makes on the DSN it is later handed, so what configuration
|
||||
accepts is exactly what will initialise. A typo in it is the one
|
||||
configuration mistake nothing downstream can ever notice — the variable
|
||||
is still set, so every later signal reports error reporting as on while
|
||||
no report is being sent — which is why it aborts rather than starting
|
||||
with reporting off. Leaving it unset is not a mistake and not affected:
|
||||
error reporting is simply off and startup is normal.
|
||||
|
||||
Boolean variables (`DEBUG`, `MAINTENANCE_MODE`) accept exactly the
|
||||
spellings Go's `strconv.ParseBool` accepts — `1`, `t`, `T`, `TRUE`,
|
||||
`true`, `True`, `0`, `f`, `F`, `FALSE`, `false`, `False` — and nothing
|
||||
@@ -543,12 +636,54 @@ decision:
|
||||
|
||||
```bash
|
||||
docker run -d \
|
||||
-p 8080:8080 \
|
||||
-p 127.0.0.1:8080:8080 \
|
||||
-v /path/to/data:/var/lib/webhooker \
|
||||
-e WEBHOOKER_ENVIRONMENT=prod \
|
||||
-e BIND_ADDRESS=0.0.0.0 \
|
||||
webhooker:latest
|
||||
```
|
||||
|
||||
**The image and the bare binary default `BIND_ADDRESS` differently, on
|
||||
purpose.** The binary defaults to `127.0.0.1`; the image ships
|
||||
`ENV BIND_ADDRESS=0.0.0.0`, so the `-e BIND_ADDRESS=0.0.0.0` above is
|
||||
belt-and-braces and the command works without it.
|
||||
|
||||
The two cases are not the same question. On a bare host, `0.0.0.0`
|
||||
puts the cleartext admin UI and the unauthenticated receiver on every
|
||||
interface of the machine, which is what the loopback default exists to
|
||||
prevent. In a container, the network namespace is already that
|
||||
boundary: nothing outside reaches `0.0.0.0:8080` because of the
|
||||
namespace, not because of the bind. What decides exposure there is the
|
||||
**publish flag**, and that is the line to get right.
|
||||
|
||||
So publish to `127.0.0.1:8080` rather than `8080`. A bare
|
||||
`-p 8080:8080` opens the port on every interface of the host — through
|
||||
firewall rules too, since Docker's forwarding rules are inserted ahead
|
||||
of most host firewalls. Publish to the host address your reverse proxy
|
||||
connects from, and nothing wider.
|
||||
|
||||
**An empty `BIND_ADDRESS` is treated as unset**, like every other
|
||||
variable here, so `-e BIND_ADDRESS=` does not mean "keep the image
|
||||
default" — it discards the image's `0.0.0.0` and falls back to the
|
||||
_binary's_ `127.0.0.1`. In a container that is the failure below, and
|
||||
nothing in the logs names the variable. A templated Compose file or a
|
||||
`.env` line with an empty value is the usual way in. Either set a
|
||||
literal or leave the variable out entirely.
|
||||
|
||||
Overriding `BIND_ADDRESS` to a loopback address in a container — by
|
||||
that route or deliberately — makes the container unreachable from
|
||||
outside its namespace even with `-p`. The published port answers
|
||||
nothing, and the health check fails too: it requests
|
||||
`http://localhost:8080`, `localhost` resolves to `::1` first, and a
|
||||
`127.0.0.1` bind is not listening there. The container then goes
|
||||
`unhealthy` about **65 seconds** after start — from `HEALTHCHECK
|
||||
--start-period=5s --interval=30s --retries=3`, so failing probes at
|
||||
5s, 35s and 65s, and `unhealthy` on the third. (Docker's probe cadence
|
||||
during the start period has changed between versions; re-derive from
|
||||
those three values rather than trusting the figure. Measured at 65s on
|
||||
Docker 29.7.2.) A container `unhealthy` with `connection refused` in
|
||||
its health log, or a published port that resets connections, is this.
|
||||
|
||||
The container runs as a non-root user (`webhooker`, UID 1000), exposes
|
||||
port 8080, and includes a health check against
|
||||
`/.well-known/healthcheck`. The `/var/lib/webhooker` volume holds all
|
||||
@@ -558,6 +693,189 @@ databases written by `database` targets (`archive-{uuid}.db`). Mount
|
||||
this as a persistent volume to preserve data across container
|
||||
restarts.
|
||||
|
||||
**The bind-mounted directory must be owned by UID 1000, or the
|
||||
container does not start.** Docker creates a `-v` source path that
|
||||
does not exist yet as `root:root`, and the process runs as UID 1000,
|
||||
so it cannot take its `DATA_DIR` lock:
|
||||
|
||||
```
|
||||
webhooker: locking data directory /var/lib/webhooker: open
|
||||
/var/lib/webhooker/webhooker.lock: permission denied
|
||||
```
|
||||
|
||||
It exits non-zero at that point, before opening any database. Create
|
||||
the directory ahead of the first `docker run`:
|
||||
|
||||
```bash
|
||||
mkdir -p /path/to/data
|
||||
chown 1000:1000 /path/to/data
|
||||
chmod 750 /path/to/data
|
||||
```
|
||||
|
||||
The same `chown` is what a restore needs — see step 4 of
|
||||
[Restore](#restore). A **named volume** does not have this problem:
|
||||
Docker copies the image's ownership onto a volume it initializes, and
|
||||
the image creates `/var/lib/webhooker` owned by `webhooker`.
|
||||
|
||||
**The file modes are not yours to set, and do not depend on the
|
||||
directory.** `webhooker.db` holds target configuration in plaintext —
|
||||
bearer tokens, API keys, Slack webhook URLs — along with the session
|
||||
encryption key, so webhooker creates every SQLite file it owns `0600`:
|
||||
each database and both of its `-wal` and `-shm` sidecars, across all
|
||||
three tiers. Files an earlier build left `0644` are tightened when
|
||||
they are opened. A `DATA_DIR` webhooker creates itself is `0750`, but
|
||||
a bind mount supplies its own directory and Docker's default for one
|
||||
it creates is `0755`; the `0600` files hold there regardless. The
|
||||
`chmod 750` above is defence in depth — it stops other local users
|
||||
listing the directory and learning your webhook UUIDs from the
|
||||
`events-{uuid}.db` filenames — not the barrier protecting the
|
||||
credentials.
|
||||
|
||||
## Deployment behind a reverse proxy
|
||||
|
||||
webhooker terminates no TLS of its own. It serves plaintext HTTP and
|
||||
expects a reverse proxy in front of it, which is the deployment it is
|
||||
built for: the proxy holds the certificate, and webhooker binds
|
||||
loopback where only the proxy can reach it.
|
||||
|
||||
Five things have to be right. Each one is silent when it is wrong —
|
||||
the service comes up, serves pages, and is broken in a way nothing
|
||||
reports.
|
||||
|
||||
1. **Bind or firewall the app port.** The binary binds `127.0.0.1` by
|
||||
default, so the cleartext listener is not published beside the
|
||||
proxy. The image binds `0.0.0.0` inside its own network namespace
|
||||
and relies on the publish address instead —
|
||||
`-p 127.0.0.1:8080:8080`. Either way the port must reach the proxy
|
||||
and nothing else; widen it only with a firewall or a publish
|
||||
address in front of it. A cleartext port answering the internet
|
||||
serves the admin login form and the unauthenticated receiver with
|
||||
no TLS at all, and the proxy in front of it changes nothing about
|
||||
that.
|
||||
2. **Set `WEBHOOKER_ENVIRONMENT=prod`, and make sure the proxy sends
|
||||
`X-Forwarded-Proto`.** These are two requirements, not one. The
|
||||
environment setting decides CORS and nothing else: the default
|
||||
`dev` answers every origin with `Access-Control-Allow-Origin: *`
|
||||
(without credentials), which a server-rendered production
|
||||
deployment has no use for. Cookie `Secure` and the strict
|
||||
Origin/Referer mode are **not** tied to it — they are decided per
|
||||
request from the transport, which behind a proxy means the
|
||||
`X-Forwarded-Proto` header. The block below sets it; without it
|
||||
every request is read as plaintext and cookies ship without
|
||||
`Secure`. See [Configuration](#configuration).
|
||||
3. **Set `TRUSTED_PROXIES` to the proxy's address.** Unset, every rate
|
||||
limiter keys on the connecting peer, which behind a proxy is the
|
||||
proxy on every request: all clients collapse into one global bucket
|
||||
per limit and the receiver's per-IP limits become service-wide
|
||||
ceilings. See [Trusted proxies](#trusted-proxies). List the proxy
|
||||
and nothing else.
|
||||
4. **Send `Host` as `$http_host`, not `$host`.** `$host` strips the
|
||||
port. webhooker's Origin/Referer check compares against the host it
|
||||
was given, so on any port other than 443 `$host` makes every form
|
||||
POST — including login — fail with `403 origin invalid`, with
|
||||
nothing in the error naming the cause.
|
||||
5. **Keep the proxy's access log.** webhooker's own access log records
|
||||
the peer address, which behind a proxy is always the proxy. The
|
||||
proxy's log is the only record of which client sent what. nginx's
|
||||
default `combined` format already logs `$remote_addr`; do not
|
||||
replace it with one that drops the client address, and retain those
|
||||
logs as long as you would want to answer a question about traffic.
|
||||
|
||||
### nginx
|
||||
|
||||
Complete server block. Replace the `server_name` and the two
|
||||
certificate paths.
|
||||
|
||||
```nginx
|
||||
server {
|
||||
listen 443 ssl;
|
||||
listen [::]:443 ssl;
|
||||
http2 on; # nginx 1.25.1+; older: listen 443 ssl http2;
|
||||
|
||||
server_name webhooker.example.com;
|
||||
|
||||
ssl_certificate /etc/ssl/certs/webhooker.example.com.crt;
|
||||
ssl_certificate_key /etc/ssl/private/webhooker.example.com.key;
|
||||
ssl_protocols TLSv1.2 TLSv1.3;
|
||||
|
||||
# webhooker caps form POST bodies at 1 MB. nginx's default happens
|
||||
# to match, so leaving this out breaks nothing today — but if you
|
||||
# ever raise webhooker's cap, this is the limit you will still be
|
||||
# hitting, and the rejection is nginx's HTML page rather than
|
||||
# webhooker's message.
|
||||
client_max_body_size 1m;
|
||||
|
||||
# $remote_addr is the client. webhooker's own log records this
|
||||
# proxy and nothing else, so this file is the only place the
|
||||
# client's address is written down.
|
||||
access_log /var/log/nginx/webhooker.access.log combined;
|
||||
|
||||
location / {
|
||||
# A literal address, not localhost: with BIND_ADDRESS at its
|
||||
# 127.0.0.1 default, a localhost that resolves to ::1 first
|
||||
# gets connection refused.
|
||||
proxy_pass http://127.0.0.1:8080;
|
||||
|
||||
# $http_host, NOT $host. $host drops the port and every form
|
||||
# POST fails with 403 origin invalid on any port but 443.
|
||||
proxy_set_header Host $http_host;
|
||||
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
|
||||
proxy_set_header X-Forwarded-Proto $scheme;
|
||||
|
||||
# Above webhooker's own 60s request timeout, so its 503
|
||||
# reaches the client instead of nginx cutting the connection
|
||||
# first and answering 504.
|
||||
proxy_read_timeout 70s;
|
||||
}
|
||||
}
|
||||
|
||||
server {
|
||||
listen 80;
|
||||
listen [::]:80;
|
||||
server_name webhooker.example.com;
|
||||
return 308 https://$host$request_uri;
|
||||
}
|
||||
```
|
||||
|
||||
`X-Forwarded-For` must be **appended** to, which
|
||||
`$proxy_add_x_forwarded_for` does. webhooker reads no other forwarded
|
||||
client header: `X-Real-IP` and `True-Client-IP` are ignored from every
|
||||
peer, so setting them has no effect. See
|
||||
[Trusted proxies](#trusted-proxies) for how the chain is walked.
|
||||
|
||||
`X-Forwarded-Proto: https` is what tells webhooker the request arrived
|
||||
over TLS, which decides the `Secure` flag on both the session and CSRF
|
||||
cookies and the strict Origin/Referer mode. Without it, requests are
|
||||
treated as plaintext and the cookies ship without `Secure`. Unlike
|
||||
`X-Forwarded-For`, this header is read from any peer and is not gated
|
||||
by `TRUSTED_PROXIES`, so the proxy must overwrite whatever a client
|
||||
sent — `$scheme` above does.
|
||||
|
||||
With that block, webhooker's environment is:
|
||||
|
||||
```sh
|
||||
WEBHOOKER_ENVIRONMENT=prod
|
||||
BIND_ADDRESS=127.0.0.1 # the default; stated here to be explicit
|
||||
TRUSTED_PROXIES=127.0.0.1
|
||||
```
|
||||
|
||||
If nginx runs on another host, `BIND_ADDRESS` becomes the address it
|
||||
connects to, `TRUSTED_PROXIES` becomes nginx's address, and the port
|
||||
must be firewalled to that address — the traffic between them is
|
||||
cleartext.
|
||||
|
||||
### HSTS is always sent, and is not configurable
|
||||
|
||||
Every response carries
|
||||
`Strict-Transport-Security: max-age=63072000; includeSubDomains; preload`.
|
||||
Two years, every subdomain, and a `preload` token. There is no setting
|
||||
that changes or suppresses it.
|
||||
|
||||
This is worth knowing before the first request reaches a browser: a
|
||||
client that sees it once will refuse plaintext HTTP to that hostname —
|
||||
and to every subdomain of it — for two years, whatever else is served
|
||||
there. Terminate TLS on a hostname you are prepared to keep on HTTPS.
|
||||
|
||||
## Backup, Restore, and Upgrades
|
||||
|
||||
### What to back up
|
||||
@@ -581,9 +899,25 @@ is both the simplest and the only complete rule:
|
||||
`events-3f2a1c9e-....db`. The only other file is `webhooker.lock`, the
|
||||
always-empty [single-instance lock](#single-instance-lock); it holds no
|
||||
state and is not part of the backup set — a copied one is stale and
|
||||
blocks nothing. No `-wal` or `-shm` files are produced (see below); a
|
||||
transient `{name}.db-journal` may exist beside a database while a write
|
||||
is in flight and is not part of the backup set either.
|
||||
blocks nothing.
|
||||
|
||||
**`-wal` and `-shm` sidecars.** Every database runs in WAL journal mode,
|
||||
so while the service is running each `{name}.db` has a `{name}.db-wal`
|
||||
and a `{name}.db-shm` beside it. **`-wal` is part of the database, not a
|
||||
scratch file**: it holds committed transactions that are not yet in the
|
||||
`.db`, so a copy of the `.db` without its `-wal` is missing data and may
|
||||
have no readable schema at all. `-shm` is regenerable, but there is no
|
||||
reason to separate the two — copy the directory and you have them.
|
||||
|
||||
A clean shutdown closes `webhooker.db` and every `events-*.db`, which
|
||||
checkpoints and removes their sidecars; a killed or crashed instance
|
||||
leaves them, and they must be carried with the `.db`. **Archive
|
||||
databases are different**: their handle is not closed at shutdown, so
|
||||
`archive-*.db-wal` and `-shm` normally survive a clean stop and the
|
||||
`-wal` can hold every row the archive has. Measured on a stopped
|
||||
instance: `archive-….db` 4096 bytes with no table, its `-wal` 157 KB
|
||||
holding all 8 archived events. Copying `DATA_DIR` in full is what makes
|
||||
this a non-issue; copying `.db` files out of it by name is not.
|
||||
|
||||
Configuration is **not** in `DATA_DIR` — it comes from the environment
|
||||
and from a `.env` file read out of the process working directory. Back
|
||||
@@ -591,17 +925,19 @@ that up with your deployment config, separately.
|
||||
|
||||
### A hot copy is not safe
|
||||
|
||||
No `journal_mode` pragma is ever issued on any database webhooker opens,
|
||||
so all of them run on SQLite's default rollback journal. There is no
|
||||
WAL. The main and event databases are also held open for the entire
|
||||
process lifetime — `WebhookDBManager` caches event database handles and
|
||||
closes them only on webhook deletion or shutdown — so "it looked idle"
|
||||
is not a guarantee that nothing was mid-transaction.
|
||||
Every database webhooker opens runs in WAL journal mode. The main and
|
||||
event databases are also held open for the entire process lifetime —
|
||||
`WebhookDBManager` caches event database handles and closes them only on
|
||||
webhook deletion or shutdown — so "it looked idle" is not a guarantee
|
||||
that nothing was mid-transaction.
|
||||
|
||||
That means `cp`, `rsync`, `tar` or a filesystem snapshot taken against a
|
||||
running instance can capture a database mid-transaction and yield a file
|
||||
that is corrupt or missing state the journal would have rolled back. Use
|
||||
one of the two procedures below instead.
|
||||
running instance can capture a database and its `-wal` at two different
|
||||
instants and yield a file that is corrupt or missing state. Copying a
|
||||
`.db` on its own is worse and fails loudly: recently written pages,
|
||||
including the schema itself on a young database, live in the `-wal`, so
|
||||
the copy reads back as an empty or table-less database. Use one of the
|
||||
two procedures below instead.
|
||||
|
||||
**Stop, copy, start.** The simplest, needs no extra tooling, and the
|
||||
only one that gives a single point in time across every file:
|
||||
@@ -620,8 +956,9 @@ for db in /path/to/data/*.db; do
|
||||
done
|
||||
```
|
||||
|
||||
`.backup` takes the proper locks and writes a consistent file. Two
|
||||
caveats. First, the runtime image is `alpine:3.21` with only
|
||||
`.backup` reads through the WAL and writes a single consistent file with
|
||||
no sidecars of its own, so the destination is complete as it stands.
|
||||
Two caveats. First, the runtime image is `alpine:3.21` with only
|
||||
`ca-certificates` added — the `sqlite3` CLI is **not** in it, so run
|
||||
this on the host against the volume path, or from a throwaway container
|
||||
that mounts the volume. Second, each file is captured at its own
|
||||
@@ -629,15 +966,37 @@ instant, so a webhook created or an event delivered between two files
|
||||
being copied lands in one and not the other. If you need the whole set
|
||||
coherent as of a single moment, stop the service.
|
||||
|
||||
Note that `sqlite3 <db> .dump` is **not** one of these procedures: it is
|
||||
an export, it holds a read transaction open for as long as it runs, and
|
||||
it pins the WAL against checkpointing for that whole time. It is safe to
|
||||
run — it does not block ingestion — but back up with `.backup` or a
|
||||
stopped copy.
|
||||
|
||||
Archive databases are the one exception the service is built for: the
|
||||
archive writer closes its handle after each write (debounced to at most
|
||||
one reopen per second), so an operator can move `archive-{uuid}.db`
|
||||
away for offline retention while the service runs, and it is recreated
|
||||
on the next write (see
|
||||
[Database Architecture](#database-architecture)). That is a
|
||||
archive writer closes and reopens its handle around writes (debounced
|
||||
to at most one reopen per second), so an operator can move
|
||||
`archive-{uuid}.db` away for offline retention while the service runs,
|
||||
and it is recreated on the next write. See
|
||||
[Database Architecture](#database-architecture). That is a
|
||||
move-the-file-away workflow, not a substitute for the backup procedures
|
||||
above.
|
||||
|
||||
**Move the sidecars with it.** Under WAL that workflow is no longer a
|
||||
single file, and the common case is the dangerous one. The reopen
|
||||
happens on the *next* write after the debounce window elapses, so after
|
||||
the last write of a burst nothing checkpoints: measured, 20 s after ten
|
||||
events the `archive-….db` was 4096 bytes — a header, no table — with
|
||||
all ten rows sitting in a 189 KB `-wal`. Copying the `.db` alone at that
|
||||
moment yields a file that opens with `no such table: archived_events`.
|
||||
The file becomes self-contained again when the handle closes, which
|
||||
happens on the next write past the debounce window, when the connection
|
||||
pool retires the idle connection (about a minute after the last write),
|
||||
or at the idle archive sweep — measured, the same file was a complete
|
||||
20 KB `.db` with no sidecars about a minute after its last write.
|
||||
Shutdown is **not** on that list: the archive handle is not closed when
|
||||
the service stops. So either move `archive-{uuid}.db` together with any
|
||||
`-wal`/`-shm` beside it, or wait until there are none.
|
||||
|
||||
### Restore
|
||||
|
||||
1. Stop the service.
|
||||
@@ -651,14 +1010,22 @@ above.
|
||||
restored without `webhooker.db` are simply orphaned; nothing
|
||||
references their UUIDs.
|
||||
|
||||
3. Do not carry `*.db-journal` files into the restore. Backups taken by
|
||||
either procedure above are self-consistent and do not need one.
|
||||
3. Carry any `*.db-wal` and `*.db-shm` files that are in the backup.
|
||||
They are part of the database, and dropping a `-wal` silently
|
||||
discards every transaction it still holds. An `.backup` set will not
|
||||
contain any: it writes a single consolidated file per database. A
|
||||
stop-and-copy set has none for `webhooker.db` or the `events-*.db`,
|
||||
because a clean stop closes those and checkpoints their sidecars
|
||||
away — but it will normally have them for `archive-*.db`, whose
|
||||
handle stays open across shutdown, and those carry the archive's
|
||||
rows. A copy salvaged from a crashed instance has them for
|
||||
everything, and needs all of them.
|
||||
|
||||
4. **Fix ownership.** The container runs as the non-root `webhooker`
|
||||
user, UID 1000 / GID 1000. Restored files must be owned by (or
|
||||
writable by) that UID, and so must the directory itself — SQLite
|
||||
creates the rollback journal beside the database, so a writable file
|
||||
inside a directory it cannot write is not enough:
|
||||
creates the `-wal` and `-shm` sidecars beside the database, so a
|
||||
writable file inside a directory it cannot write is not enough:
|
||||
|
||||
```bash
|
||||
chown -R 1000:1000 /path/to/data
|
||||
@@ -696,6 +1063,18 @@ Upgrade procedure:
|
||||
`curl -s http://host:8080/.well-known/healthcheck` reports the
|
||||
version it was stamped with (see [Version stamping](#version-stamping)).
|
||||
|
||||
**Upgrading past the introduction of `BIND_ADDRESS`:** earlier versions
|
||||
always bound every interface. **Container deployments are unaffected**
|
||||
— the image ships `ENV BIND_ADDRESS=0.0.0.0`, so a `docker run` or
|
||||
Compose service that worked before still works with nothing changed.
|
||||
|
||||
A **bare binary** is the case that changes: the listener now binds
|
||||
`127.0.0.1` unless `BIND_ADDRESS` says otherwise, so a deployment that
|
||||
relied on reaching it from another host becomes unreachable until it
|
||||
sets the address the proxy connects to. Check the `bindAddress` field
|
||||
of the startup log to see what a running process bound. See
|
||||
[Bind address](#bind-address).
|
||||
|
||||
**Downgrade is unsupported.** Once a newer binary has migrated the files
|
||||
there is no way to move them back. `AutoMigrate` is additive — it adds
|
||||
tables, columns and indexes and never drops or rewrites them — so an
|
||||
@@ -777,14 +1156,38 @@ backups at rest and restrict who can read them.
|
||||
|
||||
## The entrypoint URL is the authentication secret
|
||||
|
||||
The receiver verifies nothing about an inbound request. The UUID in an
|
||||
entrypoint's URL is its credential: anyone who holds that URL can
|
||||
submit events to it, and the receiver checks nothing else about the
|
||||
sender. Treat an entrypoint URL the way you would treat an API token.
|
||||
**The entrypoint UUID is the credential, and it is the only one.**
|
||||
webhooker mints a version 4 UUID per entrypoint and serves it at
|
||||
`/webhook/{uuid}`. Possession of that URL is the authentication:
|
||||
anyone who holds it can submit events to the entrypoint, and the
|
||||
receiver verifies nothing else about the sender.
|
||||
|
||||
There is no way to rotate the UUID in place. To retire one, delete the
|
||||
entrypoint (or deactivate it, which answers `410`) and create a new
|
||||
one, then point the sender at the new URL.
|
||||
There is no shared secret, no HMAC signature, no bearer token and no
|
||||
second factor on the receiver, and none will be added. This was
|
||||
considered and rejected; the implementation that existed was removed
|
||||
in [PR #279](https://git.eeqj.de/sneak/webhooker/pulls/279), closing
|
||||
[issue #67](https://git.eeqj.de/sneak/webhooker/issues/67) and
|
||||
[issue #241](https://git.eeqj.de/sneak/webhooker/issues/241). A
|
||||
proposal to reintroduce any of them — including as "defence in depth"
|
||||
alongside the UUID — is answered by this section. Inbound signature
|
||||
headers a sender sends anyway (`X-Hub-Signature` and its
|
||||
per-provider equivalents) are stored and forwarded as ordinary
|
||||
headers; nothing checks them.
|
||||
|
||||
What that means for an operator:
|
||||
|
||||
- **The URL is a capability, so treat it as a secret.** Keep it out of
|
||||
logs, ticket bodies, chat messages and screenshots. Anyone who reads
|
||||
it anywhere can post events as that sender.
|
||||
- **Rotating means minting a new entrypoint, not changing a key.**
|
||||
There is no way to rotate the UUID in place. To retire one, delete
|
||||
the entrypoint (or deactivate it, which answers `410`) and create a
|
||||
new one, then point the sender at the new URL.
|
||||
- **A sender that cannot be given a secret URL is a constraint on that
|
||||
integration, not a reason to change this.** If a service only
|
||||
supports signed payloads to a well-known URL, raise it as its own
|
||||
problem — pick a different integration path, or accept that it
|
||||
cannot be used. It is not grounds to reintroduce shared secrets.
|
||||
|
||||
## Entrypoints
|
||||
|
||||
@@ -862,6 +1265,16 @@ webhooker solves this by acting as a durable intermediary:
|
||||
backoff. Every delivery attempt is logged with status codes, response
|
||||
bodies, and timing.
|
||||
|
||||
**That guarantee is at-least-once, not exactly-once.** When a send
|
||||
reaches its target but the write recording that outcome fails, the
|
||||
delivery is deliberately left in a recoverable state rather than
|
||||
marked done — losing a delivery is the worse failure — so the
|
||||
pending sweep picks it up about fifteen minutes later, or the next
|
||||
restart does, and the target receives a payload it already got.
|
||||
webhooker adds no delivery identifier of its own to an outbound
|
||||
request, so **make your receiver idempotent** against whatever the
|
||||
payload itself carries.
|
||||
|
||||
3. **Observability** — Full request/response logging for every webhook
|
||||
received and every delivery attempted. Prometheus metrics expose
|
||||
volume, latency, and error rates. The web UI provides real-time
|
||||
@@ -1308,6 +1721,11 @@ webhooker uses **separate SQLite database files**: a main application
|
||||
database for configuration data and per-webhook databases for event
|
||||
storage. All database files live in the `DATA_DIR` directory.
|
||||
|
||||
Every one of them is created `0600`, and so is each `-wal` and `-shm`
|
||||
sidecar. See
|
||||
[Running with Docker](#running-with-docker) for what that does and
|
||||
does not protect.
|
||||
|
||||
**Main Application Database** (`{DATA_DIR}/webhooker.db`) — stores
|
||||
configuration and application state:
|
||||
|
||||
@@ -1349,10 +1767,12 @@ This separation provides:
|
||||
only, or disables cleanup entirely when set to `0` (retain forever).
|
||||
- **Performance** — each webhook's database has its own page cache and
|
||||
its own lock, so concurrent event ingestion across webhooks won't
|
||||
contend. No write-ahead log is involved: both DSNs are
|
||||
`file:{path}?cache=shared&mode=rwc` and no `journal_mode` pragma is
|
||||
ever issued, so every database runs on SQLite's default rollback
|
||||
journal.
|
||||
contend. Every database — main, per-webhook, and archive — is opened
|
||||
through one code path (`internal/database/sqlite_open.go`) in WAL
|
||||
journal mode, with a 10-second busy timeout, `BEGIN IMMEDIATE`
|
||||
transactions, and a bounded connection pool. Under WAL a reader never
|
||||
blocks a writer, so an operator reading a database does not stall
|
||||
event ingestion into it.
|
||||
|
||||
The **database target type** builds on this architecture to provide
|
||||
long-term archiving, separate from the per-webhook event database (which
|
||||
@@ -2237,12 +2657,15 @@ abuse limit later; they are tracked as future work.
|
||||
| `POST` | `/source/{id}/edit` | Edit webhook submission |
|
||||
| `POST` | `/source/{id}/delete` | Delete webhook |
|
||||
| `GET` | `/source/{id}/logs` | Webhook event logs |
|
||||
| `GET` | `/source/{id}/logs/{eventID}/body` | Download an event's full stored body. The log page renders each body only up to its cap, so this is the only route that serves a whole one; it is offered wherever a body is shown truncated |
|
||||
| `POST` | `/source/{id}/deliveries/{deliveryID}/replay` | Replay a finished delivery: creates a new delivery for the same event against the target's current configuration (30 per minute per bucket, then `429`) |
|
||||
| `POST` | `/source/{id}/events/{eventID}/resubmit` | Resubmit a stored event: creates a new event copying it and fans that out to every currently active target (30 per minute per bucket, then `429`) |
|
||||
| `POST` | `/source/{id}/entrypoints` | Add entrypoint to webhook |
|
||||
| `POST` | `/source/{id}/entrypoints/{entrypointID}/delete` | Delete an entrypoint |
|
||||
| `POST` | `/source/{id}/entrypoints/{entrypointID}/toggle` | Enable or disable an entrypoint |
|
||||
| `POST` | `/source/{id}/targets` | Add target to webhook |
|
||||
| `GET` | `/source/{id}/targets/{targetID}/edit` | Edit target form. The one page that renders a target's destination URL and header values in full, rather than masked |
|
||||
| `POST` | `/source/{id}/targets/{targetID}/edit` | Edit target submission |
|
||||
| `POST` | `/source/{id}/targets/{targetID}/delete` | Delete a target |
|
||||
| `POST` | `/source/{id}/targets/{targetID}/toggle` | Enable or disable a target |
|
||||
|
||||
@@ -2281,6 +2704,8 @@ webhooker/
|
||||
├── internal/
|
||||
│ ├── banner/
|
||||
│ │ └── banner.go # Ruled block for the one credential shown in the clear
|
||||
│ ├── ciscript/
|
||||
│ │ └── doc.go # Tests for the CI shell scripts in script/; no runtime code
|
||||
│ ├── resetpw/
|
||||
│ │ └── resetpw.go # `webhooker resetpw`: set an account's password, stopped deployments only
|
||||
│ ├── config/
|
||||
@@ -2350,13 +2775,17 @@ webhooker/
|
||||
│ │ ├── ratelimit.go # Per-IP rate limiting middleware (go-chi/httprate)
|
||||
│ │ ├── loginguard.go # Login failure counters and the Argon2id verification semaphore
|
||||
│ │ └── testing.go # NewForTest: Middleware without the fx lifecycle
|
||||
│ ├── reqtls/
|
||||
│ │ └── reqtls.go # IsTLS: the one TLS predicate, r.TLS or X-Forwarded-Proto
|
||||
│ ├── server/
|
||||
│ │ ├── server.go # Server struct, fx lifecycle, signal handling
|
||||
│ │ ├── http.go # HTTP server setup with timeouts
|
||||
│ │ └── routes.go # All route definitions
|
||||
│ └── session/
|
||||
│ ├── session.go # Cookie-based session management
|
||||
│ └── testing.go # NewForTest: Session without the fx lifecycle
|
||||
│ ├── session/
|
||||
│ │ ├── session.go # Cookie-based session management
|
||||
│ │ └── testing.go # NewForTest: Session without the fx lifecycle
|
||||
│ └── versionscript/
|
||||
│ └── doc.go # Tests for script/version and the build files that use it
|
||||
├── static/
|
||||
│ ├── static.go # //go:embed directive
|
||||
│ ├── css/input.css # Tailwind input, source for tailwind.css (make css)
|
||||
@@ -2469,6 +2898,10 @@ check, see [The login endpoint](#the-login-endpoint).
|
||||
|
||||
### Authentication
|
||||
|
||||
- **Webhook receiver:** the entrypoint UUID in the URL, and nothing
|
||||
else. No shared secret, no HMAC signature, no token header, and none
|
||||
will be added — see
|
||||
[The entrypoint URL is the authentication secret](#the-entrypoint-url-is-the-authentication-secret).
|
||||
- **Web UI:** Cookie-based sessions using gorilla/sessions with
|
||||
encrypted cookies. Sessions are configured with HttpOnly, SameSite
|
||||
Lax, and Secure whenever the request is on TLS — the flag follows the
|
||||
@@ -2508,7 +2941,8 @@ check, see [The login endpoint](#the-login-endpoint).
|
||||
mode
|
||||
- **The entrypoint URL is the receiver's only credential.** Nothing
|
||||
about an inbound request is verified; possession of the UUID
|
||||
authorises submission (see
|
||||
authorises submission, and no shared secret or signature check will
|
||||
be added alongside it (see
|
||||
[The entrypoint URL is the authentication secret](#the-entrypoint-url-is-the-authentication-secret))
|
||||
- **SSRF prevention** for HTTP delivery targets: private/reserved IP
|
||||
ranges (RFC 1918, loopback, link-local, cloud metadata) are blocked
|
||||
|
||||
54
TODO.md
54
TODO.md
@@ -18,18 +18,27 @@ Issue branches do NOT touch this file — the manager maintains it on
|
||||
|
||||
# Status
|
||||
|
||||
1.0.0 is open, with work remaining. The milestone
|
||||
(https://git.eeqj.de/sneak/webhooker/milestone/9) is the authoritative
|
||||
list and the only place to read a count from; this file deliberately
|
||||
carries neither, because both drift between the commits that touch it.
|
||||
`next` is ahead of `main` and a strict fast-forward. No release has
|
||||
been tagged yet.
|
||||
The milestone (https://git.eeqj.de/sneak/webhooker/milestone/9) is the
|
||||
authoritative list, and the only place to read a count or a state of
|
||||
play from. This file records where the project is, not what is in
|
||||
flight: a sentence whose truth depends on a branch being unmerged is
|
||||
wrong the moment it merges, and this file has been wrong that way
|
||||
before.
|
||||
|
||||
The tag is held on a durability defect
|
||||
(https://git.eeqj.de/sneak/webhooker/issues/256): a concurrent reader
|
||||
of a per-webhook event database strands delivered webhooks at
|
||||
`pending`, and the next restart re-delivers them. A fix is in review
|
||||
(https://git.eeqj.de/sneak/webhooker/pulls/263).
|
||||
The durability defect that held the tag has landed
|
||||
(https://git.eeqj.de/sneak/webhooker/issues/256, commit `8d64259`).
|
||||
Every SQLite handle opens with WAL journaling and a busy timeout, a
|
||||
bookkeeping write that fails leaves its delivery in a recoverable
|
||||
state rather than a lying one, and recovery skips a delivery that
|
||||
already has a successful result row. Final pre-tag verification
|
||||
exercised it and confirmed it holds. Whatever the milestone still
|
||||
shows open is what remains before `v1.0.0`.
|
||||
|
||||
Delivery is at-least-once by design, not by accident: a send whose
|
||||
result row does not land is attempted again, so a receiver can see a
|
||||
duplicate. That is deliberate — the alternative is a silent lost
|
||||
delivery — and the README says so under Rationale. It is not a defect
|
||||
to re-file.
|
||||
|
||||
One caveat on reading a green check: a docs-only commit deliberately
|
||||
replays from the layer cache
|
||||
@@ -39,16 +48,25 @@ commit invalidates the `COPY` layer and genuinely executes.
|
||||
|
||||
# Next Step
|
||||
|
||||
Land https://git.eeqj.de/sneak/webhooker/issues/256, then clear the
|
||||
rest of the open 1.0.0 milestone and tag `v1.0.0`.
|
||||
|
||||
The milestone PR (https://git.eeqj.de/sneak/webhooker/pulls/111) is
|
||||
`merge-ready` and assigned to sneak. Merging it is safe whenever sneak
|
||||
wants it — the durability defect predates the branch and exists on
|
||||
`main` too — but merging it is not the tag.
|
||||
Clear the rest of the open 1.0.0 milestone
|
||||
(https://git.eeqj.de/sneak/webhooker/milestone/9) and tag `v1.0.0`.
|
||||
Merging `next` into `main` is a separate act from tagging and waits on
|
||||
neither of those: `next` is kept mergeable at all times, which is the
|
||||
point of the branch.
|
||||
|
||||
# Completed Steps
|
||||
|
||||
- 2026-08-24 Bind the plaintext HTTP listener deliberately, via
|
||||
`BIND_ADDRESS` defaulting to `127.0.0.1`, and document the
|
||||
reverse-proxy deployment. A hostname, an empty value or a value
|
||||
carrying a port is a startup error, and the `Dockerfile` sets
|
||||
`0.0.0.0` because a loopback bind inside a container is unreachable
|
||||
(https://git.eeqj.de/sneak/webhooker/issues/268). The same commit
|
||||
removed the shutdown race: `httpServer` is built in the constructor
|
||||
rather than assigned from the serving goroutine, which orders the
|
||||
write before every fx hook and rules out the nil dereference a
|
||||
SIGTERM arriving first would have caused, and `sentryEnabled` is an
|
||||
`atomic.Bool` (https://git.eeqj.de/sneak/webhooker/issues/226)
|
||||
- 2026-08-24 Remove inbound request signature verification. The
|
||||
entrypoint UUID is the authentication secret, so the per-entrypoint
|
||||
shared secret, the `internal/signature` package, the receiver check,
|
||||
|
||||
107
cmd/webhooker/dotenv_test.go
Normal file
107
cmd/webhooker/dotenv_test.go
Normal file
@@ -0,0 +1,107 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"bytes"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
"github.com/stretchr/testify/assert"
|
||||
"github.com/stretchr/testify/require"
|
||||
"sneak.berlin/go/webhooker/internal/config"
|
||||
)
|
||||
|
||||
// dotEnvKey is a throwaway variable name these tests write and read,
|
||||
// so they cannot disturb real configuration.
|
||||
const dotEnvKey = "WEBHOOKER_TEST_DISPATCH_VALUE"
|
||||
|
||||
// writeDotEnvInWorkingDir puts contents in a .env file in a fresh
|
||||
// temporary directory and moves the process there.
|
||||
//
|
||||
// The callers are deliberately not parallel and must stay that way:
|
||||
// t.Chdir moves the whole process. Go releases parallel tests only
|
||||
// after every sequential test in the package has finished, so nothing
|
||||
// else runs while these do.
|
||||
func writeDotEnvInWorkingDir(t *testing.T, contents string) {
|
||||
t.Helper()
|
||||
|
||||
dir := t.TempDir()
|
||||
require.NoError(t, os.WriteFile(
|
||||
filepath.Join(dir, config.DotEnvPath),
|
||||
[]byte(contents), 0o600,
|
||||
))
|
||||
t.Chdir(dir)
|
||||
}
|
||||
|
||||
// TestDispatch_MalformedDotEnvRefuses pins the second half of the
|
||||
// defect. godotenv applies nothing at all when a file will not parse,
|
||||
// so one mistyped line used to revert every variable in it to its
|
||||
// default and start the server anyway, with no log line naming the
|
||||
// file. The refusal has to arrive before any subcommand runs, which
|
||||
// is why `help` — the one subcommand that touches nothing — is still
|
||||
// refused here.
|
||||
//
|
||||
//nolint:paralleltest // t.Chdir moves the whole process.
|
||||
func TestDispatch_MalformedDotEnvRefuses(t *testing.T) {
|
||||
writeDotEnvInWorkingDir(t, "PORT 19615\n")
|
||||
|
||||
var stdout, stderr bytes.Buffer
|
||||
|
||||
code := dispatch(
|
||||
[]string{helpCommand}, strings.NewReader(""), &stdout, &stderr,
|
||||
)
|
||||
|
||||
require.Equal(t, 1, code, "a broken .env must exit non-zero")
|
||||
assert.Contains(
|
||||
t, stderr.String(), config.DotEnvPath,
|
||||
"the refusal must name the file",
|
||||
)
|
||||
assert.Empty(
|
||||
t, stdout.String(),
|
||||
"the subcommand must not have run",
|
||||
)
|
||||
}
|
||||
|
||||
// TestDispatch_LoadsDotEnvBeforeSubcommands pins the ordering the
|
||||
// godotenv/autoload import used to provide for free. It ran in an
|
||||
// init(), so .env was in the environment before anything read it —
|
||||
// including config.DataDir, which both the DATA_DIR lock and resetpw
|
||||
// call outside the fx graph. Loading any later would let a .env that
|
||||
// sets DATA_DIR lock one directory while the config opened databases
|
||||
// in another.
|
||||
func TestDispatch_LoadsDotEnvBeforeSubcommands(t *testing.T) {
|
||||
t.Setenv(dotEnvKey, "placeholder")
|
||||
require.NoError(t, os.Unsetenv(dotEnvKey))
|
||||
|
||||
writeDotEnvInWorkingDir(t, dotEnvKey+"=from-dot-env\n")
|
||||
|
||||
var stdout, stderr bytes.Buffer
|
||||
|
||||
code := dispatch(
|
||||
[]string{helpCommand}, strings.NewReader(""), &stdout, &stderr,
|
||||
)
|
||||
|
||||
require.Equal(t, 0, code)
|
||||
assert.Equal(
|
||||
t, "from-dot-env", os.Getenv(dotEnvKey),
|
||||
"the file must be applied before the subcommand runs",
|
||||
)
|
||||
}
|
||||
|
||||
// TestDispatch_MissingDotEnvIsFine pins the case most deployments are
|
||||
// in: no .env at all, which must stay a normal start.
|
||||
//
|
||||
//nolint:paralleltest // t.Chdir moves the whole process.
|
||||
func TestDispatch_MissingDotEnvIsFine(t *testing.T) {
|
||||
t.Chdir(t.TempDir())
|
||||
|
||||
var stdout, stderr bytes.Buffer
|
||||
|
||||
code := dispatch(
|
||||
[]string{helpCommand}, strings.NewReader(""), &stdout, &stderr,
|
||||
)
|
||||
|
||||
require.Equal(t, 0, code)
|
||||
assert.Empty(t, stderr.String())
|
||||
}
|
||||
@@ -54,6 +54,11 @@ const stopTimeout = 5 * time.Second
|
||||
// caller can tell "called wrong" from "declined".
|
||||
const exitUsage = 2
|
||||
|
||||
// helpCommand is the subcommand that prints usage. The flag spellings
|
||||
// beside it in the switch are aliases; this is the name the usage text
|
||||
// documents and the one tests invoke.
|
||||
const helpCommand = "help"
|
||||
|
||||
// Build-time variables set via -ldflags.
|
||||
//
|
||||
//nolint:gochecknoglobals // Build-time variables injected by the linker.
|
||||
@@ -75,11 +80,27 @@ func main() {
|
||||
// every existing deployment invoke; that path is unchanged, including
|
||||
// where the DATA_DIR lock is taken relative to building the fx graph
|
||||
// and how fx propagates a non-zero exit itself.
|
||||
//
|
||||
// The optional .env file is read here, before any subcommand and so
|
||||
// before anything reads the environment — config.DataDir, which both
|
||||
// the DATA_DIR lock and resetpw call outside the fx graph, above all.
|
||||
// It used to be read from an init() in internal/config, which put it
|
||||
// earlier still but threw the error away: a single malformed line
|
||||
// applied none of the file and said nothing about it. A file that is
|
||||
// not there stays fine, since .env is optional and most deployments
|
||||
// do not have one.
|
||||
func dispatch(
|
||||
args []string,
|
||||
stdin io.Reader,
|
||||
stdout, stderr io.Writer,
|
||||
) int {
|
||||
err := config.LoadDotEnv()
|
||||
if err != nil {
|
||||
_, _ = fmt.Fprintf(stderr, "%s: %v\n", appname, err)
|
||||
|
||||
return 1
|
||||
}
|
||||
|
||||
if len(args) == 0 {
|
||||
return run(stderr)
|
||||
}
|
||||
@@ -87,7 +108,7 @@ func dispatch(
|
||||
switch args[0] {
|
||||
case resetpw.Name:
|
||||
return resetpw.Run(args[1:], stdin, stdout, stderr)
|
||||
case "help", "-h", "-help", "--help":
|
||||
case helpCommand, "-h", "-help", "--help":
|
||||
usage(stdout)
|
||||
|
||||
return 0
|
||||
|
||||
@@ -121,7 +121,7 @@ func TestDispatch_Help(t *testing.T) {
|
||||
var stdout, stderr bytes.Buffer
|
||||
|
||||
code := dispatch(
|
||||
[]string{"help"}, strings.NewReader(""), &stdout, &stderr,
|
||||
[]string{helpCommand}, strings.NewReader(""), &stdout, &stderr,
|
||||
)
|
||||
|
||||
require.Equal(t, 0, code)
|
||||
|
||||
@@ -4,6 +4,7 @@ package config
|
||||
import (
|
||||
"errors"
|
||||
"fmt"
|
||||
"io/fs"
|
||||
"log/slog"
|
||||
"net/netip"
|
||||
"os"
|
||||
@@ -11,13 +12,11 @@ import (
|
||||
"strings"
|
||||
"time"
|
||||
|
||||
"github.com/getsentry/sentry-go"
|
||||
"github.com/joho/godotenv"
|
||||
"go.uber.org/fx"
|
||||
"sneak.berlin/go/webhooker/internal/globals"
|
||||
"sneak.berlin/go/webhooker/internal/logger"
|
||||
|
||||
// Populates the environment from a ./.env file automatically for
|
||||
// development configuration. Kept in one place only (here).
|
||||
_ "github.com/joho/godotenv/autoload"
|
||||
)
|
||||
|
||||
const (
|
||||
@@ -33,6 +32,34 @@ const (
|
||||
// defaultPort is the default HTTP listen port.
|
||||
defaultPort = 8080
|
||||
|
||||
// defaultBindAddress is the interface the plaintext HTTP
|
||||
// listener claims when BIND_ADDRESS is unset.
|
||||
//
|
||||
// Loopback, because the listener speaks cleartext and serves
|
||||
// both the admin UI and the unauthenticated receiver: a
|
||||
// wildcard default publishes them on every interface of every
|
||||
// host that never configured anything, which is the failure
|
||||
// this default exists to prevent. Reaching webhooker from off
|
||||
// the host is then a deliberate act — a reverse proxy in front
|
||||
// of it, or an explicit BIND_ADDRESS.
|
||||
//
|
||||
// This is the binary's default only. The Dockerfile ships
|
||||
// ENV BIND_ADDRESS=0.0.0.0, so a container deployment needs
|
||||
// nothing set and is unaffected by this constant. The two
|
||||
// differ because they answer different questions: a container's
|
||||
// network namespace is already the boundary this default is
|
||||
// reaching for, so binding every address inside it exposes
|
||||
// nothing, and what decides exposure there is the publish flag
|
||||
// (-p 127.0.0.1:8080:8080). A loopback bind inside a container
|
||||
// buys no security and makes the process unreachable through
|
||||
// its own published port.
|
||||
//
|
||||
// The split is expressed as two explicit defaults rather than
|
||||
// container auto-detection, because a heuristic that guesses
|
||||
// wrong opens the cleartext port exactly where nobody is
|
||||
// looking.
|
||||
defaultBindAddress = "127.0.0.1"
|
||||
|
||||
// defaultRetentionSweepInterval is how often the retention
|
||||
// reaper deletes events older than each webhook's RetentionDays.
|
||||
defaultRetentionSweepInterval = time.Hour
|
||||
@@ -56,6 +83,12 @@ const (
|
||||
// IPv6 prefix spends on the ::ffff:0:0/96 wrapper, so a /104
|
||||
// covers the same addresses as an IPv4 /8.
|
||||
mappedV4Offset = 96
|
||||
|
||||
// DotEnvPath is the optional file of KEY=value lines read into the
|
||||
// environment at startup, relative to the process working
|
||||
// directory. Exported so that documentation and tests name the
|
||||
// same path the loader opens.
|
||||
DotEnvPath = ".env"
|
||||
)
|
||||
|
||||
// ErrInvalidEnvironment is returned when WEBHOOKER_ENVIRONMENT
|
||||
@@ -75,6 +108,19 @@ var ErrInvalidPort = errors.New("invalid port")
|
||||
// nor a bare IP address.
|
||||
var ErrInvalidCIDR = errors.New("invalid CIDR")
|
||||
|
||||
// ErrInvalidBindAddress is returned when BIND_ADDRESS is set to
|
||||
// something that is not an IP address literal.
|
||||
var ErrInvalidBindAddress = errors.New("invalid bind address")
|
||||
|
||||
// ErrInvalidSentryDSN is returned when SENTRY_DSN is set to something
|
||||
// the Sentry SDK cannot parse as a DSN.
|
||||
var ErrInvalidSentryDSN = errors.New("invalid Sentry DSN")
|
||||
|
||||
// ErrDotEnvUnreadable is returned when the optional .env file exists
|
||||
// but cannot be read or parsed. A file that is not there is not an
|
||||
// error; a file that is there and broken is.
|
||||
var ErrDotEnvUnreadable = errors.New("unreadable .env file")
|
||||
|
||||
// ErrIncompleteMetricsAuth is returned when exactly one of
|
||||
// METRICS_USERNAME and METRICS_PASSWORD carries a value. Neither
|
||||
// fallback is acceptable: serving /metrics on the username alone
|
||||
@@ -105,6 +151,13 @@ type Config struct {
|
||||
Port int
|
||||
SentryDSN string
|
||||
|
||||
// BindAddress is the IP address the plaintext HTTP listener
|
||||
// binds, as an address literal. It defaults to
|
||||
// defaultBindAddress and is never empty: an empty string would
|
||||
// mean the wildcard to net.Listen, which is the opposite of the
|
||||
// default this ships.
|
||||
BindAddress string
|
||||
|
||||
// RetentionSweepInterval is how often the retention reaper runs.
|
||||
// Always positive: it becomes a time.NewTicker period.
|
||||
RetentionSweepInterval time.Duration
|
||||
@@ -173,12 +226,62 @@ func (c *Config) MetricsAuthEnabled() bool {
|
||||
return c.MetricsUsername != "" && c.MetricsPassword != ""
|
||||
}
|
||||
|
||||
// SentryEnabled reports whether error reporting is shipped to Sentry.
|
||||
// It is the only answer to that question in the codebase: the SDK
|
||||
// initialisation, the sentryhttp middleware registration and the
|
||||
// startup log's sentryEnabled field all read this one method, so the
|
||||
// log cannot report reporting as on while nothing is sending.
|
||||
//
|
||||
// A non-empty DSN is enough because loadFromEnv already parsed it with
|
||||
// the SDK's own parser and refused to build a Config around one the
|
||||
// SDK would reject, and because initialising the SDK with a DSN that
|
||||
// parsed and failed anyway aborts the process rather than leaving this
|
||||
// true and the client absent.
|
||||
func (c *Config) SentryEnabled() bool {
|
||||
return c.SentryDSN != ""
|
||||
}
|
||||
|
||||
// envString returns the value of the named environment variable,
|
||||
// or an empty string if not set.
|
||||
func envString(key string) string {
|
||||
return os.Getenv(key)
|
||||
}
|
||||
|
||||
// LoadDotEnv reads DotEnvPath into the environment when that file is
|
||||
// present, and reports a file that is present but broken.
|
||||
//
|
||||
// It has to run before anything reads the environment, so that every
|
||||
// reader agrees on what the environment holds — the DATA_DIR lock
|
||||
// taken before the fx graph exists as much as loadFromEnv itself. A
|
||||
// variable already set in the real environment wins: godotenv never
|
||||
// overwrites one.
|
||||
//
|
||||
// A missing file is not an error. It is a development convenience and
|
||||
// most deployments set the environment directly.
|
||||
//
|
||||
// Any other failure is. godotenv parses the whole file before setting
|
||||
// anything, so a single malformed line applies none of it: every
|
||||
// variable in the file silently reverts to its default, which defeats
|
||||
// the fail-loud guarantee for all of them at once.
|
||||
func LoadDotEnv() error {
|
||||
return loadDotEnvFile(DotEnvPath)
|
||||
}
|
||||
|
||||
// loadDotEnvFile is LoadDotEnv over a named file, so tests can point
|
||||
// at a temporary one instead of the process working directory.
|
||||
func loadDotEnvFile(path string) error {
|
||||
err := godotenv.Load(path)
|
||||
if err == nil || errors.Is(err, fs.ErrNotExist) {
|
||||
return nil
|
||||
}
|
||||
|
||||
return fmt.Errorf(
|
||||
"%w: %s: %w; nothing in it was applied, so fix the file or "+
|
||||
"remove it",
|
||||
ErrDotEnvUnreadable, path, err,
|
||||
)
|
||||
}
|
||||
|
||||
// DataDir resolves DATA_DIR, applying DefaultDataDir when it is unset
|
||||
// or empty. It is exported so that entry points which must act on the
|
||||
// data directory before the fx graph exists — taking the exclusive
|
||||
@@ -387,6 +490,77 @@ func envPrefixList(key string) ([]netip.Prefix, error) {
|
||||
return prefixes, nil
|
||||
}
|
||||
|
||||
// envBindAddress returns the value of the named environment variable
|
||||
// parsed as an IP address literal. An unset (or empty, or
|
||||
// whitespace-only) value yields defaultValue.
|
||||
//
|
||||
// Only literals are accepted: no hostname is resolved, so `localhost`
|
||||
// is an error rather than a DNS lookup at startup whose answer could
|
||||
// be either loopback family, could change under the process, and
|
||||
// could return several addresses of which only one would be bound. A
|
||||
// value with a port in it (`127.0.0.1:8080`) is likewise an error —
|
||||
// the port is PORT's business, and silently accepting it would bind
|
||||
// something other than what was asked for.
|
||||
//
|
||||
// A set value that is not a literal is a hard error naming the key
|
||||
// and the bad value, so startup fails loudly rather than falling back
|
||||
// to a default the operator plainly did not want. A literal that is
|
||||
// not an address of this host parses here and fails at listen time
|
||||
// instead, which ends the process non-zero.
|
||||
func envBindAddress(key, defaultValue string) (string, error) {
|
||||
v := strings.TrimSpace(os.Getenv(key))
|
||||
if v == "" {
|
||||
return defaultValue, nil
|
||||
}
|
||||
|
||||
addr, err := netip.ParseAddr(v)
|
||||
if err != nil {
|
||||
return "", fmt.Errorf(
|
||||
"%w: %s: %q must be an IP address literal such as "+
|
||||
"127.0.0.1, 0.0.0.0 or ::, not a hostname and not "+
|
||||
"host:port: %w",
|
||||
ErrInvalidBindAddress, key, v, err,
|
||||
)
|
||||
}
|
||||
|
||||
return addr.String(), nil
|
||||
}
|
||||
|
||||
// envSentryDSN returns the value of the named environment variable
|
||||
// checked as a Sentry DSN. An unset (or empty, or whitespace-only)
|
||||
// value yields "", which means error reporting stays off — the common
|
||||
// case, and a normal start.
|
||||
//
|
||||
// A set value is parsed with sentry.NewDsn, which is the call
|
||||
// sentry.Init makes on the DSN it is handed, so what passes here is
|
||||
// exactly what the SDK will accept later and the two cannot disagree.
|
||||
// Reproducing the check by hand instead would cost this package its
|
||||
// dependency on the SDK — already a module dependency, already linked
|
||||
// into the binary — in exchange for a second definition of "valid DSN"
|
||||
// free to drift from the one that decides.
|
||||
//
|
||||
// A set value that does not parse is a hard error naming the key, so
|
||||
// startup fails loudly. Losing error reporting is the failure this
|
||||
// variable exists to prevent, and a typo in a DSN is silent forever:
|
||||
// nothing later in the process can notice that reports are going
|
||||
// nowhere. The bad value is quoted because it is a URL to a public
|
||||
// endpoint carrying a public key, not a secret.
|
||||
func envSentryDSN(key string) (string, error) {
|
||||
v := strings.TrimSpace(os.Getenv(key))
|
||||
if v == "" {
|
||||
return "", nil
|
||||
}
|
||||
|
||||
_, err := sentry.NewDsn(v)
|
||||
if err != nil {
|
||||
return "", fmt.Errorf(
|
||||
"%w: %s: %q: %w", ErrInvalidSentryDSN, key, v, err,
|
||||
)
|
||||
}
|
||||
|
||||
return v, nil
|
||||
}
|
||||
|
||||
// resolveMetricsAuth reads the /metrics basic-auth credentials and
|
||||
// rejects a half-set pair, naming both variables either way. The
|
||||
// error carries neither value: the password is a secret.
|
||||
@@ -431,6 +605,27 @@ func resolveEnvironment() (string, error) {
|
||||
return environment, nil
|
||||
}
|
||||
|
||||
// resolveListener reads the two variables that describe the HTTP
|
||||
// listener: which port it claims and which address it claims it on.
|
||||
// They are read together because neither is meaningful alone, and
|
||||
// because a validation failure in either has to abort startup before
|
||||
// anything binds.
|
||||
func resolveListener() (int, string, error) {
|
||||
port, err := envPort("PORT", defaultPort)
|
||||
if err != nil {
|
||||
return 0, "", err
|
||||
}
|
||||
|
||||
bindAddress, err := envBindAddress(
|
||||
"BIND_ADDRESS", defaultBindAddress,
|
||||
)
|
||||
if err != nil {
|
||||
return 0, "", err
|
||||
}
|
||||
|
||||
return port, bindAddress, nil
|
||||
}
|
||||
|
||||
// loadFromEnv builds a Config from the environment. Every value that
|
||||
// needs parsing fails loudly when it is set but unparseable: the
|
||||
// documented defaults apply only to variables that are unset (or
|
||||
@@ -442,7 +637,7 @@ func loadFromEnv() (*Config, error) {
|
||||
return nil, err
|
||||
}
|
||||
|
||||
port, err := envPort("PORT", defaultPort)
|
||||
port, bindAddress, err := resolveListener()
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
@@ -498,6 +693,11 @@ func loadFromEnv() (*Config, error) {
|
||||
return nil, err
|
||||
}
|
||||
|
||||
sentryDSN, err := envSentryDSN("SENTRY_DSN")
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
|
||||
return &Config{
|
||||
DataDir: DataDir(),
|
||||
Debug: debug,
|
||||
@@ -506,7 +706,8 @@ func loadFromEnv() (*Config, error) {
|
||||
MetricsUsername: metricsUsername,
|
||||
MetricsPassword: metricsPassword,
|
||||
Port: port,
|
||||
SentryDSN: envString("SENTRY_DSN"),
|
||||
BindAddress: bindAddress,
|
||||
SentryDSN: sentryDSN,
|
||||
RetentionSweepInterval: retentionSweepInterval,
|
||||
SessionIdleTimeout: sessionIdleTimeout,
|
||||
ReceiverRateLimit: receiverRateLimit,
|
||||
@@ -625,6 +826,11 @@ func New(lc fx.Lifecycle, params ConfigParams) (*Config, error) {
|
||||
log.Info("Configuration loaded",
|
||||
"environment", s.Environment,
|
||||
"port", s.Port,
|
||||
// Logged because which interfaces the cleartext listener
|
||||
// answers on is not otherwise observable from inside a
|
||||
// container, and it decides whether anything but the local
|
||||
// host can reach the admin UI.
|
||||
"bindAddress", s.BindAddress,
|
||||
"debug", s.Debug,
|
||||
"maintenanceMode", s.MaintenanceMode,
|
||||
"dataDir", s.DataDir,
|
||||
@@ -636,7 +842,7 @@ func New(lc fx.Lifecycle, params ConfigParams) (*Config, error) {
|
||||
"receiverRateLimit", s.ReceiverRateLimit,
|
||||
"trustedProxies", len(s.TrustedProxies),
|
||||
"allowedEgressCIDRs", len(s.AllowedEgressCIDRs),
|
||||
"hasSentryDSN", s.SentryDSN != "",
|
||||
"sentryEnabled", s.SentryEnabled(),
|
||||
"hasMetricsAuth", s.MetricsAuthEnabled(),
|
||||
)
|
||||
|
||||
|
||||
158
internal/config/dotenv_test.go
Normal file
158
internal/config/dotenv_test.go
Normal file
@@ -0,0 +1,158 @@
|
||||
package config_test
|
||||
|
||||
import (
|
||||
"os"
|
||||
"path/filepath"
|
||||
"testing"
|
||||
|
||||
"github.com/stretchr/testify/assert"
|
||||
"github.com/stretchr/testify/require"
|
||||
"sneak.berlin/go/webhooker/internal/config"
|
||||
)
|
||||
|
||||
// dotEnvKey is a throwaway variable name the .env tests write and
|
||||
// read, so they cannot disturb real configuration.
|
||||
const dotEnvKey = "WEBHOOKER_TEST_DOTENV_VALUE"
|
||||
|
||||
// malformedDotEnv is a file godotenv cannot parse. The first line is
|
||||
// the realistic typo — a space where the `=` belongs — and the rest
|
||||
// make sure nothing downstream treats the file as salvageable line by
|
||||
// line.
|
||||
const malformedDotEnv = "PORT 19615\n" +
|
||||
"this is not = valid ! syntax\n" +
|
||||
"\"unclosed\n"
|
||||
|
||||
// unsetDotEnvKey makes dotEnvKey genuinely absent for the duration of
|
||||
// the test and restores it afterwards. t.Setenv registers the restore;
|
||||
// the Unsetenv that follows is what the test actually needs, because a
|
||||
// variable set to the empty string is still present in os.Environ and
|
||||
// godotenv would refuse to overwrite it.
|
||||
func unsetDotEnvKey(t *testing.T) {
|
||||
t.Helper()
|
||||
t.Setenv(dotEnvKey, "placeholder")
|
||||
require.NoError(t, os.Unsetenv(dotEnvKey))
|
||||
}
|
||||
|
||||
// writeDotEnv writes contents to a .env file in a fresh temporary
|
||||
// directory and returns its path.
|
||||
func writeDotEnv(t *testing.T, contents string) string {
|
||||
t.Helper()
|
||||
|
||||
path := filepath.Join(t.TempDir(), config.DotEnvPath)
|
||||
require.NoError(t, os.WriteFile(path, []byte(contents), 0o600))
|
||||
|
||||
return path
|
||||
}
|
||||
|
||||
// TestLoadDotEnv_MissingFileIsFine pins the case most deployments are
|
||||
// in. The file is optional: it is a development convenience, and a
|
||||
// deployment that configures the environment directly must start
|
||||
// normally rather than be refused for a file it was never meant to
|
||||
// have.
|
||||
//
|
||||
//nolint:paralleltest // unsetDotEnvKey uses t.Setenv.
|
||||
func TestLoadDotEnv_MissingFileIsFine(t *testing.T) {
|
||||
unsetDotEnvKey(t)
|
||||
|
||||
absent := filepath.Join(t.TempDir(), config.DotEnvPath)
|
||||
require.NoError(t, config.LoadDotEnvFileForTest(absent))
|
||||
|
||||
_, present := os.LookupEnv(dotEnvKey)
|
||||
assert.False(t, present, "nothing may be set from an absent file")
|
||||
}
|
||||
|
||||
// TestLoadDotEnv_AppliesValues pins that a well-formed file still
|
||||
// reaches the environment, which is the whole reason the file is read
|
||||
// at all.
|
||||
//
|
||||
//nolint:paralleltest // unsetDotEnvKey uses t.Setenv.
|
||||
func TestLoadDotEnv_AppliesValues(t *testing.T) {
|
||||
unsetDotEnvKey(t)
|
||||
|
||||
path := writeDotEnv(t, "# a comment\n"+dotEnvKey+"=from-dot-env\n")
|
||||
|
||||
require.NoError(t, config.LoadDotEnvFileForTest(path))
|
||||
assert.Equal(t, "from-dot-env", os.Getenv(dotEnvKey))
|
||||
}
|
||||
|
||||
// TestLoadDotEnv_RealEnvironmentWins pins that the file cannot
|
||||
// override a variable the process was actually started with. A
|
||||
// deployment that sets DATA_DIR in its unit file must not have it
|
||||
// silently replaced by a stale .env left in the working directory.
|
||||
func TestLoadDotEnv_RealEnvironmentWins(t *testing.T) {
|
||||
t.Setenv(dotEnvKey, "from-environment")
|
||||
|
||||
path := writeDotEnv(t, dotEnvKey+"=from-dot-env\n")
|
||||
|
||||
require.NoError(t, config.LoadDotEnvFileForTest(path))
|
||||
assert.Equal(t, "from-environment", os.Getenv(dotEnvKey))
|
||||
}
|
||||
|
||||
// TestLoadDotEnv_MalformedFileAborts is the defect this fixes. One bad
|
||||
// line makes godotenv apply none of the file, so every variable in it
|
||||
// reverts to its default; the process used to start that way with no
|
||||
// log line naming the file at all.
|
||||
//
|
||||
//nolint:paralleltest // unsetDotEnvKey uses t.Setenv.
|
||||
func TestLoadDotEnv_MalformedFileAborts(t *testing.T) {
|
||||
unsetDotEnvKey(t)
|
||||
|
||||
path := writeDotEnv(
|
||||
t, malformedDotEnv+dotEnvKey+"=from-dot-env\n",
|
||||
)
|
||||
|
||||
err := config.LoadDotEnvFileForTest(path)
|
||||
|
||||
require.Error(t, err)
|
||||
require.ErrorIs(t, err, config.ErrDotEnvUnreadable)
|
||||
assert.Contains(
|
||||
t, err.Error(), config.DotEnvPath,
|
||||
"the failure must name the file it could not read",
|
||||
)
|
||||
|
||||
_, present := os.LookupEnv(dotEnvKey)
|
||||
assert.False(
|
||||
t, present,
|
||||
"a rejected file must apply nothing, not part of itself",
|
||||
)
|
||||
}
|
||||
|
||||
// TestLoadDotEnv_UnreadableFileAborts pins that only absence is
|
||||
// tolerated. A .env that exists but cannot be read is a file the
|
||||
// operator meant to be applied, so it fails like a malformed one
|
||||
// rather than being treated as though it were not there.
|
||||
func TestLoadDotEnv_UnreadableFileAborts(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
// A directory in the file's place: open succeeds and the read
|
||||
// fails, which no umask or root-ness can turn back into success
|
||||
// the way a chmod could.
|
||||
path := filepath.Join(t.TempDir(), config.DotEnvPath)
|
||||
require.NoError(t, os.Mkdir(path, 0o750))
|
||||
|
||||
err := config.LoadDotEnvFileForTest(path)
|
||||
|
||||
require.Error(t, err)
|
||||
require.ErrorIs(t, err, config.ErrDotEnvUnreadable)
|
||||
}
|
||||
|
||||
// TestLoadDotEnv_ReadsTheWorkingDirectory pins the path LoadDotEnv
|
||||
// itself opens, which the tests above bypass. It is relative to the
|
||||
// process working directory, as it was under godotenv/autoload and as
|
||||
// the README documents.
|
||||
//
|
||||
//nolint:paralleltest // t.Chdir moves the whole process.
|
||||
func TestLoadDotEnv_ReadsTheWorkingDirectory(t *testing.T) {
|
||||
unsetDotEnvKey(t)
|
||||
|
||||
dir := t.TempDir()
|
||||
require.NoError(t, os.WriteFile(
|
||||
filepath.Join(dir, config.DotEnvPath),
|
||||
[]byte(dotEnvKey+"=from-working-directory\n"),
|
||||
0o600,
|
||||
))
|
||||
t.Chdir(dir)
|
||||
|
||||
require.NoError(t, config.LoadDotEnv())
|
||||
assert.Equal(t, "from-working-directory", os.Getenv(dotEnvKey))
|
||||
}
|
||||
@@ -21,6 +21,22 @@ const (
|
||||
envKeyPort = "PORT"
|
||||
envKeyDebug = "DEBUG"
|
||||
envKeyMaintenanceMode = "MAINTENANCE_MODE"
|
||||
envKeyBindAddress = "BIND_ADDRESS"
|
||||
)
|
||||
|
||||
// Sample BIND_ADDRESS values used by the tables below.
|
||||
const (
|
||||
// bindAddressDefault is the shipped default. It is asserted
|
||||
// against the package's own constant in
|
||||
// TestNewUsesDefaultsWhenUnset, so the two cannot drift.
|
||||
bindAddressDefault = "127.0.0.1"
|
||||
|
||||
// bindAddressWildcard is the value a container deployment sets.
|
||||
bindAddressWildcard = "0.0.0.0"
|
||||
|
||||
// bindAddressSample is an arbitrary specific address, standing
|
||||
// for "one interface of several".
|
||||
bindAddressSample = "10.1.2.3"
|
||||
)
|
||||
|
||||
// envBoolCase is one row of the envBool table.
|
||||
@@ -291,6 +307,160 @@ func TestEnvPort(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
// TestEnvBindAddress covers BIND_ADDRESS parsing.
|
||||
//
|
||||
// Only IP address literals are accepted. Every rejection below is a
|
||||
// value an operator plausibly writes — a hostname, a host:port, a
|
||||
// CIDR block — and each has to abort startup rather than fall back to
|
||||
// the default, because falling back would bind an address other than
|
||||
// the one asked for and, in the wildcard-default case this setting
|
||||
// exists to end, publish cleartext on every interface.
|
||||
func TestEnvBindAddress(t *testing.T) {
|
||||
for _, tt := range envBindAddressCases() {
|
||||
t.Run(tt.name, func(t *testing.T) {
|
||||
// Cannot use t.Parallel() here because t.Setenv
|
||||
// is incompatible with parallel subtests.
|
||||
if tt.set {
|
||||
t.Setenv(testEnvKey, tt.value)
|
||||
} else {
|
||||
require.NoError(t, os.Unsetenv(testEnvKey))
|
||||
}
|
||||
|
||||
got, err := config.EnvBindAddressForTest(
|
||||
testEnvKey, bindAddressDefault,
|
||||
)
|
||||
|
||||
if tt.expectError {
|
||||
require.Error(t, err)
|
||||
require.ErrorIs(t, err, config.ErrInvalidBindAddress)
|
||||
assert.Contains(t, err.Error(), testEnvKey)
|
||||
assert.Contains(t, err.Error(), tt.value)
|
||||
|
||||
return
|
||||
}
|
||||
|
||||
require.NoError(t, err)
|
||||
assert.Equal(t, tt.expected, got)
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// envBindAddressCase is one row of the envBindAddress table.
|
||||
type envBindAddressCase struct {
|
||||
name string
|
||||
set bool
|
||||
value string
|
||||
expectError bool
|
||||
expected string
|
||||
}
|
||||
|
||||
// envBindAddressCases is the envBindAddress table, kept out of the
|
||||
// test body so the test itself stays readable.
|
||||
func envBindAddressCases() []envBindAddressCase {
|
||||
return append(
|
||||
envBindAddressAcceptedCases(),
|
||||
envBindAddressRejectedCases()...,
|
||||
)
|
||||
}
|
||||
|
||||
// envBindAddressAcceptedCases are the values that parse: the three
|
||||
// spellings of "unset" that take the default, and the literals.
|
||||
func envBindAddressAcceptedCases() []envBindAddressCase {
|
||||
return []envBindAddressCase{
|
||||
{
|
||||
name: "unset returns the default",
|
||||
expected: bindAddressDefault,
|
||||
},
|
||||
{
|
||||
name: "empty returns the default",
|
||||
set: true,
|
||||
value: "",
|
||||
expected: bindAddressDefault,
|
||||
},
|
||||
{
|
||||
name: "whitespace returns the default",
|
||||
set: true,
|
||||
value: " ",
|
||||
expected: bindAddressDefault,
|
||||
},
|
||||
{
|
||||
name: "ipv4 wildcard is parsed",
|
||||
set: true,
|
||||
value: bindAddressWildcard,
|
||||
expected: bindAddressWildcard,
|
||||
},
|
||||
{
|
||||
name: "ipv4 literal is parsed",
|
||||
set: true,
|
||||
value: bindAddressSample,
|
||||
expected: bindAddressSample,
|
||||
},
|
||||
{
|
||||
name: "surrounding whitespace is trimmed",
|
||||
set: true,
|
||||
value: " " + bindAddressSample + " ",
|
||||
expected: bindAddressSample,
|
||||
},
|
||||
{
|
||||
name: "ipv6 wildcard is parsed",
|
||||
set: true,
|
||||
value: "::",
|
||||
expected: "::",
|
||||
},
|
||||
{
|
||||
name: "ipv6 literal is parsed",
|
||||
set: true,
|
||||
value: "2001:db8::5",
|
||||
expected: "2001:db8::5",
|
||||
},
|
||||
}
|
||||
}
|
||||
|
||||
// envBindAddressRejectedCases are the values that abort startup.
|
||||
// Each is something an operator plausibly writes, and none may fall
|
||||
// back to the default: the default is loopback, so a silent fallback
|
||||
// would bind somewhere other than what was asked for.
|
||||
func envBindAddressRejectedCases() []envBindAddressCase {
|
||||
return []envBindAddressCase{
|
||||
{
|
||||
name: "garbage is rejected",
|
||||
set: true,
|
||||
value: "not-an-address",
|
||||
expectError: true,
|
||||
},
|
||||
{
|
||||
name: "hostname is rejected",
|
||||
set: true,
|
||||
value: "localhost",
|
||||
expectError: true,
|
||||
},
|
||||
{
|
||||
name: "unresolvable hostname is rejected",
|
||||
set: true,
|
||||
value: "no-such-host.invalid",
|
||||
expectError: true,
|
||||
},
|
||||
{
|
||||
name: "host and port is rejected",
|
||||
set: true,
|
||||
value: bindAddressDefault + ":8080",
|
||||
expectError: true,
|
||||
},
|
||||
{
|
||||
name: "bracketed ipv6 is rejected",
|
||||
set: true,
|
||||
value: "[::1]",
|
||||
expectError: true,
|
||||
},
|
||||
{
|
||||
name: "CIDR block is rejected",
|
||||
set: true,
|
||||
value: "10.0.0.0/8",
|
||||
expectError: true,
|
||||
},
|
||||
}
|
||||
}
|
||||
|
||||
// buildConfig constructs a Config through fx exactly as the
|
||||
// application does, returning the config and any construction error.
|
||||
func buildConfig(t *testing.T) (*config.Config, error) {
|
||||
@@ -312,58 +482,7 @@ func buildConfig(t *testing.T) (*config.Config, error) {
|
||||
}
|
||||
|
||||
func TestNewRejectsBadEnvValues(t *testing.T) {
|
||||
tests := []struct {
|
||||
name string
|
||||
key string
|
||||
value string
|
||||
expectError bool
|
||||
check func(t *testing.T, cfg *config.Config)
|
||||
}{
|
||||
{
|
||||
name: "valid PORT is used",
|
||||
key: envKeyPort,
|
||||
value: "9001",
|
||||
check: func(t *testing.T, cfg *config.Config) {
|
||||
t.Helper()
|
||||
assert.Equal(t, 9001, cfg.Port)
|
||||
},
|
||||
},
|
||||
{
|
||||
name: "unparseable PORT aborts startup",
|
||||
key: envKeyPort,
|
||||
value: "eighty-eighty",
|
||||
expectError: true,
|
||||
},
|
||||
{
|
||||
name: "out-of-range PORT aborts startup",
|
||||
key: envKeyPort,
|
||||
value: "70000",
|
||||
expectError: true,
|
||||
},
|
||||
{
|
||||
name: "valid DEBUG is used",
|
||||
key: envKeyDebug,
|
||||
value: "true",
|
||||
check: func(t *testing.T, cfg *config.Config) {
|
||||
t.Helper()
|
||||
assert.True(t, cfg.Debug)
|
||||
},
|
||||
},
|
||||
{
|
||||
name: "unparseable DEBUG aborts startup",
|
||||
key: envKeyDebug,
|
||||
value: "ture",
|
||||
expectError: true,
|
||||
},
|
||||
{
|
||||
name: "unparseable MAINTENANCE_MODE aborts startup",
|
||||
key: envKeyMaintenanceMode,
|
||||
value: "sometimes",
|
||||
expectError: true,
|
||||
},
|
||||
}
|
||||
|
||||
for _, tt := range tests {
|
||||
for _, tt := range badEnvValueCases() {
|
||||
t.Run(tt.name, func(t *testing.T) {
|
||||
// Cannot use t.Parallel() here because t.Setenv
|
||||
// is incompatible with parallel subtests.
|
||||
@@ -387,6 +506,149 @@ func TestNewRejectsBadEnvValues(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
// badEnvValueCase is one row of the config.New table: a variable, the
|
||||
// value it is set to, and either the assertion that startup fails
|
||||
// naming both, or a check on the Config that resulted.
|
||||
type badEnvValueCase struct {
|
||||
name string
|
||||
key string
|
||||
value string
|
||||
expectError bool
|
||||
check func(t *testing.T, cfg *config.Config)
|
||||
}
|
||||
|
||||
// badEnvValueCases is the config.New table, kept out of the test body
|
||||
// so the test itself stays readable. It is assembled from per-variable
|
||||
// groups because one literal covering every variable outgrew the
|
||||
// function-length budget.
|
||||
func badEnvValueCases() []badEnvValueCase {
|
||||
cases := listenerEnvValueCases()
|
||||
cases = append(cases, flagEnvValueCases()...)
|
||||
cases = append(cases, sentryEnvValueCases()...)
|
||||
|
||||
return cases
|
||||
}
|
||||
|
||||
// listenerEnvValueCases covers the two variables that describe the
|
||||
// HTTP listener.
|
||||
func listenerEnvValueCases() []badEnvValueCase {
|
||||
return []badEnvValueCase{
|
||||
{
|
||||
name: "valid PORT is used",
|
||||
key: envKeyPort,
|
||||
value: "9001",
|
||||
check: func(t *testing.T, cfg *config.Config) {
|
||||
t.Helper()
|
||||
assert.Equal(t, 9001, cfg.Port)
|
||||
},
|
||||
},
|
||||
{
|
||||
name: "unparseable PORT aborts startup",
|
||||
key: envKeyPort,
|
||||
value: "eighty-eighty",
|
||||
expectError: true,
|
||||
},
|
||||
{
|
||||
name: "out-of-range PORT aborts startup",
|
||||
key: envKeyPort,
|
||||
value: "70000",
|
||||
expectError: true,
|
||||
},
|
||||
{
|
||||
name: "valid BIND_ADDRESS is used",
|
||||
key: envKeyBindAddress,
|
||||
value: bindAddressWildcard,
|
||||
check: func(t *testing.T, cfg *config.Config) {
|
||||
t.Helper()
|
||||
assert.Equal(
|
||||
t, bindAddressWildcard, cfg.BindAddress,
|
||||
)
|
||||
},
|
||||
},
|
||||
{
|
||||
name: "unparseable BIND_ADDRESS aborts startup",
|
||||
key: envKeyBindAddress,
|
||||
value: "not-an-address",
|
||||
expectError: true,
|
||||
},
|
||||
{
|
||||
name: "hostname BIND_ADDRESS aborts startup",
|
||||
key: envKeyBindAddress,
|
||||
value: "localhost",
|
||||
expectError: true,
|
||||
},
|
||||
{
|
||||
name: "BIND_ADDRESS with a port aborts startup",
|
||||
key: envKeyBindAddress,
|
||||
value: bindAddressDefault + ":8080",
|
||||
expectError: true,
|
||||
},
|
||||
}
|
||||
}
|
||||
|
||||
// flagEnvValueCases covers the boolean variables.
|
||||
func flagEnvValueCases() []badEnvValueCase {
|
||||
return []badEnvValueCase{
|
||||
{
|
||||
name: "valid DEBUG is used",
|
||||
key: envKeyDebug,
|
||||
value: "true",
|
||||
check: func(t *testing.T, cfg *config.Config) {
|
||||
t.Helper()
|
||||
assert.True(t, cfg.Debug)
|
||||
},
|
||||
},
|
||||
{
|
||||
name: "unparseable DEBUG aborts startup",
|
||||
key: envKeyDebug,
|
||||
value: "ture",
|
||||
expectError: true,
|
||||
},
|
||||
{
|
||||
name: "unparseable MAINTENANCE_MODE aborts startup",
|
||||
key: envKeyMaintenanceMode,
|
||||
value: "sometimes",
|
||||
expectError: true,
|
||||
},
|
||||
}
|
||||
}
|
||||
|
||||
// sentryEnvValueCases covers SENTRY_DSN. The three rejected values are
|
||||
// the ones measured on the defect: each initialised the SDK with an
|
||||
// error and left the process serving with error reporting off.
|
||||
func sentryEnvValueCases() []badEnvValueCase {
|
||||
return []badEnvValueCase{
|
||||
{
|
||||
name: "valid SENTRY_DSN is used",
|
||||
key: envKeySentryDSN,
|
||||
value: validSentryDSN,
|
||||
check: func(t *testing.T, cfg *config.Config) {
|
||||
t.Helper()
|
||||
assert.Equal(t, validSentryDSN, cfg.SentryDSN)
|
||||
assert.True(t, cfg.SentryEnabled())
|
||||
},
|
||||
},
|
||||
{
|
||||
name: "unparseable SENTRY_DSN aborts startup",
|
||||
key: envKeySentryDSN,
|
||||
value: "not-a-dsn",
|
||||
expectError: true,
|
||||
},
|
||||
{
|
||||
name: "SENTRY_DSN that is not a URL aborts startup",
|
||||
key: envKeySentryDSN,
|
||||
value: "%%%",
|
||||
expectError: true,
|
||||
},
|
||||
{
|
||||
name: "keyless SENTRY_DSN aborts startup",
|
||||
key: envKeySentryDSN,
|
||||
value: "https://example.invalid/1",
|
||||
expectError: true,
|
||||
},
|
||||
}
|
||||
}
|
||||
|
||||
// TestNewUsesDefaultsWhenUnset proves the fail-loud behaviour did not
|
||||
// break the legitimate unset case: absent variables still get their
|
||||
// documented defaults.
|
||||
@@ -395,6 +657,7 @@ func TestNewUsesDefaultsWhenUnset(t *testing.T) {
|
||||
|
||||
for _, key := range []string{
|
||||
envKeyPort, envKeyDebug, envKeyMaintenanceMode,
|
||||
envKeyBindAddress, envKeySentryDSN,
|
||||
} {
|
||||
require.NoError(t, os.Unsetenv(key))
|
||||
}
|
||||
@@ -406,4 +669,20 @@ func TestNewUsesDefaultsWhenUnset(t *testing.T) {
|
||||
assert.Equal(t, 8080, cfg.Port)
|
||||
assert.False(t, cfg.Debug)
|
||||
assert.False(t, cfg.MaintenanceMode)
|
||||
|
||||
// Loopback, not the wildcard: the default must not publish the
|
||||
// cleartext admin UI and the unauthenticated receiver on every
|
||||
// interface of a host that configured nothing. The value is read
|
||||
// from the package rather than repeated, so the README's
|
||||
// documented default and the compiled-in one are pinned to the
|
||||
// same constant.
|
||||
assert.Equal(
|
||||
t, config.DefaultBindAddressForTest, cfg.BindAddress,
|
||||
)
|
||||
assert.Equal(t, bindAddressDefault, cfg.BindAddress)
|
||||
|
||||
// An absent SENTRY_DSN is the common case and must stay a normal
|
||||
// start with error reporting off, not a refusal.
|
||||
assert.Empty(t, cfg.SentryDSN)
|
||||
assert.False(t, cfg.SentryEnabled())
|
||||
}
|
||||
|
||||
@@ -50,3 +50,25 @@ func EnvPositiveIntForTest(key string, defaultValue int) (int, error) {
|
||||
func EnvPortForTest(key string, defaultValue int) (int, error) {
|
||||
return envPort(key, defaultValue)
|
||||
}
|
||||
|
||||
// EnvSentryDSNForTest exposes envSentryDSN.
|
||||
func EnvSentryDSNForTest(key string) (string, error) {
|
||||
return envSentryDSN(key)
|
||||
}
|
||||
|
||||
// LoadDotEnvFileForTest exposes the loader LoadDotEnv runs, over a
|
||||
// caller-named file rather than the process working directory, so
|
||||
// each .env state can be covered without moving the test process.
|
||||
func LoadDotEnvFileForTest(path string) error {
|
||||
return loadDotEnvFile(path)
|
||||
}
|
||||
|
||||
// EnvBindAddressForTest exposes envBindAddress.
|
||||
func EnvBindAddressForTest(key, defaultValue string) (string, error) {
|
||||
return envBindAddress(key, defaultValue)
|
||||
}
|
||||
|
||||
// DefaultBindAddressForTest exposes the compiled-in BIND_ADDRESS
|
||||
// default, so a test pins the documented value rather than repeating
|
||||
// a literal that could drift from it.
|
||||
const DefaultBindAddressForTest = defaultBindAddress
|
||||
|
||||
141
internal/config/sentry_test.go
Normal file
141
internal/config/sentry_test.go
Normal file
@@ -0,0 +1,141 @@
|
||||
package config_test
|
||||
|
||||
import (
|
||||
"os"
|
||||
"testing"
|
||||
|
||||
"github.com/stretchr/testify/assert"
|
||||
"github.com/stretchr/testify/require"
|
||||
"sneak.berlin/go/webhooker/internal/config"
|
||||
)
|
||||
|
||||
// envKeySentryDSN is the variable envSentryDSN reads in production.
|
||||
const envKeySentryDSN = "SENTRY_DSN"
|
||||
|
||||
// validSentryDSN is a syntactically complete DSN. The host is under
|
||||
// .invalid (RFC 2606), so nothing a test builds around it can reach a
|
||||
// real Sentry installation.
|
||||
const validSentryDSN = "https://abc123@sentry.invalid/42"
|
||||
|
||||
// envSentryDSNCase is one row of the envSentryDSN table.
|
||||
type envSentryDSNCase struct {
|
||||
name string
|
||||
set bool
|
||||
value string
|
||||
expectError bool
|
||||
expected string
|
||||
}
|
||||
|
||||
// envSentryDSNCases is the envSentryDSN table. The three invalid
|
||||
// values are the ones measured on the defect: each initialised the SDK
|
||||
// with an error and left the process serving with reporting off.
|
||||
func envSentryDSNCases() []envSentryDSNCase {
|
||||
return []envSentryDSNCase{
|
||||
{
|
||||
name: "unset means reporting off",
|
||||
expected: "",
|
||||
},
|
||||
{
|
||||
name: "empty means reporting off",
|
||||
set: true,
|
||||
value: "",
|
||||
expected: "",
|
||||
},
|
||||
{
|
||||
name: "whitespace means reporting off",
|
||||
set: true,
|
||||
value: " ",
|
||||
expected: "",
|
||||
},
|
||||
{
|
||||
name: "a valid DSN is kept",
|
||||
set: true,
|
||||
value: validSentryDSN,
|
||||
expected: validSentryDSN,
|
||||
},
|
||||
{
|
||||
name: "surrounding whitespace is trimmed",
|
||||
set: true,
|
||||
value: " " + validSentryDSN + "\t",
|
||||
expected: validSentryDSN,
|
||||
},
|
||||
{
|
||||
name: "a value that is not a URL is rejected",
|
||||
set: true,
|
||||
value: "not-a-dsn",
|
||||
expectError: true,
|
||||
},
|
||||
{
|
||||
name: "an unparseable URL is rejected",
|
||||
set: true,
|
||||
value: "%%%",
|
||||
expectError: true,
|
||||
},
|
||||
{
|
||||
name: "a DSN without a public key is rejected",
|
||||
set: true,
|
||||
value: "https://example.invalid/1",
|
||||
expectError: true,
|
||||
},
|
||||
{
|
||||
name: "a DSN without a project id is rejected",
|
||||
set: true,
|
||||
value: "https://abc123@sentry.invalid/",
|
||||
expectError: true,
|
||||
},
|
||||
{
|
||||
name: "a non-HTTP scheme is rejected",
|
||||
set: true,
|
||||
value: "ftp://abc123@sentry.invalid/42",
|
||||
expectError: true,
|
||||
},
|
||||
}
|
||||
}
|
||||
|
||||
// TestEnvSentryDSN covers the helper directly. What it pins beyond the
|
||||
// value is the failure shape: a set-but-unparseable DSN names the
|
||||
// variable and the value, exactly as the other fail-loud helpers do,
|
||||
// so an operator reads the fix off the message.
|
||||
func TestEnvSentryDSN(t *testing.T) {
|
||||
for _, tt := range envSentryDSNCases() {
|
||||
t.Run(tt.name, func(t *testing.T) {
|
||||
// Cannot use t.Parallel() here because t.Setenv
|
||||
// is incompatible with parallel subtests.
|
||||
if tt.set {
|
||||
t.Setenv(envKeySentryDSN, tt.value)
|
||||
} else {
|
||||
require.NoError(t, os.Unsetenv(envKeySentryDSN))
|
||||
}
|
||||
|
||||
got, err := config.EnvSentryDSNForTest(envKeySentryDSN)
|
||||
|
||||
if tt.expectError {
|
||||
require.Error(t, err)
|
||||
require.ErrorIs(t, err, config.ErrInvalidSentryDSN)
|
||||
assert.Contains(t, err.Error(), envKeySentryDSN)
|
||||
assert.Contains(t, err.Error(), tt.value)
|
||||
assert.Empty(t, got)
|
||||
|
||||
return
|
||||
}
|
||||
|
||||
require.NoError(t, err)
|
||||
assert.Equal(t, tt.expected, got)
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// TestSentryEnabled_TracksTheDSN pins that the one method answering
|
||||
// "is anything being reported" agrees with the DSN in every state. The
|
||||
// startup log, the SDK initialisation and the sentryhttp middleware
|
||||
// all read it, so a log field cannot report reporting as on while
|
||||
// nothing is sending.
|
||||
func TestSentryEnabled_TracksTheDSN(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
assert.False(t, (&config.Config{}).SentryEnabled())
|
||||
assert.True(
|
||||
t,
|
||||
(&config.Config{SentryDSN: validSentryDSN}).SentryEnabled(),
|
||||
)
|
||||
}
|
||||
@@ -4,7 +4,6 @@ package database
|
||||
import (
|
||||
"context"
|
||||
"crypto/rand"
|
||||
"database/sql"
|
||||
"encoding/base64"
|
||||
"errors"
|
||||
"fmt"
|
||||
@@ -16,7 +15,6 @@ import (
|
||||
"go.uber.org/fx"
|
||||
"gorm.io/driver/sqlite"
|
||||
"gorm.io/gorm"
|
||||
_ "modernc.org/sqlite" // Pure Go SQLite driver
|
||||
"sneak.berlin/go/webhooker/internal/banner"
|
||||
"sneak.berlin/go/webhooker/internal/config"
|
||||
"sneak.berlin/go/webhooker/internal/gormlog"
|
||||
@@ -198,13 +196,11 @@ func (d *Database) connectTo(dataDir string) error {
|
||||
|
||||
// Construct the main application database path inside DATA_DIR.
|
||||
dbPath := filepath.Join(dataDir, MainDBFileName)
|
||||
dbURL := fmt.Sprintf(
|
||||
"file:%s?cache=shared&mode=rwc",
|
||||
dbPath,
|
||||
)
|
||||
|
||||
// Open the database with the pure Go SQLite driver
|
||||
sqlDB, err := sql.Open("sqlite", dbURL)
|
||||
// Opened through OpenSQLite so this handle carries the same WAL
|
||||
// journaling, busy timeout, immediate-transaction locking, and pool
|
||||
// bounds as every other database file. See sqlite_open.go.
|
||||
sqlDB, err := OpenSQLite(dbPath, SQLiteModeCreate)
|
||||
if err != nil {
|
||||
d.log.Error(
|
||||
"failed to open database",
|
||||
|
||||
240
internal/database/sqlite_mode_test.go
Normal file
240
internal/database/sqlite_mode_test.go
Normal file
@@ -0,0 +1,240 @@
|
||||
package database_test
|
||||
|
||||
import (
|
||||
"context"
|
||||
"io/fs"
|
||||
"net/http"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"testing"
|
||||
|
||||
"github.com/google/uuid"
|
||||
"github.com/stretchr/testify/assert"
|
||||
"github.com/stretchr/testify/require"
|
||||
"go.uber.org/fx/fxtest"
|
||||
"sneak.berlin/go/webhooker/internal/config"
|
||||
"sneak.berlin/go/webhooker/internal/database"
|
||||
"sneak.berlin/go/webhooker/internal/globals"
|
||||
"sneak.berlin/go/webhooker/internal/logger"
|
||||
)
|
||||
|
||||
// ownerOnly is the mode every SQLite file the service owns must have.
|
||||
// Spelled out rather than referencing database.SQLiteFilePerm so the
|
||||
// test fails if the constant itself is loosened.
|
||||
const ownerOnly fs.FileMode = 0o600
|
||||
|
||||
// requireOwnerOnly asserts that path exists and is readable and
|
||||
// writable by its owner and by nobody else.
|
||||
func requireOwnerOnly(t *testing.T, path string) {
|
||||
t.Helper()
|
||||
|
||||
info, err := os.Stat(path)
|
||||
require.NoError(t, err, "%s must exist", path)
|
||||
assert.Equal(
|
||||
t,
|
||||
ownerOnly,
|
||||
info.Mode().Perm(),
|
||||
"%s holds credentials and must not be readable by "+
|
||||
"anyone but its owner",
|
||||
path,
|
||||
)
|
||||
}
|
||||
|
||||
// requireDatabaseSetOwnerOnly asserts the mode of a database file and
|
||||
// of both WAL sidecars. The sidecars carry the same rows as the
|
||||
// database, so tightening only the main file fixes nothing.
|
||||
func requireDatabaseSetOwnerOnly(t *testing.T, dbPath string) {
|
||||
t.Helper()
|
||||
|
||||
requireOwnerOnly(t, dbPath)
|
||||
requireOwnerOnly(t, dbPath+"-wal")
|
||||
requireOwnerOnly(t, dbPath+"-shm")
|
||||
}
|
||||
|
||||
// TestMainDatabaseFilesAreOwnerOnly covers the tier the defect was
|
||||
// reported against: webhooker.db holds targets.config in plaintext —
|
||||
// bearer tokens, API keys, Slack webhook URLs — and the session
|
||||
// encryption key.
|
||||
func TestMainDatabaseFilesAreOwnerOnly(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
lc := fxtest.NewLifecycle(t)
|
||||
|
||||
l, err := logger.New(lc, logger.LoggerParams{
|
||||
Globals: &globals.Globals{
|
||||
Appname: testAppname,
|
||||
Version: testVersion,
|
||||
},
|
||||
})
|
||||
require.NoError(t, err)
|
||||
|
||||
// A directory the application creates itself, not one t.TempDir
|
||||
// made at 0700, so the mode below is the application's.
|
||||
dataDir := filepath.Join(t.TempDir(), "data")
|
||||
|
||||
db, err := database.New(lc, database.DatabaseParams{
|
||||
Config: &config.Config{DataDir: dataDir},
|
||||
Logger: l,
|
||||
})
|
||||
require.NoError(t, err)
|
||||
|
||||
ctx := context.Background()
|
||||
require.NoError(t, lc.Start(ctx))
|
||||
|
||||
defer func() { require.NoError(t, lc.Stop(ctx)) }()
|
||||
|
||||
// Write through the real model so the WAL is populated and both
|
||||
// sidecars are on disk while the handle is open.
|
||||
require.NoError(t, db.DB().Create(&database.Webhook{
|
||||
Name: testWebhookName,
|
||||
}).Error)
|
||||
|
||||
requireDatabaseSetOwnerOnly(
|
||||
t, filepath.Join(dataDir, database.MainDBFileName),
|
||||
)
|
||||
|
||||
// The data directory grants nothing to `other`. Asserted as a
|
||||
// property rather than as an exact 0750, because MkdirAll applies
|
||||
// the ambient umask: the exact mode is the developer's umask as
|
||||
// much as the application's request, and pinning it would make
|
||||
// `make check` pass or fail on where it is run. The group bits are
|
||||
// deliberately left unasserted — deployments may rely on them.
|
||||
info, err := os.Stat(dataDir)
|
||||
require.NoError(t, err)
|
||||
assert.Zero(
|
||||
t,
|
||||
info.Mode().Perm()&0o007,
|
||||
"the data directory must not be world-accessible",
|
||||
)
|
||||
}
|
||||
|
||||
// TestPerWebhookEventDatabaseFilesAreOwnerOnly covers the events-*.db
|
||||
// tier. These carry no credential canaries since
|
||||
// https://git.eeqj.de/sneak/webhooker/issues/206, but they hold every
|
||||
// received request body and header.
|
||||
func TestPerWebhookEventDatabaseFilesAreOwnerOnly(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
mgr, lc := setupTestWebhookDBManager(t)
|
||||
ctx := context.Background()
|
||||
require.NoError(t, lc.Start(ctx))
|
||||
|
||||
defer func() { require.NoError(t, lc.Stop(ctx)) }()
|
||||
|
||||
webhookID := uuid.New().String()
|
||||
|
||||
db, err := mgr.GetDB(webhookID)
|
||||
require.NoError(t, err)
|
||||
|
||||
require.NoError(t, db.Create(&database.Event{
|
||||
WebhookID: webhookID,
|
||||
EntrypointID: uuid.New().String(),
|
||||
Method: http.MethodPost,
|
||||
Body: "{}",
|
||||
}).Error)
|
||||
|
||||
requireDatabaseSetOwnerOnly(t, mgr.DBPath(webhookID))
|
||||
}
|
||||
|
||||
// TestArchiveDatabaseFilesAreOwnerOnly covers the archive-*.db tier.
|
||||
// internal/delivery builds that path and opens it through OpenSQLite,
|
||||
// the same single open path exercised here, so the mode is settled for
|
||||
// all three tiers in one place.
|
||||
func TestArchiveDatabaseFilesAreOwnerOnly(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
ctx := context.Background()
|
||||
path := filepath.Join(
|
||||
t.TempDir(), "archive-"+uuid.New().String()+".db",
|
||||
)
|
||||
|
||||
sqlDB, err := database.OpenSQLite(path, database.SQLiteModeCreate)
|
||||
require.NoError(t, err)
|
||||
|
||||
defer func() { require.NoError(t, sqlDB.Close()) }()
|
||||
|
||||
_, err = sqlDB.ExecContext(ctx, "create table t (id integer)")
|
||||
require.NoError(t, err)
|
||||
|
||||
requireDatabaseSetOwnerOnly(t, path)
|
||||
}
|
||||
|
||||
// TestOpenSQLiteTightensFilesLeftWorldReadable is the upgrade case: a
|
||||
// data directory an earlier build left at 0644, including a
|
||||
// developer's own scratch directory, is fixed when it is opened rather
|
||||
// than staying exposed until it is recreated.
|
||||
func TestOpenSQLiteTightensFilesLeftWorldReadable(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
dir := t.TempDir()
|
||||
path := filepath.Join(dir, database.MainDBFileName)
|
||||
|
||||
// A database and both sidecars as the pre-fix build left them.
|
||||
for _, p := range []string{path, path + "-wal", path + "-shm"} {
|
||||
require.NoError(t, os.WriteFile(p, nil, 0o644)) //nolint:gosec // the mode under test
|
||||
}
|
||||
|
||||
sqlDB, err := database.OpenSQLite(path, database.SQLiteModeCreate)
|
||||
require.NoError(t, err)
|
||||
|
||||
require.NoError(t, sqlDB.Close())
|
||||
|
||||
requireDatabaseSetOwnerOnly(t, path)
|
||||
}
|
||||
|
||||
// TestOpenSQLiteExistingModeDoesNotCreateTheFile guards the mechanism
|
||||
// the fix uses: OpenSQLite now creates the database file itself, and
|
||||
// must not do so for a caller that asked for an existing database. An
|
||||
// empty file materialized here would turn a missing-database error
|
||||
// into a silently empty one.
|
||||
func TestOpenSQLiteExistingModeDoesNotCreateTheFile(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
ctx := context.Background()
|
||||
path := filepath.Join(t.TempDir(), "absent.db")
|
||||
|
||||
sqlDB, err := database.OpenSQLite(path, database.SQLiteModeExisting)
|
||||
if err == nil {
|
||||
// sql.Open is lazy: force the connection that fails.
|
||||
require.Error(t, sqlDB.PingContext(ctx))
|
||||
require.NoError(t, sqlDB.Close())
|
||||
}
|
||||
|
||||
_, statErr := os.Stat(path)
|
||||
assert.ErrorIs(t, statErr, fs.ErrNotExist)
|
||||
}
|
||||
|
||||
// TestReopenAfterRestartKeepsFilesOwnerOnly is the restart case: a
|
||||
// process that closed its files must be able to open them again at
|
||||
// 0600, including through a gorm handle, and the sidecars must come
|
||||
// back at 0600 too rather than at SQLite's own default.
|
||||
func TestReopenAfterRestartKeepsFilesOwnerOnly(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
ctx := context.Background()
|
||||
dir := t.TempDir()
|
||||
path := filepath.Join(dir, database.MainDBFileName)
|
||||
|
||||
first, err := database.OpenSQLite(path, database.SQLiteModeCreate)
|
||||
require.NoError(t, err)
|
||||
|
||||
_, err = first.ExecContext(ctx, "create table t (id integer)")
|
||||
require.NoError(t, err)
|
||||
require.NoError(t, first.Close())
|
||||
|
||||
second, err := database.OpenSQLite(path, database.SQLiteModeCreate)
|
||||
require.NoError(t, err)
|
||||
|
||||
defer func() { require.NoError(t, second.Close()) }()
|
||||
|
||||
_, err = second.ExecContext(ctx, "insert into t (id) values (1)")
|
||||
require.NoError(t, err)
|
||||
|
||||
requireDatabaseSetOwnerOnly(t, path)
|
||||
|
||||
var got int
|
||||
|
||||
require.NoError(t,
|
||||
second.QueryRowContext(ctx, "select id from t").Scan(&got))
|
||||
assert.Equal(t, 1, got)
|
||||
}
|
||||
252
internal/database/sqlite_open.go
Normal file
252
internal/database/sqlite_open.go
Normal file
@@ -0,0 +1,252 @@
|
||||
package database
|
||||
|
||||
import (
|
||||
"database/sql"
|
||||
"errors"
|
||||
"fmt"
|
||||
"io/fs"
|
||||
"net/url"
|
||||
"os"
|
||||
"time"
|
||||
|
||||
_ "modernc.org/sqlite" // Pure Go SQLite driver
|
||||
)
|
||||
|
||||
// Every SQLite file this service opens — the main database, the
|
||||
// per-webhook event databases, and the archive databases — is opened
|
||||
// through OpenSQLite, so the durability settings below are properties
|
||||
// of the service rather than of one call site.
|
||||
//
|
||||
// modernc.org/sqlite installs no busy handler and issues no pragmas of
|
||||
// its own: it executes only the pragmas named in explicit `_pragma=`
|
||||
// DSN parameters, and gorm.io/driver/sqlite adds none when it is
|
||||
// handed an existing *sql.DB. Every setting therefore has to be
|
||||
// spelled out here or it is simply not in effect.
|
||||
// SQLite URI open modes.
|
||||
const (
|
||||
// SQLiteModeCreate creates the database file when it is missing.
|
||||
SQLiteModeCreate = "rwc"
|
||||
|
||||
// SQLiteModeExisting requires the file to exist already.
|
||||
SQLiteModeExisting = "rw"
|
||||
)
|
||||
|
||||
const (
|
||||
// SQLiteBusyTimeout is how long SQLite retries a lock conflict
|
||||
// before returning SQLITE_BUSY.
|
||||
//
|
||||
// Under WAL a reader never blocks a writer, so the only conflict
|
||||
// left is writer against writer: this process's delivery workers
|
||||
// against each other, or against another process holding the write
|
||||
// lock. Those clear in milliseconds. Ten seconds is far above that
|
||||
// and still well inside the receiver's request budget, so an
|
||||
// inbound webhook waits rather than being rejected with a 500.
|
||||
SQLiteBusyTimeout = 10 * time.Second
|
||||
|
||||
// sqliteMaxOpenConns bounds the connection pool for one database
|
||||
// file.
|
||||
//
|
||||
// The pool needs a bound at all because database/sql cannot detect
|
||||
// a connection left mid-transaction: modernc.org/sqlite implements
|
||||
// neither driver.Validator nor driver.SessionResetter, so a
|
||||
// connection whose COMMIT failed is returned to the pool with its
|
||||
// transaction still open and handed out again indefinitely. That is
|
||||
// what turned four `database is locked` errors into 593
|
||||
// `cannot start a transaction within a transaction` in
|
||||
// https://git.eeqj.de/sneak/webhooker/issues/256.
|
||||
//
|
||||
// Four is above the one writer SQLite allows at a time, so reads
|
||||
// still proceed while a write is in flight, and low enough that
|
||||
// contention is resolved by the busy handler rather than by piling
|
||||
// up connections against a lock only one of them can hold.
|
||||
sqliteMaxOpenConns = 4
|
||||
|
||||
// sqliteMaxIdleConns keeps the pool warm without holding every
|
||||
// connection open through an idle period.
|
||||
sqliteMaxIdleConns = 2
|
||||
|
||||
// sqliteConnMaxLifetime and sqliteConnMaxIdleTime retire pooled
|
||||
// connections on a schedule. With _txlock=immediate a failed
|
||||
// COMMIT should no longer be reachable, but these bound the damage
|
||||
// if one happens anyway: a poisoned connection is closed and
|
||||
// replaced within the lifetime instead of wedging the file until
|
||||
// the process restarts.
|
||||
sqliteConnMaxLifetime = 5 * time.Minute
|
||||
sqliteConnMaxIdleTime = time.Minute
|
||||
)
|
||||
|
||||
// SQLiteFilePerm is the mode every SQLite file this service owns is
|
||||
// created with and held at: owner read/write, nothing for group or
|
||||
// other.
|
||||
//
|
||||
// These files hold credentials in plaintext. The main database stores
|
||||
// `targets.config` — bearer tokens, API keys, Slack webhook URLs — and
|
||||
// the session encryption key. SQLite left to itself creates them 0644
|
||||
// (see reserveSQLiteFile), which made the 0750 data directory the only
|
||||
// barrier; a bind-mounted directory supplied at 0755 removes it and
|
||||
// every local user on the host can read every stored credential.
|
||||
//
|
||||
// This is a file-mode fix and not encryption at rest. An unattended
|
||||
// process needs a key it can read without a human, so the key lands
|
||||
// beside the data and an attacker who can read the database can read
|
||||
// it too. See https://git.eeqj.de/sneak/webhooker/issues/212.
|
||||
const SQLiteFilePerm fs.FileMode = 0o600
|
||||
|
||||
// reserveSQLiteFile puts path at SQLiteFilePerm before the driver ever
|
||||
// touches it, and tightens any sidecar already on disk.
|
||||
//
|
||||
// The mode has to be settled here rather than by a chmod after opening,
|
||||
// because SQLite picks it: robust_open substitutes
|
||||
// SQLITE_DEFAULT_FILE_PERMISSIONS (0644) whenever it is handed mode 0,
|
||||
// and findCreateFileMode yields 0 for a main database opened by URI
|
||||
// with no `modeof` parameter. A chmod afterwards would leave a window
|
||||
// in which the credentials are on disk world-readable.
|
||||
//
|
||||
// Creating the file ourselves also settles the sidecars, which is the
|
||||
// half that could quietly not work. SQLite does not create those at a
|
||||
// mode we choose — it derives both from the main database file:
|
||||
// `-wal` through findCreateFileMode, which stats the path with the
|
||||
// suffix stripped, and `-shm` in unixOpenSharedMemory from an fstat of
|
||||
// the already-open database descriptor. A main file at 0600 therefore
|
||||
// produces sidecars at 0600. A zero-length file is a valid empty
|
||||
// database, so reserving it changes nothing else.
|
||||
//
|
||||
// create says whether the caller is opening in a mode that may create
|
||||
// the database. When it is false a missing file is left missing, so
|
||||
// SQLite still reports the absence rather than this function
|
||||
// materializing an empty database the caller asked not to create.
|
||||
//
|
||||
// Chmod of a file that already exists is what tightens a data
|
||||
// directory an earlier build left at 0644 — including a developer's
|
||||
// own scratch directory — without any migration machinery.
|
||||
func reserveSQLiteFile(path string, create bool) error {
|
||||
if create {
|
||||
// gosec G304: the path is the database file the caller asked
|
||||
// to open, and the driver is about to open the same path
|
||||
// anyway. Creating it here is what fixes its mode.
|
||||
f, err := os.OpenFile( //nolint:gosec // see above
|
||||
path, os.O_RDWR|os.O_CREATE, SQLiteFilePerm,
|
||||
)
|
||||
if err != nil {
|
||||
return fmt.Errorf("creating %s: %w", path, err)
|
||||
}
|
||||
|
||||
err = f.Close()
|
||||
if err != nil {
|
||||
return fmt.Errorf("closing %s: %w", path, err)
|
||||
}
|
||||
}
|
||||
|
||||
// O_CREATE leaves an existing file's mode alone, and umask can only
|
||||
// have narrowed a new one. Chmod settles both cases at exactly
|
||||
// SQLiteFilePerm.
|
||||
for _, p := range append(
|
||||
[]string{path}, sqliteSidecarPaths(path)...,
|
||||
) {
|
||||
err := os.Chmod(p, SQLiteFilePerm)
|
||||
if err != nil && !errors.Is(err, fs.ErrNotExist) {
|
||||
return fmt.Errorf("securing %s: %w", p, err)
|
||||
}
|
||||
}
|
||||
|
||||
return nil
|
||||
}
|
||||
|
||||
// sqliteSidecarPaths returns the files SQLite maintains beside a
|
||||
// database under WAL. They carry the same rows as the database itself,
|
||||
// so a fix that tightens only the main file has fixed nothing.
|
||||
func sqliteSidecarPaths(path string) []string {
|
||||
return []string{path + "-wal", path + "-shm"}
|
||||
}
|
||||
|
||||
// SQLiteDSN builds the connection string for one database file.
|
||||
//
|
||||
// mode is the SQLite URI open mode: "rwc" to create the file when it
|
||||
// is missing, "rw" to require that it already exists.
|
||||
//
|
||||
// Three settings carry the fix for
|
||||
// https://git.eeqj.de/sneak/webhooker/issues/256 and none of them is
|
||||
// optional:
|
||||
//
|
||||
// - journal_mode=WAL, so a reader — an operator running
|
||||
// `sqlite3 <db> .dump` over their own data — takes a snapshot
|
||||
// instead of blocking every writer behind it.
|
||||
//
|
||||
// - busy_timeout, so a writer that does meet a lock waits for it.
|
||||
// Without one SQLite gives up immediately; nothing above it
|
||||
// retries.
|
||||
//
|
||||
// - _txlock=immediate, so every transaction takes the write lock at
|
||||
// BEGIN. A deferred transaction acquires it lazily on its first
|
||||
// write, and that upgrade returns SQLITE_BUSY *without* consulting
|
||||
// the busy handler, because SQLite cannot block a transaction that
|
||||
// may already hold a read snapshot. Such a COMMIT then fails while
|
||||
// the transaction stays open on the connection. A busy timeout
|
||||
// alone does not prevent this; BEGIN IMMEDIATE does, by putting
|
||||
// the wait somewhere the handler applies.
|
||||
//
|
||||
// Note what is absent: `cache=shared`. Under a shared cache an
|
||||
// in-process conflict is reported as SQLITE_LOCKED rather than
|
||||
// SQLITE_BUSY, and the busy handler does not retry SQLITE_LOCKED — so
|
||||
// leaving it in would have defeated the busy timeout for exactly the
|
||||
// contention this service generates. Dropping it is part of the fix,
|
||||
// not housekeeping.
|
||||
//
|
||||
// synchronous is deliberately left at SQLite's default of FULL: this
|
||||
// is a webhook receiver whose one promise is that an event it answered
|
||||
// 200 for is durable.
|
||||
// The order of the _pragma parameters is load-bearing.
|
||||
// modernc.org/sqlite executes them in the order they appear, on every
|
||||
// new connection, before the connection is handed to the pool. Setting
|
||||
// journal_mode first means that pragma itself runs with no busy
|
||||
// handler installed: the pool opens connections lazily, so the moment
|
||||
// a new one is created is a moment the database is under load, and
|
||||
// PRAGMA journal_mode takes a lock. It would fail immediately with
|
||||
// SQLITE_BUSY and fail the query that caused the connection to be
|
||||
// opened. busy_timeout is therefore set first, so every pragma after
|
||||
// it — and the whole life of the connection — is covered.
|
||||
func SQLiteDSN(path, mode string) string {
|
||||
q := url.Values{}
|
||||
q.Set("mode", mode)
|
||||
q.Set("_txlock", "immediate")
|
||||
q.Add(
|
||||
"_pragma",
|
||||
fmt.Sprintf(
|
||||
"busy_timeout(%d)",
|
||||
SQLiteBusyTimeout.Milliseconds(),
|
||||
),
|
||||
)
|
||||
q.Add("_pragma", "journal_mode(WAL)")
|
||||
|
||||
return "file:" + path + "?" + q.Encode()
|
||||
}
|
||||
|
||||
// OpenSQLite opens the SQLite file at path with the service's
|
||||
// durability settings and pool bounds applied. mode is the SQLite URI
|
||||
// open mode ("rwc" or "rw").
|
||||
//
|
||||
// The file and its WAL sidecars are settled at SQLiteFilePerm before
|
||||
// the driver sees the path; see reserveSQLiteFile.
|
||||
//
|
||||
// The handle is returned rather than a *gorm.DB because the callers
|
||||
// wrap it in gorm themselves with their own logger.
|
||||
func OpenSQLite(path, mode string) (*sql.DB, error) {
|
||||
err := reserveSQLiteFile(path, mode == SQLiteModeCreate)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
|
||||
sqlDB, err := sql.Open("sqlite", SQLiteDSN(path, mode))
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf(
|
||||
"opening sqlite database %s: %w", path, err,
|
||||
)
|
||||
}
|
||||
|
||||
sqlDB.SetMaxOpenConns(sqliteMaxOpenConns)
|
||||
sqlDB.SetMaxIdleConns(sqliteMaxIdleConns)
|
||||
sqlDB.SetConnMaxLifetime(sqliteConnMaxLifetime)
|
||||
sqlDB.SetConnMaxIdleTime(sqliteConnMaxIdleTime)
|
||||
|
||||
return sqlDB, nil
|
||||
}
|
||||
178
internal/database/sqlite_open_test.go
Normal file
178
internal/database/sqlite_open_test.go
Normal file
@@ -0,0 +1,178 @@
|
||||
package database_test
|
||||
|
||||
import (
|
||||
"context"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
"github.com/google/uuid"
|
||||
"github.com/stretchr/testify/assert"
|
||||
"github.com/stretchr/testify/require"
|
||||
"gorm.io/gorm"
|
||||
"sneak.berlin/go/webhooker/internal/database"
|
||||
)
|
||||
|
||||
// livePragma reads a pragma off a live handle. Reading the DSN back
|
||||
// would prove only that the string was built; these tests assert that
|
||||
// SQLite actually applied it.
|
||||
func livePragma(t *testing.T, db *gorm.DB, name string) string {
|
||||
t.Helper()
|
||||
|
||||
var v string
|
||||
|
||||
row := db.Raw("pragma " + name).Row()
|
||||
require.NoError(t, row.Scan(&v))
|
||||
|
||||
return v
|
||||
}
|
||||
|
||||
func TestSQLiteDSNCarriesTheDurabilitySettings(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
dsn := database.SQLiteDSN(
|
||||
"/var/lib/webhooker/webhooker.db",
|
||||
database.SQLiteModeCreate,
|
||||
)
|
||||
|
||||
assert.Contains(t, dsn, "journal_mode%28WAL%29")
|
||||
assert.Contains(t, dsn, "busy_timeout%2810000%29")
|
||||
assert.Contains(t, dsn, "_txlock=immediate")
|
||||
assert.Contains(t, dsn, "mode=rwc")
|
||||
|
||||
// busy_timeout must come first. The driver runs these in order on
|
||||
// every new connection, and PRAGMA journal_mode takes a lock — a
|
||||
// connection opened while the database is busy would fail on that
|
||||
// pragma, with no busy handler yet installed to wait it out.
|
||||
assert.Less(
|
||||
t,
|
||||
strings.Index(dsn, "busy_timeout"),
|
||||
strings.Index(dsn, "journal_mode"),
|
||||
"busy_timeout must be applied before journal_mode",
|
||||
)
|
||||
|
||||
// cache=shared turns an in-process conflict into SQLITE_LOCKED,
|
||||
// which the busy handler does not retry. It must never come back.
|
||||
// See https://git.eeqj.de/sneak/webhooker/issues/256.
|
||||
assert.NotContains(t, strings.ToLower(dsn), "cache=shared")
|
||||
}
|
||||
|
||||
// TestPerWebhookDBAppliesPragmasOnALiveHandle is the check the issue
|
||||
// asks for by name: the settings are confirmed by querying the running
|
||||
// database, not by inspecting the connection string.
|
||||
func TestPerWebhookDBAppliesPragmasOnALiveHandle(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
mgr, lc := setupTestWebhookDBManager(t)
|
||||
ctx := context.Background()
|
||||
require.NoError(t, lc.Start(ctx))
|
||||
|
||||
defer func() { require.NoError(t, lc.Stop(ctx)) }()
|
||||
|
||||
webhookID := uuid.New().String()
|
||||
|
||||
db, err := mgr.GetDB(webhookID)
|
||||
require.NoError(t, err)
|
||||
|
||||
assert.Equal(
|
||||
t, "wal",
|
||||
strings.ToLower(livePragma(t, db, "journal_mode")),
|
||||
)
|
||||
assert.Equal(
|
||||
t, "10000", livePragma(t, db, "busy_timeout"),
|
||||
)
|
||||
}
|
||||
|
||||
func TestMainDBAppliesPragmasOnALiveHandle(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
ctx := context.Background()
|
||||
dir := t.TempDir()
|
||||
|
||||
sqlDB, err := database.OpenSQLite(
|
||||
filepath.Join(dir, database.MainDBFileName),
|
||||
database.SQLiteModeCreate,
|
||||
)
|
||||
require.NoError(t, err)
|
||||
|
||||
defer func() { require.NoError(t, sqlDB.Close()) }()
|
||||
|
||||
var journal string
|
||||
|
||||
require.NoError(t, sqlDB.
|
||||
QueryRowContext(ctx, "pragma journal_mode").
|
||||
Scan(&journal))
|
||||
assert.Equal(t, "wal", strings.ToLower(journal))
|
||||
|
||||
var busy string
|
||||
|
||||
require.NoError(t, sqlDB.
|
||||
QueryRowContext(ctx, "pragma busy_timeout").
|
||||
Scan(&busy))
|
||||
assert.Equal(t, "10000", busy)
|
||||
}
|
||||
|
||||
// TestConcurrentReaderDoesNotBlockWrites is the unit-scale form of the
|
||||
// reproduction in
|
||||
// https://git.eeqj.de/sneak/webhooker/issues/256: an operator's
|
||||
// long-held read of their own data used to make every concurrent write
|
||||
// fail. Under WAL the reader takes a snapshot and the writes proceed.
|
||||
func TestConcurrentReaderDoesNotBlockWrites(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
mgr, lc := setupTestWebhookDBManager(t)
|
||||
ctx := context.Background()
|
||||
require.NoError(t, lc.Start(ctx))
|
||||
|
||||
defer func() { require.NoError(t, lc.Stop(ctx)) }()
|
||||
|
||||
webhookID := uuid.New().String()
|
||||
|
||||
db, err := mgr.GetDB(webhookID)
|
||||
require.NoError(t, err)
|
||||
|
||||
// A second handle on the same file, holding a read transaction
|
||||
// open across every write below — what `sqlite3 <db> .dump` is.
|
||||
readerSQL, err := database.OpenSQLite(
|
||||
mgr.DBPath(webhookID), database.SQLiteModeExisting,
|
||||
)
|
||||
require.NoError(t, err)
|
||||
|
||||
defer func() { require.NoError(t, readerSQL.Close()) }()
|
||||
|
||||
readerConn, err := readerSQL.Conn(ctx)
|
||||
require.NoError(t, err)
|
||||
|
||||
defer func() { require.NoError(t, readerConn.Close()) }()
|
||||
|
||||
_, err = readerConn.ExecContext(ctx, "begin deferred")
|
||||
require.NoError(t, err)
|
||||
|
||||
_, err = readerConn.ExecContext(
|
||||
ctx, "select count(*) from events",
|
||||
)
|
||||
require.NoError(t, err)
|
||||
|
||||
for range 25 {
|
||||
err = db.Transaction(func(tx *gorm.DB) error {
|
||||
return tx.Create(&database.Event{
|
||||
WebhookID: webhookID,
|
||||
EntrypointID: uuid.New().String(),
|
||||
Method: "POST",
|
||||
Body: "{}",
|
||||
}).Error
|
||||
})
|
||||
require.NoError(t, err)
|
||||
}
|
||||
|
||||
_, err = readerConn.ExecContext(ctx, "commit")
|
||||
require.NoError(t, err)
|
||||
|
||||
var count int64
|
||||
|
||||
require.NoError(
|
||||
t,
|
||||
db.Model(&database.Event{}).Count(&count).Error,
|
||||
)
|
||||
assert.Equal(t, int64(25), count)
|
||||
}
|
||||
@@ -2,7 +2,6 @@ package database
|
||||
|
||||
import (
|
||||
"context"
|
||||
"database/sql"
|
||||
"errors"
|
||||
"fmt"
|
||||
"log/slog"
|
||||
@@ -234,12 +233,11 @@ func (m *WebhookDBManager) openDB(
|
||||
webhookID string,
|
||||
) (*gorm.DB, error) {
|
||||
path := m.dbPath(webhookID)
|
||||
dbURL := fmt.Sprintf(
|
||||
"file:%s?cache=shared&mode=rwc",
|
||||
path,
|
||||
)
|
||||
|
||||
sqlDB, err := sql.Open("sqlite", dbURL)
|
||||
// See sqlite_open.go: WAL, a busy timeout, immediate-transaction
|
||||
// locking, and a bounded pool, all of which this file needs most —
|
||||
// it is the one every delivery worker writes to concurrently.
|
||||
sqlDB, err := OpenSQLite(path, SQLiteModeCreate)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf(
|
||||
"opening webhook database %s: %w",
|
||||
|
||||
File diff suppressed because it is too large
Load Diff
@@ -2,7 +2,6 @@ package delivery_test
|
||||
|
||||
import (
|
||||
"context"
|
||||
"database/sql"
|
||||
"encoding/json"
|
||||
"fmt"
|
||||
"io"
|
||||
@@ -70,11 +69,12 @@ func iMainDB(t *testing.T) *gorm.DB {
|
||||
t.TempDir(), "main-test.db",
|
||||
)
|
||||
|
||||
dsn := fmt.Sprintf(
|
||||
"file:%s?cache=shared&mode=rwc", dbPath,
|
||||
// Opened the way the service opens the main database, so these
|
||||
// tests cannot pass against journal and locking settings
|
||||
// production does not use.
|
||||
sqlDB, err := database.OpenSQLite(
|
||||
dbPath, database.SQLiteModeCreate,
|
||||
)
|
||||
|
||||
sqlDB, err := sql.Open("sqlite", dsn)
|
||||
require.NoError(t, err)
|
||||
|
||||
t.Cleanup(func() { _ = sqlDB.Close() })
|
||||
@@ -377,6 +377,17 @@ func TestProcessRetryTask_SuccessfulRetry(t *testing.T) {
|
||||
|
||||
bodyStr := event.Body
|
||||
cfg := iHTTPConfig(ts.URL)
|
||||
|
||||
// The target row exists because the engine confirms a scheduled
|
||||
// retry's target has not been deleted before it runs it. A retry
|
||||
// task whose target id names no row at all is a state the service
|
||||
// does not produce: the handler read that target to build the
|
||||
// task. See https://git.eeqj.de/sneak/webhooker/issues/107.
|
||||
iCreateTarget(
|
||||
t, s.MainDB, targetID, s.WebhookID, "retry-target",
|
||||
database.TargetTypeHTTP, cfg, 5,
|
||||
)
|
||||
|
||||
task := iTask(
|
||||
d, event, s.WebhookID, targetID,
|
||||
"retry-target", cfg, 5, 2, &bodyStr,
|
||||
@@ -456,6 +467,12 @@ func TestProcessRetryTask_LargeBody_FetchFromDB(
|
||||
)
|
||||
|
||||
cfg := iHTTPConfig(ts.URL)
|
||||
|
||||
iCreateTarget(
|
||||
t, s.MainDB, targetID, s.WebhookID, "retry-large",
|
||||
database.TargetTypeHTTP, cfg, 5,
|
||||
)
|
||||
|
||||
task := iTask(
|
||||
d, event, s.WebhookID, targetID,
|
||||
"retry-large", cfg, 5, 2, nil,
|
||||
@@ -558,6 +575,12 @@ func TestWorkerLifecycle_ProcessesRetryChannel(
|
||||
|
||||
bodyStr := event.Body
|
||||
cfg := iHTTPConfig(ts.URL)
|
||||
|
||||
iCreateTarget(
|
||||
t, s.MainDB, targetID, s.WebhookID, "retry-chan-test",
|
||||
database.TargetTypeHTTP, cfg, 5,
|
||||
)
|
||||
|
||||
task := iTask(
|
||||
d, event, s.WebhookID, targetID,
|
||||
"retry-chan-test", cfg, 5, 2, &bodyStr,
|
||||
|
||||
@@ -3,7 +3,6 @@ package delivery_test
|
||||
import (
|
||||
"bytes"
|
||||
"context"
|
||||
"database/sql"
|
||||
"encoding/json"
|
||||
"fmt"
|
||||
"log/slog"
|
||||
@@ -37,11 +36,12 @@ func testWebhookDB(t *testing.T) *gorm.DB {
|
||||
t.TempDir(), "events-test.db",
|
||||
)
|
||||
|
||||
dsn := fmt.Sprintf(
|
||||
"file:%s?cache=shared&mode=rwc", dbPath,
|
||||
// Opened the way the service opens a per-webhook database, so
|
||||
// these tests cannot pass against journal and locking settings
|
||||
// production does not use.
|
||||
sqlDB, err := database.OpenSQLite(
|
||||
dbPath, database.SQLiteModeCreate,
|
||||
)
|
||||
|
||||
sqlDB, err := sql.Open("sqlite", dsn)
|
||||
require.NoError(t, err)
|
||||
|
||||
t.Cleanup(func() { _ = sqlDB.Close() })
|
||||
|
||||
@@ -96,7 +96,15 @@ func TestEventDBHoldsNoTargetRows(t *testing.T) {
|
||||
)
|
||||
assertNoTargetRows(t, dbPath)
|
||||
|
||||
// A retry.
|
||||
// A retry. Its target exists in the main database, because the
|
||||
// engine confirms a scheduled retry's target has not been
|
||||
// deleted before running it; see
|
||||
// https://git.eeqj.de/sneak/webhooker/issues/107.
|
||||
iCreateTarget(
|
||||
t, s.MainDB, targetID, s.WebhookID, "leaky-target",
|
||||
database.TargetTypeHTTP, cfg, 5,
|
||||
)
|
||||
|
||||
rd := iSeedDelivery(
|
||||
t, s.WebhookDB, event.ID, targetID,
|
||||
database.DeliveryStatusRetrying,
|
||||
|
||||
442
internal/delivery/event_timestamp_test.go
Normal file
442
internal/delivery/event_timestamp_test.go
Normal file
@@ -0,0 +1,442 @@
|
||||
package delivery_test
|
||||
|
||||
import (
|
||||
"context"
|
||||
"encoding/json"
|
||||
"io"
|
||||
"net/http"
|
||||
"net/http/httptest"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"github.com/google/uuid"
|
||||
"github.com/stretchr/testify/assert"
|
||||
"github.com/stretchr/testify/require"
|
||||
"gorm.io/gorm"
|
||||
"sneak.berlin/go/webhooker/internal/database"
|
||||
"sneak.berlin/go/webhooker/internal/delivery"
|
||||
)
|
||||
|
||||
// tsEventCreatedAt is the receipt time seeded on the events these
|
||||
// tests deliver. It is far enough from both the zero time and from
|
||||
// now that neither can be mistaken for it.
|
||||
func tsEventCreatedAt() time.Time {
|
||||
return time.Date(
|
||||
2026, time.March, 4, 5, 6, 7, 0, time.UTC,
|
||||
)
|
||||
}
|
||||
|
||||
// tsZeroStamp is what a Slack message renders when the event handed
|
||||
// to FormatSlackMessage carries no CreatedAt.
|
||||
const tsZeroStamp = "*Timestamp:* `0001-01-01T00:00:00Z`"
|
||||
|
||||
// tsEventBody is the body seeded on every event in this file. It is
|
||||
// small enough that a Task can inline it.
|
||||
const tsEventBody = `{"hello":"world"}`
|
||||
|
||||
// tsUndeliverableHook stands in for a Slack incoming webhook on the
|
||||
// tests that never send: the config parser requires a URL, but no
|
||||
// request is made.
|
||||
const tsUndeliverableHook = "https://hooks.slack.com/services/T/B/x"
|
||||
|
||||
// tsSink is a stand-in Slack incoming webhook that records the raw
|
||||
// body posted to it.
|
||||
type tsSink struct {
|
||||
*httptest.Server
|
||||
|
||||
bodies chan []byte
|
||||
}
|
||||
|
||||
func newTSSink(t *testing.T) *tsSink {
|
||||
t.Helper()
|
||||
|
||||
s := &tsSink{bodies: make(chan []byte, 8)}
|
||||
|
||||
s.Server = httptest.NewServer(http.HandlerFunc(
|
||||
func(w http.ResponseWriter, r *http.Request) {
|
||||
body, _ := io.ReadAll(r.Body)
|
||||
|
||||
select {
|
||||
case s.bodies <- body:
|
||||
default:
|
||||
}
|
||||
|
||||
w.WriteHeader(http.StatusOK)
|
||||
},
|
||||
))
|
||||
|
||||
t.Cleanup(s.Close)
|
||||
|
||||
return s
|
||||
}
|
||||
|
||||
// text returns the Slack message text from the single payload the
|
||||
// sink received.
|
||||
func (s *tsSink) text(t *testing.T) string {
|
||||
t.Helper()
|
||||
|
||||
select {
|
||||
case raw := <-s.bodies:
|
||||
t.Logf("raw slack payload: %s", raw)
|
||||
|
||||
var payload struct {
|
||||
Text string `json:"text"`
|
||||
}
|
||||
|
||||
require.NoError(t, json.Unmarshal(raw, &payload))
|
||||
|
||||
return payload.Text
|
||||
case <-time.After(5 * time.Second):
|
||||
t.Fatal("slack sink received no payload")
|
||||
|
||||
return ""
|
||||
}
|
||||
}
|
||||
|
||||
func tsSlackConfig(t *testing.T, url string) string {
|
||||
t.Helper()
|
||||
|
||||
data, err := json.Marshal(
|
||||
delivery.SlackTargetConfig{WebhookURL: url},
|
||||
)
|
||||
require.NoError(t, err)
|
||||
|
||||
return string(data)
|
||||
}
|
||||
|
||||
// tsSeedEvent writes an event whose CreatedAt is tsEventCreatedAt
|
||||
// rather than the write time, so an assertion on the rendered
|
||||
// timestamp cannot pass by accident against "roughly now".
|
||||
func tsSeedEvent(
|
||||
t *testing.T, db *gorm.DB, webhookID string,
|
||||
) database.Event {
|
||||
t.Helper()
|
||||
|
||||
event := database.Event{
|
||||
WebhookID: webhookID,
|
||||
EntrypointID: uuid.New().String(),
|
||||
Method: http.MethodPost,
|
||||
Headers: `{}`,
|
||||
Body: tsEventBody,
|
||||
ContentType: "application/json",
|
||||
}
|
||||
event.ID = uuid.New().String()
|
||||
event.CreatedAt = tsEventCreatedAt()
|
||||
event.UpdatedAt = tsEventCreatedAt()
|
||||
|
||||
require.NoError(t, db.Create(&event).Error)
|
||||
|
||||
var stored database.Event
|
||||
|
||||
require.NoError(t,
|
||||
db.First(&stored, "id = ?", event.ID).Error,
|
||||
)
|
||||
require.Equal(t,
|
||||
tsEventCreatedAt().UTC(), stored.CreatedAt.UTC(),
|
||||
"seeded created_at did not round-trip",
|
||||
)
|
||||
|
||||
return event
|
||||
}
|
||||
|
||||
// tsSeedTarget writes the slack target row into the main database.
|
||||
// The retry path confirms the target still exists before sending.
|
||||
func tsSeedTarget(
|
||||
t *testing.T, mainDB *gorm.DB, webhookID, config string,
|
||||
) database.Target {
|
||||
t.Helper()
|
||||
|
||||
target := database.Target{
|
||||
WebhookID: webhookID,
|
||||
Name: "slack-sink",
|
||||
Type: database.TargetTypeSlack,
|
||||
Config: config,
|
||||
Active: true,
|
||||
}
|
||||
|
||||
require.NoError(t, mainDB.Create(&target).Error)
|
||||
|
||||
return target
|
||||
}
|
||||
|
||||
func tsTask(
|
||||
d database.Delivery,
|
||||
event database.Event,
|
||||
webhookID string,
|
||||
target database.Target,
|
||||
attemptNum int,
|
||||
body *string,
|
||||
) delivery.Task {
|
||||
return delivery.Task{
|
||||
DeliveryID: d.ID,
|
||||
EventID: event.ID,
|
||||
WebhookID: webhookID,
|
||||
EntrypointID: event.EntrypointID,
|
||||
TargetID: target.ID,
|
||||
TargetName: target.Name,
|
||||
TargetType: database.TargetTypeSlack,
|
||||
TargetConfig: target.Config,
|
||||
MaxRetries: 0,
|
||||
Method: event.Method,
|
||||
Headers: event.Headers,
|
||||
ContentType: event.ContentType,
|
||||
Body: body,
|
||||
AttemptNum: attemptNum,
|
||||
}
|
||||
}
|
||||
|
||||
func tsAssertRealTimestamp(t *testing.T, text string) {
|
||||
t.Helper()
|
||||
|
||||
assert.NotContains(t, text, tsZeroStamp,
|
||||
"slack message carries the zero timestamp",
|
||||
)
|
||||
assert.Contains(t, text,
|
||||
"*Timestamp:* `"+
|
||||
tsEventCreatedAt().UTC().Format(time.RFC3339)+"`",
|
||||
"slack message does not carry the event's receipt time",
|
||||
)
|
||||
}
|
||||
|
||||
// tsCase is one end-to-end delivery of a seeded event to a slack
|
||||
// sink, over whichever engine path `process` names.
|
||||
type tsCase struct {
|
||||
// status is the delivery row's status before the engine runs.
|
||||
// The retry path refuses a delivery that is not retrying.
|
||||
status database.DeliveryStatus
|
||||
|
||||
// inlineBody mirrors a Task built for a body under
|
||||
// MaxInlineBodySize. When false the engine reads the body back
|
||||
// from the stored row.
|
||||
inlineBody bool
|
||||
|
||||
attemptNum int
|
||||
|
||||
process func(
|
||||
ctx context.Context, e *delivery.Engine, task *delivery.Task,
|
||||
)
|
||||
}
|
||||
|
||||
// run delivers one event through the named path and returns the
|
||||
// Slack message text the sink received.
|
||||
func (c tsCase) run(t *testing.T) (iSetup, database.Delivery, string) {
|
||||
t.Helper()
|
||||
|
||||
s := newISetup(t)
|
||||
sink := newTSSink(t)
|
||||
|
||||
cfg := tsSlackConfig(t, sink.URL)
|
||||
target := tsSeedTarget(t, s.MainDB, s.WebhookID, cfg)
|
||||
event := tsSeedEvent(t, s.WebhookDB, s.WebhookID)
|
||||
|
||||
d := iSeedDelivery(
|
||||
t, s.WebhookDB, event.ID, target.ID, c.status,
|
||||
)
|
||||
|
||||
var body *string
|
||||
|
||||
if c.inlineBody {
|
||||
bodyStr := event.Body
|
||||
body = &bodyStr
|
||||
}
|
||||
|
||||
task := tsTask(
|
||||
d, event, s.WebhookID, target, c.attemptNum, body,
|
||||
)
|
||||
|
||||
c.process(context.TODO(), s.Engine, &task)
|
||||
|
||||
return s, d, sink.text(t)
|
||||
}
|
||||
|
||||
// TestSlackFirstAttemptCarriesEventTimestamp covers the path an
|
||||
// event takes on its first delivery: the task comes from the
|
||||
// receiver and the engine reconstructs the event from it.
|
||||
func TestSlackFirstAttemptCarriesEventTimestamp(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
s, d, text := tsCase{
|
||||
status: database.DeliveryStatusPending,
|
||||
inlineBody: true,
|
||||
attemptNum: 1,
|
||||
process: func(
|
||||
ctx context.Context,
|
||||
e *delivery.Engine,
|
||||
task *delivery.Task,
|
||||
) {
|
||||
e.ExportProcessNewTask(ctx, task)
|
||||
},
|
||||
}.run(t)
|
||||
|
||||
tsAssertRealTimestamp(t, text)
|
||||
|
||||
iAssertStatus(t, s.WebhookDB, d.ID,
|
||||
database.DeliveryStatusDelivered,
|
||||
)
|
||||
}
|
||||
|
||||
// TestSlackFirstAttemptLargeBodyCarriesEventTimestamp covers the
|
||||
// first-attempt path for an event whose body exceeded
|
||||
// MaxInlineBodySize, so the task carries no body and the engine
|
||||
// reads it back from the stored row.
|
||||
func TestSlackFirstAttemptLargeBodyCarriesEventTimestamp(
|
||||
t *testing.T,
|
||||
) {
|
||||
t.Parallel()
|
||||
|
||||
_, _, text := tsCase{
|
||||
status: database.DeliveryStatusPending,
|
||||
inlineBody: false,
|
||||
attemptNum: 1,
|
||||
process: func(
|
||||
ctx context.Context,
|
||||
e *delivery.Engine,
|
||||
task *delivery.Task,
|
||||
) {
|
||||
e.ExportProcessNewTask(ctx, task)
|
||||
},
|
||||
}.run(t)
|
||||
|
||||
tsAssertRealTimestamp(t, text)
|
||||
}
|
||||
|
||||
// TestSlackRetryCarriesEventTimestamp covers the retry path, which
|
||||
// reconstructs the event from the same task the first attempt used.
|
||||
func TestSlackRetryCarriesEventTimestamp(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
s, d, text := tsCase{
|
||||
status: database.DeliveryStatusRetrying,
|
||||
inlineBody: true,
|
||||
attemptNum: 2,
|
||||
process: func(
|
||||
ctx context.Context,
|
||||
e *delivery.Engine,
|
||||
task *delivery.Task,
|
||||
) {
|
||||
e.ExportProcessRetryTask(ctx, task)
|
||||
},
|
||||
}.run(t)
|
||||
|
||||
tsAssertRealTimestamp(t, text)
|
||||
|
||||
iAssertStatus(t, s.WebhookDB, d.ID,
|
||||
database.DeliveryStatusDelivered,
|
||||
)
|
||||
}
|
||||
|
||||
// TestFormatSlackMessageOverTaskReconstructedEvent asserts on the
|
||||
// formatted message directly, over the event the delivery paths
|
||||
// reconstruct from a Task. It is the unit-level guard under the
|
||||
// end-to-end tests: revert the CreatedAt population in hydrateEvent
|
||||
// and this fails on the zero timestamp.
|
||||
func TestFormatSlackMessageOverTaskReconstructedEvent(
|
||||
t *testing.T,
|
||||
) {
|
||||
t.Parallel()
|
||||
|
||||
s := newISetup(t)
|
||||
|
||||
cfg := tsSlackConfig(t, tsUndeliverableHook)
|
||||
target := tsSeedTarget(t, s.MainDB, s.WebhookID, cfg)
|
||||
event := tsSeedEvent(t, s.WebhookDB, s.WebhookID)
|
||||
|
||||
d := iSeedDelivery(
|
||||
t, s.WebhookDB, event.ID, target.ID,
|
||||
database.DeliveryStatusPending,
|
||||
)
|
||||
|
||||
bodyStr := event.Body
|
||||
task := tsTask(d, event, s.WebhookID, target, 1, &bodyStr)
|
||||
|
||||
rebuilt, err := s.Engine.ExportEventForTask(
|
||||
s.WebhookDB, &task,
|
||||
)
|
||||
require.NoError(t, err)
|
||||
assert.False(t, rebuilt.CreatedAt.IsZero(),
|
||||
"reconstructed event carries the zero time",
|
||||
)
|
||||
assert.Equal(t,
|
||||
tsEventCreatedAt().UTC(), rebuilt.CreatedAt.UTC(),
|
||||
)
|
||||
|
||||
tsAssertRealTimestamp(
|
||||
t, delivery.FormatSlackMessage(&rebuilt),
|
||||
)
|
||||
}
|
||||
|
||||
// TestFormatSlackMessageZeroTimestamp asserts the rendering choice
|
||||
// directly, without going through the engine: a zero CreatedAt (the
|
||||
// shape a reaped-row fallback produces) renders as "unknown" rather
|
||||
// than the year-1 zero time, while a real CreatedAt still renders as
|
||||
// RFC3339.
|
||||
func TestFormatSlackMessageZeroTimestamp(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
zeroEvent := database.Event{
|
||||
Method: http.MethodPost,
|
||||
ContentType: testContentType,
|
||||
Body: tsEventBody,
|
||||
}
|
||||
|
||||
zeroText := delivery.FormatSlackMessage(&zeroEvent)
|
||||
|
||||
assert.NotContains(t, zeroText, "0001-01-01",
|
||||
"slack message carries the zero-time year",
|
||||
)
|
||||
assert.Contains(t, zeroText, "*Timestamp:* `unknown`",
|
||||
"slack message does not mark an unset receipt time as unknown",
|
||||
)
|
||||
|
||||
nonZeroEvent := zeroEvent
|
||||
nonZeroEvent.CreatedAt = tsEventCreatedAt()
|
||||
|
||||
nonZeroText := delivery.FormatSlackMessage(&nonZeroEvent)
|
||||
|
||||
assert.Contains(t, nonZeroText,
|
||||
"*Timestamp:* `"+
|
||||
tsEventCreatedAt().UTC().Format(time.RFC3339)+"`",
|
||||
"slack message does not render a real receipt time as RFC3339",
|
||||
)
|
||||
}
|
||||
|
||||
// TestEventReconstructionSurvivesAReapedRow pins the fallback: an
|
||||
// event row reaped by retention while its delivery still holds the
|
||||
// body inline is still delivered, with the receipt time unset,
|
||||
// rather than dropped.
|
||||
func TestEventReconstructionSurvivesAReapedRow(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
s := newISetup(t)
|
||||
|
||||
cfg := tsSlackConfig(t, tsUndeliverableHook)
|
||||
target := tsSeedTarget(t, s.MainDB, s.WebhookID, cfg)
|
||||
event := tsSeedEvent(t, s.WebhookDB, s.WebhookID)
|
||||
|
||||
d := iSeedDelivery(
|
||||
t, s.WebhookDB, event.ID, target.ID,
|
||||
database.DeliveryStatusPending,
|
||||
)
|
||||
|
||||
bodyStr := event.Body
|
||||
task := tsTask(d, event, s.WebhookID, target, 1, &bodyStr)
|
||||
|
||||
require.NoError(t, s.WebhookDB.Unscoped().Delete(
|
||||
&database.Event{}, "id = ?", event.ID,
|
||||
).Error)
|
||||
|
||||
rebuilt, err := s.Engine.ExportEventForTask(
|
||||
s.WebhookDB, &task,
|
||||
)
|
||||
require.NoError(t, err)
|
||||
assert.Equal(t, bodyStr, rebuilt.Body)
|
||||
assert.True(t, rebuilt.CreatedAt.IsZero())
|
||||
|
||||
// A task with no inlined body has nothing left to deliver, so
|
||||
// the same reaped row is an error there.
|
||||
noBody := task
|
||||
noBody.Body = nil
|
||||
|
||||
_, err = s.Engine.ExportEventForTask(s.WebhookDB, &noBody)
|
||||
require.Error(t, err)
|
||||
}
|
||||
@@ -33,6 +33,11 @@ const (
|
||||
// response is written against this number, so a test has to
|
||||
// be able to name it.
|
||||
ExportMaxBodyLog = maxBodyLog
|
||||
|
||||
// ExportPendingSweepMinAge is how long a delivery must sit at
|
||||
// pending before the sweep treats it as stranded. A test has to
|
||||
// name it to age a row past the bound.
|
||||
ExportPendingSweepMinAge = pendingSweepMinAge
|
||||
)
|
||||
|
||||
// ExportIsBlockedIP exposes isBlockedIP for testing.
|
||||
@@ -146,6 +151,16 @@ func (e *Engine) ExportProcessRetryTask(
|
||||
e.processRetryTask(ctx, task)
|
||||
}
|
||||
|
||||
// ExportEventForTask exposes the event reconstruction the delivery
|
||||
// paths run: buildEventFromTask followed by hydrateEvent.
|
||||
func (e *Engine) ExportEventForTask(
|
||||
webhookDB *gorm.DB, task *Task,
|
||||
) (database.Event, error) {
|
||||
return e.hydrateEvent(
|
||||
webhookDB, buildEventFromTask(task), task,
|
||||
)
|
||||
}
|
||||
|
||||
// ExportProcessDelivery exposes processDelivery.
|
||||
func (e *Engine) ExportProcessDelivery(
|
||||
ctx context.Context,
|
||||
@@ -286,6 +301,26 @@ func (e *Engine) ExportWedgeWorker(release <-chan struct{}) {
|
||||
})
|
||||
}
|
||||
|
||||
// ExportInflightHeld reports how many deliveries the engine currently
|
||||
// owns, so a test can prove ownership is released rather than leaked.
|
||||
func (e *Engine) ExportInflightHeld() int {
|
||||
return e.inflight.held()
|
||||
}
|
||||
|
||||
// ExportRetainDelivery takes the first reference on a delivery, as the
|
||||
// queueing side does. It lets a test put a delivery into the state a
|
||||
// worker or a full channel would, without running the pool.
|
||||
func (e *Engine) ExportRetainDelivery(deliveryID string) bool {
|
||||
return e.inflight.retainIdle(deliveryID)
|
||||
}
|
||||
|
||||
// ExportRecoverRetryingDeliveries exposes recoverRetryingDeliveries.
|
||||
func (e *Engine) ExportRecoverRetryingDeliveries(
|
||||
webhookDB *gorm.DB, webhookID string,
|
||||
) {
|
||||
e.recoverRetryingDeliveries(webhookDB, webhookID)
|
||||
}
|
||||
|
||||
// ExportDeliveryCh returns the delivery channel.
|
||||
func (e *Engine) ExportDeliveryCh() chan Task {
|
||||
return e.deliveryCh
|
||||
|
||||
110
internal/delivery/inflight.go
Normal file
110
internal/delivery/inflight.go
Normal file
@@ -0,0 +1,110 @@
|
||||
package delivery
|
||||
|
||||
import "sync"
|
||||
|
||||
// inflightSet records which deliveries the engine currently owns.
|
||||
//
|
||||
// A delivery is owned from the moment a task for it is handed to a
|
||||
// channel or to a retry timer until the engine has no further plan for
|
||||
// it in memory. Restart recovery and both arms of the periodic sweep
|
||||
// re-dispatch only deliveries the set does not hold, which is what
|
||||
// makes them exact rather than a guess about how long a row has sat at
|
||||
// pending.
|
||||
//
|
||||
// This replaces reasoning from timestamps. A delivery's row says
|
||||
// pending from creation until its outcome is written, which covers
|
||||
// four different situations — never dispatched, waiting in a channel,
|
||||
// being attempted right now, and genuinely stranded — and no column
|
||||
// distinguishes them. Only the engine knows which, and it knows
|
||||
// exactly. `deliveryChannelSize` is 10000 against 10 workers, so a
|
||||
// perfectly healthy delivery can wait far longer than any age bound
|
||||
// worth setting before its attempt even begins; an age bound alone
|
||||
// re-sends it. See
|
||||
// https://git.eeqj.de/sneak/webhooker/issues/256.
|
||||
//
|
||||
// In-memory state is sufficient because a data directory admits one
|
||||
// process: internal/datadir takes an flock on it at startup and a
|
||||
// second instance refuses to run. Deliveries owned by a process that
|
||||
// died are not in any successor's set, and restart recovery is what
|
||||
// picks those up.
|
||||
//
|
||||
// References are counted rather than held as a plain set because
|
||||
// ownership outlives the worker that took it. A target that schedules
|
||||
// a retry from inside Deliver adds a reference while the worker still
|
||||
// holds one, so the delivery stays owned across the gap between the
|
||||
// worker returning and the timer firing — the window in which a sweep
|
||||
// would otherwise find the row at retrying and send it again.
|
||||
//
|
||||
// The zero value is ready to use, and the Engine holds one by value.
|
||||
// That is deliberate: an engine built by a constructor that forgot to
|
||||
// initialise this would not refuse to re-dispatch anything, and the
|
||||
// symptom would be duplicate deliveries rather than a failure anybody
|
||||
// notices.
|
||||
type inflightSet struct {
|
||||
mu sync.Mutex
|
||||
ids map[string]int
|
||||
}
|
||||
|
||||
// retain adds a reference to a delivery the caller already knows the
|
||||
// engine owns, so that ownership survives the current holder letting
|
||||
// go. It cannot fail.
|
||||
func (s *inflightSet) retain(deliveryID string) {
|
||||
s.mu.Lock()
|
||||
defer s.mu.Unlock()
|
||||
|
||||
if s.ids == nil {
|
||||
s.ids = make(map[string]int)
|
||||
}
|
||||
|
||||
s.ids[deliveryID]++
|
||||
}
|
||||
|
||||
// retainIdle takes the first reference on a delivery, and reports
|
||||
// whether it got it. It fails when the engine already owns the
|
||||
// delivery, which is what makes two claimants — restart recovery and
|
||||
// the sweep run concurrently, or two sweep arms — mutually exclusive
|
||||
// rather than merely atomic.
|
||||
func (s *inflightSet) retainIdle(deliveryID string) bool {
|
||||
s.mu.Lock()
|
||||
defer s.mu.Unlock()
|
||||
|
||||
if s.ids[deliveryID] > 0 {
|
||||
return false
|
||||
}
|
||||
|
||||
if s.ids == nil {
|
||||
s.ids = make(map[string]int)
|
||||
}
|
||||
|
||||
s.ids[deliveryID] = 1
|
||||
|
||||
return true
|
||||
}
|
||||
|
||||
// release drops one reference. The delivery becomes eligible for
|
||||
// re-dispatch again once the last one goes.
|
||||
func (s *inflightSet) release(deliveryID string) {
|
||||
s.mu.Lock()
|
||||
defer s.mu.Unlock()
|
||||
|
||||
n := s.ids[deliveryID] - 1
|
||||
if n <= 0 {
|
||||
delete(s.ids, deliveryID)
|
||||
|
||||
return
|
||||
}
|
||||
|
||||
s.ids[deliveryID] = n
|
||||
}
|
||||
|
||||
// held reports how many deliveries the engine currently owns. It
|
||||
// exists so a test can assert that ownership is released rather than
|
||||
// leaked: a reference that is never dropped hides its delivery from
|
||||
// every sweep for the life of the process, which is the one way this
|
||||
// mechanism can fail silently.
|
||||
func (s *inflightSet) held() int {
|
||||
s.mu.Lock()
|
||||
defer s.mu.Unlock()
|
||||
|
||||
return len(s.ids)
|
||||
}
|
||||
428
internal/delivery/inflight_test.go
Normal file
428
internal/delivery/inflight_test.go
Normal file
@@ -0,0 +1,428 @@
|
||||
package delivery_test
|
||||
|
||||
import (
|
||||
"context"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"sync"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"github.com/google/uuid"
|
||||
"github.com/stretchr/testify/assert"
|
||||
"github.com/stretchr/testify/require"
|
||||
"sneak.berlin/go/webhooker/internal/database"
|
||||
"sneak.berlin/go/webhooker/internal/delivery"
|
||||
)
|
||||
|
||||
// These tests pin the rule that decides whether a delivery may be
|
||||
// handed back to a worker: the engine re-dispatches only what it does
|
||||
// not already own. Age alone is not that rule — a healthy delivery
|
||||
// waiting in a 10000-deep channel is old and must not be re-sent. See
|
||||
// https://git.eeqj.de/sneak/webhooker/issues/256.
|
||||
|
||||
// fSweepSetup seeds the main database with the webhook row the sweep
|
||||
// enumerates, and returns the setup.
|
||||
func fSweepSetup(
|
||||
t *testing.T, targetID, name string,
|
||||
) iSetup {
|
||||
t.Helper()
|
||||
|
||||
s := newISetup(t)
|
||||
|
||||
iCreateTarget(t, s.MainDB, targetID,
|
||||
s.WebhookID, name,
|
||||
database.TargetTypeLog, "", 0,
|
||||
)
|
||||
|
||||
require.NoError(t, s.MainDB.Create(&database.Webhook{
|
||||
BaseModel: database.BaseModel{ID: s.WebhookID},
|
||||
UserID: uuid.New().String(),
|
||||
Name: name,
|
||||
}).Error)
|
||||
|
||||
return s
|
||||
}
|
||||
|
||||
// fDrain collects every task the engine has queued.
|
||||
//
|
||||
// Every caller drives the dispatch paths synchronously and has already
|
||||
// waited for them to return, so anything they queued is in the channel
|
||||
// by now. The short grace covers nothing but scheduler jitter, and is
|
||||
// kept small because one of these tests runs the drain forty times.
|
||||
func fDrain(e *delivery.Engine) []delivery.Task {
|
||||
var out []delivery.Task
|
||||
|
||||
for {
|
||||
select {
|
||||
case task := <-e.ExportDeliveryCh():
|
||||
out = append(out, task)
|
||||
case task := <-e.ExportRetryCh():
|
||||
out = append(out, task)
|
||||
case <-time.After(25 * time.Millisecond):
|
||||
return out
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestArchiveHandleIsWAL closes the last gap in the durability
|
||||
// evidence: the main and per-webhook tiers each assert their journal
|
||||
// mode on a live handle, and the archive tier gets its settings from
|
||||
// the same code path but nothing checked the running file.
|
||||
func TestArchiveHandleIsWAL(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
w := delivery.NewExportArchiveWriter(
|
||||
filepath.Join(t.TempDir(), "archive-wal.db"),
|
||||
archiveTestLogger(), 0,
|
||||
)
|
||||
|
||||
require.NoError(t, w.Open(0))
|
||||
|
||||
var mode string
|
||||
|
||||
row := w.DB().Raw("pragma journal_mode").Row()
|
||||
require.NoError(t, row.Scan(&mode))
|
||||
assert.Equal(t, "wal", strings.ToLower(mode))
|
||||
|
||||
var busy string
|
||||
|
||||
row = w.DB().Raw("pragma busy_timeout").Row()
|
||||
require.NoError(t, row.Scan(&busy))
|
||||
assert.Equal(t, "10000", busy)
|
||||
}
|
||||
|
||||
// TestSweepLeavesAQueuedDeliveryAlone is the case the age bound cannot
|
||||
// see. The delivery is queued and untouched, so its row is arbitrarily
|
||||
// old and still perfectly healthy; only ownership distinguishes it
|
||||
// from a stranded one.
|
||||
func TestSweepLeavesAQueuedDeliveryAlone(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
targetID := uuid.New().String()
|
||||
s := fSweepSetup(t, targetID, "queued")
|
||||
|
||||
event := iSeedEvent(
|
||||
t, s.WebhookDB, s.WebhookID, `{"queued":true}`,
|
||||
)
|
||||
|
||||
d := iSeedDelivery(
|
||||
t, s.WebhookDB, event.ID, targetID,
|
||||
database.DeliveryStatusPending,
|
||||
)
|
||||
rAgePending(t, s.WebhookDB, d.ID)
|
||||
|
||||
// Queued exactly as the receiver queues it, and never dequeued:
|
||||
// no workers are running in this engine.
|
||||
s.Engine.Notify([]delivery.Task{{
|
||||
DeliveryID: d.ID,
|
||||
EventID: event.ID,
|
||||
WebhookID: s.WebhookID,
|
||||
TargetID: targetID,
|
||||
}})
|
||||
|
||||
require.Equal(t, 1, s.Engine.ExportInflightHeld())
|
||||
|
||||
s.Engine.ExportSweepWebhookRetries(
|
||||
context.Background(), s.WebhookID,
|
||||
)
|
||||
|
||||
tasks := fDrain(s.Engine)
|
||||
assert.Len(
|
||||
t, tasks, 1,
|
||||
"the sweep must not queue a delivery that is "+
|
||||
"already waiting for a worker",
|
||||
)
|
||||
}
|
||||
|
||||
// TestRecoveryAndSweepDoNotDoubleDispatch drives the two entry points
|
||||
// the engine starts concurrently against one aged pending row. Before
|
||||
// ownership they both dispatched it.
|
||||
func TestRecoveryAndSweepDoNotDoubleDispatch(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
targetID := uuid.New().String()
|
||||
s := fSweepSetup(t, targetID, "racing")
|
||||
|
||||
event := iSeedEvent(
|
||||
t, s.WebhookDB, s.WebhookID, `{"racing":true}`,
|
||||
)
|
||||
|
||||
ctx := context.Background()
|
||||
|
||||
for range 40 {
|
||||
d := iSeedDelivery(
|
||||
t, s.WebhookDB, event.ID, targetID,
|
||||
database.DeliveryStatusPending,
|
||||
)
|
||||
rAgePending(t, s.WebhookDB, d.ID)
|
||||
|
||||
var wg sync.WaitGroup
|
||||
|
||||
wg.Go(func() {
|
||||
s.Engine.ExportRecoverPendingDeliveries(
|
||||
ctx, s.WebhookDB, s.WebhookID,
|
||||
)
|
||||
})
|
||||
wg.Go(func() {
|
||||
s.Engine.ExportSweepWebhookRetries(
|
||||
ctx, s.WebhookID,
|
||||
)
|
||||
})
|
||||
wg.Wait()
|
||||
|
||||
tasks := fDrain(s.Engine)
|
||||
require.Len(
|
||||
t, tasks, 1,
|
||||
"delivery %s dispatched %d times",
|
||||
d.ID, len(tasks),
|
||||
)
|
||||
|
||||
// No worker runs in this engine, so the reference the winner
|
||||
// took is never released and earlier iterations' deliveries
|
||||
// stay owned — which is itself the property under test, since
|
||||
// both paths see them on every subsequent pass.
|
||||
}
|
||||
}
|
||||
|
||||
// TestConcurrentClaimsOfOneDeliveryYieldOneOwner exercises the
|
||||
// exclusion directly, rather than arguing it from a SQL predicate.
|
||||
func TestConcurrentClaimsOfOneDeliveryYieldOneOwner(
|
||||
t *testing.T,
|
||||
) {
|
||||
t.Parallel()
|
||||
|
||||
eng := newISetup(t).Engine
|
||||
deliveryID := uuid.New().String()
|
||||
|
||||
var (
|
||||
wg sync.WaitGroup
|
||||
mu sync.Mutex
|
||||
won int
|
||||
)
|
||||
|
||||
for range 64 {
|
||||
wg.Go(func() {
|
||||
if eng.ExportRetainDelivery(deliveryID) {
|
||||
mu.Lock()
|
||||
won++
|
||||
mu.Unlock()
|
||||
}
|
||||
})
|
||||
}
|
||||
|
||||
wg.Wait()
|
||||
|
||||
assert.Equal(t, 1, won)
|
||||
assert.Equal(t, 1, eng.ExportInflightHeld())
|
||||
}
|
||||
|
||||
// TestOwnershipIsReleasedAfterDelivery guards the other direction: a
|
||||
// leaked reference hides a delivery from every sweep for the life of
|
||||
// the process.
|
||||
func TestOwnershipIsReleasedAfterDelivery(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
s := newISetup(t)
|
||||
targetID := uuid.New().String()
|
||||
|
||||
iCreateTarget(t, s.MainDB, targetID,
|
||||
s.WebhookID, "released",
|
||||
database.TargetTypeLog, "", 0,
|
||||
)
|
||||
|
||||
event := iSeedEvent(
|
||||
t, s.WebhookDB, s.WebhookID, `{"released":true}`,
|
||||
)
|
||||
|
||||
d := iSeedDelivery(
|
||||
t, s.WebhookDB, event.ID, targetID,
|
||||
database.DeliveryStatusPending,
|
||||
)
|
||||
|
||||
s.Engine.ExportStart()
|
||||
|
||||
defer func() {
|
||||
require.NoError(
|
||||
t, s.Engine.ExportStop(context.Background()),
|
||||
)
|
||||
}()
|
||||
|
||||
body := `{"released":true}`
|
||||
|
||||
s.Engine.Notify([]delivery.Task{{
|
||||
DeliveryID: d.ID,
|
||||
EventID: event.ID,
|
||||
WebhookID: s.WebhookID,
|
||||
TargetID: targetID,
|
||||
TargetName: "released",
|
||||
TargetType: database.TargetTypeLog,
|
||||
Body: &body,
|
||||
EntrypointID: event.EntrypointID,
|
||||
}})
|
||||
|
||||
iWaitForDelivered(t, s.WebhookDB, d.ID)
|
||||
|
||||
assert.Eventually(
|
||||
t,
|
||||
func() bool {
|
||||
return s.Engine.ExportInflightHeld() == 0
|
||||
},
|
||||
2*time.Second, 20*time.Millisecond,
|
||||
"the delivery stayed owned after it was delivered",
|
||||
)
|
||||
}
|
||||
|
||||
// TestRetryingRecoverySkipsASuccessfulResult is the retrying-side twin
|
||||
// of the pending reconcile. A second attempt that reached the receiver
|
||||
// and whose status write then failed sits at retrying holding a
|
||||
// successful result, and re-sending it is the same duplicate.
|
||||
func TestRetryingRecoverySkipsASuccessfulResult(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
targetID := uuid.New().String()
|
||||
s := fSweepSetup(t, targetID, "retry-settled")
|
||||
|
||||
event := iSeedEvent(
|
||||
t, s.WebhookDB, s.WebhookID, `{"retry":true}`,
|
||||
)
|
||||
|
||||
d := iSeedDelivery(
|
||||
t, s.WebhookDB, event.ID, targetID,
|
||||
database.DeliveryStatusRetrying,
|
||||
)
|
||||
rSeedResult(t, s.WebhookDB, d.ID, 1, false)
|
||||
rSeedResult(t, s.WebhookDB, d.ID, 2, true)
|
||||
|
||||
s.Engine.ExportRecoverRetryingDeliveries(
|
||||
s.WebhookDB, s.WebhookID,
|
||||
)
|
||||
|
||||
assert.Empty(
|
||||
t, fDrain(s.Engine),
|
||||
"a retrying delivery holding a successful result "+
|
||||
"must not be sent again",
|
||||
)
|
||||
|
||||
iAssertStatus(
|
||||
t, s.WebhookDB, d.ID,
|
||||
database.DeliveryStatusDelivered,
|
||||
)
|
||||
}
|
||||
|
||||
// TestRetryingSweepSkipsASuccessfulResult is the same rule on the
|
||||
// periodic sweep's retrying arm.
|
||||
func TestRetryingSweepSkipsASuccessfulResult(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
targetID := uuid.New().String()
|
||||
s := fSweepSetup(t, targetID, "retry-swept")
|
||||
|
||||
event := iSeedEvent(
|
||||
t, s.WebhookDB, s.WebhookID, `{"swept":true}`,
|
||||
)
|
||||
|
||||
d := iSeedDelivery(
|
||||
t, s.WebhookDB, event.ID, targetID,
|
||||
database.DeliveryStatusRetrying,
|
||||
)
|
||||
rSeedResult(t, s.WebhookDB, d.ID, 1, false)
|
||||
rSeedResult(t, s.WebhookDB, d.ID, 2, true)
|
||||
|
||||
s.Engine.ExportSweepWebhookRetries(
|
||||
context.Background(), s.WebhookID,
|
||||
)
|
||||
|
||||
assert.Empty(t, fDrain(s.Engine))
|
||||
|
||||
iAssertStatus(
|
||||
t, s.WebhookDB, d.ID,
|
||||
database.DeliveryStatusDelivered,
|
||||
)
|
||||
|
||||
var attempts int64
|
||||
|
||||
require.NoError(t, s.WebhookDB.
|
||||
Model(&database.DeliveryResult{}).
|
||||
Where("delivery_id = ?", d.ID).
|
||||
Count(&attempts).Error)
|
||||
assert.Equal(
|
||||
t, int64(2), attempts,
|
||||
"settling must not invent an attempt",
|
||||
)
|
||||
}
|
||||
|
||||
// TestScheduledRetryIsNotSweptDuringBackoff closes the window between
|
||||
// a target scheduling a retry and the timer firing. The row says
|
||||
// retrying and nothing is running, which is exactly what an orphaned
|
||||
// retry looks like from the database.
|
||||
func TestScheduledRetryIsNotSweptDuringBackoff(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
targetID := uuid.New().String()
|
||||
s := fSweepSetup(t, targetID, "backoff")
|
||||
|
||||
event := iSeedEvent(
|
||||
t, s.WebhookDB, s.WebhookID, `{"backoff":true}`,
|
||||
)
|
||||
|
||||
d := iSeedDelivery(
|
||||
t, s.WebhookDB, event.ID, targetID,
|
||||
database.DeliveryStatusRetrying,
|
||||
)
|
||||
|
||||
s.Engine.ExportScheduleRetry(delivery.Task{
|
||||
DeliveryID: d.ID,
|
||||
EventID: event.ID,
|
||||
WebhookID: s.WebhookID,
|
||||
TargetID: targetID,
|
||||
AttemptNum: 2,
|
||||
}, time.Hour)
|
||||
|
||||
require.Equal(t, 1, s.Engine.ExportInflightHeld())
|
||||
|
||||
s.Engine.ExportSweepWebhookRetries(
|
||||
context.Background(), s.WebhookID,
|
||||
)
|
||||
|
||||
assert.Empty(
|
||||
t, fDrain(s.Engine),
|
||||
"the sweep must not duplicate a retry that is "+
|
||||
"already scheduled",
|
||||
)
|
||||
}
|
||||
|
||||
// TestRedispatchStampsTheRow pins the cadence control: a stranded
|
||||
// delivery that has just been handed out is not selected again by the
|
||||
// next tick a minute later.
|
||||
func TestRedispatchStampsTheRow(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
targetID := uuid.New().String()
|
||||
s := fSweepSetup(t, targetID, "stamped")
|
||||
|
||||
event := iSeedEvent(
|
||||
t, s.WebhookDB, s.WebhookID, `{"stamped":true}`,
|
||||
)
|
||||
|
||||
d := iSeedDelivery(
|
||||
t, s.WebhookDB, event.ID, targetID,
|
||||
database.DeliveryStatusPending,
|
||||
)
|
||||
rAgePending(t, s.WebhookDB, d.ID)
|
||||
|
||||
ctx := context.Background()
|
||||
|
||||
s.Engine.ExportSweepWebhookRetries(ctx, s.WebhookID)
|
||||
require.Len(t, fDrain(s.Engine), 1)
|
||||
|
||||
var row database.Delivery
|
||||
|
||||
require.NoError(t, s.WebhookDB.
|
||||
First(&row, "id = ?", d.ID).Error)
|
||||
assert.WithinDuration(
|
||||
t, time.Now(), row.UpdatedAt, time.Minute,
|
||||
"a re-dispatched delivery must be stamped so the "+
|
||||
"next tick does not select it again",
|
||||
)
|
||||
}
|
||||
@@ -229,6 +229,13 @@ func mExhaustRetries(t *testing.T, s iSetup) {
|
||||
body := event.Body
|
||||
cfg := iHTTPConfig(ts.URL)
|
||||
|
||||
// The retry below is only run if its target still exists; see
|
||||
// https://git.eeqj.de/sneak/webhooker/issues/107.
|
||||
iCreateTarget(
|
||||
t, s.MainDB, targetID, s.WebhookID, "metrics-fail",
|
||||
database.TargetTypeHTTP, cfg, 2,
|
||||
)
|
||||
|
||||
first := iTask(
|
||||
d, event, s.WebhookID, targetID,
|
||||
"metrics-fail", cfg, 2, 1, &body,
|
||||
@@ -289,6 +296,13 @@ func TestDeliveryMetrics_CircuitBreakerGauge(t *testing.T) {
|
||||
// rather than the budget is what stops the delivery.
|
||||
maxRetries := delivery.ExportDefaultFailureThreshold + 5
|
||||
|
||||
// The retries below are only run if their target still exists;
|
||||
// see https://git.eeqj.de/sneak/webhooker/issues/107.
|
||||
iCreateTarget(
|
||||
t, s.MainDB, targetID, s.WebhookID, "metrics-trip",
|
||||
database.TargetTypeHTTP, cfg, maxRetries,
|
||||
)
|
||||
|
||||
first := iTask(
|
||||
d, event, s.WebhookID, targetID,
|
||||
"metrics-trip", cfg, maxRetries, 1, &body,
|
||||
@@ -353,6 +367,11 @@ func TestDeliveryMetrics_BreakerBlockedIsNotAnAttempt(
|
||||
cfg := iHTTPConfig(ts.URL)
|
||||
maxRetries := delivery.ExportDefaultFailureThreshold + 5
|
||||
|
||||
iCreateTarget(
|
||||
t, s.MainDB, targetID, s.WebhookID, "metrics-blocked",
|
||||
database.TargetTypeHTTP, cfg, maxRetries,
|
||||
)
|
||||
|
||||
first := iTask(
|
||||
d, event, s.WebhookID, targetID,
|
||||
"metrics-blocked", cfg, maxRetries, 1, &body,
|
||||
|
||||
@@ -3,8 +3,6 @@ package delivery_test
|
||||
import (
|
||||
"bytes"
|
||||
"context"
|
||||
"database/sql"
|
||||
"fmt"
|
||||
"log/slog"
|
||||
"net/http"
|
||||
"path/filepath"
|
||||
@@ -54,12 +52,10 @@ func (q *qdSyncBuf) String() string {
|
||||
func qdMainDB(t *testing.T, log *slog.Logger) *gorm.DB {
|
||||
t.Helper()
|
||||
|
||||
dsn := fmt.Sprintf(
|
||||
"file:%s?cache=shared&mode=rwc",
|
||||
sqlDB, err := database.OpenSQLite(
|
||||
filepath.Join(t.TempDir(), "main-gormlog.db"),
|
||||
database.SQLiteModeCreate,
|
||||
)
|
||||
|
||||
sqlDB, err := sql.Open("sqlite", dsn)
|
||||
require.NoError(t, err)
|
||||
|
||||
t.Cleanup(func() { _ = sqlDB.Close() })
|
||||
|
||||
378
internal/delivery/recovery_durability_test.go
Normal file
378
internal/delivery/recovery_durability_test.go
Normal file
@@ -0,0 +1,378 @@
|
||||
package delivery_test
|
||||
|
||||
import (
|
||||
"context"
|
||||
"net/http"
|
||||
"net/http/httptest"
|
||||
"sync/atomic"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"github.com/google/uuid"
|
||||
"github.com/stretchr/testify/assert"
|
||||
"github.com/stretchr/testify/require"
|
||||
"gorm.io/gorm"
|
||||
"sneak.berlin/go/webhooker/internal/database"
|
||||
"sneak.berlin/go/webhooker/internal/delivery"
|
||||
)
|
||||
|
||||
// These tests cover the delivery half of
|
||||
// https://git.eeqj.de/sneak/webhooker/issues/256: a delivery that
|
||||
// reached its receiver but whose bookkeeping write failed used to be
|
||||
// left at pending and re-sent on the next restart, giving the receiver
|
||||
// a second copy while the event log recorded one attempt.
|
||||
|
||||
// rSeedResult records a DeliveryResult against a delivery, standing in
|
||||
// for the attempt row the send path writes before the status.
|
||||
func rSeedResult(
|
||||
t *testing.T,
|
||||
db *gorm.DB,
|
||||
deliveryID string,
|
||||
attemptNum int,
|
||||
success bool,
|
||||
) {
|
||||
t.Helper()
|
||||
|
||||
require.NoError(t, db.Create(&database.DeliveryResult{
|
||||
DeliveryID: deliveryID,
|
||||
AttemptNum: attemptNum,
|
||||
Success: success,
|
||||
}).Error)
|
||||
}
|
||||
|
||||
// rAgePending backdates a delivery past the sweep's age bound, which is
|
||||
// what separates a stranded delivery from one a worker still holds.
|
||||
func rAgePending(
|
||||
t *testing.T, db *gorm.DB, deliveryID string,
|
||||
) {
|
||||
t.Helper()
|
||||
|
||||
old := time.Now().Add(
|
||||
-2 * delivery.ExportPendingSweepMinAge,
|
||||
)
|
||||
|
||||
require.NoError(t, db.Model(&database.Delivery{}).
|
||||
Where("id = ?", deliveryID).
|
||||
UpdateColumn("updated_at", old).Error)
|
||||
}
|
||||
|
||||
func TestRecoverySkipsPendingWithSuccessfulResult(
|
||||
t *testing.T,
|
||||
) {
|
||||
t.Parallel()
|
||||
|
||||
s := newISetup(t)
|
||||
targetID := uuid.New().String()
|
||||
|
||||
iCreateTarget(t, s.MainDB, targetID,
|
||||
s.WebhookID, "already-delivered",
|
||||
database.TargetTypeLog, "", 0,
|
||||
)
|
||||
|
||||
event := iSeedEvent(
|
||||
t, s.WebhookDB, s.WebhookID, `{"delivered":true}`,
|
||||
)
|
||||
|
||||
// The delivery whose send succeeded and whose result row landed:
|
||||
// only the status write failed, so it sits at pending.
|
||||
done := iSeedDelivery(
|
||||
t, s.WebhookDB, event.ID, targetID,
|
||||
database.DeliveryStatusPending,
|
||||
)
|
||||
rSeedResult(t, s.WebhookDB, done.ID, 1, true)
|
||||
|
||||
// A delivery that was genuinely never attempted.
|
||||
fresh := iSeedDelivery(
|
||||
t, s.WebhookDB, event.ID, targetID,
|
||||
database.DeliveryStatusPending,
|
||||
)
|
||||
|
||||
s.Engine.ExportRecoverPendingDeliveries(
|
||||
context.Background(), s.WebhookDB, s.WebhookID,
|
||||
)
|
||||
|
||||
select {
|
||||
case task := <-s.Engine.ExportDeliveryCh():
|
||||
assert.Equal(
|
||||
t, fresh.ID, task.DeliveryID,
|
||||
"only the unattempted delivery may be re-sent",
|
||||
)
|
||||
case <-time.After(2 * time.Second):
|
||||
t.Fatal("expected the unattempted delivery")
|
||||
}
|
||||
|
||||
select {
|
||||
case task := <-s.Engine.ExportDeliveryCh():
|
||||
t.Fatalf(
|
||||
"re-sent an already delivered delivery: %s",
|
||||
task.DeliveryID,
|
||||
)
|
||||
case <-time.After(200 * time.Millisecond):
|
||||
}
|
||||
|
||||
// It is settled rather than merely skipped: leaving it pending
|
||||
// would strand it again on the next sweep.
|
||||
iAssertStatus(
|
||||
t, s.WebhookDB, done.ID,
|
||||
database.DeliveryStatusDelivered,
|
||||
)
|
||||
}
|
||||
|
||||
// TestRecoveryContinuesTheAttemptNumbering pins the audit trail: a
|
||||
// recovered delivery that already recorded two attempts is re-sent as
|
||||
// attempt three, not as attempt one again.
|
||||
func TestRecoveryContinuesTheAttemptNumbering(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
s := newISetup(t)
|
||||
targetID := uuid.New().String()
|
||||
|
||||
iCreateTarget(t, s.MainDB, targetID,
|
||||
s.WebhookID, "numbering",
|
||||
database.TargetTypeLog, "", 0,
|
||||
)
|
||||
|
||||
event := iSeedEvent(
|
||||
t, s.WebhookDB, s.WebhookID, `{"numbering":true}`,
|
||||
)
|
||||
|
||||
d := iSeedDelivery(
|
||||
t, s.WebhookDB, event.ID, targetID,
|
||||
database.DeliveryStatusPending,
|
||||
)
|
||||
|
||||
rSeedResult(t, s.WebhookDB, d.ID, 1, false)
|
||||
rSeedResult(t, s.WebhookDB, d.ID, 2, false)
|
||||
|
||||
s.Engine.ExportRecoverPendingDeliveries(
|
||||
context.Background(), s.WebhookDB, s.WebhookID,
|
||||
)
|
||||
|
||||
select {
|
||||
case task := <-s.Engine.ExportDeliveryCh():
|
||||
assert.Equal(t, d.ID, task.DeliveryID)
|
||||
assert.Equal(t, 3, task.AttemptNum)
|
||||
case <-time.After(2 * time.Second):
|
||||
t.Fatal("expected the delivery to be recovered")
|
||||
}
|
||||
}
|
||||
|
||||
// TestSweepRecoversStrandedPending is the half that removes the
|
||||
// restart requirement: a delivery left at pending is picked up by the
|
||||
// periodic sweep.
|
||||
func TestSweepRecoversStrandedPending(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
targetID := uuid.New().String()
|
||||
s := fSweepSetup(t, targetID, "stranded")
|
||||
|
||||
event := iSeedEvent(
|
||||
t, s.WebhookDB, s.WebhookID, `{"stranded":true}`,
|
||||
)
|
||||
|
||||
stranded := iSeedDelivery(
|
||||
t, s.WebhookDB, event.ID, targetID,
|
||||
database.DeliveryStatusPending,
|
||||
)
|
||||
rAgePending(t, s.WebhookDB, stranded.ID)
|
||||
|
||||
// A delivery a worker may still be holding: young, and therefore
|
||||
// none of the sweep's business.
|
||||
inFlight := iSeedDelivery(
|
||||
t, s.WebhookDB, event.ID, targetID,
|
||||
database.DeliveryStatusPending,
|
||||
)
|
||||
|
||||
s.Engine.ExportSweepWebhookRetries(
|
||||
context.Background(), s.WebhookID,
|
||||
)
|
||||
|
||||
select {
|
||||
case task := <-s.Engine.ExportDeliveryCh():
|
||||
assert.Equal(t, stranded.ID, task.DeliveryID)
|
||||
case <-time.After(2 * time.Second):
|
||||
t.Fatal("expected the stranded delivery")
|
||||
}
|
||||
|
||||
select {
|
||||
case task := <-s.Engine.ExportDeliveryCh():
|
||||
t.Fatalf(
|
||||
"swept an in-flight delivery: %s",
|
||||
task.DeliveryID,
|
||||
)
|
||||
case <-time.After(200 * time.Millisecond):
|
||||
}
|
||||
|
||||
iAssertStatus(
|
||||
t, s.WebhookDB, inFlight.ID,
|
||||
database.DeliveryStatusPending,
|
||||
)
|
||||
}
|
||||
|
||||
// TestSweepClaimsAStrandedDeliveryOnlyOnce guards the repeat the sweep
|
||||
// would otherwise be: the row stays pending for as long as the attempt
|
||||
// runs, and a sweep a minute later must not send it a second time.
|
||||
func TestSweepClaimsAStrandedDeliveryOnlyOnce(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
targetID := uuid.New().String()
|
||||
s := fSweepSetup(t, targetID, "claimed")
|
||||
|
||||
event := iSeedEvent(
|
||||
t, s.WebhookDB, s.WebhookID, `{"claimed":true}`,
|
||||
)
|
||||
|
||||
d := iSeedDelivery(
|
||||
t, s.WebhookDB, event.ID, targetID,
|
||||
database.DeliveryStatusPending,
|
||||
)
|
||||
rAgePending(t, s.WebhookDB, d.ID)
|
||||
|
||||
ctx := context.Background()
|
||||
|
||||
s.Engine.ExportSweepWebhookRetries(ctx, s.WebhookID)
|
||||
|
||||
select {
|
||||
case task := <-s.Engine.ExportDeliveryCh():
|
||||
assert.Equal(t, d.ID, task.DeliveryID)
|
||||
case <-time.After(2 * time.Second):
|
||||
t.Fatal("expected the stranded delivery")
|
||||
}
|
||||
|
||||
// The delivery is still pending — nothing has run it yet — but
|
||||
// the claim must keep the next sweep off it.
|
||||
iAssertStatus(
|
||||
t, s.WebhookDB, d.ID,
|
||||
database.DeliveryStatusPending,
|
||||
)
|
||||
|
||||
s.Engine.ExportSweepWebhookRetries(ctx, s.WebhookID)
|
||||
|
||||
select {
|
||||
case task := <-s.Engine.ExportDeliveryCh():
|
||||
t.Fatalf(
|
||||
"sent a claimed delivery again: %s",
|
||||
task.DeliveryID,
|
||||
)
|
||||
case <-time.After(200 * time.Millisecond):
|
||||
}
|
||||
}
|
||||
|
||||
// TestSweepSettlesStrandedPendingWithoutResending is the sweep's own
|
||||
// version of the reconcile: a stranded delivery holding a successful
|
||||
// result is settled where it stands, and the receiver hears nothing.
|
||||
func TestSweepSettlesStrandedPendingWithoutResending(
|
||||
t *testing.T,
|
||||
) {
|
||||
t.Parallel()
|
||||
|
||||
targetID := uuid.New().String()
|
||||
s := fSweepSetup(t, targetID, "settled")
|
||||
|
||||
event := iSeedEvent(
|
||||
t, s.WebhookDB, s.WebhookID, `{"settled":true}`,
|
||||
)
|
||||
|
||||
d := iSeedDelivery(
|
||||
t, s.WebhookDB, event.ID, targetID,
|
||||
database.DeliveryStatusPending,
|
||||
)
|
||||
rSeedResult(t, s.WebhookDB, d.ID, 1, true)
|
||||
rAgePending(t, s.WebhookDB, d.ID)
|
||||
|
||||
s.Engine.ExportSweepWebhookRetries(
|
||||
context.Background(), s.WebhookID,
|
||||
)
|
||||
|
||||
select {
|
||||
case task := <-s.Engine.ExportDeliveryCh():
|
||||
t.Fatalf(
|
||||
"re-sent a delivery that already succeeded: %s",
|
||||
task.DeliveryID,
|
||||
)
|
||||
case <-time.After(200 * time.Millisecond):
|
||||
}
|
||||
|
||||
iAssertStatus(
|
||||
t, s.WebhookDB, d.ID,
|
||||
database.DeliveryStatusDelivered,
|
||||
)
|
||||
|
||||
var attempts int64
|
||||
|
||||
require.NoError(t, s.WebhookDB.
|
||||
Model(&database.DeliveryResult{}).
|
||||
Where("delivery_id = ?", d.ID).
|
||||
Count(&attempts).Error)
|
||||
assert.Equal(
|
||||
t, int64(1), attempts,
|
||||
"settling must not invent an attempt",
|
||||
)
|
||||
}
|
||||
|
||||
// TestFailedResultWriteLeavesDeliveryRecoverable is the rule the
|
||||
// targets now follow: a bookkeeping write that fails must not advance
|
||||
// the status, because pending and retrying are the states the sweeps
|
||||
// recover and delivered is a claim the database refused to record.
|
||||
func TestFailedResultWriteLeavesDeliveryRecoverable(
|
||||
t *testing.T,
|
||||
) {
|
||||
t.Parallel()
|
||||
|
||||
s := newISetup(t)
|
||||
targetID := uuid.New().String()
|
||||
|
||||
var hits atomic.Int64
|
||||
|
||||
ts := httptest.NewServer(http.HandlerFunc(
|
||||
func(w http.ResponseWriter, _ *http.Request) {
|
||||
hits.Add(1)
|
||||
w.WriteHeader(http.StatusOK)
|
||||
},
|
||||
))
|
||||
defer ts.Close()
|
||||
|
||||
event := iSeedEvent(
|
||||
t, s.WebhookDB, s.WebhookID, `{"unwritable":true}`,
|
||||
)
|
||||
|
||||
d := iSeedDelivery(
|
||||
t, s.WebhookDB, event.ID, targetID,
|
||||
database.DeliveryStatusPending,
|
||||
)
|
||||
|
||||
// Drop the table the attempt row goes in, so the send succeeds
|
||||
// and only the bookkeeping write fails.
|
||||
require.NoError(
|
||||
t,
|
||||
s.WebhookDB.Exec("drop table delivery_results").Error,
|
||||
)
|
||||
|
||||
full := &database.Delivery{
|
||||
EventID: event.ID,
|
||||
TargetID: targetID,
|
||||
Status: database.DeliveryStatusPending,
|
||||
Event: event,
|
||||
Target: database.Target{
|
||||
Name: "unwritable",
|
||||
Type: database.TargetTypeHTTP,
|
||||
Config: iHTTPConfig(ts.URL),
|
||||
},
|
||||
}
|
||||
full.ID = d.ID
|
||||
|
||||
s.Engine.ExportDeliverHTTP(
|
||||
context.Background(), s.WebhookDB, full,
|
||||
&delivery.Task{DeliveryID: d.ID, AttemptNum: 1},
|
||||
)
|
||||
|
||||
assert.Equal(
|
||||
t, int64(1), hits.Load(),
|
||||
"the send itself must still happen",
|
||||
)
|
||||
|
||||
iAssertStatus(
|
||||
t, s.WebhookDB, d.ID,
|
||||
database.DeliveryStatusPending,
|
||||
)
|
||||
}
|
||||
@@ -170,7 +170,8 @@ func TestDelivery_CrossOriginRedirectDropsOriginScopedHeaders(
|
||||
// Stripping must not fire within the configured origin, or every
|
||||
// destination that redirects its own path would lose its
|
||||
// credential and start answering 401 — and would lose the inbound
|
||||
// signature the receiver verifies.
|
||||
// signature header the target endpoint verifies. webhooker's own
|
||||
// receiver verifies no signature; it only forwards the header.
|
||||
func TestDelivery_SameOriginRedirectKeepsOriginScopedHeaders(
|
||||
t *testing.T,
|
||||
) {
|
||||
|
||||
@@ -58,12 +58,17 @@ func (t *databaseTarget) Deliver(
|
||||
"error", err,
|
||||
)
|
||||
|
||||
t.eng.recordResult(
|
||||
recErr := t.eng.recordResult(
|
||||
webhookDB, d, 1, false, 0, "",
|
||||
err.Error(), elapsed.Milliseconds(),
|
||||
)
|
||||
if recErr != nil {
|
||||
t.eng.bookkeepingFailed(d, recErr)
|
||||
|
||||
t.eng.updateDeliveryStatus(
|
||||
return
|
||||
}
|
||||
|
||||
t.eng.settleStatus(
|
||||
webhookDB, d, d.Target.Type,
|
||||
database.DeliveryStatusFailed,
|
||||
)
|
||||
@@ -71,12 +76,17 @@ func (t *databaseTarget) Deliver(
|
||||
return
|
||||
}
|
||||
|
||||
t.eng.recordResult(
|
||||
recErr := t.eng.recordResult(
|
||||
webhookDB, d, 1, true, 0, "", "",
|
||||
elapsed.Milliseconds(),
|
||||
)
|
||||
if recErr != nil {
|
||||
t.eng.bookkeepingFailed(d, recErr)
|
||||
|
||||
t.eng.updateDeliveryStatus(
|
||||
return
|
||||
}
|
||||
|
||||
t.eng.settleStatus(
|
||||
webhookDB, d, d.Target.Type,
|
||||
database.DeliveryStatusDelivered,
|
||||
)
|
||||
|
||||
@@ -1,7 +1,6 @@
|
||||
package delivery
|
||||
|
||||
import (
|
||||
"database/sql"
|
||||
"encoding/json"
|
||||
"errors"
|
||||
"fmt"
|
||||
@@ -12,6 +11,7 @@ import (
|
||||
|
||||
"gorm.io/driver/sqlite"
|
||||
"gorm.io/gorm"
|
||||
"sneak.berlin/go/webhooker/internal/database"
|
||||
"sneak.berlin/go/webhooker/internal/gormlog"
|
||||
)
|
||||
|
||||
@@ -30,13 +30,13 @@ const (
|
||||
// path: open the archive file, creating it if missing, so a
|
||||
// first write (or a write after the operator moved the file
|
||||
// away) recreates it.
|
||||
archiveModeCreate = "rwc"
|
||||
archiveModeCreate = database.SQLiteModeCreate
|
||||
|
||||
// archiveModeExisting is the SQLite URI mode used by the idle
|
||||
// sweep: open read-write but never create. A sweep must never
|
||||
// conjure an empty archive file for a webhook that has a
|
||||
// database target but has never received an event.
|
||||
archiveModeExisting = "rw"
|
||||
archiveModeExisting = database.SQLiteModeExisting
|
||||
)
|
||||
|
||||
var (
|
||||
@@ -273,9 +273,11 @@ func (w *archiveWriter) open(expiry time.Duration) error {
|
||||
func (w *archiveWriter) openMode(
|
||||
mode string, expiry time.Duration,
|
||||
) error {
|
||||
dbURL := fmt.Sprintf("file:%s?mode=%s", w.path, mode)
|
||||
|
||||
sqlDB, err := sql.Open("sqlite", dbURL)
|
||||
// Opened through database.OpenSQLite so an archive file carries
|
||||
// the same WAL journaling, busy timeout, immediate-transaction
|
||||
// locking, and pool bounds as every other database file. See
|
||||
// internal/database/sqlite_open.go.
|
||||
sqlDB, err := database.OpenSQLite(w.path, mode)
|
||||
if err != nil {
|
||||
return fmt.Errorf(
|
||||
"opening archive database %s: %w", w.path, err,
|
||||
|
||||
@@ -77,14 +77,19 @@ func (c *httpCore) fireAndForget(
|
||||
) {
|
||||
c.eng.observeAttempt(d.Target.Type, res.elapsed())
|
||||
|
||||
c.eng.recordResult(
|
||||
err := c.eng.recordResult(
|
||||
webhookDB, d, 1, res.success,
|
||||
res.statusCode, res.respBody, res.errMsg,
|
||||
res.duration,
|
||||
)
|
||||
if err != nil {
|
||||
c.eng.bookkeepingFailed(d, err)
|
||||
|
||||
return
|
||||
}
|
||||
|
||||
if res.success {
|
||||
c.eng.updateDeliveryStatus(
|
||||
c.eng.settleStatus(
|
||||
webhookDB, d, d.Target.Type,
|
||||
database.DeliveryStatusDelivered,
|
||||
)
|
||||
@@ -92,7 +97,7 @@ func (c *httpCore) fireAndForget(
|
||||
return
|
||||
}
|
||||
|
||||
c.eng.updateDeliveryStatus(
|
||||
c.eng.settleStatus(
|
||||
webhookDB, d, d.Target.Type,
|
||||
database.DeliveryStatusFailed,
|
||||
)
|
||||
@@ -122,16 +127,25 @@ func (c *httpCore) withRetry(
|
||||
|
||||
c.eng.observeAttempt(d.Target.Type, res.elapsed())
|
||||
|
||||
c.eng.recordResult(
|
||||
err := c.eng.recordResult(
|
||||
webhookDB, d, attemptNum, res.success,
|
||||
res.statusCode, res.respBody, res.errMsg,
|
||||
res.duration,
|
||||
)
|
||||
if err != nil {
|
||||
// The breaker still learns the outcome: it describes the
|
||||
// target's health, which is unaffected by this database's.
|
||||
c.recordCircuitOutcome(cb, res.success)
|
||||
|
||||
c.eng.bookkeepingFailed(d, err)
|
||||
|
||||
return
|
||||
}
|
||||
|
||||
if res.success {
|
||||
cb.RecordSuccess()
|
||||
|
||||
c.eng.updateDeliveryStatus(
|
||||
c.eng.settleStatus(
|
||||
webhookDB, d, d.Target.Type,
|
||||
database.DeliveryStatusDelivered,
|
||||
)
|
||||
@@ -146,6 +160,20 @@ func (c *httpCore) withRetry(
|
||||
)
|
||||
}
|
||||
|
||||
// recordCircuitOutcome feeds one attempt's outcome to the target's
|
||||
// circuit breaker.
|
||||
func (c *httpCore) recordCircuitOutcome(
|
||||
cb *CircuitBreaker, success bool,
|
||||
) {
|
||||
if success {
|
||||
cb.RecordSuccess()
|
||||
|
||||
return
|
||||
}
|
||||
|
||||
cb.RecordFailure()
|
||||
}
|
||||
|
||||
func (c *httpCore) circuitBreakerBlock(
|
||||
webhookDB *gorm.DB,
|
||||
d *database.Delivery,
|
||||
@@ -169,7 +197,7 @@ func (c *httpCore) circuitBreakerBlock(
|
||||
"cooldown_remaining", remaining,
|
||||
)
|
||||
|
||||
c.eng.updateDeliveryStatus(
|
||||
c.eng.settleStatus(
|
||||
webhookDB, d, d.Target.Type,
|
||||
database.DeliveryStatusRetrying,
|
||||
)
|
||||
@@ -189,7 +217,7 @@ func (c *httpCore) handleRetry(
|
||||
attemptNum int,
|
||||
) {
|
||||
if attemptNum >= maxRetries {
|
||||
c.eng.updateDeliveryStatus(
|
||||
c.eng.settleStatus(
|
||||
webhookDB, d, d.Target.Type,
|
||||
database.DeliveryStatusFailed,
|
||||
)
|
||||
@@ -197,7 +225,7 @@ func (c *httpCore) handleRetry(
|
||||
return
|
||||
}
|
||||
|
||||
c.eng.updateDeliveryStatus(
|
||||
c.eng.settleStatus(
|
||||
webhookDB, d, d.Target.Type,
|
||||
database.DeliveryStatusRetrying,
|
||||
)
|
||||
@@ -332,12 +360,17 @@ func (t *httpTarget) Deliver(
|
||||
"error", err,
|
||||
)
|
||||
|
||||
t.eng.recordResult(
|
||||
recErr := t.eng.recordResult(
|
||||
webhookDB, d, task.AttemptNum,
|
||||
false, 0, "", err.Error(), 0,
|
||||
)
|
||||
if recErr != nil {
|
||||
t.eng.bookkeepingFailed(d, recErr)
|
||||
|
||||
t.eng.updateDeliveryStatus(
|
||||
return
|
||||
}
|
||||
|
||||
t.eng.settleStatus(
|
||||
webhookDB, d, d.Target.Type,
|
||||
database.DeliveryStatusFailed,
|
||||
)
|
||||
|
||||
@@ -55,12 +55,17 @@ func (t *logTarget) Deliver(
|
||||
|
||||
t.eng.observeAttempt(d.Target.Type, elapsed)
|
||||
|
||||
t.eng.recordResult(
|
||||
err := t.eng.recordResult(
|
||||
webhookDB, d, 1, true, 0, "", "",
|
||||
elapsed.Milliseconds(),
|
||||
)
|
||||
if err != nil {
|
||||
t.eng.bookkeepingFailed(d, err)
|
||||
|
||||
t.eng.updateDeliveryStatus(
|
||||
return
|
||||
}
|
||||
|
||||
t.eng.settleStatus(
|
||||
webhookDB, d, d.Target.Type,
|
||||
database.DeliveryStatusDelivered,
|
||||
)
|
||||
|
||||
@@ -95,12 +95,17 @@ func (t *slackTarget) failConfig(
|
||||
d *database.Delivery,
|
||||
err error,
|
||||
) {
|
||||
t.eng.recordResult(
|
||||
recErr := t.eng.recordResult(
|
||||
webhookDB, d, 1,
|
||||
false, 0, "", err.Error(), 0,
|
||||
)
|
||||
if recErr != nil {
|
||||
t.eng.bookkeepingFailed(d, recErr)
|
||||
|
||||
t.eng.updateDeliveryStatus(
|
||||
return
|
||||
}
|
||||
|
||||
t.eng.settleStatus(
|
||||
webhookDB, d, d.Target.Type,
|
||||
database.DeliveryStatusFailed,
|
||||
)
|
||||
@@ -226,10 +231,15 @@ func FormatSlackMessage(
|
||||
event.ContentType,
|
||||
)
|
||||
|
||||
timestamp := "unknown"
|
||||
if !event.CreatedAt.IsZero() {
|
||||
timestamp = event.CreatedAt.UTC().Format(time.RFC3339)
|
||||
}
|
||||
|
||||
fmt.Fprintf(
|
||||
&b,
|
||||
"*Timestamp:* `%s`\n",
|
||||
event.CreatedAt.UTC().Format(time.RFC3339),
|
||||
timestamp,
|
||||
)
|
||||
|
||||
fmt.Fprintf(
|
||||
|
||||
531
internal/delivery/terminal_state_test.go
Normal file
531
internal/delivery/terminal_state_test.go
Normal file
@@ -0,0 +1,531 @@
|
||||
package delivery_test
|
||||
|
||||
import (
|
||||
"context"
|
||||
"net/http"
|
||||
"net/http/httptest"
|
||||
"sync/atomic"
|
||||
"testing"
|
||||
|
||||
"github.com/google/uuid"
|
||||
"github.com/stretchr/testify/assert"
|
||||
"github.com/stretchr/testify/require"
|
||||
"sneak.berlin/go/webhooker/internal/database"
|
||||
"sneak.berlin/go/webhooker/internal/delivery"
|
||||
)
|
||||
|
||||
// The two terminal-state gaps of
|
||||
// https://git.eeqj.de/sneak/webhooker/issues/107: a delivery failed
|
||||
// with nothing in its event log to say why, and a retrying delivery
|
||||
// whose target was deleted, which used to keep sending and then never
|
||||
// terminalise.
|
||||
|
||||
// tUnknownType is a target type no build implements. It stands in for
|
||||
// a target whose type was written by a build that knew a type this one
|
||||
// does not.
|
||||
const tUnknownType = database.TargetType("pubsub")
|
||||
|
||||
// tSeedDeletedTarget creates a target, a retrying delivery against it
|
||||
// with one recorded failed attempt, and then deletes the target the
|
||||
// way the source page does.
|
||||
//
|
||||
// It asserts the delete is soft, because that is the whole reason the
|
||||
// engine could not tell a deleted target from a target id that never
|
||||
// named a row: the surviving row is invisible to a scoped read.
|
||||
func tSeedDeletedTarget(
|
||||
t *testing.T,
|
||||
s iSetup,
|
||||
name, url string,
|
||||
) string {
|
||||
t.Helper()
|
||||
|
||||
targetID := uuid.New().String()
|
||||
|
||||
iCreateTarget(
|
||||
t, s.MainDB, targetID, s.WebhookID, name,
|
||||
database.TargetTypeHTTP, iHTTPConfig(url), 5,
|
||||
)
|
||||
|
||||
event := iSeedEvent(
|
||||
t, s.WebhookDB, s.WebhookID, `{"target":"deleted"}`,
|
||||
)
|
||||
|
||||
d := iSeedDelivery(
|
||||
t, s.WebhookDB, event.ID, targetID,
|
||||
database.DeliveryStatusRetrying,
|
||||
)
|
||||
|
||||
iSeedFailedResult(t, s.WebhookDB, d.ID)
|
||||
|
||||
require.NoError(t, s.MainDB.Delete(
|
||||
&database.Target{}, "id = ?", targetID,
|
||||
).Error)
|
||||
|
||||
var scoped, unscoped int64
|
||||
|
||||
require.NoError(t, s.MainDB.
|
||||
Model(&database.Target{}).
|
||||
Where("id = ?", targetID).
|
||||
Count(&scoped).Error)
|
||||
|
||||
require.NoError(t, s.MainDB.Unscoped().
|
||||
Model(&database.Target{}).
|
||||
Where("id = ?", targetID).
|
||||
Count(&unscoped).Error)
|
||||
|
||||
require.Zero(t, scoped,
|
||||
"the deleted target is still visible to a scoped read",
|
||||
)
|
||||
require.Equal(t, int64(1), unscoped,
|
||||
"the delete was hard, so this test proves nothing about "+
|
||||
"the soft-delete case it exists for",
|
||||
)
|
||||
|
||||
return d.ID
|
||||
}
|
||||
|
||||
// tLastResult returns a delivery's final recorded attempt, asserting
|
||||
// the expected number of them.
|
||||
func tLastResult(
|
||||
t *testing.T,
|
||||
s iSetup,
|
||||
deliveryID string,
|
||||
want int,
|
||||
) database.DeliveryResult {
|
||||
t.Helper()
|
||||
|
||||
results := iResults(t, s.WebhookDB, deliveryID)
|
||||
require.Len(t, results, want)
|
||||
|
||||
return results[want-1]
|
||||
}
|
||||
|
||||
// --- 1. A failure with nothing recorded ---
|
||||
|
||||
func TestProcessDelivery_UnknownTargetType_RecordsWhy(
|
||||
t *testing.T,
|
||||
) {
|
||||
t.Parallel()
|
||||
|
||||
s := newISetup(t)
|
||||
|
||||
targetID := uuid.New().String()
|
||||
|
||||
event := iSeedEvent(
|
||||
t, s.WebhookDB, s.WebhookID, `{"unknown":"type"}`,
|
||||
)
|
||||
|
||||
seeded := iSeedDelivery(
|
||||
t, s.WebhookDB, event.ID, targetID,
|
||||
database.DeliveryStatusPending,
|
||||
)
|
||||
|
||||
target := database.Target{
|
||||
Name: "mystery",
|
||||
Type: tUnknownType,
|
||||
Config: iHTTPConfig("http://example.com/hook"),
|
||||
}
|
||||
target.ID = targetID
|
||||
|
||||
d := database.Delivery{
|
||||
EventID: event.ID,
|
||||
TargetID: targetID,
|
||||
Status: database.DeliveryStatusPending,
|
||||
Event: event,
|
||||
Target: target,
|
||||
}
|
||||
d.ID = seeded.ID
|
||||
|
||||
body := event.Body
|
||||
task := iTask(
|
||||
seeded, event, s.WebhookID, targetID, "mystery",
|
||||
target.Config, 0, 1, &body,
|
||||
)
|
||||
task.TargetType = tUnknownType
|
||||
|
||||
s.Engine.ExportProcessDelivery(
|
||||
context.Background(), s.WebhookDB, &d, &task,
|
||||
)
|
||||
|
||||
iAssertStatus(
|
||||
t, s.WebhookDB, d.ID, database.DeliveryStatusFailed,
|
||||
)
|
||||
|
||||
last := tLastResult(t, s, d.ID, 1)
|
||||
|
||||
assert.False(t, last.Success)
|
||||
assert.Equal(t, 1, last.AttemptNum)
|
||||
assert.Contains(t, last.Error, string(tUnknownType),
|
||||
"the recorded reason does not name the offending type",
|
||||
)
|
||||
}
|
||||
|
||||
// --- 2. A retrying delivery whose target is gone ---
|
||||
|
||||
func TestRecoverSingleRetry_TargetDeleted(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
s := newISetup(t)
|
||||
|
||||
iCreateWebhook(
|
||||
t, s.MainDB, s.WebhookID, "deleted-target-recovery",
|
||||
)
|
||||
|
||||
deliveryID := tSeedDeletedTarget(
|
||||
t, s, "gone-on-recovery", "http://example.com/hook",
|
||||
)
|
||||
|
||||
s.Engine.ExportRecoverWebhookDeliveries(
|
||||
context.Background(), s.WebhookID,
|
||||
)
|
||||
|
||||
iAssertStatus(
|
||||
t, s.WebhookDB, deliveryID,
|
||||
database.DeliveryStatusFailed,
|
||||
)
|
||||
|
||||
last := tLastResult(t, s, deliveryID, 2)
|
||||
|
||||
assert.False(t, last.Success)
|
||||
assert.Equal(t, 2, last.AttemptNum)
|
||||
assert.Contains(t, last.Error, "gone-on-recovery")
|
||||
assert.Contains(t, last.Error, "was deleted")
|
||||
|
||||
assert.Empty(t, s.Engine.ExportRetryCh(),
|
||||
"a delivery whose target is gone was rescheduled",
|
||||
)
|
||||
assert.Zero(t, s.Engine.ExportInflightHeld(),
|
||||
"the terminal path leaked its ownership reference",
|
||||
)
|
||||
}
|
||||
|
||||
func TestSweepSingleRetry_TargetDeleted(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
s := newISetup(t)
|
||||
|
||||
iCreateWebhook(
|
||||
t, s.MainDB, s.WebhookID, "deleted-target-sweep",
|
||||
)
|
||||
|
||||
deliveryID := tSeedDeletedTarget(
|
||||
t, s, "gone-on-sweep", "http://example.com/hook",
|
||||
)
|
||||
|
||||
// Twice, because the bug was an error the sweep repeated every
|
||||
// minute for the life of the database: the second sweep must
|
||||
// find nothing left to do.
|
||||
s.Engine.ExportSweepWebhookRetries(
|
||||
context.Background(), s.WebhookID,
|
||||
)
|
||||
s.Engine.ExportSweepWebhookRetries(
|
||||
context.Background(), s.WebhookID,
|
||||
)
|
||||
|
||||
iAssertStatus(
|
||||
t, s.WebhookDB, deliveryID,
|
||||
database.DeliveryStatusFailed,
|
||||
)
|
||||
|
||||
last := tLastResult(t, s, deliveryID, 2)
|
||||
|
||||
assert.Contains(t, last.Error, "gone-on-sweep")
|
||||
assert.Contains(t, last.Error, "was deleted")
|
||||
|
||||
assert.Empty(t, s.Engine.ExportRetryCh())
|
||||
assert.Zero(t, s.Engine.ExportInflightHeld())
|
||||
}
|
||||
|
||||
// TestSweepSingleRetry_TargetNeverExisted covers the other half of the
|
||||
// soft-delete distinction: an id with no row at all, deleted or
|
||||
// otherwise, must not be reported as something the operator deleted.
|
||||
func TestSweepSingleRetry_TargetNeverExisted(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
s := newISetup(t)
|
||||
|
||||
iCreateWebhook(
|
||||
t, s.MainDB, s.WebhookID, "target-never-existed",
|
||||
)
|
||||
|
||||
targetID := uuid.New().String()
|
||||
|
||||
event := iSeedEvent(
|
||||
t, s.WebhookDB, s.WebhookID, `{"target":"absent"}`,
|
||||
)
|
||||
|
||||
d := iSeedDelivery(
|
||||
t, s.WebhookDB, event.ID, targetID,
|
||||
database.DeliveryStatusRetrying,
|
||||
)
|
||||
|
||||
iSeedFailedResult(t, s.WebhookDB, d.ID)
|
||||
|
||||
s.Engine.ExportSweepWebhookRetries(
|
||||
context.Background(), s.WebhookID,
|
||||
)
|
||||
|
||||
iAssertStatus(
|
||||
t, s.WebhookDB, d.ID, database.DeliveryStatusFailed,
|
||||
)
|
||||
|
||||
last := tLastResult(t, s, d.ID, 2)
|
||||
|
||||
assert.Contains(t, last.Error, targetID)
|
||||
assert.Contains(t, last.Error, "no longer exists")
|
||||
assert.NotContains(t, last.Error, "was deleted",
|
||||
"an id that never named a row was reported as a deletion",
|
||||
)
|
||||
}
|
||||
|
||||
// TestFailMissingTargetRetry_WritesNoTargetRow holds the new terminal
|
||||
// path to the same rule as the existing one: no target row, and so no
|
||||
// plaintext target config, may be written into the per-webhook event
|
||||
// database. See https://git.eeqj.de/sneak/webhooker/issues/206.
|
||||
func TestFailMissingTargetRetry_WritesNoTargetRow(
|
||||
t *testing.T,
|
||||
) {
|
||||
t.Parallel()
|
||||
|
||||
s := newISetup(t)
|
||||
|
||||
iCreateWebhook(
|
||||
t, s.MainDB, s.WebhookID, "no-target-row-deleted",
|
||||
)
|
||||
|
||||
hookURL := "https://hooks.slack.com/services/T00/B00/x"
|
||||
|
||||
deliveryID := tSeedDeletedTarget(
|
||||
t, s, "credential-bearing", hookURL,
|
||||
)
|
||||
|
||||
s.Engine.ExportSweepWebhookRetries(
|
||||
context.Background(), s.WebhookID,
|
||||
)
|
||||
|
||||
iAssertStatus(
|
||||
t, s.WebhookDB, deliveryID,
|
||||
database.DeliveryStatusFailed,
|
||||
)
|
||||
|
||||
var configs []string
|
||||
|
||||
require.NoError(t, s.WebhookDB.
|
||||
Table("targets").
|
||||
Pluck("config", &configs).Error)
|
||||
|
||||
assert.Empty(t, configs,
|
||||
"the deleted-target terminal path wrote a target row "+
|
||||
"into the per-webhook event database",
|
||||
)
|
||||
}
|
||||
|
||||
// --- 3. The scheduled retry chain ---
|
||||
|
||||
// tRetryChainSetup wires a counting sink and a retrying delivery
|
||||
// against a live target pointing at it, and returns the task a
|
||||
// scheduled retry would carry — config and all, snapshotted as
|
||||
// ScheduleRetry snapshots it.
|
||||
func tRetryChainSetup(
|
||||
t *testing.T,
|
||||
s iSetup,
|
||||
name string,
|
||||
hits *atomic.Int64,
|
||||
) (delivery.Task, string) {
|
||||
t.Helper()
|
||||
|
||||
ts := httptest.NewServer(http.HandlerFunc(
|
||||
func(w http.ResponseWriter, _ *http.Request) {
|
||||
hits.Add(1)
|
||||
w.WriteHeader(http.StatusOK)
|
||||
},
|
||||
))
|
||||
t.Cleanup(ts.Close)
|
||||
|
||||
iCreateWebhook(t, s.MainDB, s.WebhookID, name)
|
||||
|
||||
targetID := uuid.New().String()
|
||||
cfg := iHTTPConfig(ts.URL)
|
||||
|
||||
iCreateTarget(
|
||||
t, s.MainDB, targetID, s.WebhookID, name,
|
||||
database.TargetTypeHTTP, cfg, 5,
|
||||
)
|
||||
|
||||
event := iSeedEvent(
|
||||
t, s.WebhookDB, s.WebhookID, `{"chain":"retry"}`,
|
||||
)
|
||||
|
||||
d := iSeedDelivery(
|
||||
t, s.WebhookDB, event.ID, targetID,
|
||||
database.DeliveryStatusRetrying,
|
||||
)
|
||||
|
||||
iSeedFailedResult(t, s.WebhookDB, d.ID)
|
||||
|
||||
body := event.Body
|
||||
|
||||
return iTask(
|
||||
d, event, s.WebhookID, targetID, name, cfg, 5, 2, &body,
|
||||
), targetID
|
||||
}
|
||||
|
||||
// TestProcessRetryTask_TargetDeleted_MakesNoAttempt is the half the
|
||||
// deployability audit found worse than filed: terminalising on
|
||||
// recovery and sweep alone leaves the already-scheduled timer chain
|
||||
// running, and it holds the target's configuration from before the
|
||||
// deletion, so it goes on sending to a destination that was removed.
|
||||
func TestProcessRetryTask_TargetDeleted_MakesNoAttempt(
|
||||
t *testing.T,
|
||||
) {
|
||||
t.Parallel()
|
||||
|
||||
s := newISetup(t)
|
||||
|
||||
var hits atomic.Int64
|
||||
|
||||
task, targetID := tRetryChainSetup(
|
||||
t, s, "gone-mid-chain", &hits,
|
||||
)
|
||||
|
||||
require.NoError(t, s.MainDB.Delete(
|
||||
&database.Target{}, "id = ?", targetID,
|
||||
).Error)
|
||||
|
||||
s.Engine.ExportProcessRetryTask(
|
||||
context.Background(), &task,
|
||||
)
|
||||
|
||||
assert.Zero(t, hits.Load(),
|
||||
"a scheduled retry fired at a target the operator "+
|
||||
"had already deleted",
|
||||
)
|
||||
|
||||
iAssertStatus(
|
||||
t, s.WebhookDB, task.DeliveryID,
|
||||
database.DeliveryStatusFailed,
|
||||
)
|
||||
|
||||
last := tLastResult(t, s, task.DeliveryID, 2)
|
||||
|
||||
assert.False(t, last.Success)
|
||||
assert.Contains(t, last.Error, "was deleted")
|
||||
|
||||
assert.Zero(t, s.Engine.ExportInflightHeld())
|
||||
}
|
||||
|
||||
// TestProcessRetryTask_TargetPresent_StillDelivers is the guard's
|
||||
// mutation check: a liveness check that refused every retry would pass
|
||||
// the test above and break every retry there is.
|
||||
func TestProcessRetryTask_TargetPresent_StillDelivers(
|
||||
t *testing.T,
|
||||
) {
|
||||
t.Parallel()
|
||||
|
||||
s := newISetup(t)
|
||||
|
||||
var hits atomic.Int64
|
||||
|
||||
task, _ := tRetryChainSetup(t, s, "still-there", &hits)
|
||||
|
||||
s.Engine.ExportProcessRetryTask(
|
||||
context.Background(), &task,
|
||||
)
|
||||
|
||||
assert.Equal(t, int64(1), hits.Load())
|
||||
|
||||
iAssertStatus(
|
||||
t, s.WebhookDB, task.DeliveryID,
|
||||
database.DeliveryStatusDelivered,
|
||||
)
|
||||
}
|
||||
|
||||
// TestProcessRetryTask_TargetUnreadable_StillDelivers pins the other
|
||||
// half of the guard: only a target that is confirmed gone stops a
|
||||
// retry. A main database that cannot be read is a transient fault, and
|
||||
// a guard that abandoned deliveries on one would be a worse bug than
|
||||
// the one it fixes.
|
||||
func TestProcessRetryTask_TargetUnreadable_StillDelivers(
|
||||
t *testing.T,
|
||||
) {
|
||||
t.Parallel()
|
||||
|
||||
s := newISetup(t)
|
||||
|
||||
var hits atomic.Int64
|
||||
|
||||
task, _ := tRetryChainSetup(t, s, "unreadable-main", &hits)
|
||||
|
||||
sqlDB, err := s.MainDB.DB()
|
||||
require.NoError(t, err)
|
||||
require.NoError(t, sqlDB.Close())
|
||||
|
||||
s.Engine.ExportProcessRetryTask(
|
||||
context.Background(), &task,
|
||||
)
|
||||
|
||||
assert.Equal(t, int64(1), hits.Load(),
|
||||
"a retry was abandoned because the main database "+
|
||||
"could not be read, not because its target was gone",
|
||||
)
|
||||
|
||||
iAssertStatus(
|
||||
t, s.WebhookDB, task.DeliveryID,
|
||||
database.DeliveryStatusDelivered,
|
||||
)
|
||||
}
|
||||
|
||||
// TestRecoverSingleRetry_TargetUnreadable_LeavesDeliveryAlone is the
|
||||
// same rule on the recovery path. A read failure that is not
|
||||
// "record not found" must leave every retrying delivery of every
|
||||
// webhook exactly as it was.
|
||||
func TestRecoverSingleRetry_TargetUnreadable_LeavesDeliveryAlone(
|
||||
t *testing.T,
|
||||
) {
|
||||
t.Parallel()
|
||||
|
||||
s := newISetup(t)
|
||||
|
||||
iCreateWebhook(
|
||||
t, s.MainDB, s.WebhookID, "unreadable-on-recovery",
|
||||
)
|
||||
|
||||
targetID := uuid.New().String()
|
||||
|
||||
iCreateTarget(
|
||||
t, s.MainDB, targetID, s.WebhookID, "healthy",
|
||||
database.TargetTypeHTTP,
|
||||
iHTTPConfig("http://example.com/hook"), 5,
|
||||
)
|
||||
|
||||
event := iSeedEvent(
|
||||
t, s.WebhookDB, s.WebhookID, `{"still":"retrying"}`,
|
||||
)
|
||||
|
||||
d := iSeedDelivery(
|
||||
t, s.WebhookDB, event.ID, targetID,
|
||||
database.DeliveryStatusRetrying,
|
||||
)
|
||||
|
||||
iSeedFailedResult(t, s.WebhookDB, d.ID)
|
||||
|
||||
sqlDB, err := s.MainDB.DB()
|
||||
require.NoError(t, err)
|
||||
require.NoError(t, sqlDB.Close())
|
||||
|
||||
s.Engine.ExportRecoverRetryingDeliveries(
|
||||
s.WebhookDB, s.WebhookID,
|
||||
)
|
||||
|
||||
iAssertStatus(
|
||||
t, s.WebhookDB, d.ID,
|
||||
database.DeliveryStatusRetrying,
|
||||
)
|
||||
|
||||
assert.Len(t, iResults(t, s.WebhookDB, d.ID), 1,
|
||||
"an unreadable main database produced a terminal "+
|
||||
"failure row",
|
||||
)
|
||||
|
||||
assert.Zero(t, s.Engine.ExportInflightHeld())
|
||||
}
|
||||
@@ -92,11 +92,10 @@ func (h *Handlers) HandleEventBodyDownload() http.HandlerFunc {
|
||||
// once per range.
|
||||
//
|
||||
// One consequence is worth keeping in view: the read finishes
|
||||
// before the client is written to, so no read lock is held for
|
||||
// the length of a slow download. These per-webhook databases
|
||||
// run in SQLite's default journal mode rather than WAL, so a
|
||||
// lock held that long would block the receiver from recording
|
||||
// new events.
|
||||
// before the client is written to, so nothing is held open for
|
||||
// the length of a slow download. Under WAL a read no longer
|
||||
// blocks the receiver, but it does pin the WAL against
|
||||
// checkpointing, and a download can last minutes.
|
||||
func (h *Handlers) serveEventBody(
|
||||
w http.ResponseWriter,
|
||||
r *http.Request,
|
||||
|
||||
@@ -139,10 +139,10 @@ func (h *Handlers) resubmitEvent(
|
||||
}
|
||||
|
||||
// Read before the write transaction is opened. The body can be up
|
||||
// to the 1 MB ingest cap, and holding a read of it inside the
|
||||
// transaction would extend how long the per-webhook database is
|
||||
// locked against the receiver, which runs these files in
|
||||
// SQLite's default journal mode rather than WAL.
|
||||
// to the 1 MB ingest cap, and every transaction on these files
|
||||
// takes the write lock at BEGIN (_txlock=immediate, see
|
||||
// internal/database/sqlite_open.go), so reading inside it would
|
||||
// hold that lock against the receiver for the length of the read.
|
||||
src, found, err := loadResubmitSource(
|
||||
webhookDB, webhook.ID, eventID.String(),
|
||||
)
|
||||
|
||||
283
internal/handlers/source_detail_baseurl_test.go
Normal file
283
internal/handlers/source_detail_baseurl_test.go
Normal file
@@ -0,0 +1,283 @@
|
||||
package handlers_test
|
||||
|
||||
import (
|
||||
"context"
|
||||
"crypto/tls"
|
||||
"net/http"
|
||||
"net/http/httptest"
|
||||
"regexp"
|
||||
"testing"
|
||||
|
||||
"github.com/go-chi/chi"
|
||||
"github.com/stretchr/testify/assert"
|
||||
"github.com/stretchr/testify/require"
|
||||
"sneak.berlin/go/webhooker/internal/database"
|
||||
"sneak.berlin/go/webhooker/internal/handlers"
|
||||
"sneak.berlin/go/webhooker/internal/session"
|
||||
)
|
||||
|
||||
// The only two schemes a rendered entrypoint URL may carry,
|
||||
// whatever the request claimed.
|
||||
const (
|
||||
schemeHTTPS = "https"
|
||||
schemeHTTP = "http"
|
||||
)
|
||||
|
||||
// entrypointURLPattern captures the entrypoint URL the source
|
||||
// detail page renders, which is the operator-visible product of
|
||||
// BaseURL. Asserting on the extracted string rather than on a
|
||||
// substring of the page proves the raw header value cannot reach
|
||||
// the scheme by any route.
|
||||
var entrypointURLPattern = regexp.MustCompile(
|
||||
`<code id="entrypoint-url-[^"]*"[^>]*>([^<]*)</code>`,
|
||||
)
|
||||
|
||||
// baseURLFixture is one started app plus the webhook whose
|
||||
// entrypoint URL the BaseURL cases read.
|
||||
type baseURLFixture struct {
|
||||
handlers *handlers.Handlers
|
||||
session *session.Session
|
||||
webhook string
|
||||
path string
|
||||
}
|
||||
|
||||
// newBaseURLFixture starts the app and seeds a webhook with one
|
||||
// entrypoint.
|
||||
func newBaseURLFixture(t *testing.T) *baseURLFixture {
|
||||
t.Helper()
|
||||
|
||||
var (
|
||||
h *handlers.Handlers
|
||||
sess *session.Session
|
||||
db *database.Database
|
||||
)
|
||||
|
||||
app := newTestApp(t, &h, &sess, &db)
|
||||
app.RequireStart()
|
||||
|
||||
t.Cleanup(app.RequireStop)
|
||||
|
||||
wh := seedWebhook(t, db)
|
||||
seedEntrypoint(t, db, wh.ID)
|
||||
|
||||
return &baseURLFixture{
|
||||
handlers: h,
|
||||
session: sess,
|
||||
webhook: wh.ID,
|
||||
path: "ep-" + wh.ID,
|
||||
}
|
||||
}
|
||||
|
||||
// entrypointURL renders the source detail page for the fixture's
|
||||
// webhook over a request the caller shapes, and returns the
|
||||
// entrypoint URL as an operator would copy it.
|
||||
func (f *baseURLFixture) entrypointURL(
|
||||
t *testing.T,
|
||||
host string,
|
||||
shape func(*http.Request),
|
||||
) string {
|
||||
t.Helper()
|
||||
|
||||
req := httptest.NewRequestWithContext(
|
||||
context.Background(),
|
||||
http.MethodGet,
|
||||
"/source/"+f.webhook,
|
||||
nil,
|
||||
)
|
||||
req.Host = host
|
||||
|
||||
shape(req)
|
||||
|
||||
for _, c := range authenticatedCookies(
|
||||
t, f.session, deleteTestUserID, deleteTestUsername,
|
||||
) {
|
||||
req.AddCookie(c)
|
||||
}
|
||||
|
||||
rctx := chi.NewRouteContext()
|
||||
rctx.URLParams.Add(paramSourceID, f.webhook)
|
||||
req = req.WithContext(
|
||||
context.WithValue(
|
||||
req.Context(), chi.RouteCtxKey, rctx,
|
||||
),
|
||||
)
|
||||
|
||||
w := httptest.NewRecorder()
|
||||
f.handlers.HandleSourceDetail().ServeHTTP(w, req)
|
||||
|
||||
require.Equal(t, http.StatusOK, w.Code)
|
||||
|
||||
match := entrypointURLPattern.FindStringSubmatch(w.Body.String())
|
||||
require.Len(
|
||||
t, match, 2,
|
||||
"the page must render exactly one entrypoint URL",
|
||||
)
|
||||
|
||||
return match[1]
|
||||
}
|
||||
|
||||
// forwardedProto returns a request shaper setting
|
||||
// X-Forwarded-Proto, or leaving the request alone for "".
|
||||
func forwardedProto(value string) func(*http.Request) {
|
||||
return func(r *http.Request) {
|
||||
if value == "" {
|
||||
return
|
||||
}
|
||||
|
||||
r.Header.Set("X-Forwarded-Proto", value)
|
||||
}
|
||||
}
|
||||
|
||||
// baseURLCase is one X-Forwarded-Proto spelling and the scheme
|
||||
// the rendered entrypoint URL owes it.
|
||||
type baseURLCase struct {
|
||||
name string
|
||||
header string
|
||||
scheme string
|
||||
why string
|
||||
}
|
||||
|
||||
// baseURLCases enumerate the spellings a proxy really emits. The
|
||||
// scheme is only ever http or https: the header value itself is
|
||||
// never a scheme, however it is spelled.
|
||||
func baseURLCases() []baseURLCase {
|
||||
return []baseURLCase{
|
||||
{
|
||||
name: "lowercase",
|
||||
header: schemeHTTPS,
|
||||
scheme: schemeHTTPS,
|
||||
why: "the ordinary spelling",
|
||||
},
|
||||
{
|
||||
name: "uppercase",
|
||||
header: "HTTPS",
|
||||
scheme: schemeHTTPS,
|
||||
why: "the token is case-insensitive; the scheme " +
|
||||
"in a copyable URL is not",
|
||||
},
|
||||
{
|
||||
name: "chain with plaintext inner hop",
|
||||
header: "https, http",
|
||||
scheme: schemeHTTPS,
|
||||
why: "a chained proxy appends its hop; the " +
|
||||
"leftmost element faces the client",
|
||||
},
|
||||
{
|
||||
name: "chain of two TLS hops",
|
||||
header: "https,https",
|
||||
scheme: schemeHTTPS,
|
||||
why: "appended chain with no space after the comma",
|
||||
},
|
||||
{
|
||||
name: "trailing space",
|
||||
header: "https ",
|
||||
scheme: schemeHTTPS,
|
||||
why: "whitespace is not part of the token",
|
||||
},
|
||||
{
|
||||
name: "plaintext",
|
||||
header: schemeHTTP,
|
||||
scheme: schemeHTTP,
|
||||
why: "the negative control: the proxy reports plaintext",
|
||||
},
|
||||
{
|
||||
name: "no header",
|
||||
header: "",
|
||||
scheme: schemeHTTP,
|
||||
why: "a plaintext request asserting nothing is http",
|
||||
},
|
||||
{
|
||||
name: "garbage token",
|
||||
header: "javascript:alert(1)//",
|
||||
scheme: schemeHTTP,
|
||||
why: "anything that is not https is not TLS, and " +
|
||||
"the token never becomes the scheme",
|
||||
},
|
||||
}
|
||||
}
|
||||
|
||||
// TestSourceDetailBaseURL_ForwardedProtoSpellings is the
|
||||
// regression test for the entrypoint URL an operator pastes into
|
||||
// the sending system: a header spelling that used to land in the
|
||||
// scheme verbatim produced a URL no sender could deliver to.
|
||||
func TestSourceDetailBaseURL_ForwardedProtoSpellings(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
const host = "hooks.example.com"
|
||||
|
||||
fixture := newBaseURLFixture(t)
|
||||
|
||||
for _, tc := range baseURLCases() {
|
||||
t.Run(tc.name, func(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
assert.Equal(
|
||||
t,
|
||||
tc.scheme+"://"+host+"/webhook/"+fixture.path,
|
||||
fixture.entrypointURL(
|
||||
t, host, forwardedProto(tc.header),
|
||||
),
|
||||
"X-Forwarded-Proto %q: %s", tc.header, tc.why,
|
||||
)
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// TestSourceDetailBaseURL_DirectTLSBeatsPlaintextHeader pins the
|
||||
// precedence the old code had backwards: it let any present
|
||||
// header overwrite what the connection itself proved, so a
|
||||
// direct-TLS request behind a proxy reporting http rendered an
|
||||
// http URL.
|
||||
func TestSourceDetailBaseURL_DirectTLSBeatsPlaintextHeader(
|
||||
t *testing.T,
|
||||
) {
|
||||
t.Parallel()
|
||||
|
||||
const host = "hooks.example.com"
|
||||
|
||||
fixture := newBaseURLFixture(t)
|
||||
|
||||
got := fixture.entrypointURL(t, host, func(r *http.Request) {
|
||||
r.TLS = &tls.ConnectionState{}
|
||||
r.Header.Set("X-Forwarded-Proto", "http")
|
||||
})
|
||||
|
||||
assert.Equal(
|
||||
t,
|
||||
"https://"+host+"/webhook/"+fixture.path,
|
||||
got,
|
||||
"a connection this process terminated with TLS "+
|
||||
"outranks a header claiming plaintext",
|
||||
)
|
||||
}
|
||||
|
||||
// TestSourceDetailBaseURL_KeepsHostAuthority pins the host half
|
||||
// of the URL: it is taken from the request unchanged, so the
|
||||
// deployments that do not sit on port 443 still get a URL that
|
||||
// works. Constraining the host would break exactly these.
|
||||
func TestSourceDetailBaseURL_KeepsHostAuthority(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
fixture := newBaseURLFixture(t)
|
||||
|
||||
hosts := []string{
|
||||
"hooks.example.com:8443",
|
||||
"[2001:db8::1]:8443",
|
||||
"internal-host",
|
||||
}
|
||||
|
||||
for _, host := range hosts {
|
||||
t.Run(host, func(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
assert.Equal(
|
||||
t,
|
||||
"https://"+host+"/webhook/"+fixture.path,
|
||||
fixture.entrypointURL(
|
||||
t, host, forwardedProto("HTTPS"),
|
||||
),
|
||||
"the authority must survive verbatim, port and all",
|
||||
)
|
||||
})
|
||||
}
|
||||
}
|
||||
@@ -13,6 +13,7 @@ import (
|
||||
"gorm.io/gorm"
|
||||
"sneak.berlin/go/webhooker/internal/database"
|
||||
"sneak.berlin/go/webhooker/internal/delivery"
|
||||
"sneak.berlin/go/webhooker/internal/reqtls"
|
||||
)
|
||||
|
||||
// WebhookListItem holds data for the webhook list view.
|
||||
@@ -427,16 +428,16 @@ func (h *Handlers) renderSourceDetail(
|
||||
}
|
||||
}
|
||||
|
||||
host := r.Host
|
||||
scheme := "https"
|
||||
|
||||
if r.TLS == nil {
|
||||
scheme = "http"
|
||||
scheme := "http"
|
||||
if reqtls.IsTLS(r) {
|
||||
scheme = "https"
|
||||
}
|
||||
|
||||
if fwdProto := r.Header.Get("X-Forwarded-Proto"); fwdProto != "" {
|
||||
scheme = fwdProto
|
||||
}
|
||||
// The host is the client's Host header, unvalidated. It is
|
||||
// inert only because source_detail.html renders BaseURL as
|
||||
// text inside a <code> element; putting it in an href or any
|
||||
// other URL context needs it constrained first.
|
||||
baseURL := scheme + "://" + r.Host
|
||||
|
||||
// The template calls Webhook methods, which take pointer
|
||||
// receivers; html/template cannot address a value stored in a map.
|
||||
@@ -448,7 +449,7 @@ func (h *Handlers) renderSourceDetail(
|
||||
"Entrypoints": NewEntrypointViews(entrypoints),
|
||||
"Targets": delivery.NewTargetViews(targets),
|
||||
"Events": events,
|
||||
"BaseURL": scheme + "://" + host,
|
||||
"BaseURL": baseURL,
|
||||
}
|
||||
|
||||
h.renderTemplate(w, r, "source_detail.html", data)
|
||||
|
||||
336
internal/server/bind_address_test.go
Normal file
336
internal/server/bind_address_test.go
Normal file
@@ -0,0 +1,336 @@
|
||||
package server_test
|
||||
|
||||
import (
|
||||
"context"
|
||||
"net"
|
||||
"net/http"
|
||||
"strconv"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"github.com/stretchr/testify/assert"
|
||||
"github.com/stretchr/testify/require"
|
||||
"go.uber.org/fx"
|
||||
"sneak.berlin/go/webhooker/internal/config"
|
||||
"sneak.berlin/go/webhooker/internal/globals"
|
||||
"sneak.berlin/go/webhooker/internal/server"
|
||||
)
|
||||
|
||||
const (
|
||||
// loopbackV4 is the shipped BIND_ADDRESS default.
|
||||
loopbackV4 = "127.0.0.1"
|
||||
|
||||
// wildcardV4 is the value a container deployment must set,
|
||||
// where a loopback-bound process is unreachable from outside
|
||||
// its network namespace even with a published port.
|
||||
wildcardV4 = "0.0.0.0"
|
||||
|
||||
// unavailableAddr is a TEST-NET-1 address (RFC 5737). It is a
|
||||
// well-formed literal that no host is assigned, so binding it
|
||||
// fails with EADDRNOTAVAIL rather than succeeding somewhere
|
||||
// unexpected.
|
||||
unavailableAddr = "192.0.2.1"
|
||||
|
||||
// listenReadyTimeout bounds the wait for the listener to accept
|
||||
// connections. The bind itself is immediate; this only covers
|
||||
// goroutine scheduling.
|
||||
listenReadyTimeout = 3 * time.Second
|
||||
|
||||
// listenPollInterval is how often the readiness wait retries.
|
||||
listenPollInterval = 10 * time.Millisecond
|
||||
|
||||
// dialTimeout bounds a single connection attempt in these
|
||||
// tests. Everything dialled here is on this host, so a dial
|
||||
// that is not answered immediately is a failure, not slowness.
|
||||
dialTimeout = time.Second
|
||||
)
|
||||
|
||||
// freePort returns a TCP port that is free on every local address at
|
||||
// the moment it returns, by taking one on the wildcard and releasing
|
||||
// it. The window between release and re-bind is the standard one
|
||||
// every "pick a free port" helper carries.
|
||||
func freePort(t *testing.T) int {
|
||||
t.Helper()
|
||||
|
||||
var listenCfg net.ListenConfig
|
||||
|
||||
l, err := listenCfg.Listen(t.Context(), "tcp", "0.0.0.0:0")
|
||||
require.NoError(t, err)
|
||||
|
||||
addr, ok := l.Addr().(*net.TCPAddr)
|
||||
require.True(t, ok, "listener is not TCP")
|
||||
require.NoError(t, l.Close())
|
||||
|
||||
return addr.Port
|
||||
}
|
||||
|
||||
// otherLocalAddr returns a local IPv4 address that is not
|
||||
// loopbackV4, or skips the test when the host has none.
|
||||
//
|
||||
// The bind-address tests need a second address of this host to stand
|
||||
// in for "another interface": what a wildcard bind claims and a
|
||||
// loopback bind does not. 127.0.0.2 is that address on Linux, where
|
||||
// the whole 127.0.0.0/8 is local; elsewhere an interface address is
|
||||
// used instead. Each candidate is proven bindable before it is
|
||||
// returned, so a host that offers neither skips rather than fails on
|
||||
// something that was never about the code under test.
|
||||
func otherLocalAddr(t *testing.T) string {
|
||||
t.Helper()
|
||||
|
||||
candidates := []string{"127.0.0.2"}
|
||||
|
||||
ifaceAddrs, err := net.InterfaceAddrs()
|
||||
require.NoError(t, err)
|
||||
|
||||
for _, a := range ifaceAddrs {
|
||||
ipNet, ok := a.(*net.IPNet)
|
||||
if !ok {
|
||||
continue
|
||||
}
|
||||
|
||||
ip4 := ipNet.IP.To4()
|
||||
if ip4 == nil || ip4.String() == loopbackV4 {
|
||||
continue
|
||||
}
|
||||
|
||||
candidates = append(candidates, ip4.String())
|
||||
}
|
||||
|
||||
var listenCfg net.ListenConfig
|
||||
|
||||
for _, candidate := range candidates {
|
||||
l, listenErr := listenCfg.Listen(
|
||||
t.Context(), "tcp", net.JoinHostPort(candidate, "0"),
|
||||
)
|
||||
if listenErr != nil {
|
||||
continue
|
||||
}
|
||||
|
||||
require.NoError(t, l.Close())
|
||||
|
||||
return candidate
|
||||
}
|
||||
|
||||
t.Skip("host has no second local IPv4 address to bind")
|
||||
|
||||
return ""
|
||||
}
|
||||
|
||||
// startBoundServer starts the wired app with the given bind address
|
||||
// on a free port and returns that port. The app is stopped on
|
||||
// cleanup.
|
||||
func startBoundServer(t *testing.T, bindAddress string) int {
|
||||
t.Helper()
|
||||
|
||||
port := freePort(t)
|
||||
|
||||
env := newTestEnv(t)
|
||||
env.cfg.BindAddress = bindAddress
|
||||
env.cfg.Port = port
|
||||
|
||||
app := fx.New(
|
||||
fx.NopLogger,
|
||||
fx.Supply(env.log, env.cfg, env.mw, env.hnd),
|
||||
fx.Provide(globals.New, server.New),
|
||||
fx.Invoke(func(*server.Server) {}),
|
||||
)
|
||||
|
||||
startCtx, cancelStart := context.WithTimeout(
|
||||
context.Background(), lifecycleTimeout,
|
||||
)
|
||||
defer cancelStart()
|
||||
|
||||
require.NoError(t, app.Start(startCtx))
|
||||
|
||||
t.Cleanup(func() {
|
||||
stopCtx, cancelStop := context.WithTimeout(
|
||||
context.Background(), lifecycleTimeout,
|
||||
)
|
||||
defer cancelStop()
|
||||
|
||||
require.NoError(t, app.Stop(stopCtx))
|
||||
})
|
||||
|
||||
return port
|
||||
}
|
||||
|
||||
// dialable reports whether a TCP connection to addr succeeds.
|
||||
func dialable(ctx context.Context, addr string) bool {
|
||||
dialer := net.Dialer{Timeout: dialTimeout}
|
||||
|
||||
conn, err := dialer.DialContext(ctx, "tcp", addr)
|
||||
if err != nil {
|
||||
return false
|
||||
}
|
||||
|
||||
_ = conn.Close()
|
||||
|
||||
return true
|
||||
}
|
||||
|
||||
// requireDialable waits for addr to accept connections, failing the
|
||||
// test if it never does.
|
||||
func requireDialable(t *testing.T, addr string) {
|
||||
t.Helper()
|
||||
|
||||
deadline := time.Now().Add(listenReadyTimeout)
|
||||
for time.Now().Before(deadline) {
|
||||
if dialable(t.Context(), addr) {
|
||||
return
|
||||
}
|
||||
|
||||
time.Sleep(listenPollInterval)
|
||||
}
|
||||
|
||||
t.Fatalf("nothing accepted connections on %s", addr)
|
||||
}
|
||||
|
||||
// TestListenAddr pins how BindAddress and Port are rendered into the
|
||||
// listen address.
|
||||
//
|
||||
// The defect this covers was a bare fmt.Sprintf(":%d", port), which
|
||||
// binds every interface with no way to say otherwise. The IPv6 rows
|
||||
// are here because an unbracketed IPv6 host would produce an address
|
||||
// net.Listen rejects, turning a valid configuration into a startup
|
||||
// failure.
|
||||
func TestListenAddr(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
tests := []struct {
|
||||
name string
|
||||
bindAddress string
|
||||
port int
|
||||
expected string
|
||||
}{
|
||||
{
|
||||
name: "loopback default",
|
||||
bindAddress: loopbackV4,
|
||||
port: 8080,
|
||||
expected: "127.0.0.1:8080",
|
||||
},
|
||||
{
|
||||
name: "ipv4 wildcard",
|
||||
bindAddress: wildcardV4,
|
||||
port: 8080,
|
||||
expected: "0.0.0.0:8080",
|
||||
},
|
||||
{
|
||||
name: "ipv6 wildcard is bracketed",
|
||||
bindAddress: "::",
|
||||
port: 8080,
|
||||
expected: "[::]:8080",
|
||||
},
|
||||
{
|
||||
name: "ipv6 literal is bracketed",
|
||||
bindAddress: "2001:db8::5",
|
||||
port: 9001,
|
||||
expected: "[2001:db8::5]:9001",
|
||||
},
|
||||
}
|
||||
|
||||
for _, tt := range tests {
|
||||
t.Run(tt.name, func(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
assert.Equal(t, tt.expected, server.ListenAddrForTest(
|
||||
&config.Config{
|
||||
BindAddress: tt.bindAddress,
|
||||
Port: tt.port,
|
||||
},
|
||||
))
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// TestBindAddress_LoopbackIsNotOnOtherAddresses proves the fix end to
|
||||
// end: with BIND_ADDRESS at its loopback default, the cleartext
|
||||
// listener answers on loopback and has not claimed any other address
|
||||
// of this host.
|
||||
//
|
||||
// The second address is proven free by binding it on the same port
|
||||
// while the server runs. That is the assertion that fails against the
|
||||
// old wildcard bind — a wildcard listener owns the port on every
|
||||
// address, so this bind would return EADDRINUSE. Dialling from
|
||||
// another machine is what the operator cares about, and this is the
|
||||
// in-process form of it: the socket the remote host would connect to
|
||||
// does not exist.
|
||||
func TestBindAddress_LoopbackIsNotOnOtherAddresses(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
other := otherLocalAddr(t)
|
||||
port := startBoundServer(t, loopbackV4)
|
||||
|
||||
// Positive control: the service really is up and serving.
|
||||
requireDialable(t, net.JoinHostPort(loopbackV4, strconv.Itoa(port)))
|
||||
|
||||
var listenCfg net.ListenConfig
|
||||
|
||||
l, err := listenCfg.Listen(
|
||||
t.Context(), "tcp",
|
||||
net.JoinHostPort(other, strconv.Itoa(port)),
|
||||
)
|
||||
require.NoError(
|
||||
t, err,
|
||||
"port %d on %s is taken while bound to %s: the listener "+
|
||||
"claimed more than its configured address",
|
||||
port, other, loopbackV4,
|
||||
)
|
||||
|
||||
require.NoError(t, l.Close())
|
||||
}
|
||||
|
||||
// TestBindAddress_WildcardReachesOtherAddresses is the counterpart:
|
||||
// the value a container deployment sets does reach the addresses the
|
||||
// default withholds. Without this, a loopback-only bind would pass
|
||||
// the test above by never listening at all.
|
||||
func TestBindAddress_WildcardReachesOtherAddresses(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
other := otherLocalAddr(t)
|
||||
port := startBoundServer(t, wildcardV4)
|
||||
|
||||
requireDialable(t, net.JoinHostPort(other, strconv.Itoa(port)))
|
||||
}
|
||||
|
||||
// TestBindAddress_ServesRequestsOnConfiguredAddress proves the bound
|
||||
// listener serves the application rather than merely accepting TCP,
|
||||
// so a bind address that is honoured cannot be mistaken for one that
|
||||
// is honoured and broken.
|
||||
func TestBindAddress_ServesRequestsOnConfiguredAddress(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
port := startBoundServer(t, loopbackV4)
|
||||
addr := net.JoinHostPort(loopbackV4, strconv.Itoa(port))
|
||||
requireDialable(t, addr)
|
||||
|
||||
req, err := http.NewRequestWithContext(
|
||||
t.Context(), http.MethodGet,
|
||||
"http://"+addr+"/.well-known/healthcheck", nil,
|
||||
)
|
||||
require.NoError(t, err)
|
||||
|
||||
client := &http.Client{Timeout: dialTimeout}
|
||||
|
||||
resp, err := client.Do(req)
|
||||
require.NoError(t, err)
|
||||
|
||||
defer func() { _ = resp.Body.Close() }()
|
||||
|
||||
assert.Equal(t, http.StatusOK, resp.StatusCode)
|
||||
}
|
||||
|
||||
// TestBindAddress_UnavailableAddressShutsDownTheApp covers the half
|
||||
// of the fail-loud rule that configuration parsing cannot reach. A
|
||||
// syntactically valid address that is not assigned to this host
|
||||
// parses fine and fails at bind time, after fx has already reported
|
||||
// RUNNING. It must end the process non-zero rather than leave it
|
||||
// alive with nothing listening.
|
||||
func TestBindAddress_UnavailableAddressShutsDownTheApp(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
env := newTestEnv(t)
|
||||
env.cfg.BindAddress = unavailableAddr
|
||||
env.cfg.Port = freePort(t)
|
||||
|
||||
requireListenFailureExit(t, env)
|
||||
}
|
||||
91
internal/server/early_shutdown_test.go
Normal file
91
internal/server/early_shutdown_test.go
Normal file
@@ -0,0 +1,91 @@
|
||||
package server_test
|
||||
|
||||
import (
|
||||
"context"
|
||||
"testing"
|
||||
|
||||
"github.com/stretchr/testify/require"
|
||||
"go.uber.org/fx"
|
||||
"sneak.berlin/go/webhooker/internal/globals"
|
||||
"sneak.berlin/go/webhooker/internal/server"
|
||||
)
|
||||
|
||||
// earlyStopIterations is how many start/stop cycles the race test
|
||||
// runs. The window it aims at is the gap between the OnStart hook
|
||||
// returning and the serving goroutine reaching its first field
|
||||
// access, which is microseconds wide. The race detector reports an
|
||||
// unsynchronised pair whenever it observes one, but it has to observe
|
||||
// one, so a single cycle can miss purely on scheduling. Repetition
|
||||
// makes the observation reliable; the collaborators are built once,
|
||||
// so the cycles themselves are cheap.
|
||||
const earlyStopIterations = 25
|
||||
|
||||
// TestEarlyShutdown_NoPanicAndNoRace stops the application
|
||||
// immediately after starting it, before the serving goroutine has
|
||||
// necessarily run at all.
|
||||
//
|
||||
// Two defects live in that window. The OnStart hook returns as soon
|
||||
// as it has spawned the serving goroutine, so fx runs the stop
|
||||
// sequence against a Server whose serving goroutine may not have
|
||||
// executed a single line. cleanShutdown called Shutdown on an
|
||||
// httpServer that goroutine was supposed to assign, which was a nil
|
||||
// dereference on an early SIGTERM; and it read httpServer and
|
||||
// sentryEnabled with nothing ordering those reads against the
|
||||
// goroutine's writes, which is a data race that only surfaces once
|
||||
// something both starts and stops the server. Nothing did before this
|
||||
// test: the listen-failure test never binds, and the router tests
|
||||
// bypass the lifecycle entirely.
|
||||
//
|
||||
// httpServer is now built in New, on the constructing goroutine, so
|
||||
// it is written before any hook exists and can never be nil.
|
||||
// sentryEnabled is atomic. This test is what catches either one
|
||||
// coming back — under -race, which is how the suite runs.
|
||||
func TestEarlyShutdown_NoPanicAndNoRace(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
// Built once: the collaborators are not what is under test, and
|
||||
// standing up a database per iteration would make repetition too
|
||||
// expensive to be worth having.
|
||||
env := newTestEnv(t)
|
||||
env.cfg.BindAddress = loopbackV4
|
||||
|
||||
for range earlyStopIterations {
|
||||
requireStartStopIsClean(t, env)
|
||||
}
|
||||
}
|
||||
|
||||
// requireStartStopIsClean runs one start/stop cycle with no wait in
|
||||
// between, failing the test if either half errors.
|
||||
//
|
||||
// Each cycle gets a fresh fx app, so the Server under test is
|
||||
// constructed anew every time — that construction is where the
|
||||
// httpServer write now happens, and reusing one Server would test it
|
||||
// only once.
|
||||
func requireStartStopIsClean(t *testing.T, env *testEnv) {
|
||||
t.Helper()
|
||||
|
||||
env.cfg.Port = freePort(t)
|
||||
|
||||
app := fx.New(
|
||||
fx.NopLogger,
|
||||
fx.Supply(env.log, env.cfg, env.mw, env.hnd),
|
||||
fx.Provide(globals.New, server.New),
|
||||
fx.Invoke(func(*server.Server) {}),
|
||||
)
|
||||
|
||||
startCtx, cancelStart := context.WithTimeout(
|
||||
context.Background(), lifecycleTimeout,
|
||||
)
|
||||
defer cancelStart()
|
||||
|
||||
require.NoError(t, app.Start(startCtx))
|
||||
|
||||
// No sleep and no readiness wait: stopping while the serving
|
||||
// goroutine is still in flight is the whole point.
|
||||
stopCtx, cancelStop := context.WithTimeout(
|
||||
context.Background(), lifecycleTimeout,
|
||||
)
|
||||
defer cancelStop()
|
||||
|
||||
require.NoError(t, app.Stop(stopCtx))
|
||||
}
|
||||
@@ -55,6 +55,16 @@ func NewRouterForTest(
|
||||
return s.router
|
||||
}
|
||||
|
||||
// ListenAddrForTest exposes the address the HTTP listener binds for
|
||||
// a given Config, so the rendering of host and port — IPv6
|
||||
// bracketing above all — can be pinned without standing up a
|
||||
// listener.
|
||||
func ListenAddrForTest(cfg *config.Config) string {
|
||||
s := &Server{params: ServerParams{Config: cfg}}
|
||||
|
||||
return s.listenAddr()
|
||||
}
|
||||
|
||||
// ProbePattern is the route NewRouterWithProbeForTest adds to the
|
||||
// production route tree.
|
||||
const ProbePattern = "/probe"
|
||||
@@ -84,8 +94,8 @@ func NewRouterWithProbeForTest(
|
||||
mw: mw,
|
||||
h: h,
|
||||
params: ServerParams{Config: cfg},
|
||||
sentryEnabled: sentryEnabled,
|
||||
}
|
||||
s.sentryEnabled.Store(sentryEnabled)
|
||||
s.SetupRoutes()
|
||||
s.router.Handle(ProbePattern, probe)
|
||||
|
||||
|
||||
@@ -2,8 +2,9 @@ package server
|
||||
|
||||
import (
|
||||
"errors"
|
||||
"fmt"
|
||||
"net"
|
||||
"net/http"
|
||||
"strconv"
|
||||
"time"
|
||||
)
|
||||
|
||||
@@ -24,26 +25,52 @@ const (
|
||||
httpMaxHeaderBytes = 1 << 20
|
||||
)
|
||||
|
||||
func (s *Server) serveUntilShutdown() {
|
||||
listenAddr := fmt.Sprintf(":%d", s.params.Config.Port)
|
||||
s.httpServer = &http.Server{
|
||||
Addr: listenAddr,
|
||||
// listenAddr renders the address the HTTP listener binds.
|
||||
//
|
||||
// The host half is always present: an empty host would be the
|
||||
// wildcard, and the whole point of BIND_ADDRESS is that binding every
|
||||
// interface is a choice the operator makes rather than one the
|
||||
// process makes for them. Config guarantees a literal, so
|
||||
// JoinHostPort's bracketing is enough to make IPv6 well formed.
|
||||
func (s *Server) listenAddr() string {
|
||||
return net.JoinHostPort(
|
||||
s.params.Config.BindAddress,
|
||||
strconv.Itoa(s.params.Config.Port),
|
||||
)
|
||||
}
|
||||
|
||||
// newHTTPServer builds the HTTP server for this Server's
|
||||
// configuration.
|
||||
//
|
||||
// It is called from New, on the constructing goroutine, rather than
|
||||
// from the serving goroutine that used to assign s.httpServer
|
||||
// directly. Two goroutines reach that field — the serving goroutine
|
||||
// and the fx stop hook, which calls Shutdown on it — with nothing
|
||||
// ordering them. Constructing it during New puts the write before
|
||||
// every hook fx will later run, which both removes the race and rules
|
||||
// out the nil dereference a stop that arrived before the serving
|
||||
// goroutine had run would have caused.
|
||||
func (s *Server) newHTTPServer() *http.Server {
|
||||
return &http.Server{
|
||||
Addr: s.listenAddr(),
|
||||
ReadTimeout: httpReadTimeout,
|
||||
WriteTimeout: httpWriteTimeout,
|
||||
MaxHeaderBytes: httpMaxHeaderBytes,
|
||||
Handler: s,
|
||||
}
|
||||
}
|
||||
|
||||
func (s *Server) serveUntilShutdown() {
|
||||
// add routes
|
||||
// this does any necessary setup in each handler
|
||||
s.SetupRoutes()
|
||||
|
||||
s.log.Info("http begin listen", "listenaddr", listenAddr)
|
||||
s.log.Info("http begin listen", "listenaddr", s.httpServer.Addr)
|
||||
|
||||
err := s.httpServer.ListenAndServe()
|
||||
if err != nil && !errors.Is(err, http.ErrServerClosed) {
|
||||
s.log.Error("listen error", "error", err)
|
||||
s.shutdownOnListenFailure()
|
||||
s.shutdownWithFailure()
|
||||
}
|
||||
}
|
||||
|
||||
|
||||
@@ -36,9 +36,9 @@ const lifecycleTimeout = 15 * time.Second
|
||||
//
|
||||
// The port is occupied by a listener this test holds open, on a
|
||||
// kernel-chosen port, so the failure is the real EADDRINUSE the
|
||||
// operator hits when a second instance starts. Loopback is enough to
|
||||
// collide with the server's wildcard bind: a listening socket on a
|
||||
// specific address blocks the wildcard from claiming the same port.
|
||||
// operator hits when a second instance starts. The server is pointed
|
||||
// at the same loopback address, so the collision is a direct one on
|
||||
// the exact address it asks the kernel for.
|
||||
func TestListenFailure_ShutsDownTheApp(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
@@ -55,11 +55,27 @@ func TestListenFailure_ShutsDownTheApp(t *testing.T) {
|
||||
require.True(t, ok, "listener is not TCP")
|
||||
|
||||
// The collaborators come from the wired graph rather than stubs,
|
||||
// so the Server under test is the one that ships. Only the port
|
||||
// is test-specific.
|
||||
// so the Server under test is the one that ships. Only the
|
||||
// listen address is test-specific.
|
||||
env := newTestEnv(t)
|
||||
env.cfg.BindAddress = loopbackV4
|
||||
env.cfg.Port = addr.Port
|
||||
|
||||
requireListenFailureExit(t, env)
|
||||
}
|
||||
|
||||
// requireListenFailureExit starts the wired app over env and asserts
|
||||
// that it gives up on its own with the listen-failure status, then
|
||||
// completes its stop sequence.
|
||||
//
|
||||
// Two different listen failures share it — a port already in use and
|
||||
// an address that is not on this host — because what has to hold for
|
||||
// both is the same: the failure is discovered after fx has already
|
||||
// reported RUNNING, so the only thing that can turn it into a visible
|
||||
// exit is the shutdown path under test.
|
||||
func requireListenFailureExit(t *testing.T, env *testEnv) {
|
||||
t.Helper()
|
||||
|
||||
app := fx.New(
|
||||
fx.NopLogger,
|
||||
fx.Supply(env.log, env.cfg, env.mw, env.hnd),
|
||||
@@ -77,7 +93,7 @@ func TestListenFailure_ShutsDownTheApp(t *testing.T) {
|
||||
select {
|
||||
case sig := <-app.Wait():
|
||||
require.Equal(
|
||||
t, server.ListenFailureExitCode, sig.ExitCode,
|
||||
t, server.StartupFailureExitCode, sig.ExitCode,
|
||||
"listen failure must exit non-zero",
|
||||
)
|
||||
case <-time.After(listenFailureDeadline):
|
||||
|
||||
@@ -81,7 +81,7 @@ func (s *Server) setupGlobalMiddleware() {
|
||||
// Sentry error reporting (if SENTRY_DSN is set). Repanic is
|
||||
// true so panics still bubble up to the Recoverer middleware
|
||||
// registered immediately above.
|
||||
if s.sentryEnabled {
|
||||
if s.sentryEnabled.Load() {
|
||||
sentryHandler := sentryhttp.New(sentryhttp.Options{
|
||||
Repanic: true,
|
||||
})
|
||||
|
||||
99
internal/server/sentry_failure_test.go
Normal file
99
internal/server/sentry_failure_test.go
Normal file
@@ -0,0 +1,99 @@
|
||||
package server_test
|
||||
|
||||
import (
|
||||
"context"
|
||||
"net"
|
||||
"strconv"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"github.com/stretchr/testify/require"
|
||||
"go.uber.org/fx"
|
||||
"sneak.berlin/go/webhooker/internal/config"
|
||||
"sneak.berlin/go/webhooker/internal/globals"
|
||||
"sneak.berlin/go/webhooker/internal/server"
|
||||
)
|
||||
|
||||
// TestSentryInitFailure_ShutsDownTheApp pins that error reporting
|
||||
// which is configured and cannot be started ends the application
|
||||
// instead of serving without it.
|
||||
//
|
||||
// The measured defect logged `sentry init failure` and kept running,
|
||||
// so the deployment served traffic with reporting off while every
|
||||
// other signal — SENTRY_DSN still set, the startup summary's own
|
||||
// field — said it was on. Nothing later in the process can notice
|
||||
// that reports are going nowhere, which is why this exits rather than
|
||||
// degrades.
|
||||
//
|
||||
// The DSN is placed on a hand-built Config, which is the only way to
|
||||
// reach this branch at all: loadFromEnv now parses SENTRY_DSN with
|
||||
// sentry.NewDsn, the same call sentry.Init makes, so a DSN that
|
||||
// survives configuration cannot fail initialisation in the SDK
|
||||
// version this pins. The branch stays because that is a property of
|
||||
// the SDK's current implementation rather than of its contract.
|
||||
func TestSentryInitFailure_ShutsDownTheApp(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
port := freePort(t)
|
||||
|
||||
env := newTestEnvWithConfig(t, &config.Config{
|
||||
DataDir: t.TempDir(),
|
||||
Environment: config.EnvironmentDev,
|
||||
BindAddress: loopbackV4,
|
||||
Port: port,
|
||||
SentryDSN: "not-a-dsn",
|
||||
})
|
||||
|
||||
app := fx.New(
|
||||
fx.NopLogger,
|
||||
fx.Supply(env.log, env.cfg, env.mw, env.hnd),
|
||||
fx.Provide(globals.New, server.New),
|
||||
fx.Invoke(func(*server.Server) {}),
|
||||
)
|
||||
|
||||
startCtx, cancelStart := context.WithTimeout(
|
||||
context.Background(), lifecycleTimeout,
|
||||
)
|
||||
defer cancelStart()
|
||||
|
||||
require.NoError(t, app.Start(startCtx))
|
||||
|
||||
select {
|
||||
case sig := <-app.Wait():
|
||||
require.Equal(
|
||||
t, server.StartupFailureExitCode, sig.ExitCode,
|
||||
"a sentry failure must exit non-zero",
|
||||
)
|
||||
case <-time.After(listenFailureDeadline):
|
||||
t.Fatal("a sentry failure left the app running")
|
||||
}
|
||||
|
||||
// The stop sequence still has to complete: the failure must reach
|
||||
// shutdown through fx rather than around it.
|
||||
stopCtx, cancelStop := context.WithTimeout(
|
||||
context.Background(), lifecycleTimeout,
|
||||
)
|
||||
defer cancelStop()
|
||||
|
||||
require.NoError(t, app.Stop(stopCtx))
|
||||
|
||||
// And it must give up before it listens. A process that bound the
|
||||
// port and then exited would have accepted requests it could not
|
||||
// report on, which is the state under test in miniature.
|
||||
requireBindable(t, port)
|
||||
}
|
||||
|
||||
// requireBindable asserts that the port is free, which it is only if
|
||||
// the server under test never claimed it.
|
||||
func requireBindable(t *testing.T, port int) {
|
||||
t.Helper()
|
||||
|
||||
var listenCfg net.ListenConfig
|
||||
|
||||
listener, err := listenCfg.Listen(
|
||||
t.Context(), "tcp",
|
||||
net.JoinHostPort(loopbackV4, strconv.Itoa(port)),
|
||||
)
|
||||
require.NoError(t, err, "the server bound a port it then gave up")
|
||||
require.NoError(t, listener.Close())
|
||||
}
|
||||
@@ -9,6 +9,7 @@ import (
|
||||
"net/http"
|
||||
"os"
|
||||
"os/signal"
|
||||
"sync/atomic"
|
||||
"syscall"
|
||||
"time"
|
||||
|
||||
@@ -50,12 +51,13 @@ const (
|
||||
minSentryFlush = 250 * time.Millisecond
|
||||
)
|
||||
|
||||
// ListenFailureExitCode is the status the process exits with when the
|
||||
// HTTP listener cannot be established, or dies for a reason other
|
||||
// than a requested shutdown. It must stay non-zero: systemd
|
||||
// `Restart=on-failure` and Docker's restart policies key off it, and a
|
||||
// zero exit would read as a deliberate stop.
|
||||
const ListenFailureExitCode = 1
|
||||
// StartupFailureExitCode is the status the process exits with when
|
||||
// the serving goroutine gives up: the HTTP listener cannot be
|
||||
// established or dies for a reason other than a requested shutdown, or
|
||||
// error reporting is configured and cannot be started. It must stay
|
||||
// non-zero: systemd `Restart=on-failure` and Docker's restart policies
|
||||
// key off it, and a zero exit would read as a deliberate stop.
|
||||
const StartupFailureExitCode = 1
|
||||
|
||||
// SentryFlushBudget reports how long the Sentry flush may run when
|
||||
// remaining is the time left on the fx stop context after the HTTP
|
||||
@@ -89,7 +91,14 @@ type ServerParams struct {
|
||||
// graceful shutdown.
|
||||
type Server struct {
|
||||
startupTime time.Time
|
||||
sentryEnabled bool
|
||||
|
||||
// sentryEnabled is written by the serving goroutine, in
|
||||
// enableSentry, and read by the fx stop hook in cleanShutdown.
|
||||
// Nothing orders those two: the OnStart hook returns as soon as
|
||||
// the goroutine is spawned, so a stop can be running while
|
||||
// enableSentry is still deciding. It is atomic to supply the
|
||||
// edge the goroutines do not.
|
||||
sentryEnabled atomic.Bool
|
||||
log *slog.Logger
|
||||
cancelFunc context.CancelFunc
|
||||
httpServer *http.Server
|
||||
@@ -107,6 +116,7 @@ func New(lc fx.Lifecycle, params ServerParams) (*Server, error) {
|
||||
s.mw = params.Middleware
|
||||
s.h = params.Handlers
|
||||
s.log = params.Logger.Get()
|
||||
s.httpServer = s.newHTTPServer()
|
||||
|
||||
lc.Append(fx.Hook{
|
||||
OnStart: func(_ context.Context) error {
|
||||
@@ -126,11 +136,25 @@ func New(lc fx.Lifecycle, params ServerParams) (*Server, error) {
|
||||
}
|
||||
|
||||
// Run configures Sentry and starts serving HTTP requests.
|
||||
//
|
||||
// A Sentry failure ends the application instead of listening. It runs
|
||||
// before the listener rather than after it so that the process never
|
||||
// binds a port it is about to give up.
|
||||
func (s *Server) Run() {
|
||||
s.configure()
|
||||
|
||||
// logging before sentry, because sentry logs
|
||||
s.enableSentry()
|
||||
err := s.enableSentry()
|
||||
if err != nil {
|
||||
s.log.Error(
|
||||
"SENTRY_DSN is set but error reporting could not be "+
|
||||
"started; refusing to serve with it off",
|
||||
"error", err,
|
||||
)
|
||||
s.shutdownWithFailure()
|
||||
|
||||
return
|
||||
}
|
||||
|
||||
s.serve()
|
||||
}
|
||||
@@ -141,11 +165,23 @@ func (s *Server) MaintenanceMode() bool {
|
||||
return s.params.Config.MaintenanceMode
|
||||
}
|
||||
|
||||
func (s *Server) enableSentry() {
|
||||
s.sentryEnabled = false
|
||||
// enableSentry initialises the Sentry SDK when error reporting is
|
||||
// configured, and reports the failure when it is configured and cannot
|
||||
// be initialised. A DSN that is not set is not a failure: reporting
|
||||
// stays off and the server starts normally.
|
||||
//
|
||||
// There is no fallback to running with reporting off. An operator who
|
||||
// set SENTRY_DSN asked for failures to be visible, and serving traffic
|
||||
// with reporting quietly off is the one state nothing can ever tell
|
||||
// them about — the DSN is still set, so every later signal says it is
|
||||
// on. Config already refused a DSN the SDK cannot parse, which is what
|
||||
// a typo produces, so reaching this branch means the SDK refused
|
||||
// something that parsed: not a condition to guess at either.
|
||||
func (s *Server) enableSentry() error {
|
||||
s.sentryEnabled.Store(false)
|
||||
|
||||
if s.params.Config.SentryDSN == "" {
|
||||
return
|
||||
if !s.params.Config.SentryEnabled() {
|
||||
return nil
|
||||
}
|
||||
|
||||
err := sentry.Init(sentryClientOptions(
|
||||
@@ -157,19 +193,19 @@ func (s *Server) enableSentry() {
|
||||
),
|
||||
))
|
||||
if err != nil {
|
||||
s.log.Error("sentry init failure", "error", err)
|
||||
// Don't use fatal since we still want the service to run
|
||||
return
|
||||
return fmt.Errorf("initialising sentry: %w", err)
|
||||
}
|
||||
|
||||
s.log.Info("sentry error reporting activated")
|
||||
s.sentryEnabled = true
|
||||
s.sentryEnabled.Store(true)
|
||||
|
||||
return nil
|
||||
}
|
||||
|
||||
// serve installs the signal watcher, starts the listener and blocks
|
||||
// until the server's context is cancelled. The process exit status is
|
||||
// fx's to decide — from a signal, or from the code
|
||||
// shutdownOnListenFailure hands the Shutdowner — so this reports
|
||||
// shutdownWithFailure hands the Shutdowner — so this reports
|
||||
// nothing back to its caller.
|
||||
func (s *Server) serve() {
|
||||
ctx, cancelFunc := context.WithCancel(context.Background())
|
||||
@@ -199,20 +235,24 @@ func (s *Server) serve() {
|
||||
// Do not call cleanShutdown() here to avoid double invocation.
|
||||
}
|
||||
|
||||
// shutdownOnListenFailure ends the application after the HTTP
|
||||
// listener failed. The fx OnStart hook returns as soon as the serving
|
||||
// goroutine is spawned, so nothing downstream of it ever learns that
|
||||
// the listen failed: fx reports RUNNING and the process sits alive
|
||||
// with nothing bound, which is invisible to systemd and Docker
|
||||
// restart policies. Asking the Shutdowner to stop the app with a
|
||||
// non-zero code is what turns that into a visible failure.
|
||||
// shutdownWithFailure ends the application non-zero from the serving
|
||||
// goroutine. It is how anything on that goroutine fails fatally: the
|
||||
// fx OnStart hook returns as soon as the goroutine is spawned, so
|
||||
// nothing downstream of it ever learns that the goroutine gave up. fx
|
||||
// reports RUNNING and the process sits alive having done neither what
|
||||
// it was asked nor anything visible instead, which systemd and
|
||||
// Docker restart policies cannot see. Asking the Shutdowner to stop
|
||||
// the app with a non-zero code is what turns that into a visible
|
||||
// failure, and it is the whole of "fatal" here — no panic, no
|
||||
// os.Exit, and every stop hook still runs.
|
||||
//
|
||||
// The context cancel that follows only unwinds serve()'s own wait.
|
||||
// The shutdown itself runs through fx's normal stop sequence, so the
|
||||
// clean-shutdown drain in cleanShutdown is reached unchanged.
|
||||
func (s *Server) shutdownOnListenFailure() {
|
||||
// The context cancel that follows only unwinds serve()'s own wait,
|
||||
// and is skipped before serve has installed one. The shutdown itself
|
||||
// runs through fx's normal stop sequence, so the clean-shutdown drain
|
||||
// in cleanShutdown is reached unchanged.
|
||||
func (s *Server) shutdownWithFailure() {
|
||||
err := s.params.Shutdowner.Shutdown(
|
||||
fx.ExitCode(ListenFailureExitCode),
|
||||
fx.ExitCode(StartupFailureExitCode),
|
||||
)
|
||||
if err != nil {
|
||||
s.log.Error("shutdown request failed", "error", err)
|
||||
@@ -242,7 +282,7 @@ func (s *Server) cleanShutdown(ctx context.Context) {
|
||||
|
||||
s.cleanupForExit()
|
||||
|
||||
if s.sentryEnabled {
|
||||
if s.sentryEnabled.Load() {
|
||||
s.flushSentry(ctx)
|
||||
}
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user