Compare commits
33
Commits
5ce8fb57bc
..
prod
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
1647b43aa6 | ||
|
|
f7151f0168 | ||
|
|
3cc05a36eb | ||
|
|
7316f0a7e2 | ||
|
|
5f84d891cf | ||
|
|
9cf9cdd8eb | ||
|
|
f0adeafde3 | ||
|
|
f755c03110 | ||
|
|
4a724130ca | ||
|
|
d4f4ddf51f | ||
|
|
51580a2bc6 | ||
|
|
e0b211f960 | ||
|
|
978eb01b29 | ||
|
|
3cdab97930 | ||
|
|
83740b1de1 | ||
|
|
aeeeca5ea1 | ||
|
|
6ebac4fa71 | ||
|
|
237f131367 | ||
|
|
7ed1588443 | ||
|
|
b051821370 | ||
|
|
39afa69bfc | ||
|
|
888eaf526b | ||
|
|
251cb3d3d3 | ||
|
|
d61d9dc1c1 | ||
|
|
b0a011f6b4 | ||
|
|
5976a4a98f | ||
|
|
b2c9acdaa6 | ||
|
|
af3703d748 | ||
|
|
322d9a6d6b | ||
|
|
b9f7db6901 | ||
|
|
48cf93ec7e | ||
|
|
8d64259283 | ||
|
|
bde32d3ee6 |
+9
-2
@@ -88,7 +88,9 @@ RUN CGO_ENABLED=1 make build VERSION="$VERSION" GO_LDFLAGS='-extldflags "-static
|
||||
# alpine:3.21, 2026-03-17
|
||||
FROM alpine:3.21@sha256:c3f8e73fdb79deaebaa2037150150191b9dcbfba68b4a46d70103204c53f4709
|
||||
|
||||
RUN apk --no-cache add ca-certificates
|
||||
# su-exec 0.2-r3 (Alpine 3.21), 2026-09-29: the entrypoint runs the app
|
||||
# as webhooker with it.
|
||||
RUN apk --no-cache add ca-certificates su-exec=0.2-r3
|
||||
|
||||
# Create non-root user
|
||||
RUN addgroup -g 1000 -S webhooker && \
|
||||
@@ -99,13 +101,17 @@ WORKDIR /app
|
||||
# Copy binary from builder
|
||||
COPY --from=builder /build/bin/webhooker /app/webhooker
|
||||
|
||||
# Not under /app, which belongs to webhooker: this script runs as root.
|
||||
COPY deploy/docker-entrypoint.sh /usr/local/bin/docker-entrypoint.sh
|
||||
|
||||
# Create data directory for all SQLite databases (main app DB +
|
||||
# per-webhook event DBs). DATA_DIR defaults to /var/lib/webhooker.
|
||||
RUN mkdir -p /var/lib/webhooker
|
||||
|
||||
RUN chown -R webhooker:webhooker /app /var/lib/webhooker
|
||||
|
||||
USER webhooker
|
||||
# No USER: the entrypoint starts as root to make the data directory
|
||||
# webhooker's, then runs the app as webhooker.
|
||||
|
||||
EXPOSE 8080
|
||||
|
||||
@@ -124,4 +130,5 @@ ENV BIND_ADDRESS=0.0.0.0
|
||||
HEALTHCHECK --interval=30s --timeout=3s --start-period=5s --retries=3 \
|
||||
CMD wget --no-verbose --tries=1 --spider http://localhost:8080/.well-known/healthcheck || exit 1
|
||||
|
||||
ENTRYPOINT ["/usr/local/bin/docker-entrypoint.sh"]
|
||||
CMD ["/app/webhooker"]
|
||||
|
||||
@@ -7,6 +7,13 @@ services, durably stores them, and delivers them to configured targets
|
||||
with retry support, logging, and observability. Category: infrastructure
|
||||
/ web service. License: MIT.
|
||||
|
||||
Each entrypoint is a version 4 UUID served at `/webhook/{uuid}`, and
|
||||
that UUID is the entrypoint's only credential. webhooker does not use
|
||||
shared secrets, HMAC signatures or token headers on the receiver, and
|
||||
will not add them — read
|
||||
[The entrypoint URL is the authentication secret](#the-entrypoint-url-is-the-authentication-secret)
|
||||
before deploying one.
|
||||
|
||||
## Getting Started
|
||||
|
||||
### Prerequisites
|
||||
@@ -37,9 +44,9 @@ make bootstrap
|
||||
# Run all checks (test, lint, format check)
|
||||
make check
|
||||
|
||||
# Run in development mode. DATA_DIR defaults to /var/lib/webhooker in
|
||||
# every environment, so set it (in .env or the shell) to a writable
|
||||
# directory when running from a clone.
|
||||
# Run the server from the clone. DATA_DIR defaults to
|
||||
# /var/lib/webhooker in every environment, so set it (in .env or the
|
||||
# shell) to a writable directory.
|
||||
DATA_DIR=./data make dev
|
||||
|
||||
# Build Docker image
|
||||
@@ -71,11 +78,22 @@ make clean # Remove bin/
|
||||
### Configuration
|
||||
|
||||
All configuration is via environment variables. For local development,
|
||||
you can place variables in a `.env` file in the project root (loaded
|
||||
automatically via `godotenv/autoload`).
|
||||
you can place variables in a `.env` file in the process working
|
||||
directory, read once at startup before anything else looks at the
|
||||
environment.
|
||||
|
||||
The file is optional and having none is the normal case for a
|
||||
deployment. A file that is there but cannot be parsed aborts startup
|
||||
with a message naming it, because a single malformed line makes none
|
||||
of the file apply: every variable in it silently reverts to its
|
||||
default, which is exactly the failure [Invalid values abort
|
||||
startup](#invalid-values-abort-startup) exists to prevent, for all of
|
||||
them at once. A variable already present in the real environment wins
|
||||
over the file's value for the same name.
|
||||
|
||||
The environment is selected by setting `WEBHOOKER_ENVIRONMENT` to `dev`
|
||||
or `prod` (default: `dev`). The setting controls exactly one behavior:
|
||||
or `prod` (default: `prod`; `dev` must be set explicitly). The setting
|
||||
controls exactly one behavior:
|
||||
|
||||
| Behavior | `dev` | `prod` |
|
||||
| -------- | ----------------------- | ---------------- |
|
||||
@@ -117,7 +135,7 @@ TTY detection, and security headers are always applied.
|
||||
|
||||
| Variable | Description | Default |
|
||||
| ----------------------- | ----------------------------------- | -------- |
|
||||
| `WEBHOOKER_ENVIRONMENT` | `dev` or `prod` | `dev` |
|
||||
| `WEBHOOKER_ENVIRONMENT` | `dev` or `prod` | `prod` |
|
||||
| `PORT` | HTTP listen port | `8080` |
|
||||
| `BIND_ADDRESS` | IP address the HTTP listener binds. Loopback by default, so the cleartext listener is not published on every interface. The Docker image ships `0.0.0.0` instead. See [Bind address](#bind-address) | `127.0.0.1` (image: `0.0.0.0`) |
|
||||
| `DATA_DIR` | Directory for all SQLite databases | `/var/lib/webhooker` |
|
||||
@@ -125,7 +143,7 @@ TTY detection, and security headers are always applied.
|
||||
| `MAINTENANCE_MODE` | Report `maintenanceMode: true` in the healthcheck JSON. It does not change how any request is served — no maintenance page exists | `false` |
|
||||
| `METRICS_USERNAME` | Basic auth username for `/metrics`. Must be set together with `METRICS_PASSWORD`; one without the other fails startup | `""` |
|
||||
| `METRICS_PASSWORD` | Basic auth password for `/metrics`. Must be set together with `METRICS_USERNAME`; one without the other fails startup | `""` |
|
||||
| `SENTRY_DSN` | Sentry error reporting DSN | `""` |
|
||||
| `SENTRY_DSN` | Sentry error reporting DSN. Unset leaves error reporting off; a value the Sentry SDK cannot parse fails startup rather than serving with reporting silently off | `""` |
|
||||
| `RETENTION_SWEEP_INTERVAL` | How often the retention reaper and archive sweeper run (Go duration, must be positive) | `1h` |
|
||||
| `SESSION_IDLE_TIMEOUT` | Idle session timeout (Go duration) | `24h` |
|
||||
| `RECEIVER_RATE_LIMIT` | Receiver requests/minute per IP per entrypoint (10x that per IP across the route) | `120` |
|
||||
@@ -139,6 +157,11 @@ private and reserved ranges — RFC 1918, loopback, CGNAT, link-local and
|
||||
the rest — are refused, which stops a target from being used to make
|
||||
webhooker probe the network it sits in.
|
||||
|
||||
Besides the private and reserved ranges, the default blocklist refuses
|
||||
public cloud metadata addresses: currently only `168.63.129.16`, Azure's
|
||||
WireServer, which serves an Azure VM its credentials. Because it is a
|
||||
public address, listing it in `ALLOWED_EGRESS_CIDRS` reopens it.
|
||||
|
||||
That default is also inconvenient for the thing webhooker is mostly
|
||||
for: taking a public webhook and forwarding it to something on your own
|
||||
network. A container on the same Docker network, a box on `10.x`, a
|
||||
@@ -177,15 +200,16 @@ Two things this setting cannot do:
|
||||
the list is always an allowlist; an empty list (the default) means
|
||||
every private and reserved range stays refused. Note that
|
||||
`0.0.0.0/0` gets you most of the way there anyway, per above.
|
||||
- **It cannot open link-local, or a cloud metadata endpoint that
|
||||
discloses credentials or user data.** An address is on the list below
|
||||
when both of these hold: the provider fixes it, so it cannot collide
|
||||
with anything you run; and reaching it hands out credentials, user
|
||||
data or bootstrap material. Those stay blocked no matter what you
|
||||
list, including when you list them outright or list a supernet such
|
||||
as `0.0.0.0/0`, `::/0`, `fd00::/8` or `100.64.0.0/10`. Treat this as
|
||||
best effort rather than a guarantee — it is a hand-maintained list
|
||||
and the caveat below the table applies:
|
||||
- **It cannot open link-local, or a cloud metadata endpoint at a
|
||||
non-public address that discloses credentials or user data.** An
|
||||
address is on the list below when it is not a public address and both
|
||||
of these hold: the provider fixes it, so it cannot collide with
|
||||
anything you run; and reaching it hands out credentials, user data or
|
||||
bootstrap material. Those stay blocked no matter what you list,
|
||||
including when you list them outright or list a supernet such as
|
||||
`0.0.0.0/0`, `::/0`, `fd00::/8` or `100.64.0.0/10`. Treat this as best
|
||||
effort rather than a guarantee — it is a hand-maintained list and the
|
||||
caveat below the table applies:
|
||||
|
||||
| Blocked unconditionally | What it is |
|
||||
| ----------------------- | ---------- |
|
||||
@@ -224,7 +248,8 @@ Two things this setting cannot do:
|
||||
encodings, which the default blocklist does not match. A publicly
|
||||
routable metadata address is not listed here, because nothing on this
|
||||
list can be reopened and blocking one that way would leave you no
|
||||
escape hatch at all.
|
||||
escape hatch at all; Azure's `168.63.129.16` is refused by the default
|
||||
blocklist instead, as described above.
|
||||
|
||||
This list is not exhaustive of every cloud's metadata address — if
|
||||
yours is not here, do not allowlist the block that contains it.
|
||||
@@ -382,10 +407,9 @@ the bucket is. See [Rate Limiting](#rate-limiting).
|
||||
|
||||
The remedy is to set `TRUSTED_PROXIES` to your reverse proxy's
|
||||
address, which restores per-client buckets. webhooker logs a warning
|
||||
at startup whenever `TRUSTED_PROXIES` is empty, in every environment —
|
||||
not only when `WEBHOOKER_ENVIRONMENT=prod`, because that variable
|
||||
defaults to `dev` and an operator who never set it is precisely the
|
||||
one at risk. The warning is informational when nothing proxies to the
|
||||
at startup whenever `TRUSTED_PROXIES` is empty, in every environment,
|
||||
because behind a proxy every client shares one bucket in `dev` and
|
||||
`prod` alike. The warning is informational when nothing proxies to the
|
||||
process: with no proxy in front, the peer address is the client's own
|
||||
and the buckets are already per-client. See
|
||||
[Rate Limiting](#rate-limiting) for what each limit shares.
|
||||
@@ -464,10 +488,20 @@ startup), every entry in `TRUSTED_PROXIES` and
|
||||
`ALLOWED_EGRESS_CIDRS` must be a CIDR block or a bare IP address, and
|
||||
`BIND_ADDRESS` must be an IP address literal — `localhost`,
|
||||
`127.0.0.1:8080` and `10.0.0.0/8` are each rejected rather than
|
||||
resolved, split, or narrowed to something they do not say.
|
||||
resolved, split, or narrowed to something they do not say — and
|
||||
`SENTRY_DSN` must parse as a Sentry DSN.
|
||||
`SESSION_IDLE_TIMEOUT` is the exception: a
|
||||
non-positive value there means idle expiry is disabled, not invalid.
|
||||
|
||||
`SENTRY_DSN` is checked with the Sentry SDK's own parser, the same call
|
||||
the SDK makes on the DSN it is later handed, so what configuration
|
||||
accepts is exactly what will initialise. A typo in it is the one
|
||||
configuration mistake nothing downstream can ever notice — the variable
|
||||
is still set, so every later signal reports error reporting as on while
|
||||
no report is being sent — which is why it aborts rather than starting
|
||||
with reporting off. Leaving it unset is not a mistake and not affected:
|
||||
error reporting is simply off and startup is normal.
|
||||
|
||||
Boolean variables (`DEBUG`, `MAINTENANCE_MODE`) accept exactly the
|
||||
spellings Go's `strconv.ParseBool` accepts — `1`, `t`, `T`, `TRUE`,
|
||||
`true`, `True`, `0`, `f`, `F`, `FALSE`, `false`, `False` — and nothing
|
||||
@@ -504,6 +538,12 @@ its Argon2id hash. There is no second account and no forgot-password
|
||||
flow, so the banner and the reset command below are the only two ways
|
||||
in.
|
||||
|
||||
A start that finds no `webhooker.db` in `DATA_DIR` also logs
|
||||
`created a new, empty database` at `WARN`, with the file's path,
|
||||
shortly before the banner. On a deployment that has run before, that
|
||||
line means `DATA_DIR` was empty, most often because its volume is not
|
||||
mounted.
|
||||
|
||||
#### Recovering a lost admin password
|
||||
|
||||
`webhooker resetpw` sets an existing account's password from the
|
||||
@@ -518,8 +558,9 @@ printf '%s' "$NEW_PASSWORD" | \
|
||||
DATA_DIR=/var/lib/webhooker webhooker resetpw admin
|
||||
```
|
||||
|
||||
In a container it is the same binary, which the image sets as `CMD`
|
||||
rather than `ENTRYPOINT`, so the whole command has to be given:
|
||||
In a container it is the same binary. The image's `CMD` is
|
||||
`/app/webhooker`, and a command given to `docker run` replaces all of
|
||||
it, so the whole command has to be given:
|
||||
|
||||
```bash
|
||||
docker run --rm -v webhooker-data:/var/lib/webhooker \
|
||||
@@ -611,7 +652,6 @@ decision:
|
||||
docker run -d \
|
||||
-p 127.0.0.1:8080:8080 \
|
||||
-v /path/to/data:/var/lib/webhooker \
|
||||
-e WEBHOOKER_ENVIRONMENT=prod \
|
||||
-e BIND_ADDRESS=0.0.0.0 \
|
||||
webhooker:latest
|
||||
```
|
||||
@@ -657,14 +697,82 @@ those three values rather than trusting the figure. Measured at 65s on
|
||||
Docker 29.7.2.) A container `unhealthy` with `connection refused` in
|
||||
its health log, or a published port that resets connections, is this.
|
||||
|
||||
The container runs as a non-root user (`webhooker`, UID 1000), exposes
|
||||
port 8080, and includes a health check against
|
||||
`/.well-known/healthcheck`. The `/var/lib/webhooker` volume holds all
|
||||
SQLite databases: the main application database (`webhooker.db`), the
|
||||
per-webhook event databases (`events-{uuid}.db`), and any archive
|
||||
databases written by `database` targets (`archive-{uuid}.db`). Mount
|
||||
this as a persistent volume to preserve data across container
|
||||
restarts.
|
||||
The app runs as a non-root user (`webhooker`, UID 1000), exposes port
|
||||
8080, and includes a health check against `/.well-known/healthcheck`.
|
||||
The `/var/lib/webhooker` volume holds all SQLite databases: the main
|
||||
application database (`webhooker.db`), the per-webhook event databases
|
||||
(`events-{uuid}.db`), and any archive databases written by `database`
|
||||
targets (`archive-{uuid}.db`). Mount this as a persistent volume to
|
||||
preserve data across container restarts.
|
||||
|
||||
**The container sets its data directory's owner and mode itself
|
||||
before the app starts**, so a host directory can be mounted as it is,
|
||||
whoever owns it. The image's `ENTRYPOINT`,
|
||||
`deploy/docker-entrypoint.sh`, starts as root, creates `DATA_DIR` if
|
||||
it is missing, gives the directory and anything in it that belongs to
|
||||
another user to `webhooker`, sets the directory to `0750`, and only
|
||||
then runs the app as `webhooker`. Started with `--user`, it changes
|
||||
nothing and runs the app as that user.
|
||||
|
||||
**The file modes are not yours to set, and do not depend on the
|
||||
directory.** `webhooker.db` holds target configuration in plaintext —
|
||||
bearer tokens, API keys, Slack webhook URLs — along with the session
|
||||
encryption key, so webhooker creates every SQLite file it owns `0600`:
|
||||
each database and both of its `-wal` and `-shm` sidecars, across all
|
||||
three tiers. Files an earlier build left `0644` are tightened when
|
||||
they are opened. The directory's `0750` is defence in depth — it stops
|
||||
other local users listing the directory and learning your webhook
|
||||
UUIDs from the `events-{uuid}.db` filenames — not the barrier
|
||||
protecting the credentials.
|
||||
|
||||
### Running under upaas
|
||||
|
||||
[upaas](https://git.eeqj.de/sneak/upaas) builds the image from this
|
||||
repository's `Dockerfile` and runs it. The app needs:
|
||||
|
||||
- **Network and port:** add no port mapping in upaas. upaas publishes
|
||||
every mapped port on all interfaces of the host
|
||||
([upaas issue 113](https://git.eeqj.de/sneak/upaas/issues/113)),
|
||||
which would put the plain-HTTP admin UI and receiver there. Instead,
|
||||
set the app's Docker Network in upaas to your reverse proxy's Docker
|
||||
network; the proxy then reaches the app at `upaas-` followed by the
|
||||
app name, port `8080`. Leave `PORT` unset: the image's health check
|
||||
probes `8080`.
|
||||
- **Volume:** one host directory mounted at `/var/lib/webhooker`.
|
||||
- **Environment variables:**
|
||||
- `WEBHOOKER_ENVIRONMENT=prod`
|
||||
- `TRUSTED_PROXIES`: your reverse proxy's address on that Docker
|
||||
network. The `remoteIP` field of the `http request` log line for a
|
||||
request that came through the proxy shows it; the health check's
|
||||
own lines show `::1`. See [Trusted proxies](#trusted-proxies).
|
||||
- Leave `BIND_ADDRESS` and `DATA_DIR` unset: the image sets
|
||||
`BIND_ADDRESS` to `0.0.0.0`, and `DATA_DIR` defaults to
|
||||
`/var/lib/webhooker`.
|
||||
- Everything else is optional; see [Configuration](#configuration).
|
||||
- **Health check:** the image's own, which requests
|
||||
`/.well-known/healthcheck`. upaas reads the container's health 60
|
||||
seconds after a deploy and marks the deploy failed unless it is
|
||||
`healthy`.
|
||||
- **First run:** the first start prints the `admin` password once, in
|
||||
the banner described under [The admin account](#the-admin-account),
|
||||
to the container's log. upaas names the container `upaas-` followed
|
||||
by the app name, so for an app named `webhooker`:
|
||||
|
||||
```bash
|
||||
docker logs upaas-webhooker
|
||||
```
|
||||
|
||||
If the password is lost, stop the container, set a new password with
|
||||
the app's own image and volume, and start it again (see
|
||||
[Recovering a lost admin password](#recovering-a-lost-admin-password)):
|
||||
|
||||
```bash
|
||||
docker stop upaas-webhooker
|
||||
docker run --rm --volumes-from upaas-webhooker \
|
||||
"$(docker inspect -f '{{.Image}}' upaas-webhooker)" \
|
||||
/app/webhooker resetpw -generate admin
|
||||
docker start upaas-webhooker
|
||||
```
|
||||
|
||||
## Deployment behind a reverse proxy
|
||||
|
||||
@@ -687,17 +795,18 @@ reports.
|
||||
serves the admin login form and the unauthenticated receiver with
|
||||
no TLS at all, and the proxy in front of it changes nothing about
|
||||
that.
|
||||
2. **Set `WEBHOOKER_ENVIRONMENT=prod`, and make sure the proxy sends
|
||||
`X-Forwarded-Proto`.** These are two requirements, not one. The
|
||||
environment setting decides CORS and nothing else: the default
|
||||
2. **Make sure the environment is not `dev` (leave
|
||||
`WEBHOOKER_ENVIRONMENT` unset or set it to `prod`), and make sure
|
||||
the proxy sends `X-Forwarded-Proto`.** These are two requirements,
|
||||
not one. The environment setting decides CORS and nothing else:
|
||||
`dev` answers every origin with `Access-Control-Allow-Origin: *`
|
||||
(without credentials), which a server-rendered production
|
||||
deployment has no use for. Cookie `Secure` and the strict
|
||||
Origin/Referer mode are **not** tied to it — they are decided per
|
||||
request from the transport, which behind a proxy means the
|
||||
`X-Forwarded-Proto` header. The block below sets it; without it
|
||||
every request is read as plaintext and cookies ship without
|
||||
`Secure`. See [Configuration](#configuration).
|
||||
deployment has no use for, and `prod` — the default — disables it.
|
||||
Cookie `Secure` and the strict Origin/Referer mode are **not** tied
|
||||
to it — they are decided per request from the transport, which
|
||||
behind a proxy means the `X-Forwarded-Proto` header. The block below
|
||||
sets it; without it every request is read as plaintext and cookies
|
||||
ship without `Secure`. See [Configuration](#configuration).
|
||||
3. **Set `TRUSTED_PROXIES` to the proxy's address.** Unset, every rate
|
||||
limiter keys on the connecting peer, which behind a proxy is the
|
||||
proxy on every request: all clients collapse into one global bucket
|
||||
@@ -789,7 +898,6 @@ sent — `$scheme` above does.
|
||||
With that block, webhooker's environment is:
|
||||
|
||||
```sh
|
||||
WEBHOOKER_ENVIRONMENT=prod
|
||||
BIND_ADDRESS=127.0.0.1 # the default; stated here to be explicit
|
||||
TRUSTED_PROXIES=127.0.0.1
|
||||
```
|
||||
@@ -834,9 +942,20 @@ is both the simplest and the only complete rule:
|
||||
`events-3f2a1c9e-....db`. The only other file is `webhooker.lock`, the
|
||||
always-empty [single-instance lock](#single-instance-lock); it holds no
|
||||
state and is not part of the backup set — a copied one is stale and
|
||||
blocks nothing. No `-wal` or `-shm` files are produced (see below); a
|
||||
transient `{name}.db-journal` may exist beside a database while a write
|
||||
is in flight and is not part of the backup set either.
|
||||
blocks nothing.
|
||||
|
||||
**`-wal` and `-shm` sidecars.** Every database runs in WAL journal mode,
|
||||
so while the service is running each `{name}.db` has a `{name}.db-wal`
|
||||
and a `{name}.db-shm` beside it. **`-wal` is part of the database, not a
|
||||
scratch file**: it holds committed transactions that are not yet in the
|
||||
`.db`, so a copy of the `.db` without its `-wal` is missing data and may
|
||||
have no readable schema at all. `-shm` is regenerable, but there is no
|
||||
reason to separate the two — copy the directory and you have them.
|
||||
|
||||
A clean shutdown closes every database, which checkpoints and removes
|
||||
its sidecars; a killed or crashed instance leaves them, and they must be
|
||||
carried with the `.db`. An archive the service has not opened since a
|
||||
crash keeps that crash's sidecars, even across a later clean stop.
|
||||
|
||||
Configuration is **not** in `DATA_DIR` — it comes from the environment
|
||||
and from a `.env` file read out of the process working directory. Back
|
||||
@@ -844,17 +963,19 @@ that up with your deployment config, separately.
|
||||
|
||||
### A hot copy is not safe
|
||||
|
||||
No `journal_mode` pragma is ever issued on any database webhooker opens,
|
||||
so all of them run on SQLite's default rollback journal. There is no
|
||||
WAL. The main and event databases are also held open for the entire
|
||||
process lifetime — `WebhookDBManager` caches event database handles and
|
||||
closes them only on webhook deletion or shutdown — so "it looked idle"
|
||||
is not a guarantee that nothing was mid-transaction.
|
||||
Every database webhooker opens runs in WAL journal mode. The main and
|
||||
event databases are also held open for the entire process lifetime —
|
||||
`WebhookDBManager` caches event database handles and closes them only on
|
||||
webhook deletion or shutdown — so "it looked idle" is not a guarantee
|
||||
that nothing was mid-transaction.
|
||||
|
||||
That means `cp`, `rsync`, `tar` or a filesystem snapshot taken against a
|
||||
running instance can capture a database mid-transaction and yield a file
|
||||
that is corrupt or missing state the journal would have rolled back. Use
|
||||
one of the two procedures below instead.
|
||||
running instance can capture a database and its `-wal` at two different
|
||||
instants and yield a file that is corrupt or missing state. Copying a
|
||||
`.db` on its own is worse and fails loudly: recently written pages,
|
||||
including the schema itself on a young database, live in the `-wal`, so
|
||||
the copy reads back as an empty or table-less database. Use one of the
|
||||
two procedures below instead.
|
||||
|
||||
**Stop, copy, start.** The simplest, needs no extra tooling, and the
|
||||
only one that gives a single point in time across every file:
|
||||
@@ -873,24 +994,46 @@ for db in /path/to/data/*.db; do
|
||||
done
|
||||
```
|
||||
|
||||
`.backup` takes the proper locks and writes a consistent file. Two
|
||||
caveats. First, the runtime image is `alpine:3.21` with only
|
||||
`ca-certificates` added — the `sqlite3` CLI is **not** in it, so run
|
||||
this on the host against the volume path, or from a throwaway container
|
||||
that mounts the volume. Second, each file is captured at its own
|
||||
instant, so a webhook created or an event delivered between two files
|
||||
being copied lands in one and not the other. If you need the whole set
|
||||
coherent as of a single moment, stop the service.
|
||||
`.backup` reads through the WAL and writes a single consistent file with
|
||||
no sidecars of its own, so the destination is complete as it stands.
|
||||
Two caveats. First, the runtime image is `alpine:3.21` with only
|
||||
`ca-certificates` and `su-exec` added — the `sqlite3` CLI is **not** in
|
||||
it, so run this on the host against the volume path, or from a
|
||||
throwaway container that mounts the volume. Second, each file is
|
||||
captured at its own instant, so a webhook created or an event delivered
|
||||
between two files being copied lands in one and not the other. If you
|
||||
need the whole set coherent as of a single moment, stop the service.
|
||||
|
||||
Note that `sqlite3 <db> .dump` is **not** one of these procedures: it is
|
||||
an export, it holds a read transaction open for as long as it runs, and
|
||||
it pins the WAL against checkpointing for that whole time. It is safe to
|
||||
run — it does not block ingestion — but back up with `.backup` or a
|
||||
stopped copy.
|
||||
|
||||
Archive databases are the one exception the service is built for: the
|
||||
archive writer closes its handle after each write (debounced to at most
|
||||
one reopen per second), so an operator can move `archive-{uuid}.db`
|
||||
away for offline retention while the service runs, and it is recreated
|
||||
on the next write (see
|
||||
[Database Architecture](#database-architecture)). That is a
|
||||
archive writer closes and reopens its handle around writes (debounced
|
||||
to at most one reopen per second), so an operator can move
|
||||
`archive-{uuid}.db` away for offline retention while the service runs,
|
||||
and it is recreated on the next write. See
|
||||
[Database Architecture](#database-architecture). That is a
|
||||
move-the-file-away workflow, not a substitute for the backup procedures
|
||||
above.
|
||||
|
||||
**Move the sidecars with it.** Under WAL that workflow is no longer a
|
||||
single file, and the common case is the dangerous one. The reopen
|
||||
happens on the *next* write after the debounce window elapses, so after
|
||||
the last write of a burst nothing checkpoints: measured, 20 s after ten
|
||||
events the `archive-….db` was 4096 bytes — a header, no table — with
|
||||
all ten rows sitting in a 189 KB `-wal`. Copying the `.db` alone at that
|
||||
moment yields a file that opens with `no such table: archived_events`.
|
||||
The file becomes self-contained again when the handle closes, which
|
||||
happens on the next write past the debounce window, when the connection
|
||||
pool retires the idle connection (about a minute after the last write),
|
||||
or at the idle archive sweep — measured, the same file was a complete
|
||||
20 KB `.db` with no sidecars about a minute after its last write. A
|
||||
clean stop closes it too. So either move `archive-{uuid}.db` together
|
||||
with any `-wal`/`-shm` beside it, or wait until there are none.
|
||||
|
||||
### Restore
|
||||
|
||||
1. Stop the service.
|
||||
@@ -904,24 +1047,20 @@ above.
|
||||
restored without `webhooker.db` are simply orphaned; nothing
|
||||
references their UUIDs.
|
||||
|
||||
3. Do not carry `*.db-journal` files into the restore. Backups taken by
|
||||
either procedure above are self-consistent and do not need one.
|
||||
3. Carry any `*.db-wal` and `*.db-shm` files that are in the backup.
|
||||
They are part of the database, and dropping a `-wal` silently
|
||||
discards every transaction it still holds. An `.backup` set will not
|
||||
contain any: it writes a single consolidated file per database. A
|
||||
stop-and-copy set normally has none, because a clean stop closes
|
||||
every database and checkpoints its sidecars away; the exception is an
|
||||
archive not opened since a crash. A copy salvaged from a crashed
|
||||
instance has them for everything, and needs all of them.
|
||||
|
||||
4. **Fix ownership.** The container runs as the non-root `webhooker`
|
||||
user, UID 1000 / GID 1000. Restored files must be owned by (or
|
||||
writable by) that UID, and so must the directory itself — SQLite
|
||||
creates the rollback journal beside the database, so a writable file
|
||||
inside a directory it cannot write is not enough:
|
||||
|
||||
```bash
|
||||
chown -R 1000:1000 /path/to/data
|
||||
```
|
||||
|
||||
Restoring as `root` on the host and forgetting this step is the
|
||||
usual way a restore fails.
|
||||
|
||||
5. Start the service. `AutoMigrate` runs against each restored database
|
||||
as it is opened.
|
||||
4. Start the service. The container gives the directory and the
|
||||
restored files to the `webhooker` user before the app starts,
|
||||
whoever restored them (see
|
||||
[Running with Docker](#running-with-docker)). `AutoMigrate` runs
|
||||
against each restored database as it is opened.
|
||||
|
||||
### Upgrades
|
||||
|
||||
@@ -1042,14 +1181,38 @@ backups at rest and restrict who can read them.
|
||||
|
||||
## The entrypoint URL is the authentication secret
|
||||
|
||||
The receiver verifies nothing about an inbound request. The UUID in an
|
||||
entrypoint's URL is its credential: anyone who holds that URL can
|
||||
submit events to it, and the receiver checks nothing else about the
|
||||
sender. Treat an entrypoint URL the way you would treat an API token.
|
||||
**The entrypoint UUID is the credential, and it is the only one.**
|
||||
webhooker mints a version 4 UUID per entrypoint and serves it at
|
||||
`/webhook/{uuid}`. Possession of that URL is the authentication:
|
||||
anyone who holds it can submit events to the entrypoint, and the
|
||||
receiver verifies nothing else about the sender.
|
||||
|
||||
There is no way to rotate the UUID in place. To retire one, delete the
|
||||
entrypoint (or deactivate it, which answers `410`) and create a new
|
||||
one, then point the sender at the new URL.
|
||||
There is no shared secret, no HMAC signature, no bearer token and no
|
||||
second factor on the receiver, and none will be added. This was
|
||||
considered and rejected; the implementation that existed was removed
|
||||
in [PR #279](https://git.eeqj.de/sneak/webhooker/pulls/279), closing
|
||||
[issue #67](https://git.eeqj.de/sneak/webhooker/issues/67) and
|
||||
[issue #241](https://git.eeqj.de/sneak/webhooker/issues/241). A
|
||||
proposal to reintroduce any of them — including as "defence in depth"
|
||||
alongside the UUID — is answered by this section. Inbound signature
|
||||
headers a sender sends anyway (`X-Hub-Signature` and its
|
||||
per-provider equivalents) are stored and forwarded as ordinary
|
||||
headers; nothing checks them.
|
||||
|
||||
What that means for an operator:
|
||||
|
||||
- **The URL is a capability, so treat it as a secret.** Keep it out of
|
||||
logs, ticket bodies, chat messages and screenshots. Anyone who reads
|
||||
it anywhere can post events as that sender.
|
||||
- **Rotating means minting a new entrypoint, not changing a key.**
|
||||
There is no way to rotate the UUID in place. To retire one, delete
|
||||
the entrypoint (or deactivate it, which answers `410`) and create a
|
||||
new one, then point the sender at the new URL.
|
||||
- **A sender that cannot be given a secret URL is a constraint on that
|
||||
integration, not a reason to change this.** If a service only
|
||||
supports signed payloads to a well-known URL, raise it as its own
|
||||
problem — pick a different integration path, or accept that it
|
||||
cannot be used. It is not grounds to reintroduce shared secrets.
|
||||
|
||||
## Entrypoints
|
||||
|
||||
@@ -1059,8 +1222,15 @@ standard: normalized scripts in `script/` are the entrypoints for the
|
||||
development workflow. Ten of the Makefile's seventeen targets are thin
|
||||
shims that call them; `build`, `run`, `dev`, `deps`, `clean`, `css` and
|
||||
`version` are inline commands with no script behind them, though
|
||||
`build` and `version` both take their value from `script/version`. We
|
||||
provide:
|
||||
`build` and `version` both take their value from `script/version`.
|
||||
|
||||
`make check` needs the third-party browser assets in `static/`, which
|
||||
are not committed, so run `make bootstrap` (or just `make assets`) once
|
||||
after cloning. Without them the tests fail with a message naming that
|
||||
remedy. `make check` does not fetch them itself because it must not
|
||||
change any files in the repo.
|
||||
|
||||
We provide:
|
||||
|
||||
- `script/bootstrap` — install all dependencies (idempotent)
|
||||
- `script/setup` — make a fresh clone ready for development
|
||||
@@ -1127,6 +1297,16 @@ webhooker solves this by acting as a durable intermediary:
|
||||
backoff. Every delivery attempt is logged with status codes, response
|
||||
bodies, and timing.
|
||||
|
||||
**That guarantee is at-least-once, not exactly-once.** When a send
|
||||
reaches its target but the write recording that outcome fails, the
|
||||
delivery is deliberately left in a recoverable state rather than
|
||||
marked done — losing a delivery is the worse failure — so the
|
||||
pending sweep picks it up about fifteen minutes later, or the next
|
||||
restart does, and the target receives a payload it already got.
|
||||
webhooker adds no delivery identifier of its own to an outbound
|
||||
request, so **make your receiver idempotent** against whatever the
|
||||
payload itself carries.
|
||||
|
||||
3. **Observability** — Full request/response logging for every webhook
|
||||
received and every delivery attempted. Prometheus metrics expose
|
||||
volume, latency, and error rates. The web UI provides real-time
|
||||
@@ -1359,7 +1539,7 @@ events should be forwarded.
|
||||
| `type` | TargetType | One of: `http`, `slack`, `database`, `log` |
|
||||
| `active` | boolean | Whether deliveries are enabled (default: true) |
|
||||
| `config` | JSON text | Type-specific configuration |
|
||||
| `max_retries` | integer | Maximum retry attempts for `http` and `slack` targets (0 = fire-and-forget, >0 = retries with backoff and a circuit breaker). Ignored by `database` and `log` targets |
|
||||
| `max_retries` | integer | Total delivery attempts for `http` and `slack` targets, not retries on top of the first: 0 is a single fire-and-forget attempt with no retries and no circuit breaker, and a value of N makes N attempts in all, with exponential backoff and a per-target circuit breaker. Ignored by `database` and `log` targets |
|
||||
| `max_queue_size` | integer | Stored and shown on the target's detail view, but not enforced anywhere yet: nothing in the delivery engine consults it. Queue depth is set by the two fixed 10,000-entry channels |
|
||||
|
||||
**Relations:** Belongs to Webhook. Has many Deliveries.
|
||||
@@ -1367,12 +1547,12 @@ events should be forwarded.
|
||||
**Target types:**
|
||||
|
||||
- **`http`** — Forward the event as an HTTP POST to a configured URL.
|
||||
Behavior depends on `max_retries`: when `max_retries` is 0 (the
|
||||
default), the target operates in fire-and-forget mode — a single
|
||||
attempt with no retries and no circuit breaker. When `max_retries` is
|
||||
greater than 0, failed deliveries are retried with exponential backoff
|
||||
up to `max_retries` attempts, protected by a per-target circuit
|
||||
breaker.
|
||||
`max_retries` is the total number of delivery attempts, not retries on
|
||||
top of the first: when `max_retries` is 0 (the default), the target
|
||||
operates in fire-and-forget mode, a single attempt with no retries and
|
||||
no circuit breaker; a value of N makes up to N attempts in all,
|
||||
retrying failed deliveries with exponential backoff and protecting them
|
||||
with a per-target circuit breaker.
|
||||
- **`slack`** — Post the event as a formatted message to a
|
||||
Slack-compatible incoming webhook URL (`webhookUrl` in `config`). It
|
||||
is built on the same HTTP core as `http` and honours `max_retries`
|
||||
@@ -1552,6 +1732,28 @@ retries) is individually logged for full observability.
|
||||
|
||||
**Relations:** Belongs to Delivery.
|
||||
|
||||
#### Event-tier indexes
|
||||
|
||||
These indexes on the per-webhook event databases are declared in the model
|
||||
tags, so `AutoMigrate` creates them on a fresh and on an existing database:
|
||||
|
||||
| Table | Columns | Serves |
|
||||
| ------------------ | --------------------------- | ------ |
|
||||
| `deliveries` | `status`, `deleted_at` | Startup recovery, the retry and pending sweeps every 60 seconds and the queue-depth sampler every 30 seconds, which select deliveries by status |
|
||||
| `deliveries` | `event_id`, `deleted_at` | The event log, which loads each event's deliveries, and retention, which selects and deletes the deliveries of expired events |
|
||||
| `delivery_results` | `delivery_id`, `deleted_at` | The event log, which loads the attempts of a page's deliveries, and retention, which deletes the attempts of expired events |
|
||||
| `events` | `deleted_at`, `created_at` | Retention, which selects expired events by age |
|
||||
| `events` | `created_at` | Retention's delete of the expired events themselves |
|
||||
|
||||
GORM's soft delete adds `deleted_at IS NULL` to these queries; retention's
|
||||
deletes leave it out, but their lookups of expired rows keep it. SQLite keeps
|
||||
no statistics on these tables, and without them it rates the `deleted_at`
|
||||
index, which every live row matches, above an index on a column matched
|
||||
against several values or compared with `<`. So every index but the last also
|
||||
covers `deleted_at`. It comes second, so that retention's deletes can use the
|
||||
index without it, except in `events`, where `created_at` is compared with `<`
|
||||
and SQLite narrows by a `<` only on the last column it uses.
|
||||
|
||||
#### Common Fields
|
||||
|
||||
Every entity except `Setting` includes these fields from `BaseModel`.
|
||||
@@ -1573,6 +1775,11 @@ webhooker uses **separate SQLite database files**: a main application
|
||||
database for configuration data and per-webhook databases for event
|
||||
storage. All database files live in the `DATA_DIR` directory.
|
||||
|
||||
Every one of them is created `0600`, and so is each `-wal` and `-shm`
|
||||
sidecar. See
|
||||
[Running with Docker](#running-with-docker) for what that does and
|
||||
does not protect.
|
||||
|
||||
**Main Application Database** (`{DATA_DIR}/webhooker.db`) — stores
|
||||
configuration and application state:
|
||||
|
||||
@@ -1614,10 +1821,12 @@ This separation provides:
|
||||
only, or disables cleanup entirely when set to `0` (retain forever).
|
||||
- **Performance** — each webhook's database has its own page cache and
|
||||
its own lock, so concurrent event ingestion across webhooks won't
|
||||
contend. No write-ahead log is involved: both DSNs are
|
||||
`file:{path}?cache=shared&mode=rwc` and no `journal_mode` pragma is
|
||||
ever issued, so every database runs on SQLite's default rollback
|
||||
journal.
|
||||
contend. Every database — main, per-webhook, and archive — is opened
|
||||
through one code path (`internal/database/sqlite_open.go`) in WAL
|
||||
journal mode, with a 10-second busy timeout, `BEGIN IMMEDIATE`
|
||||
transactions, and a bounded connection pool. Under WAL a reader never
|
||||
blocks a writer, so an operator reading a database does not stall
|
||||
event ingestion into it.
|
||||
|
||||
The **database target type** builds on this architecture to provide
|
||||
long-term archiving, separate from the per-webhook event database (which
|
||||
@@ -1808,7 +2017,7 @@ rescans the database anyway).
|
||||
| ----------- | -------- |
|
||||
| **Closed** | Normal operation. Deliveries flow through. Consecutive failures are counted. |
|
||||
| **Open** | Target appears down. Deliveries are skipped and rescheduled for after the cooldown. |
|
||||
| **Half-Open** | Cooldown expired. One probe delivery is allowed to test if the target has recovered. |
|
||||
| **Half-Open** | Cooldown expired. One probe delivery is allowed to test if the target has recovered. Other deliveries are rescheduled for one whole cooldown later. |
|
||||
|
||||
**Transitions:**
|
||||
|
||||
@@ -1846,7 +2055,9 @@ operations), and log targets (stdout) do not use circuit breakers.
|
||||
When a circuit is open and a new delivery arrives, the engine marks the
|
||||
delivery as `retrying` and schedules a retry timer for after the
|
||||
remaining cooldown period. This ensures no deliveries are lost — they're
|
||||
just delayed until the target is healthy again.
|
||||
just delayed until the target is healthy again. A delivery already in
|
||||
`retrying` keeps that status without another database write each time
|
||||
the breaker turns it away.
|
||||
|
||||
### Metrics
|
||||
|
||||
@@ -1860,7 +2071,7 @@ arriving and being stored, they are just not getting anywhere.
|
||||
| Metric | Type | Meaning |
|
||||
| ------ | ---- | ------- |
|
||||
| `webhooker_events_received_total` | counter | Events received and durably stored. Compare against the delivery counters on one dashboard |
|
||||
| `webhooker_delivery_attempts_total` | counter | Delivery attempts actually dispatched to a target. A delivery an open circuit breaker refused is not one: it is counted as a retry instead |
|
||||
| `webhooker_delivery_attempts_total` | counter | Delivery attempts actually dispatched to a target. A delivery a circuit breaker refused is not one: it is counted as a retry instead, but only when the refusal moves it into `retrying` |
|
||||
| `webhooker_deliveries_succeeded_total` | counter | Deliveries that reached `delivered` |
|
||||
| `webhooker_deliveries_failed_total` | counter | Deliveries that failed terminally and will not be retried |
|
||||
| `webhooker_delivery_retries_total` | counter | Deliveries put back into `retrying` |
|
||||
@@ -2477,7 +2688,7 @@ abuse limit later; they are tracked as future work.
|
||||
| ------ | --------------------------- | ----------- |
|
||||
| `GET` | `/` | Root redirect, 303 (authenticated → `/sources`, unauthenticated → `/pages/login`) |
|
||||
| `GET` | `/.well-known/healthcheck` | Health check (JSON: `status`, `now`, `uptimeSeconds`, `uptimeHuman`, `version`, `appname`, `maintenanceMode`) |
|
||||
| any | `/s/*` | Static file serving (embedded CSS, JS). Mounted for every method, not just `GET`/`HEAD`: chi's `Mount` registers all methods and `http.FileServer` special-cases only `HEAD` (by omitting the body), so a `POST` or `DELETE` to an asset is answered `200` with the file. Pinned by `TestStaticServesEveryMethod` |
|
||||
| `GET`, `HEAD` | `/s/*` | Static file serving (embedded CSS, JS). `GET` and `HEAD` only — `POST`, `PUT`, `PATCH`, `DELETE`, `OPTIONS`, `TRACE` and `CONNECT` are answered `405 Method Not Allowed` with `Allow: GET, HEAD`. Any other method (such as `PROPFIND`) is refused by chi before it reaches this route, and gets `405` without an `Allow` header. Pinned by `TestStaticServesOnlyGetAndHead` |
|
||||
| `POST` | `/webhook/{uuid}` | Webhook receiver endpoint. `POST` only — every other method is answered `405 Method Not Allowed` with `Allow: POST`. Rate limited (see [Rate Limiting](#rate-limiting)) |
|
||||
|
||||
#### Authentication Endpoints
|
||||
@@ -2502,12 +2713,15 @@ abuse limit later; they are tracked as future work.
|
||||
| `POST` | `/source/{id}/edit` | Edit webhook submission |
|
||||
| `POST` | `/source/{id}/delete` | Delete webhook |
|
||||
| `GET` | `/source/{id}/logs` | Webhook event logs |
|
||||
| `GET` | `/source/{id}/logs/{eventID}/body` | Download an event's full stored body. The log page renders each body only up to its cap, so this is the only route that serves a whole one; it is offered wherever a body is shown truncated |
|
||||
| `POST` | `/source/{id}/deliveries/{deliveryID}/replay` | Replay a finished delivery: creates a new delivery for the same event against the target's current configuration (30 per minute per bucket, then `429`) |
|
||||
| `POST` | `/source/{id}/events/{eventID}/resubmit` | Resubmit a stored event: creates a new event copying it and fans that out to every currently active target (30 per minute per bucket, then `429`) |
|
||||
| `POST` | `/source/{id}/entrypoints` | Add entrypoint to webhook |
|
||||
| `POST` | `/source/{id}/entrypoints/{entrypointID}/delete` | Delete an entrypoint |
|
||||
| `POST` | `/source/{id}/entrypoints/{entrypointID}/toggle` | Enable or disable an entrypoint |
|
||||
| `POST` | `/source/{id}/targets` | Add target to webhook |
|
||||
| `GET` | `/source/{id}/targets/{targetID}/edit` | Edit target form. The one page that renders a target's destination URL and header values in full, rather than masked |
|
||||
| `POST` | `/source/{id}/targets/{targetID}/edit` | Edit target submission |
|
||||
| `POST` | `/source/{id}/targets/{targetID}/delete` | Delete a target |
|
||||
| `POST` | `/source/{id}/targets/{targetID}/toggle` | Enable or disable a target |
|
||||
|
||||
@@ -2546,6 +2760,8 @@ webhooker/
|
||||
├── internal/
|
||||
│ ├── banner/
|
||||
│ │ └── banner.go # Ruled block for the one credential shown in the clear
|
||||
│ ├── ciscript/
|
||||
│ │ └── doc.go # Tests for the CI shell scripts in script/; no runtime code
|
||||
│ ├── resetpw/
|
||||
│ │ └── resetpw.go # `webhooker resetpw`: set an account's password, stopped deployments only
|
||||
│ ├── config/
|
||||
@@ -2615,13 +2831,17 @@ webhooker/
|
||||
│ │ ├── ratelimit.go # Per-IP rate limiting middleware (go-chi/httprate)
|
||||
│ │ ├── loginguard.go # Login failure counters and the Argon2id verification semaphore
|
||||
│ │ └── testing.go # NewForTest: Middleware without the fx lifecycle
|
||||
│ ├── reqtls/
|
||||
│ │ └── reqtls.go # IsTLS: the one TLS predicate, r.TLS or X-Forwarded-Proto
|
||||
│ ├── server/
|
||||
│ │ ├── server.go # Server struct, fx lifecycle, signal handling
|
||||
│ │ ├── http.go # HTTP server setup with timeouts
|
||||
│ │ └── routes.go # All route definitions
|
||||
│ └── session/
|
||||
│ ├── session.go # Cookie-based session management
|
||||
│ └── testing.go # NewForTest: Session without the fx lifecycle
|
||||
│ ├── session/
|
||||
│ │ ├── session.go # Cookie-based session management
|
||||
│ │ └── testing.go # NewForTest: Session without the fx lifecycle
|
||||
│ └── versionscript/
|
||||
│ └── doc.go # Tests for script/version and the build files that use it
|
||||
├── static/
|
||||
│ ├── static.go # //go:embed directive
|
||||
│ ├── css/input.css # Tailwind input, source for tailwind.css (make css)
|
||||
@@ -2734,6 +2954,10 @@ check, see [The login endpoint](#the-login-endpoint).
|
||||
|
||||
### Authentication
|
||||
|
||||
- **Webhook receiver:** the entrypoint UUID in the URL, and nothing
|
||||
else. No shared secret, no HMAC signature, no token header, and none
|
||||
will be added — see
|
||||
[The entrypoint URL is the authentication secret](#the-entrypoint-url-is-the-authentication-secret).
|
||||
- **Web UI:** Cookie-based sessions using gorilla/sessions with
|
||||
encrypted cookies. Sessions are configured with HttpOnly, SameSite
|
||||
Lax, and Secure whenever the request is on TLS — the flag follows the
|
||||
@@ -2773,7 +2997,8 @@ check, see [The login endpoint](#the-login-endpoint).
|
||||
mode
|
||||
- **The entrypoint URL is the receiver's only credential.** Nothing
|
||||
about an inbound request is verified; possession of the UUID
|
||||
authorises submission (see
|
||||
authorises submission, and no shared secret or signature check will
|
||||
be added alongside it (see
|
||||
[The entrypoint URL is the authentication secret](#the-entrypoint-url-is-the-authentication-secret))
|
||||
- **SSRF prevention** for HTTP delivery targets: private/reserved IP
|
||||
ranges (RFC 1918, loopback, link-local, cloud metadata) are blocked
|
||||
@@ -2813,7 +3038,11 @@ check, see [The login endpoint](#the-login-endpoint).
|
||||
- Prometheus metrics behind basic auth
|
||||
- Static assets embedded in binary (no filesystem access needed at
|
||||
runtime)
|
||||
- Container runs as non-root user (UID 1000)
|
||||
- The app runs as the non-root `webhooker` user (UID 1000) in the
|
||||
container. The image sets no `USER`, so these run as root: the
|
||||
`ENTRYPOINT` script, which sets the data directory's owner and mode
|
||||
before the app starts; the image's health check; and `docker exec`,
|
||||
unless given `--user`
|
||||
- GORM soft deletes on every entity that carries `BaseModel`, which is
|
||||
all of them but `Setting` (data preserved for audit)
|
||||
|
||||
@@ -2829,7 +3058,8 @@ each hook. The order, read off the fx stop-hook log:
|
||||
3. `server` — the HTTP drain, bounded separately by
|
||||
`server.ShutdownTimeout` (**3 seconds**), then a Sentry flush if
|
||||
`SENTRY_DSN` is set
|
||||
4. `delivery.Engine`
|
||||
4. `delivery.Engine` — waits for its workers, then closes the archive
|
||||
databases
|
||||
5. `healthcheck`
|
||||
6. `WebhookDBManager`
|
||||
7. the database close
|
||||
@@ -2939,10 +3169,13 @@ version is fixed independently of the compiler's:
|
||||
`GO_LDFLAGS`, so neither can drop the `-X` that stamps the version.
|
||||
The version arrives as the `VERSION` build arg, since the context
|
||||
has no `.git` (see [Version stamping](#version-stamping)).
|
||||
3. **Runtime stage** (`alpine:3.21`) — copies the static binary,
|
||||
creates the `/var/lib/webhooker` directory for all SQLite databases,
|
||||
runs as the non-root `webhooker` user (UID 1000), exposes port 8080,
|
||||
and includes a health check against `/.well-known/healthcheck`.
|
||||
3. **Runtime stage** (`alpine:3.21`) — copies the static binary and
|
||||
`deploy/docker-entrypoint.sh`, creates the `/var/lib/webhooker`
|
||||
directory for all SQLite databases, exposes port 8080, and includes
|
||||
a health check against `/.well-known/healthcheck`. It sets no
|
||||
`USER`: the `ENTRYPOINT` script starts as root, sets the data
|
||||
directory's owner and mode, and runs the app as the non-root
|
||||
`webhooker` user (UID 1000) through `su-exec`.
|
||||
|
||||
The lint stage invokes `golangci-lint` directly rather than `make lint`:
|
||||
it is already the pinned linter image, and `make lint` builds
|
||||
@@ -3021,3 +3254,5 @@ MIT
|
||||
## Author
|
||||
|
||||
[@sneak](https://sneak.berlin)
|
||||
|
||||
|
||||
|
||||
@@ -18,18 +18,27 @@ Issue branches do NOT touch this file — the manager maintains it on
|
||||
|
||||
# Status
|
||||
|
||||
1.0.0 is open, with work remaining. The milestone
|
||||
(https://git.eeqj.de/sneak/webhooker/milestone/9) is the authoritative
|
||||
list, and the only place to read a count or a state of play from. This
|
||||
file records where the project is, not what is in flight: a sentence
|
||||
whose truth depends on a branch being unmerged is wrong the moment it
|
||||
merges, and this file has been wrong that way before.
|
||||
The milestone (https://git.eeqj.de/sneak/webhooker/milestone/9) is the
|
||||
authoritative list, and the only place to read a count or a state of
|
||||
play from. This file records where the project is, not what is in
|
||||
flight: a sentence whose truth depends on a branch being unmerged is
|
||||
wrong the moment it merges, and this file has been wrong that way
|
||||
before.
|
||||
|
||||
The tag is held on a durability defect
|
||||
(https://git.eeqj.de/sneak/webhooker/issues/256): a concurrent reader
|
||||
of a per-webhook event database strands delivered webhooks at
|
||||
`pending`, and the next restart re-delivers them. That issue gates
|
||||
`v1.0.0`, and is where the fix's own state is tracked.
|
||||
The durability defect that held the tag has landed
|
||||
(https://git.eeqj.de/sneak/webhooker/issues/256, commit `8d64259`).
|
||||
Every SQLite handle opens with WAL journaling and a busy timeout, a
|
||||
bookkeeping write that fails leaves its delivery in a recoverable
|
||||
state rather than a lying one, and recovery skips a delivery that
|
||||
already has a successful result row. Final pre-tag verification
|
||||
exercised it and confirmed it holds. Whatever the milestone still
|
||||
shows open is what remains before `v1.0.0`.
|
||||
|
||||
Delivery is at-least-once by design, not by accident: a send whose
|
||||
result row does not land is attempted again, so a receiver can see a
|
||||
duplicate. That is deliberate — the alternative is a silent lost
|
||||
delivery — and the README says so under Rationale. It is not a defect
|
||||
to re-file.
|
||||
|
||||
One caveat on reading a green check: a docs-only commit deliberately
|
||||
replays from the layer cache
|
||||
@@ -39,11 +48,11 @@ commit invalidates the `COPY` layer and genuinely executes.
|
||||
|
||||
# Next Step
|
||||
|
||||
Land https://git.eeqj.de/sneak/webhooker/issues/256, then clear the
|
||||
rest of the open 1.0.0 milestone and tag `v1.0.0`. Merging `next` into
|
||||
`main` is a separate act from tagging and waits on neither of those:
|
||||
`next` is kept mergeable at all times, which is the point of the
|
||||
branch.
|
||||
Clear the rest of the open 1.0.0 milestone
|
||||
(https://git.eeqj.de/sneak/webhooker/milestone/9) and tag `v1.0.0`.
|
||||
Merging `next` into `main` is a separate act from tagging and waits on
|
||||
neither of those: `next` is kept mergeable at all times, which is the
|
||||
point of the branch.
|
||||
|
||||
# Completed Steps
|
||||
|
||||
|
||||
@@ -0,0 +1,107 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"bytes"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
"github.com/stretchr/testify/assert"
|
||||
"github.com/stretchr/testify/require"
|
||||
"sneak.berlin/go/webhooker/internal/config"
|
||||
)
|
||||
|
||||
// dotEnvKey is a throwaway variable name these tests write and read,
|
||||
// so they cannot disturb real configuration.
|
||||
const dotEnvKey = "WEBHOOKER_TEST_DISPATCH_VALUE"
|
||||
|
||||
// writeDotEnvInWorkingDir puts contents in a .env file in a fresh
|
||||
// temporary directory and moves the process there.
|
||||
//
|
||||
// The callers are deliberately not parallel and must stay that way:
|
||||
// t.Chdir moves the whole process. Go releases parallel tests only
|
||||
// after every sequential test in the package has finished, so nothing
|
||||
// else runs while these do.
|
||||
func writeDotEnvInWorkingDir(t *testing.T, contents string) {
|
||||
t.Helper()
|
||||
|
||||
dir := t.TempDir()
|
||||
require.NoError(t, os.WriteFile(
|
||||
filepath.Join(dir, config.DotEnvPath),
|
||||
[]byte(contents), 0o600,
|
||||
))
|
||||
t.Chdir(dir)
|
||||
}
|
||||
|
||||
// TestDispatch_MalformedDotEnvRefuses pins the second half of the
|
||||
// defect. godotenv applies nothing at all when a file will not parse,
|
||||
// so one mistyped line used to revert every variable in it to its
|
||||
// default and start the server anyway, with no log line naming the
|
||||
// file. The refusal has to arrive before any subcommand runs, which
|
||||
// is why `help` — the one subcommand that touches nothing — is still
|
||||
// refused here.
|
||||
//
|
||||
//nolint:paralleltest // t.Chdir moves the whole process.
|
||||
func TestDispatch_MalformedDotEnvRefuses(t *testing.T) {
|
||||
writeDotEnvInWorkingDir(t, "PORT 19615\n")
|
||||
|
||||
var stdout, stderr bytes.Buffer
|
||||
|
||||
code := dispatch(
|
||||
[]string{helpCommand}, strings.NewReader(""), &stdout, &stderr,
|
||||
)
|
||||
|
||||
require.Equal(t, 1, code, "a broken .env must exit non-zero")
|
||||
assert.Contains(
|
||||
t, stderr.String(), config.DotEnvPath,
|
||||
"the refusal must name the file",
|
||||
)
|
||||
assert.Empty(
|
||||
t, stdout.String(),
|
||||
"the subcommand must not have run",
|
||||
)
|
||||
}
|
||||
|
||||
// TestDispatch_LoadsDotEnvBeforeSubcommands pins the ordering the
|
||||
// godotenv/autoload import used to provide for free. It ran in an
|
||||
// init(), so .env was in the environment before anything read it —
|
||||
// including config.DataDir, which both the DATA_DIR lock and resetpw
|
||||
// call outside the fx graph. Loading any later would let a .env that
|
||||
// sets DATA_DIR lock one directory while the config opened databases
|
||||
// in another.
|
||||
func TestDispatch_LoadsDotEnvBeforeSubcommands(t *testing.T) {
|
||||
t.Setenv(dotEnvKey, "placeholder")
|
||||
require.NoError(t, os.Unsetenv(dotEnvKey))
|
||||
|
||||
writeDotEnvInWorkingDir(t, dotEnvKey+"=from-dot-env\n")
|
||||
|
||||
var stdout, stderr bytes.Buffer
|
||||
|
||||
code := dispatch(
|
||||
[]string{helpCommand}, strings.NewReader(""), &stdout, &stderr,
|
||||
)
|
||||
|
||||
require.Equal(t, 0, code)
|
||||
assert.Equal(
|
||||
t, "from-dot-env", os.Getenv(dotEnvKey),
|
||||
"the file must be applied before the subcommand runs",
|
||||
)
|
||||
}
|
||||
|
||||
// TestDispatch_MissingDotEnvIsFine pins the case most deployments are
|
||||
// in: no .env at all, which must stay a normal start.
|
||||
//
|
||||
//nolint:paralleltest // t.Chdir moves the whole process.
|
||||
func TestDispatch_MissingDotEnvIsFine(t *testing.T) {
|
||||
t.Chdir(t.TempDir())
|
||||
|
||||
var stdout, stderr bytes.Buffer
|
||||
|
||||
code := dispatch(
|
||||
[]string{helpCommand}, strings.NewReader(""), &stdout, &stderr,
|
||||
)
|
||||
|
||||
require.Equal(t, 0, code)
|
||||
assert.Empty(t, stderr.String())
|
||||
}
|
||||
+22
-1
@@ -54,6 +54,11 @@ const stopTimeout = 5 * time.Second
|
||||
// caller can tell "called wrong" from "declined".
|
||||
const exitUsage = 2
|
||||
|
||||
// helpCommand is the subcommand that prints usage. The flag spellings
|
||||
// beside it in the switch are aliases; this is the name the usage text
|
||||
// documents and the one tests invoke.
|
||||
const helpCommand = "help"
|
||||
|
||||
// Build-time variables set via -ldflags.
|
||||
//
|
||||
//nolint:gochecknoglobals // Build-time variables injected by the linker.
|
||||
@@ -75,11 +80,27 @@ func main() {
|
||||
// every existing deployment invoke; that path is unchanged, including
|
||||
// where the DATA_DIR lock is taken relative to building the fx graph
|
||||
// and how fx propagates a non-zero exit itself.
|
||||
//
|
||||
// The optional .env file is read here, before any subcommand and so
|
||||
// before anything reads the environment — config.DataDir, which both
|
||||
// the DATA_DIR lock and resetpw call outside the fx graph, above all.
|
||||
// It used to be read from an init() in internal/config, which put it
|
||||
// earlier still but threw the error away: a single malformed line
|
||||
// applied none of the file and said nothing about it. A file that is
|
||||
// not there stays fine, since .env is optional and most deployments
|
||||
// do not have one.
|
||||
func dispatch(
|
||||
args []string,
|
||||
stdin io.Reader,
|
||||
stdout, stderr io.Writer,
|
||||
) int {
|
||||
err := config.LoadDotEnv()
|
||||
if err != nil {
|
||||
_, _ = fmt.Fprintf(stderr, "%s: %v\n", appname, err)
|
||||
|
||||
return 1
|
||||
}
|
||||
|
||||
if len(args) == 0 {
|
||||
return run(stderr)
|
||||
}
|
||||
@@ -87,7 +108,7 @@ func dispatch(
|
||||
switch args[0] {
|
||||
case resetpw.Name:
|
||||
return resetpw.Run(args[1:], stdin, stdout, stderr)
|
||||
case "help", "-h", "-help", "--help":
|
||||
case helpCommand, "-h", "-help", "--help":
|
||||
usage(stdout)
|
||||
|
||||
return 0
|
||||
|
||||
@@ -121,7 +121,7 @@ func TestDispatch_Help(t *testing.T) {
|
||||
var stdout, stderr bytes.Buffer
|
||||
|
||||
code := dispatch(
|
||||
[]string{"help"}, strings.NewReader(""), &stdout, &stderr,
|
||||
[]string{helpCommand}, strings.NewReader(""), &stdout, &stderr,
|
||||
)
|
||||
|
||||
require.Equal(t, 0, code)
|
||||
|
||||
Executable
+22
@@ -0,0 +1,22 @@
|
||||
#!/bin/sh
|
||||
# deploy/docker-entrypoint.sh: the image's ENTRYPOINT. A bind-mounted
|
||||
# data directory keeps its owner from the host, often root, and the app
|
||||
# could not write to it. Started as root, this creates DATA_DIR if
|
||||
# needed, gives it and everything in it to webhooker, sets its mode, and
|
||||
# runs the command as webhooker, so the app never runs as root. Started
|
||||
# as another user, it only runs the command.
|
||||
set -eu
|
||||
|
||||
main() {
|
||||
if [ "$(id -u)" != 0 ]; then
|
||||
exec "$@"
|
||||
fi
|
||||
|
||||
dir="${DATA_DIR:-/var/lib/webhooker}"
|
||||
mkdir -p "$dir"
|
||||
find "$dir" ! -user webhooker -exec chown -h webhooker:webhooker {} +
|
||||
chmod 750 "$dir"
|
||||
exec su-exec webhooker "$@"
|
||||
}
|
||||
|
||||
main "$@"
|
||||
+129
-22
@@ -4,6 +4,7 @@ package config
|
||||
import (
|
||||
"errors"
|
||||
"fmt"
|
||||
"io/fs"
|
||||
"log/slog"
|
||||
"net/netip"
|
||||
"os"
|
||||
@@ -11,13 +12,11 @@ import (
|
||||
"strings"
|
||||
"time"
|
||||
|
||||
"github.com/getsentry/sentry-go"
|
||||
"github.com/joho/godotenv"
|
||||
"go.uber.org/fx"
|
||||
"sneak.berlin/go/webhooker/internal/globals"
|
||||
"sneak.berlin/go/webhooker/internal/logger"
|
||||
|
||||
// Populates the environment from a ./.env file automatically for
|
||||
// development configuration. Kept in one place only (here).
|
||||
_ "github.com/joho/godotenv/autoload"
|
||||
)
|
||||
|
||||
const (
|
||||
@@ -84,6 +83,12 @@ const (
|
||||
// IPv6 prefix spends on the ::ffff:0:0/96 wrapper, so a /104
|
||||
// covers the same addresses as an IPv4 /8.
|
||||
mappedV4Offset = 96
|
||||
|
||||
// DotEnvPath is the optional file of KEY=value lines read into the
|
||||
// environment at startup, relative to the process working
|
||||
// directory. Exported so that documentation and tests name the
|
||||
// same path the loader opens.
|
||||
DotEnvPath = ".env"
|
||||
)
|
||||
|
||||
// ErrInvalidEnvironment is returned when WEBHOOKER_ENVIRONMENT
|
||||
@@ -107,6 +112,15 @@ var ErrInvalidCIDR = errors.New("invalid CIDR")
|
||||
// something that is not an IP address literal.
|
||||
var ErrInvalidBindAddress = errors.New("invalid bind address")
|
||||
|
||||
// ErrInvalidSentryDSN is returned when SENTRY_DSN is set to something
|
||||
// the Sentry SDK cannot parse as a DSN.
|
||||
var ErrInvalidSentryDSN = errors.New("invalid Sentry DSN")
|
||||
|
||||
// ErrDotEnvUnreadable is returned when the optional .env file exists
|
||||
// but cannot be read or parsed. A file that is not there is not an
|
||||
// error; a file that is there and broken is.
|
||||
var ErrDotEnvUnreadable = errors.New("unreadable .env file")
|
||||
|
||||
// ErrIncompleteMetricsAuth is returned when exactly one of
|
||||
// METRICS_USERNAME and METRICS_PASSWORD carries a value. Neither
|
||||
// fallback is acceptable: serving /metrics on the username alone
|
||||
@@ -178,9 +192,10 @@ type Config struct {
|
||||
// alwaysBlockedNetworks stays blocked no matter what is listed
|
||||
// here. That set is link-local plus the cloud metadata
|
||||
// endpoints outside it that disclose credentials or user data
|
||||
// at a provider-fixed address; it is not exhaustive of every
|
||||
// cloud's metadata address. See alwaysBlockedNetworks for the
|
||||
// authoritative list and the criterion it is built from.
|
||||
// at a provider-fixed, non-public address; it is not
|
||||
// exhaustive of every cloud's metadata address. See
|
||||
// alwaysBlockedNetworks for the authoritative list and the
|
||||
// criterion it is built from.
|
||||
AllowedEgressCIDRs []netip.Prefix
|
||||
|
||||
params *ConfigParams
|
||||
@@ -212,12 +227,62 @@ func (c *Config) MetricsAuthEnabled() bool {
|
||||
return c.MetricsUsername != "" && c.MetricsPassword != ""
|
||||
}
|
||||
|
||||
// SentryEnabled reports whether error reporting is shipped to Sentry.
|
||||
// It is the only answer to that question in the codebase: the SDK
|
||||
// initialisation, the sentryhttp middleware registration and the
|
||||
// startup log's sentryEnabled field all read this one method, so the
|
||||
// log cannot report reporting as on while nothing is sending.
|
||||
//
|
||||
// A non-empty DSN is enough because loadFromEnv already parsed it with
|
||||
// the SDK's own parser and refused to build a Config around one the
|
||||
// SDK would reject, and because initialising the SDK with a DSN that
|
||||
// parsed and failed anyway aborts the process rather than leaving this
|
||||
// true and the client absent.
|
||||
func (c *Config) SentryEnabled() bool {
|
||||
return c.SentryDSN != ""
|
||||
}
|
||||
|
||||
// envString returns the value of the named environment variable,
|
||||
// or an empty string if not set.
|
||||
func envString(key string) string {
|
||||
return os.Getenv(key)
|
||||
}
|
||||
|
||||
// LoadDotEnv reads DotEnvPath into the environment when that file is
|
||||
// present, and reports a file that is present but broken.
|
||||
//
|
||||
// It has to run before anything reads the environment, so that every
|
||||
// reader agrees on what the environment holds — the DATA_DIR lock
|
||||
// taken before the fx graph exists as much as loadFromEnv itself. A
|
||||
// variable already set in the real environment wins: godotenv never
|
||||
// overwrites one.
|
||||
//
|
||||
// A missing file is not an error. It is a development convenience and
|
||||
// most deployments set the environment directly.
|
||||
//
|
||||
// Any other failure is. godotenv parses the whole file before setting
|
||||
// anything, so a single malformed line applies none of it: every
|
||||
// variable in the file silently reverts to its default, which defeats
|
||||
// the fail-loud guarantee for all of them at once.
|
||||
func LoadDotEnv() error {
|
||||
return loadDotEnvFile(DotEnvPath)
|
||||
}
|
||||
|
||||
// loadDotEnvFile is LoadDotEnv over a named file, so tests can point
|
||||
// at a temporary one instead of the process working directory.
|
||||
func loadDotEnvFile(path string) error {
|
||||
err := godotenv.Load(path)
|
||||
if err == nil || errors.Is(err, fs.ErrNotExist) {
|
||||
return nil
|
||||
}
|
||||
|
||||
return fmt.Errorf(
|
||||
"%w: %s: %w; nothing in it was applied, so fix the file or "+
|
||||
"remove it",
|
||||
ErrDotEnvUnreadable, path, err,
|
||||
)
|
||||
}
|
||||
|
||||
// DataDir resolves DATA_DIR, applying DefaultDataDir when it is unset
|
||||
// or empty. It is exported so that entry points which must act on the
|
||||
// data directory before the fx graph exists — taking the exclusive
|
||||
@@ -462,6 +527,41 @@ func envBindAddress(key, defaultValue string) (string, error) {
|
||||
return addr.String(), nil
|
||||
}
|
||||
|
||||
// envSentryDSN returns the value of the named environment variable
|
||||
// checked as a Sentry DSN. An unset (or empty, or whitespace-only)
|
||||
// value yields "", which means error reporting stays off — the common
|
||||
// case, and a normal start.
|
||||
//
|
||||
// A set value is parsed with sentry.NewDsn, which is the call
|
||||
// sentry.Init makes on the DSN it is handed, so what passes here is
|
||||
// exactly what the SDK will accept later and the two cannot disagree.
|
||||
// Reproducing the check by hand instead would cost this package its
|
||||
// dependency on the SDK — already a module dependency, already linked
|
||||
// into the binary — in exchange for a second definition of "valid DSN"
|
||||
// free to drift from the one that decides.
|
||||
//
|
||||
// A set value that does not parse is a hard error naming the key, so
|
||||
// startup fails loudly. Losing error reporting is the failure this
|
||||
// variable exists to prevent, and a typo in a DSN is silent forever:
|
||||
// nothing later in the process can notice that reports are going
|
||||
// nowhere. The bad value is quoted because it is a URL to a public
|
||||
// endpoint carrying a public key, not a secret.
|
||||
func envSentryDSN(key string) (string, error) {
|
||||
v := strings.TrimSpace(os.Getenv(key))
|
||||
if v == "" {
|
||||
return "", nil
|
||||
}
|
||||
|
||||
_, err := sentry.NewDsn(v)
|
||||
if err != nil {
|
||||
return "", fmt.Errorf(
|
||||
"%w: %s: %q: %w", ErrInvalidSentryDSN, key, v, err,
|
||||
)
|
||||
}
|
||||
|
||||
return v, nil
|
||||
}
|
||||
|
||||
// resolveMetricsAuth reads the /metrics basic-auth credentials and
|
||||
// rejects a half-set pair, naming both variables either way. The
|
||||
// error carries neither value: the password is a secret.
|
||||
@@ -486,12 +586,14 @@ func resolveMetricsAuth() (string, string, error) {
|
||||
)
|
||||
}
|
||||
|
||||
// resolveEnvironment reads WEBHOOKER_ENVIRONMENT, defaulting to
|
||||
// dev, and rejects unrecognised values.
|
||||
// resolveEnvironment reads WEBHOOKER_ENVIRONMENT, defaulting to prod
|
||||
// when it is unset so a deployment that forgets the variable is not
|
||||
// silently permissive; dev must be set explicitly. It rejects
|
||||
// unrecognised values.
|
||||
func resolveEnvironment() (string, error) {
|
||||
environment := os.Getenv("WEBHOOKER_ENVIRONMENT")
|
||||
if environment == "" {
|
||||
environment = EnvironmentDev
|
||||
environment = EnvironmentProd
|
||||
}
|
||||
|
||||
if environment != EnvironmentDev &&
|
||||
@@ -594,6 +696,11 @@ func loadFromEnv() (*Config, error) {
|
||||
return nil, err
|
||||
}
|
||||
|
||||
sentryDSN, err := envSentryDSN("SENTRY_DSN")
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
|
||||
return &Config{
|
||||
DataDir: DataDir(),
|
||||
Debug: debug,
|
||||
@@ -603,7 +710,7 @@ func loadFromEnv() (*Config, error) {
|
||||
MetricsPassword: metricsPassword,
|
||||
Port: port,
|
||||
BindAddress: bindAddress,
|
||||
SentryDSN: envString("SENTRY_DSN"),
|
||||
SentryDSN: sentryDSN,
|
||||
RetentionSweepInterval: retentionSweepInterval,
|
||||
SessionIdleTimeout: sessionIdleTimeout,
|
||||
ReceiverRateLimit: receiverRateLimit,
|
||||
@@ -640,12 +747,14 @@ func (c *Config) warnEgressAllowlist(log *slog.Logger) {
|
||||
|
||||
log.Warn(
|
||||
"ALLOWED_EGRESS_CIDRS lets delivery targets reach these "+
|
||||
"otherwise-blocked private/reserved networks. Anyone "+
|
||||
"who can create a delivery target can now make this "+
|
||||
"process issue requests into them, and read back the "+
|
||||
"response. Link-local and the known cloud instance "+
|
||||
"metadata endpoints outside it stay blocked "+
|
||||
"regardless of what is listed here.",
|
||||
"otherwise-blocked networks. Anyone who can create a "+
|
||||
"delivery target can now make this process issue "+
|
||||
"requests into them, and read back the response. Only "+
|
||||
"the addresses the README lists as blocked "+
|
||||
"unconditionally stay blocked regardless of what is "+
|
||||
"listed here; a public cloud metadata address such as "+
|
||||
"168.63.129.16 is reachable once it, or a block "+
|
||||
"covering it, is listed.",
|
||||
"allowedEgressCIDRs",
|
||||
strings.Join(PrefixStrings(c.AllowedEgressCIDRs), ","),
|
||||
)
|
||||
@@ -668,10 +777,8 @@ func (c *Config) warnEgressAllowlist(log *slog.Logger) {
|
||||
// everyone else's wrong passwords, and the receiver's limits become
|
||||
// service-wide ceilings.
|
||||
//
|
||||
// The warning is deliberately not gated on WEBHOOKER_ENVIRONMENT. That
|
||||
// variable defaults to dev, so gating on it would silence the warning
|
||||
// for exactly the operator who forgot to configure the deployment —
|
||||
// the case it exists to catch.
|
||||
// The warning is deliberately not gated on WEBHOOKER_ENVIRONMENT:
|
||||
// behind a proxy every client shares one bucket in dev and prod alike.
|
||||
//
|
||||
// The default of trusting nobody is deliberate — trusting forwarded
|
||||
// headers from arbitrary peers lets any client choose its own bucket —
|
||||
@@ -738,7 +845,7 @@ func New(lc fx.Lifecycle, params ConfigParams) (*Config, error) {
|
||||
"receiverRateLimit", s.ReceiverRateLimit,
|
||||
"trustedProxies", len(s.TrustedProxies),
|
||||
"allowedEgressCIDRs", len(s.AllowedEgressCIDRs),
|
||||
"hasSentryDSN", s.SentryDSN != "",
|
||||
"sentryEnabled", s.SentryEnabled(),
|
||||
"hasMetricsAuth", s.MetricsAuthEnabled(),
|
||||
)
|
||||
|
||||
|
||||
@@ -44,9 +44,9 @@ func TestEnvironmentConfig(t *testing.T) {
|
||||
isProd bool
|
||||
}{
|
||||
{
|
||||
name: "default is dev",
|
||||
isDev: true,
|
||||
isProd: false,
|
||||
name: "default is prod",
|
||||
isDev: false,
|
||||
isProd: true,
|
||||
},
|
||||
{
|
||||
name: "explicit dev",
|
||||
@@ -834,12 +834,13 @@ func TestEgressAllowlistWarning(t *testing.T) {
|
||||
// to be able to read back which networks are open.
|
||||
assert.Contains(t, logged, "10.0.0.0/8")
|
||||
assert.Contains(t, logged, "127.0.0.0/8")
|
||||
// What stays shut. Asserted on the clause naming the
|
||||
// wider set rather than on "Link-local" alone, so the
|
||||
// string cannot narrow back to link-local only while
|
||||
// the always-blocked set covers ULA, CGNAT and two
|
||||
// public metadata addresses as well.
|
||||
assert.Contains(t, logged, "metadata endpoints outside it")
|
||||
// What stays shut is the whole unconditional set, not
|
||||
// link-local alone; a public metadata address is not in
|
||||
// it, so a listed block covering it opens it.
|
||||
assert.Contains(t, logged, "blocked unconditionally")
|
||||
assert.Contains(t, logged, "168.63.129.16 is reachable")
|
||||
// The listed blocks need not be private or reserved.
|
||||
assert.NotContains(t, logged, "private/reserved")
|
||||
})
|
||||
}
|
||||
}
|
||||
@@ -848,10 +849,9 @@ func TestEgressAllowlistWarning(t *testing.T) {
|
||||
// tells an operator a deployment behind a reverse proxy shares one
|
||||
// rate-limit bucket between every client, which turns the receiver
|
||||
// limits into service-wide ceilings and collapses login failure
|
||||
// counting. It must fire whenever TRUSTED_PROXIES is empty,
|
||||
// in any environment: WEBHOOKER_ENVIRONMENT defaults to dev, so gating
|
||||
// on it would silence the warning for exactly the operator who never
|
||||
// configured the deployment. It stays quiet once proxies are named.
|
||||
// counting. It must fire whenever TRUSTED_PROXIES is empty, in any
|
||||
// environment, because behind a proxy every client shares one bucket
|
||||
// in dev and prod alike. It stays quiet once proxies are named.
|
||||
func TestSharedRateLimitBucketWarning(t *testing.T) {
|
||||
tests := []struct {
|
||||
name string
|
||||
@@ -871,10 +871,6 @@ func TestSharedRateLimitBucketWarning(t *testing.T) {
|
||||
expectWarning: false,
|
||||
},
|
||||
{
|
||||
// The default environment. An internet-exposed
|
||||
// deployment whose operator never set
|
||||
// WEBHOOKER_ENVIRONMENT lands here and has exactly
|
||||
// the exposure the warning announces.
|
||||
name: "dev without trusted proxies warns",
|
||||
environment: config.EnvironmentDev,
|
||||
expectWarning: true,
|
||||
|
||||
@@ -0,0 +1,158 @@
|
||||
package config_test
|
||||
|
||||
import (
|
||||
"os"
|
||||
"path/filepath"
|
||||
"testing"
|
||||
|
||||
"github.com/stretchr/testify/assert"
|
||||
"github.com/stretchr/testify/require"
|
||||
"sneak.berlin/go/webhooker/internal/config"
|
||||
)
|
||||
|
||||
// dotEnvKey is a throwaway variable name the .env tests write and
|
||||
// read, so they cannot disturb real configuration.
|
||||
const dotEnvKey = "WEBHOOKER_TEST_DOTENV_VALUE"
|
||||
|
||||
// malformedDotEnv is a file godotenv cannot parse. The first line is
|
||||
// the realistic typo — a space where the `=` belongs — and the rest
|
||||
// make sure nothing downstream treats the file as salvageable line by
|
||||
// line.
|
||||
const malformedDotEnv = "PORT 19615\n" +
|
||||
"this is not = valid ! syntax\n" +
|
||||
"\"unclosed\n"
|
||||
|
||||
// unsetDotEnvKey makes dotEnvKey genuinely absent for the duration of
|
||||
// the test and restores it afterwards. t.Setenv registers the restore;
|
||||
// the Unsetenv that follows is what the test actually needs, because a
|
||||
// variable set to the empty string is still present in os.Environ and
|
||||
// godotenv would refuse to overwrite it.
|
||||
func unsetDotEnvKey(t *testing.T) {
|
||||
t.Helper()
|
||||
t.Setenv(dotEnvKey, "placeholder")
|
||||
require.NoError(t, os.Unsetenv(dotEnvKey))
|
||||
}
|
||||
|
||||
// writeDotEnv writes contents to a .env file in a fresh temporary
|
||||
// directory and returns its path.
|
||||
func writeDotEnv(t *testing.T, contents string) string {
|
||||
t.Helper()
|
||||
|
||||
path := filepath.Join(t.TempDir(), config.DotEnvPath)
|
||||
require.NoError(t, os.WriteFile(path, []byte(contents), 0o600))
|
||||
|
||||
return path
|
||||
}
|
||||
|
||||
// TestLoadDotEnv_MissingFileIsFine pins the case most deployments are
|
||||
// in. The file is optional: it is a development convenience, and a
|
||||
// deployment that configures the environment directly must start
|
||||
// normally rather than be refused for a file it was never meant to
|
||||
// have.
|
||||
//
|
||||
//nolint:paralleltest // unsetDotEnvKey uses t.Setenv.
|
||||
func TestLoadDotEnv_MissingFileIsFine(t *testing.T) {
|
||||
unsetDotEnvKey(t)
|
||||
|
||||
absent := filepath.Join(t.TempDir(), config.DotEnvPath)
|
||||
require.NoError(t, config.LoadDotEnvFileForTest(absent))
|
||||
|
||||
_, present := os.LookupEnv(dotEnvKey)
|
||||
assert.False(t, present, "nothing may be set from an absent file")
|
||||
}
|
||||
|
||||
// TestLoadDotEnv_AppliesValues pins that a well-formed file still
|
||||
// reaches the environment, which is the whole reason the file is read
|
||||
// at all.
|
||||
//
|
||||
//nolint:paralleltest // unsetDotEnvKey uses t.Setenv.
|
||||
func TestLoadDotEnv_AppliesValues(t *testing.T) {
|
||||
unsetDotEnvKey(t)
|
||||
|
||||
path := writeDotEnv(t, "# a comment\n"+dotEnvKey+"=from-dot-env\n")
|
||||
|
||||
require.NoError(t, config.LoadDotEnvFileForTest(path))
|
||||
assert.Equal(t, "from-dot-env", os.Getenv(dotEnvKey))
|
||||
}
|
||||
|
||||
// TestLoadDotEnv_RealEnvironmentWins pins that the file cannot
|
||||
// override a variable the process was actually started with. A
|
||||
// deployment that sets DATA_DIR in its unit file must not have it
|
||||
// silently replaced by a stale .env left in the working directory.
|
||||
func TestLoadDotEnv_RealEnvironmentWins(t *testing.T) {
|
||||
t.Setenv(dotEnvKey, "from-environment")
|
||||
|
||||
path := writeDotEnv(t, dotEnvKey+"=from-dot-env\n")
|
||||
|
||||
require.NoError(t, config.LoadDotEnvFileForTest(path))
|
||||
assert.Equal(t, "from-environment", os.Getenv(dotEnvKey))
|
||||
}
|
||||
|
||||
// TestLoadDotEnv_MalformedFileAborts is the defect this fixes. One bad
|
||||
// line makes godotenv apply none of the file, so every variable in it
|
||||
// reverts to its default; the process used to start that way with no
|
||||
// log line naming the file at all.
|
||||
//
|
||||
//nolint:paralleltest // unsetDotEnvKey uses t.Setenv.
|
||||
func TestLoadDotEnv_MalformedFileAborts(t *testing.T) {
|
||||
unsetDotEnvKey(t)
|
||||
|
||||
path := writeDotEnv(
|
||||
t, malformedDotEnv+dotEnvKey+"=from-dot-env\n",
|
||||
)
|
||||
|
||||
err := config.LoadDotEnvFileForTest(path)
|
||||
|
||||
require.Error(t, err)
|
||||
require.ErrorIs(t, err, config.ErrDotEnvUnreadable)
|
||||
assert.Contains(
|
||||
t, err.Error(), config.DotEnvPath,
|
||||
"the failure must name the file it could not read",
|
||||
)
|
||||
|
||||
_, present := os.LookupEnv(dotEnvKey)
|
||||
assert.False(
|
||||
t, present,
|
||||
"a rejected file must apply nothing, not part of itself",
|
||||
)
|
||||
}
|
||||
|
||||
// TestLoadDotEnv_UnreadableFileAborts pins that only absence is
|
||||
// tolerated. A .env that exists but cannot be read is a file the
|
||||
// operator meant to be applied, so it fails like a malformed one
|
||||
// rather than being treated as though it were not there.
|
||||
func TestLoadDotEnv_UnreadableFileAborts(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
// A directory in the file's place: open succeeds and the read
|
||||
// fails, which no umask or root-ness can turn back into success
|
||||
// the way a chmod could.
|
||||
path := filepath.Join(t.TempDir(), config.DotEnvPath)
|
||||
require.NoError(t, os.Mkdir(path, 0o750))
|
||||
|
||||
err := config.LoadDotEnvFileForTest(path)
|
||||
|
||||
require.Error(t, err)
|
||||
require.ErrorIs(t, err, config.ErrDotEnvUnreadable)
|
||||
}
|
||||
|
||||
// TestLoadDotEnv_ReadsTheWorkingDirectory pins the path LoadDotEnv
|
||||
// itself opens, which the tests above bypass. It is relative to the
|
||||
// process working directory, as it was under godotenv/autoload and as
|
||||
// the README documents.
|
||||
//
|
||||
//nolint:paralleltest // t.Chdir moves the whole process.
|
||||
func TestLoadDotEnv_ReadsTheWorkingDirectory(t *testing.T) {
|
||||
unsetDotEnvKey(t)
|
||||
|
||||
dir := t.TempDir()
|
||||
require.NoError(t, os.WriteFile(
|
||||
filepath.Join(dir, config.DotEnvPath),
|
||||
[]byte(dotEnvKey+"=from-working-directory\n"),
|
||||
0o600,
|
||||
))
|
||||
t.Chdir(dir)
|
||||
|
||||
require.NoError(t, config.LoadDotEnv())
|
||||
assert.Equal(t, "from-working-directory", os.Getenv(dotEnvKey))
|
||||
}
|
||||
+82
-23
@@ -518,8 +518,20 @@ type badEnvValueCase struct {
|
||||
}
|
||||
|
||||
// badEnvValueCases is the config.New table, kept out of the test body
|
||||
// so the test itself stays readable.
|
||||
// so the test itself stays readable. It is assembled from per-variable
|
||||
// groups because one literal covering every variable outgrew the
|
||||
// function-length budget.
|
||||
func badEnvValueCases() []badEnvValueCase {
|
||||
cases := listenerEnvValueCases()
|
||||
cases = append(cases, flagEnvValueCases()...)
|
||||
cases = append(cases, sentryEnvValueCases()...)
|
||||
|
||||
return cases
|
||||
}
|
||||
|
||||
// listenerEnvValueCases covers the two variables that describe the
|
||||
// HTTP listener.
|
||||
func listenerEnvValueCases() []badEnvValueCase {
|
||||
return []badEnvValueCase{
|
||||
{
|
||||
name: "valid PORT is used",
|
||||
@@ -542,27 +554,6 @@ func badEnvValueCases() []badEnvValueCase {
|
||||
value: "70000",
|
||||
expectError: true,
|
||||
},
|
||||
{
|
||||
name: "valid DEBUG is used",
|
||||
key: envKeyDebug,
|
||||
value: "true",
|
||||
check: func(t *testing.T, cfg *config.Config) {
|
||||
t.Helper()
|
||||
assert.True(t, cfg.Debug)
|
||||
},
|
||||
},
|
||||
{
|
||||
name: "unparseable DEBUG aborts startup",
|
||||
key: envKeyDebug,
|
||||
value: "ture",
|
||||
expectError: true,
|
||||
},
|
||||
{
|
||||
name: "unparseable MAINTENANCE_MODE aborts startup",
|
||||
key: envKeyMaintenanceMode,
|
||||
value: "sometimes",
|
||||
expectError: true,
|
||||
},
|
||||
{
|
||||
name: "valid BIND_ADDRESS is used",
|
||||
key: envKeyBindAddress,
|
||||
@@ -595,6 +586,69 @@ func badEnvValueCases() []badEnvValueCase {
|
||||
}
|
||||
}
|
||||
|
||||
// flagEnvValueCases covers the boolean variables.
|
||||
func flagEnvValueCases() []badEnvValueCase {
|
||||
return []badEnvValueCase{
|
||||
{
|
||||
name: "valid DEBUG is used",
|
||||
key: envKeyDebug,
|
||||
value: "true",
|
||||
check: func(t *testing.T, cfg *config.Config) {
|
||||
t.Helper()
|
||||
assert.True(t, cfg.Debug)
|
||||
},
|
||||
},
|
||||
{
|
||||
name: "unparseable DEBUG aborts startup",
|
||||
key: envKeyDebug,
|
||||
value: "ture",
|
||||
expectError: true,
|
||||
},
|
||||
{
|
||||
name: "unparseable MAINTENANCE_MODE aborts startup",
|
||||
key: envKeyMaintenanceMode,
|
||||
value: "sometimes",
|
||||
expectError: true,
|
||||
},
|
||||
}
|
||||
}
|
||||
|
||||
// sentryEnvValueCases covers SENTRY_DSN. The three rejected values are
|
||||
// the ones measured on the defect: each initialised the SDK with an
|
||||
// error and left the process serving with error reporting off.
|
||||
func sentryEnvValueCases() []badEnvValueCase {
|
||||
return []badEnvValueCase{
|
||||
{
|
||||
name: "valid SENTRY_DSN is used",
|
||||
key: envKeySentryDSN,
|
||||
value: validSentryDSN,
|
||||
check: func(t *testing.T, cfg *config.Config) {
|
||||
t.Helper()
|
||||
assert.Equal(t, validSentryDSN, cfg.SentryDSN)
|
||||
assert.True(t, cfg.SentryEnabled())
|
||||
},
|
||||
},
|
||||
{
|
||||
name: "unparseable SENTRY_DSN aborts startup",
|
||||
key: envKeySentryDSN,
|
||||
value: "not-a-dsn",
|
||||
expectError: true,
|
||||
},
|
||||
{
|
||||
name: "SENTRY_DSN that is not a URL aborts startup",
|
||||
key: envKeySentryDSN,
|
||||
value: "%%%",
|
||||
expectError: true,
|
||||
},
|
||||
{
|
||||
name: "keyless SENTRY_DSN aborts startup",
|
||||
key: envKeySentryDSN,
|
||||
value: "https://example.invalid/1",
|
||||
expectError: true,
|
||||
},
|
||||
}
|
||||
}
|
||||
|
||||
// TestNewUsesDefaultsWhenUnset proves the fail-loud behaviour did not
|
||||
// break the legitimate unset case: absent variables still get their
|
||||
// documented defaults.
|
||||
@@ -603,7 +657,7 @@ func TestNewUsesDefaultsWhenUnset(t *testing.T) {
|
||||
|
||||
for _, key := range []string{
|
||||
envKeyPort, envKeyDebug, envKeyMaintenanceMode,
|
||||
envKeyBindAddress,
|
||||
envKeyBindAddress, envKeySentryDSN,
|
||||
} {
|
||||
require.NoError(t, os.Unsetenv(key))
|
||||
}
|
||||
@@ -626,4 +680,9 @@ func TestNewUsesDefaultsWhenUnset(t *testing.T) {
|
||||
t, config.DefaultBindAddressForTest, cfg.BindAddress,
|
||||
)
|
||||
assert.Equal(t, bindAddressDefault, cfg.BindAddress)
|
||||
|
||||
// An absent SENTRY_DSN is the common case and must stay a normal
|
||||
// start with error reporting off, not a refusal.
|
||||
assert.Empty(t, cfg.SentryDSN)
|
||||
assert.False(t, cfg.SentryEnabled())
|
||||
}
|
||||
|
||||
@@ -51,6 +51,18 @@ func EnvPortForTest(key string, defaultValue int) (int, error) {
|
||||
return envPort(key, defaultValue)
|
||||
}
|
||||
|
||||
// EnvSentryDSNForTest exposes envSentryDSN.
|
||||
func EnvSentryDSNForTest(key string) (string, error) {
|
||||
return envSentryDSN(key)
|
||||
}
|
||||
|
||||
// LoadDotEnvFileForTest exposes the loader LoadDotEnv runs, over a
|
||||
// caller-named file rather than the process working directory, so
|
||||
// each .env state can be covered without moving the test process.
|
||||
func LoadDotEnvFileForTest(path string) error {
|
||||
return loadDotEnvFile(path)
|
||||
}
|
||||
|
||||
// EnvBindAddressForTest exposes envBindAddress.
|
||||
func EnvBindAddressForTest(key, defaultValue string) (string, error) {
|
||||
return envBindAddress(key, defaultValue)
|
||||
|
||||
@@ -0,0 +1,141 @@
|
||||
package config_test
|
||||
|
||||
import (
|
||||
"os"
|
||||
"testing"
|
||||
|
||||
"github.com/stretchr/testify/assert"
|
||||
"github.com/stretchr/testify/require"
|
||||
"sneak.berlin/go/webhooker/internal/config"
|
||||
)
|
||||
|
||||
// envKeySentryDSN is the variable envSentryDSN reads in production.
|
||||
const envKeySentryDSN = "SENTRY_DSN"
|
||||
|
||||
// validSentryDSN is a syntactically complete DSN. The host is under
|
||||
// .invalid (RFC 2606), so nothing a test builds around it can reach a
|
||||
// real Sentry installation.
|
||||
const validSentryDSN = "https://abc123@sentry.invalid/42"
|
||||
|
||||
// envSentryDSNCase is one row of the envSentryDSN table.
|
||||
type envSentryDSNCase struct {
|
||||
name string
|
||||
set bool
|
||||
value string
|
||||
expectError bool
|
||||
expected string
|
||||
}
|
||||
|
||||
// envSentryDSNCases is the envSentryDSN table. The three invalid
|
||||
// values are the ones measured on the defect: each initialised the SDK
|
||||
// with an error and left the process serving with reporting off.
|
||||
func envSentryDSNCases() []envSentryDSNCase {
|
||||
return []envSentryDSNCase{
|
||||
{
|
||||
name: "unset means reporting off",
|
||||
expected: "",
|
||||
},
|
||||
{
|
||||
name: "empty means reporting off",
|
||||
set: true,
|
||||
value: "",
|
||||
expected: "",
|
||||
},
|
||||
{
|
||||
name: "whitespace means reporting off",
|
||||
set: true,
|
||||
value: " ",
|
||||
expected: "",
|
||||
},
|
||||
{
|
||||
name: "a valid DSN is kept",
|
||||
set: true,
|
||||
value: validSentryDSN,
|
||||
expected: validSentryDSN,
|
||||
},
|
||||
{
|
||||
name: "surrounding whitespace is trimmed",
|
||||
set: true,
|
||||
value: " " + validSentryDSN + "\t",
|
||||
expected: validSentryDSN,
|
||||
},
|
||||
{
|
||||
name: "a value that is not a URL is rejected",
|
||||
set: true,
|
||||
value: "not-a-dsn",
|
||||
expectError: true,
|
||||
},
|
||||
{
|
||||
name: "an unparseable URL is rejected",
|
||||
set: true,
|
||||
value: "%%%",
|
||||
expectError: true,
|
||||
},
|
||||
{
|
||||
name: "a DSN without a public key is rejected",
|
||||
set: true,
|
||||
value: "https://example.invalid/1",
|
||||
expectError: true,
|
||||
},
|
||||
{
|
||||
name: "a DSN without a project id is rejected",
|
||||
set: true,
|
||||
value: "https://abc123@sentry.invalid/",
|
||||
expectError: true,
|
||||
},
|
||||
{
|
||||
name: "a non-HTTP scheme is rejected",
|
||||
set: true,
|
||||
value: "ftp://abc123@sentry.invalid/42",
|
||||
expectError: true,
|
||||
},
|
||||
}
|
||||
}
|
||||
|
||||
// TestEnvSentryDSN covers the helper directly. What it pins beyond the
|
||||
// value is the failure shape: a set-but-unparseable DSN names the
|
||||
// variable and the value, exactly as the other fail-loud helpers do,
|
||||
// so an operator reads the fix off the message.
|
||||
func TestEnvSentryDSN(t *testing.T) {
|
||||
for _, tt := range envSentryDSNCases() {
|
||||
t.Run(tt.name, func(t *testing.T) {
|
||||
// Cannot use t.Parallel() here because t.Setenv
|
||||
// is incompatible with parallel subtests.
|
||||
if tt.set {
|
||||
t.Setenv(envKeySentryDSN, tt.value)
|
||||
} else {
|
||||
require.NoError(t, os.Unsetenv(envKeySentryDSN))
|
||||
}
|
||||
|
||||
got, err := config.EnvSentryDSNForTest(envKeySentryDSN)
|
||||
|
||||
if tt.expectError {
|
||||
require.Error(t, err)
|
||||
require.ErrorIs(t, err, config.ErrInvalidSentryDSN)
|
||||
assert.Contains(t, err.Error(), envKeySentryDSN)
|
||||
assert.Contains(t, err.Error(), tt.value)
|
||||
assert.Empty(t, got)
|
||||
|
||||
return
|
||||
}
|
||||
|
||||
require.NoError(t, err)
|
||||
assert.Equal(t, tt.expected, got)
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// TestSentryEnabled_TracksTheDSN pins that the one method answering
|
||||
// "is anything being reported" agrees with the DSN in every state. The
|
||||
// startup log, the SDK initialisation and the sentryhttp middleware
|
||||
// all read it, so a log field cannot report reporting as on while
|
||||
// nothing is sending.
|
||||
func TestSentryEnabled_TracksTheDSN(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
assert.False(t, (&config.Config{}).SentryEnabled())
|
||||
assert.True(
|
||||
t,
|
||||
(&config.Config{SentryDSN: validSentryDSN}).SentryEnabled(),
|
||||
)
|
||||
}
|
||||
@@ -3,6 +3,8 @@ package database_test
|
||||
import (
|
||||
"bytes"
|
||||
"context"
|
||||
"log/slog"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
@@ -83,3 +85,37 @@ func TestFirstBoot_PrintsTheAdminPasswordAsABanner(t *testing.T) {
|
||||
t, ok, "the printed password must open the seeded account",
|
||||
)
|
||||
}
|
||||
|
||||
// TestNewDatabase_IsLoggedWithItsPath is the log half of
|
||||
// https://git.eeqj.de/sneak/webhooker/issues/359. A DATA_DIR that is
|
||||
// unexpectedly empty boots exactly like a first start, so the start
|
||||
// that creates the database must say so, and where. Opening that
|
||||
// database again must not.
|
||||
func TestNewDatabase_IsLoggedWithItsPath(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
dir := t.TempDir()
|
||||
|
||||
open := func() string {
|
||||
var out bytes.Buffer
|
||||
|
||||
db, err := database.Open(dir, slog.New(slog.NewTextHandler(&out, nil)))
|
||||
require.NoError(t, err)
|
||||
require.NoError(t, db.Close())
|
||||
|
||||
return out.String()
|
||||
}
|
||||
|
||||
const created = `level=WARN msg="created a new, empty database"`
|
||||
|
||||
first := open()
|
||||
second := open()
|
||||
|
||||
assert.Contains(
|
||||
t, first,
|
||||
created+" path="+filepath.Join(dir, database.MainDBFileName),
|
||||
)
|
||||
assert.NotContains(
|
||||
t, second, created, "an existing database is not new",
|
||||
)
|
||||
}
|
||||
|
||||
@@ -4,11 +4,11 @@ package database
|
||||
import (
|
||||
"context"
|
||||
"crypto/rand"
|
||||
"database/sql"
|
||||
"encoding/base64"
|
||||
"errors"
|
||||
"fmt"
|
||||
"io"
|
||||
"io/fs"
|
||||
"log/slog"
|
||||
"os"
|
||||
"path/filepath"
|
||||
@@ -16,15 +16,14 @@ import (
|
||||
"go.uber.org/fx"
|
||||
"gorm.io/driver/sqlite"
|
||||
"gorm.io/gorm"
|
||||
_ "modernc.org/sqlite" // Pure Go SQLite driver
|
||||
"sneak.berlin/go/webhooker/internal/banner"
|
||||
"sneak.berlin/go/webhooker/internal/config"
|
||||
"sneak.berlin/go/webhooker/internal/datadir"
|
||||
"sneak.berlin/go/webhooker/internal/gormlog"
|
||||
"sneak.berlin/go/webhooker/internal/logger"
|
||||
)
|
||||
|
||||
const (
|
||||
dataDirPerm = 0750
|
||||
randomPasswordLen = 16
|
||||
sessionKeyLen = 32
|
||||
)
|
||||
@@ -187,7 +186,9 @@ func (d *Database) connect() error {
|
||||
// caller's decision.
|
||||
func (d *Database) connectTo(dataDir string) error {
|
||||
// Ensure the data directory exists before opening the database.
|
||||
err := os.MkdirAll(dataDir, dataDirPerm)
|
||||
// datadir.DirPerm is the single source of the directory mode; this
|
||||
// package creates the directory too, since either may run first.
|
||||
err := os.MkdirAll(dataDir, datadir.DirPerm)
|
||||
if err != nil {
|
||||
return fmt.Errorf(
|
||||
"creating data directory %s: %w",
|
||||
@@ -198,13 +199,17 @@ func (d *Database) connectTo(dataDir string) error {
|
||||
|
||||
// Construct the main application database path inside DATA_DIR.
|
||||
dbPath := filepath.Join(dataDir, MainDBFileName)
|
||||
dbURL := fmt.Sprintf(
|
||||
"file:%s?cache=shared&mode=rwc",
|
||||
dbPath,
|
||||
)
|
||||
|
||||
// Open the database with the pure Go SQLite driver
|
||||
sqlDB, err := sql.Open("sqlite", dbURL)
|
||||
// Checked before opening, which creates the file. A DATA_DIR that
|
||||
// is unexpectedly empty -- its volume not mounted, say -- looks
|
||||
// exactly like a first start, so a new database is a warning.
|
||||
_, statErr := os.Stat(dbPath)
|
||||
created := errors.Is(statErr, fs.ErrNotExist)
|
||||
|
||||
// Opened through OpenSQLite so this handle carries the same WAL
|
||||
// journaling, busy timeout, immediate-transaction locking, and pool
|
||||
// bounds as every other database file. See sqlite_open.go.
|
||||
sqlDB, err := OpenSQLite(dbPath, SQLiteModeCreate)
|
||||
if err != nil {
|
||||
d.log.Error(
|
||||
"failed to open database",
|
||||
@@ -231,7 +236,12 @@ func (d *Database) connectTo(dataDir string) error {
|
||||
}
|
||||
|
||||
d.db = db
|
||||
d.log.Info("connected to database", "path", dbPath)
|
||||
|
||||
if created {
|
||||
d.log.Warn("created a new, empty database", "path", dbPath)
|
||||
} else {
|
||||
d.log.Info("connected to database", "path", dbPath)
|
||||
}
|
||||
|
||||
// Run migrations
|
||||
return d.migrate()
|
||||
|
||||
@@ -0,0 +1,166 @@
|
||||
package database_test
|
||||
|
||||
import (
|
||||
"context"
|
||||
"fmt"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"github.com/google/uuid"
|
||||
"github.com/stretchr/testify/assert"
|
||||
"github.com/stretchr/testify/require"
|
||||
"gorm.io/gorm"
|
||||
"sneak.berlin/go/webhooker/internal/database"
|
||||
)
|
||||
|
||||
// TestWebhookDBManager_OpenAddsEventTierIndexes verifies that opening a
|
||||
// per-webhook database that predates these indexes creates them. It
|
||||
// stands in for an older database file by dropping the indexes
|
||||
// AutoMigrate just created, then reopening the same file.
|
||||
func TestWebhookDBManager_OpenAddsEventTierIndexes(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
indexes := []struct {
|
||||
model any
|
||||
name string
|
||||
}{
|
||||
{&database.Delivery{}, "idx_deliveries_status"},
|
||||
{&database.Delivery{}, "idx_deliveries_event_id"},
|
||||
{&database.DeliveryResult{}, "idx_delivery_results_delivery_id"},
|
||||
{&database.Event{}, "idx_events_deleted_at_created_at"},
|
||||
{&database.Event{}, "idx_events_created_at"},
|
||||
}
|
||||
|
||||
mgr, lc := setupTestWebhookDBManager(t)
|
||||
ctx := context.Background()
|
||||
require.NoError(t, lc.Start(ctx))
|
||||
|
||||
defer func() { require.NoError(t, lc.Stop(ctx)) }()
|
||||
|
||||
webhookID := uuid.New().String()
|
||||
|
||||
db, err := mgr.GetDB(webhookID)
|
||||
require.NoError(t, err)
|
||||
|
||||
// A fresh database has them.
|
||||
for _, ix := range indexes {
|
||||
require.True(t, db.Migrator().HasIndex(ix.model, ix.name))
|
||||
}
|
||||
|
||||
// Stand in for a database file created before the indexes existed.
|
||||
for _, ix := range indexes {
|
||||
require.NoError(t, db.Migrator().DropIndex(ix.model, ix.name))
|
||||
require.False(t, db.Migrator().HasIndex(ix.model, ix.name))
|
||||
}
|
||||
|
||||
// Drop the cached connection so the next open reopens the file and
|
||||
// runs AutoMigrate against it, as a restart would.
|
||||
require.NoError(t, mgr.CloseAll())
|
||||
|
||||
db, err = mgr.GetDB(webhookID)
|
||||
require.NoError(t, err)
|
||||
|
||||
for _, ix := range indexes {
|
||||
assert.True(t, db.Migrator().HasIndex(ix.model, ix.name),
|
||||
"opening the existing database should create %s", ix.name)
|
||||
}
|
||||
}
|
||||
|
||||
// TestEventTierQueriesUseTheirIndexes verifies that the statements the
|
||||
// indexes are for use them. GORM builds each statement in a dry run as
|
||||
// the code named above it does, soft-delete condition included, and
|
||||
// SQLite, which keeps no statistics on these tables, must plan to seek
|
||||
// on each index listed by the columns in parentheses.
|
||||
func TestEventTierQueriesUseTheirIndexes(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
mgr, lc := setupTestWebhookDBManager(t)
|
||||
ctx := context.Background()
|
||||
require.NoError(t, lc.Start(ctx))
|
||||
|
||||
defer func() { require.NoError(t, lc.Stop(ctx)) }()
|
||||
|
||||
db, err := mgr.GetDB(uuid.New().String())
|
||||
require.NoError(t, err)
|
||||
|
||||
dry := db.Session(&gorm.Session{DryRun: true})
|
||||
ids := []string{
|
||||
uuid.New().String(), uuid.New().String(), uuid.New().String(),
|
||||
}
|
||||
cutoff := time.Now()
|
||||
|
||||
var (
|
||||
deliveries []database.Delivery
|
||||
results []database.DeliveryResult
|
||||
depths []struct{ Depth int }
|
||||
)
|
||||
|
||||
byStatus := "idx_deliveries_status (status=? AND deleted_at=?)"
|
||||
byEvent := "idx_deliveries_event_id (event_id=? AND deleted_at=?)"
|
||||
byAge := "idx_events_deleted_at_created_at (deleted_at=? AND created_at<?)"
|
||||
|
||||
// The delivery engine: recovery and the retry sweep, the sweep for
|
||||
// stranded pending deliveries, and the queue depth count.
|
||||
assertPlanUses(t, db, dry.Where(
|
||||
"status = ?", database.DeliveryStatusRetrying,
|
||||
).Find(&deliveries), byStatus)
|
||||
assertPlanUses(t, db, dry.Where(
|
||||
"status = ? AND updated_at < ?",
|
||||
database.DeliveryStatusPending, cutoff,
|
||||
).Limit(500).Find(&deliveries), byStatus)
|
||||
assertPlanUses(t, db, dry.Model(&database.Delivery{}).
|
||||
Select("target_id", "status", "count(*) as depth").
|
||||
Where("status IN ?", []database.DeliveryStatus{
|
||||
database.DeliveryStatusPending,
|
||||
database.DeliveryStatusRetrying,
|
||||
}).Group("target_id, status").Find(&depths), byStatus)
|
||||
|
||||
// The event log: each event's deliveries, then their attempts
|
||||
// (loadEventsWithDeliveries, loadDeliveryResults).
|
||||
assertPlanUses(t, db, dry.Where("event_id = ?", ids[0]).
|
||||
Find(&deliveries), byEvent)
|
||||
assertPlanUses(t, db, dry.Where("delivery_id IN ?", ids).
|
||||
Order("attempt_num ASC").Find(&results),
|
||||
"idx_delivery_results_delivery_id (delivery_id=? AND deleted_at=?)")
|
||||
|
||||
// Retention's three deletes (reapExpired), whose subqueries are built
|
||||
// afresh for each statement as it builds them.
|
||||
expiredEventIDs := func() *gorm.DB {
|
||||
return dry.Model(&database.Event{}).Select("id").
|
||||
Where("created_at < ?", cutoff)
|
||||
}
|
||||
|
||||
assertPlanUses(t, db, dry.Unscoped().Where(
|
||||
"delivery_id IN (?)", dry.Model(&database.Delivery{}).
|
||||
Select("id").Where("event_id IN (?)", expiredEventIDs()),
|
||||
).Delete(&database.DeliveryResult{}),
|
||||
"idx_delivery_results_delivery_id (delivery_id=?)", byEvent, byAge)
|
||||
assertPlanUses(t, db, dry.Unscoped().Where(
|
||||
"event_id IN (?)", expiredEventIDs(),
|
||||
).Delete(&database.Delivery{}),
|
||||
"idx_deliveries_event_id (event_id=?)", byAge)
|
||||
assertPlanUses(t, db, dry.Unscoped().Where(
|
||||
"created_at < ?", cutoff,
|
||||
).Delete(&database.Event{}), "idx_events_created_at (created_at<?)")
|
||||
}
|
||||
|
||||
// assertPlanUses asserts that SQLite's plan for a statement GORM built
|
||||
// in a dry run, run with the same SQL and arguments GORM would send,
|
||||
// names each of the given indexes.
|
||||
func assertPlanUses(
|
||||
t *testing.T, db, built *gorm.DB, indexes ...string,
|
||||
) {
|
||||
t.Helper()
|
||||
|
||||
var plan []struct{ Detail string }
|
||||
|
||||
require.NoError(t, db.Raw(
|
||||
"EXPLAIN QUERY PLAN "+built.Statement.SQL.String(),
|
||||
built.Statement.Vars...,
|
||||
).Scan(&plan).Error)
|
||||
|
||||
for _, index := range indexes {
|
||||
assert.Contains(t, fmt.Sprint(plan), index,
|
||||
built.Statement.SQL.String())
|
||||
}
|
||||
}
|
||||
@@ -1,5 +1,7 @@
|
||||
package database
|
||||
|
||||
import "gorm.io/gorm"
|
||||
|
||||
// DeliveryStatus represents the status of a delivery
|
||||
type DeliveryStatus string
|
||||
|
||||
@@ -29,12 +31,19 @@ func (s DeliveryStatus) Terminal() bool {
|
||||
}
|
||||
|
||||
// Delivery represents a delivery attempt for an event to a target
|
||||
//
|
||||
//nolint:lll // a struct tag cannot wrap
|
||||
type Delivery struct {
|
||||
BaseModel
|
||||
|
||||
EventID string `gorm:"type:uuid;not null" json:"eventId"`
|
||||
TargetID string `gorm:"type:uuid;not null" json:"targetId"`
|
||||
Status DeliveryStatus `gorm:"not null;default:'pending'" json:"status"`
|
||||
EventID string `gorm:"type:uuid;not null;index:idx_deliveries_event_id,priority:1" json:"eventId"`
|
||||
TargetID string `gorm:"type:uuid;not null" json:"targetId"`
|
||||
Status DeliveryStatus `gorm:"not null;default:'pending';index:idx_deliveries_status,priority:1" json:"status"`
|
||||
|
||||
// DeletedAt repeats the BaseModel field only to be the second column
|
||||
// of the event_id and status indexes, for the reason DeliveryResult
|
||||
// gives.
|
||||
DeletedAt gorm.DeletedAt `gorm:"index:idx_deliveries_event_id,priority:2;index:idx_deliveries_status,priority:2" json:"deletedAt,omitzero"`
|
||||
|
||||
// Relations
|
||||
Event Event `json:"event,omitzero"`
|
||||
|
||||
@@ -1,10 +1,21 @@
|
||||
package database
|
||||
|
||||
import "gorm.io/gorm"
|
||||
|
||||
// DeliveryResult represents the result of a delivery attempt
|
||||
//
|
||||
//nolint:lll // a struct tag cannot wrap
|
||||
type DeliveryResult struct {
|
||||
BaseModel
|
||||
|
||||
DeliveryID string `gorm:"type:uuid;not null" json:"deliveryId"`
|
||||
// DeliveryID and DeletedAt make up one index, in that order.
|
||||
// DeletedAt repeats the BaseModel field only to join it: GORM adds
|
||||
// "deleted_at IS NULL" to almost every query, and where a column is
|
||||
// matched against several values SQLite otherwise reads through the
|
||||
// deleted_at index, which every live row matches.
|
||||
DeliveryID string `gorm:"type:uuid;not null;index:idx_delivery_results_delivery_id,priority:1" json:"deliveryId"`
|
||||
DeletedAt gorm.DeletedAt `gorm:"index:idx_delivery_results_delivery_id,priority:2" json:"deletedAt,omitzero"`
|
||||
|
||||
AttemptNum int `gorm:"not null" json:"attemptNum"`
|
||||
Success bool `json:"success"`
|
||||
StatusCode int `json:"statusCode,omitempty"`
|
||||
|
||||
@@ -1,9 +1,27 @@
|
||||
package database
|
||||
|
||||
import (
|
||||
"time"
|
||||
|
||||
"gorm.io/gorm"
|
||||
)
|
||||
|
||||
// Event represents a captured webhook event
|
||||
//
|
||||
//nolint:lll // a struct tag cannot wrap
|
||||
type Event struct {
|
||||
BaseModel
|
||||
|
||||
// CreatedAt and DeletedAt repeat the BaseModel fields only to index
|
||||
// them for retention, which finds events by age. Its lookups carry
|
||||
// GORM's "deleted_at IS NULL" (see DeliveryResult) and compare
|
||||
// created_at with <, so their index has deleted_at first: SQLite
|
||||
// narrows by a < only on the last column it uses. Its final delete
|
||||
// has no deleted_at condition and uses the index on created_at
|
||||
// alone. The other tables keep the unindexed BaseModel created_at.
|
||||
CreatedAt time.Time `gorm:"index;index:idx_events_deleted_at_created_at,priority:2" json:"createdAt"`
|
||||
DeletedAt gorm.DeletedAt `gorm:"index:idx_events_deleted_at_created_at,priority:1" json:"deletedAt,omitzero"`
|
||||
|
||||
WebhookID string `gorm:"type:uuid;not null" json:"webhookId"`
|
||||
EntrypointID string `gorm:"type:uuid;not null" json:"entrypointId"`
|
||||
|
||||
|
||||
@@ -0,0 +1,240 @@
|
||||
package database_test
|
||||
|
||||
import (
|
||||
"context"
|
||||
"io/fs"
|
||||
"net/http"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"testing"
|
||||
|
||||
"github.com/google/uuid"
|
||||
"github.com/stretchr/testify/assert"
|
||||
"github.com/stretchr/testify/require"
|
||||
"go.uber.org/fx/fxtest"
|
||||
"sneak.berlin/go/webhooker/internal/config"
|
||||
"sneak.berlin/go/webhooker/internal/database"
|
||||
"sneak.berlin/go/webhooker/internal/globals"
|
||||
"sneak.berlin/go/webhooker/internal/logger"
|
||||
)
|
||||
|
||||
// ownerOnly is the mode every SQLite file the service owns must have.
|
||||
// Spelled out rather than referencing database.SQLiteFilePerm so the
|
||||
// test fails if the constant itself is loosened.
|
||||
const ownerOnly fs.FileMode = 0o600
|
||||
|
||||
// requireOwnerOnly asserts that path exists and is readable and
|
||||
// writable by its owner and by nobody else.
|
||||
func requireOwnerOnly(t *testing.T, path string) {
|
||||
t.Helper()
|
||||
|
||||
info, err := os.Stat(path)
|
||||
require.NoError(t, err, "%s must exist", path)
|
||||
assert.Equal(
|
||||
t,
|
||||
ownerOnly,
|
||||
info.Mode().Perm(),
|
||||
"%s holds credentials and must not be readable by "+
|
||||
"anyone but its owner",
|
||||
path,
|
||||
)
|
||||
}
|
||||
|
||||
// requireDatabaseSetOwnerOnly asserts the mode of a database file and
|
||||
// of both WAL sidecars. The sidecars carry the same rows as the
|
||||
// database, so tightening only the main file fixes nothing.
|
||||
func requireDatabaseSetOwnerOnly(t *testing.T, dbPath string) {
|
||||
t.Helper()
|
||||
|
||||
requireOwnerOnly(t, dbPath)
|
||||
requireOwnerOnly(t, dbPath+"-wal")
|
||||
requireOwnerOnly(t, dbPath+"-shm")
|
||||
}
|
||||
|
||||
// TestMainDatabaseFilesAreOwnerOnly covers the tier the defect was
|
||||
// reported against: webhooker.db holds targets.config in plaintext —
|
||||
// bearer tokens, API keys, Slack webhook URLs — and the session
|
||||
// encryption key.
|
||||
func TestMainDatabaseFilesAreOwnerOnly(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
lc := fxtest.NewLifecycle(t)
|
||||
|
||||
l, err := logger.New(lc, logger.LoggerParams{
|
||||
Globals: &globals.Globals{
|
||||
Appname: testAppname,
|
||||
Version: testVersion,
|
||||
},
|
||||
})
|
||||
require.NoError(t, err)
|
||||
|
||||
// A directory the application creates itself, not one t.TempDir
|
||||
// made at 0700, so the mode below is the application's.
|
||||
dataDir := filepath.Join(t.TempDir(), "data")
|
||||
|
||||
db, err := database.New(lc, database.DatabaseParams{
|
||||
Config: &config.Config{DataDir: dataDir},
|
||||
Logger: l,
|
||||
})
|
||||
require.NoError(t, err)
|
||||
|
||||
ctx := context.Background()
|
||||
require.NoError(t, lc.Start(ctx))
|
||||
|
||||
defer func() { require.NoError(t, lc.Stop(ctx)) }()
|
||||
|
||||
// Write through the real model so the WAL is populated and both
|
||||
// sidecars are on disk while the handle is open.
|
||||
require.NoError(t, db.DB().Create(&database.Webhook{
|
||||
Name: testWebhookName,
|
||||
}).Error)
|
||||
|
||||
requireDatabaseSetOwnerOnly(
|
||||
t, filepath.Join(dataDir, database.MainDBFileName),
|
||||
)
|
||||
|
||||
// The data directory grants nothing to `other`. Asserted as a
|
||||
// property rather than as an exact 0750, because MkdirAll applies
|
||||
// the ambient umask: the exact mode is the developer's umask as
|
||||
// much as the application's request, and pinning it would make
|
||||
// `make check` pass or fail on where it is run. The group bits are
|
||||
// deliberately left unasserted — deployments may rely on them.
|
||||
info, err := os.Stat(dataDir)
|
||||
require.NoError(t, err)
|
||||
assert.Zero(
|
||||
t,
|
||||
info.Mode().Perm()&0o007,
|
||||
"the data directory must not be world-accessible",
|
||||
)
|
||||
}
|
||||
|
||||
// TestPerWebhookEventDatabaseFilesAreOwnerOnly covers the events-*.db
|
||||
// tier. These carry no credential canaries since
|
||||
// https://git.eeqj.de/sneak/webhooker/issues/206, but they hold every
|
||||
// received request body and header.
|
||||
func TestPerWebhookEventDatabaseFilesAreOwnerOnly(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
mgr, lc := setupTestWebhookDBManager(t)
|
||||
ctx := context.Background()
|
||||
require.NoError(t, lc.Start(ctx))
|
||||
|
||||
defer func() { require.NoError(t, lc.Stop(ctx)) }()
|
||||
|
||||
webhookID := uuid.New().String()
|
||||
|
||||
db, err := mgr.GetDB(webhookID)
|
||||
require.NoError(t, err)
|
||||
|
||||
require.NoError(t, db.Create(&database.Event{
|
||||
WebhookID: webhookID,
|
||||
EntrypointID: uuid.New().String(),
|
||||
Method: http.MethodPost,
|
||||
Body: "{}",
|
||||
}).Error)
|
||||
|
||||
requireDatabaseSetOwnerOnly(t, mgr.DBPath(webhookID))
|
||||
}
|
||||
|
||||
// TestArchiveDatabaseFilesAreOwnerOnly covers the archive-*.db tier.
|
||||
// internal/delivery builds that path and opens it through OpenSQLite,
|
||||
// the same single open path exercised here, so the mode is settled for
|
||||
// all three tiers in one place.
|
||||
func TestArchiveDatabaseFilesAreOwnerOnly(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
ctx := context.Background()
|
||||
path := filepath.Join(
|
||||
t.TempDir(), "archive-"+uuid.New().String()+".db",
|
||||
)
|
||||
|
||||
sqlDB, err := database.OpenSQLite(path, database.SQLiteModeCreate)
|
||||
require.NoError(t, err)
|
||||
|
||||
defer func() { require.NoError(t, sqlDB.Close()) }()
|
||||
|
||||
_, err = sqlDB.ExecContext(ctx, "create table t (id integer)")
|
||||
require.NoError(t, err)
|
||||
|
||||
requireDatabaseSetOwnerOnly(t, path)
|
||||
}
|
||||
|
||||
// TestOpenSQLiteTightensFilesLeftWorldReadable is the upgrade case: a
|
||||
// data directory an earlier build left at 0644, including a
|
||||
// developer's own scratch directory, is fixed when it is opened rather
|
||||
// than staying exposed until it is recreated.
|
||||
func TestOpenSQLiteTightensFilesLeftWorldReadable(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
dir := t.TempDir()
|
||||
path := filepath.Join(dir, database.MainDBFileName)
|
||||
|
||||
// A database and both sidecars as the pre-fix build left them.
|
||||
for _, p := range []string{path, path + "-wal", path + "-shm"} {
|
||||
require.NoError(t, os.WriteFile(p, nil, 0o644)) //nolint:gosec // the mode under test
|
||||
}
|
||||
|
||||
sqlDB, err := database.OpenSQLite(path, database.SQLiteModeCreate)
|
||||
require.NoError(t, err)
|
||||
|
||||
require.NoError(t, sqlDB.Close())
|
||||
|
||||
requireDatabaseSetOwnerOnly(t, path)
|
||||
}
|
||||
|
||||
// TestOpenSQLiteExistingModeDoesNotCreateTheFile guards the mechanism
|
||||
// the fix uses: OpenSQLite now creates the database file itself, and
|
||||
// must not do so for a caller that asked for an existing database. An
|
||||
// empty file materialized here would turn a missing-database error
|
||||
// into a silently empty one.
|
||||
func TestOpenSQLiteExistingModeDoesNotCreateTheFile(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
ctx := context.Background()
|
||||
path := filepath.Join(t.TempDir(), "absent.db")
|
||||
|
||||
sqlDB, err := database.OpenSQLite(path, database.SQLiteModeExisting)
|
||||
if err == nil {
|
||||
// sql.Open is lazy: force the connection that fails.
|
||||
require.Error(t, sqlDB.PingContext(ctx))
|
||||
require.NoError(t, sqlDB.Close())
|
||||
}
|
||||
|
||||
_, statErr := os.Stat(path)
|
||||
assert.ErrorIs(t, statErr, fs.ErrNotExist)
|
||||
}
|
||||
|
||||
// TestReopenAfterRestartKeepsFilesOwnerOnly is the restart case: a
|
||||
// process that closed its files must be able to open them again at
|
||||
// 0600, including through a gorm handle, and the sidecars must come
|
||||
// back at 0600 too rather than at SQLite's own default.
|
||||
func TestReopenAfterRestartKeepsFilesOwnerOnly(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
ctx := context.Background()
|
||||
dir := t.TempDir()
|
||||
path := filepath.Join(dir, database.MainDBFileName)
|
||||
|
||||
first, err := database.OpenSQLite(path, database.SQLiteModeCreate)
|
||||
require.NoError(t, err)
|
||||
|
||||
_, err = first.ExecContext(ctx, "create table t (id integer)")
|
||||
require.NoError(t, err)
|
||||
require.NoError(t, first.Close())
|
||||
|
||||
second, err := database.OpenSQLite(path, database.SQLiteModeCreate)
|
||||
require.NoError(t, err)
|
||||
|
||||
defer func() { require.NoError(t, second.Close()) }()
|
||||
|
||||
_, err = second.ExecContext(ctx, "insert into t (id) values (1)")
|
||||
require.NoError(t, err)
|
||||
|
||||
requireDatabaseSetOwnerOnly(t, path)
|
||||
|
||||
var got int
|
||||
|
||||
require.NoError(t,
|
||||
second.QueryRowContext(ctx, "select id from t").Scan(&got))
|
||||
assert.Equal(t, 1, got)
|
||||
}
|
||||
@@ -0,0 +1,252 @@
|
||||
package database
|
||||
|
||||
import (
|
||||
"database/sql"
|
||||
"errors"
|
||||
"fmt"
|
||||
"io/fs"
|
||||
"net/url"
|
||||
"os"
|
||||
"time"
|
||||
|
||||
_ "modernc.org/sqlite" // Pure Go SQLite driver
|
||||
)
|
||||
|
||||
// Every SQLite file this service opens — the main database, the
|
||||
// per-webhook event databases, and the archive databases — is opened
|
||||
// through OpenSQLite, so the durability settings below are properties
|
||||
// of the service rather than of one call site.
|
||||
//
|
||||
// modernc.org/sqlite installs no busy handler and issues no pragmas of
|
||||
// its own: it executes only the pragmas named in explicit `_pragma=`
|
||||
// DSN parameters, and gorm.io/driver/sqlite adds none when it is
|
||||
// handed an existing *sql.DB. Every setting therefore has to be
|
||||
// spelled out here or it is simply not in effect.
|
||||
// SQLite URI open modes.
|
||||
const (
|
||||
// SQLiteModeCreate creates the database file when it is missing.
|
||||
SQLiteModeCreate = "rwc"
|
||||
|
||||
// SQLiteModeExisting requires the file to exist already.
|
||||
SQLiteModeExisting = "rw"
|
||||
)
|
||||
|
||||
const (
|
||||
// SQLiteBusyTimeout is how long SQLite retries a lock conflict
|
||||
// before returning SQLITE_BUSY.
|
||||
//
|
||||
// Under WAL a reader never blocks a writer, so the only conflict
|
||||
// left is writer against writer: this process's delivery workers
|
||||
// against each other, or against another process holding the write
|
||||
// lock. Those clear in milliseconds. Ten seconds is far above that
|
||||
// and still well inside the receiver's request budget, so an
|
||||
// inbound webhook waits rather than being rejected with a 500.
|
||||
SQLiteBusyTimeout = 10 * time.Second
|
||||
|
||||
// sqliteMaxOpenConns bounds the connection pool for one database
|
||||
// file.
|
||||
//
|
||||
// The pool needs a bound at all because database/sql cannot detect
|
||||
// a connection left mid-transaction: modernc.org/sqlite implements
|
||||
// neither driver.Validator nor driver.SessionResetter, so a
|
||||
// connection whose COMMIT failed is returned to the pool with its
|
||||
// transaction still open and handed out again indefinitely. That is
|
||||
// what turned four `database is locked` errors into 593
|
||||
// `cannot start a transaction within a transaction` in
|
||||
// https://git.eeqj.de/sneak/webhooker/issues/256.
|
||||
//
|
||||
// Four is above the one writer SQLite allows at a time, so reads
|
||||
// still proceed while a write is in flight, and low enough that
|
||||
// contention is resolved by the busy handler rather than by piling
|
||||
// up connections against a lock only one of them can hold.
|
||||
sqliteMaxOpenConns = 4
|
||||
|
||||
// sqliteMaxIdleConns keeps the pool warm without holding every
|
||||
// connection open through an idle period.
|
||||
sqliteMaxIdleConns = 2
|
||||
|
||||
// sqliteConnMaxLifetime and sqliteConnMaxIdleTime retire pooled
|
||||
// connections on a schedule. With _txlock=immediate a failed
|
||||
// COMMIT should no longer be reachable, but these bound the damage
|
||||
// if one happens anyway: a poisoned connection is closed and
|
||||
// replaced within the lifetime instead of wedging the file until
|
||||
// the process restarts.
|
||||
sqliteConnMaxLifetime = 5 * time.Minute
|
||||
sqliteConnMaxIdleTime = time.Minute
|
||||
)
|
||||
|
||||
// SQLiteFilePerm is the mode every SQLite file this service owns is
|
||||
// created with and held at: owner read/write, nothing for group or
|
||||
// other.
|
||||
//
|
||||
// These files hold credentials in plaintext. The main database stores
|
||||
// `targets.config` — bearer tokens, API keys, Slack webhook URLs — and
|
||||
// the session encryption key. SQLite left to itself creates them 0644
|
||||
// (see reserveSQLiteFile), which made the 0750 data directory the only
|
||||
// barrier; a bind-mounted directory supplied at 0755 removes it and
|
||||
// every local user on the host can read every stored credential.
|
||||
//
|
||||
// This is a file-mode fix and not encryption at rest. An unattended
|
||||
// process needs a key it can read without a human, so the key lands
|
||||
// beside the data and an attacker who can read the database can read
|
||||
// it too. See https://git.eeqj.de/sneak/webhooker/issues/212.
|
||||
const SQLiteFilePerm fs.FileMode = 0o600
|
||||
|
||||
// reserveSQLiteFile puts path at SQLiteFilePerm before the driver ever
|
||||
// touches it, and tightens any sidecar already on disk.
|
||||
//
|
||||
// The mode has to be settled here rather than by a chmod after opening,
|
||||
// because SQLite picks it: robust_open substitutes
|
||||
// SQLITE_DEFAULT_FILE_PERMISSIONS (0644) whenever it is handed mode 0,
|
||||
// and findCreateFileMode yields 0 for a main database opened by URI
|
||||
// with no `modeof` parameter. A chmod afterwards would leave a window
|
||||
// in which the credentials are on disk world-readable.
|
||||
//
|
||||
// Creating the file ourselves also settles the sidecars, which is the
|
||||
// half that could quietly not work. SQLite does not create those at a
|
||||
// mode we choose — it derives both from the main database file:
|
||||
// `-wal` through findCreateFileMode, which stats the path with the
|
||||
// suffix stripped, and `-shm` in unixOpenSharedMemory from an fstat of
|
||||
// the already-open database descriptor. A main file at 0600 therefore
|
||||
// produces sidecars at 0600. A zero-length file is a valid empty
|
||||
// database, so reserving it changes nothing else.
|
||||
//
|
||||
// create says whether the caller is opening in a mode that may create
|
||||
// the database. When it is false a missing file is left missing, so
|
||||
// SQLite still reports the absence rather than this function
|
||||
// materializing an empty database the caller asked not to create.
|
||||
//
|
||||
// Chmod of a file that already exists is what tightens a data
|
||||
// directory an earlier build left at 0644 — including a developer's
|
||||
// own scratch directory — without any migration machinery.
|
||||
func reserveSQLiteFile(path string, create bool) error {
|
||||
if create {
|
||||
// gosec G304: the path is the database file the caller asked
|
||||
// to open, and the driver is about to open the same path
|
||||
// anyway. Creating it here is what fixes its mode.
|
||||
f, err := os.OpenFile( //nolint:gosec // see above
|
||||
path, os.O_RDWR|os.O_CREATE, SQLiteFilePerm,
|
||||
)
|
||||
if err != nil {
|
||||
return fmt.Errorf("creating %s: %w", path, err)
|
||||
}
|
||||
|
||||
err = f.Close()
|
||||
if err != nil {
|
||||
return fmt.Errorf("closing %s: %w", path, err)
|
||||
}
|
||||
}
|
||||
|
||||
// O_CREATE leaves an existing file's mode alone, and umask can only
|
||||
// have narrowed a new one. Chmod settles both cases at exactly
|
||||
// SQLiteFilePerm.
|
||||
for _, p := range append(
|
||||
[]string{path}, sqliteSidecarPaths(path)...,
|
||||
) {
|
||||
err := os.Chmod(p, SQLiteFilePerm)
|
||||
if err != nil && !errors.Is(err, fs.ErrNotExist) {
|
||||
return fmt.Errorf("securing %s: %w", p, err)
|
||||
}
|
||||
}
|
||||
|
||||
return nil
|
||||
}
|
||||
|
||||
// sqliteSidecarPaths returns the files SQLite maintains beside a
|
||||
// database under WAL. They carry the same rows as the database itself,
|
||||
// so a fix that tightens only the main file has fixed nothing.
|
||||
func sqliteSidecarPaths(path string) []string {
|
||||
return []string{path + "-wal", path + "-shm"}
|
||||
}
|
||||
|
||||
// SQLiteDSN builds the connection string for one database file.
|
||||
//
|
||||
// mode is the SQLite URI open mode: "rwc" to create the file when it
|
||||
// is missing, "rw" to require that it already exists.
|
||||
//
|
||||
// Three settings carry the fix for
|
||||
// https://git.eeqj.de/sneak/webhooker/issues/256 and none of them is
|
||||
// optional:
|
||||
//
|
||||
// - journal_mode=WAL, so a reader — an operator running
|
||||
// `sqlite3 <db> .dump` over their own data — takes a snapshot
|
||||
// instead of blocking every writer behind it.
|
||||
//
|
||||
// - busy_timeout, so a writer that does meet a lock waits for it.
|
||||
// Without one SQLite gives up immediately; nothing above it
|
||||
// retries.
|
||||
//
|
||||
// - _txlock=immediate, so every transaction takes the write lock at
|
||||
// BEGIN. A deferred transaction acquires it lazily on its first
|
||||
// write, and that upgrade returns SQLITE_BUSY *without* consulting
|
||||
// the busy handler, because SQLite cannot block a transaction that
|
||||
// may already hold a read snapshot. Such a COMMIT then fails while
|
||||
// the transaction stays open on the connection. A busy timeout
|
||||
// alone does not prevent this; BEGIN IMMEDIATE does, by putting
|
||||
// the wait somewhere the handler applies.
|
||||
//
|
||||
// Note what is absent: `cache=shared`. Under a shared cache an
|
||||
// in-process conflict is reported as SQLITE_LOCKED rather than
|
||||
// SQLITE_BUSY, and the busy handler does not retry SQLITE_LOCKED — so
|
||||
// leaving it in would have defeated the busy timeout for exactly the
|
||||
// contention this service generates. Dropping it is part of the fix,
|
||||
// not housekeeping.
|
||||
//
|
||||
// synchronous is deliberately left at SQLite's default of FULL: this
|
||||
// is a webhook receiver whose one promise is that an event it answered
|
||||
// 200 for is durable.
|
||||
// The order of the _pragma parameters is load-bearing.
|
||||
// modernc.org/sqlite executes them in the order they appear, on every
|
||||
// new connection, before the connection is handed to the pool. Setting
|
||||
// journal_mode first means that pragma itself runs with no busy
|
||||
// handler installed: the pool opens connections lazily, so the moment
|
||||
// a new one is created is a moment the database is under load, and
|
||||
// PRAGMA journal_mode takes a lock. It would fail immediately with
|
||||
// SQLITE_BUSY and fail the query that caused the connection to be
|
||||
// opened. busy_timeout is therefore set first, so every pragma after
|
||||
// it — and the whole life of the connection — is covered.
|
||||
func SQLiteDSN(path, mode string) string {
|
||||
q := url.Values{}
|
||||
q.Set("mode", mode)
|
||||
q.Set("_txlock", "immediate")
|
||||
q.Add(
|
||||
"_pragma",
|
||||
fmt.Sprintf(
|
||||
"busy_timeout(%d)",
|
||||
SQLiteBusyTimeout.Milliseconds(),
|
||||
),
|
||||
)
|
||||
q.Add("_pragma", "journal_mode(WAL)")
|
||||
|
||||
return "file:" + path + "?" + q.Encode()
|
||||
}
|
||||
|
||||
// OpenSQLite opens the SQLite file at path with the service's
|
||||
// durability settings and pool bounds applied. mode is the SQLite URI
|
||||
// open mode ("rwc" or "rw").
|
||||
//
|
||||
// The file and its WAL sidecars are settled at SQLiteFilePerm before
|
||||
// the driver sees the path; see reserveSQLiteFile.
|
||||
//
|
||||
// The handle is returned rather than a *gorm.DB because the callers
|
||||
// wrap it in gorm themselves with their own logger.
|
||||
func OpenSQLite(path, mode string) (*sql.DB, error) {
|
||||
err := reserveSQLiteFile(path, mode == SQLiteModeCreate)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
|
||||
sqlDB, err := sql.Open("sqlite", SQLiteDSN(path, mode))
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf(
|
||||
"opening sqlite database %s: %w", path, err,
|
||||
)
|
||||
}
|
||||
|
||||
sqlDB.SetMaxOpenConns(sqliteMaxOpenConns)
|
||||
sqlDB.SetMaxIdleConns(sqliteMaxIdleConns)
|
||||
sqlDB.SetConnMaxLifetime(sqliteConnMaxLifetime)
|
||||
sqlDB.SetConnMaxIdleTime(sqliteConnMaxIdleTime)
|
||||
|
||||
return sqlDB, nil
|
||||
}
|
||||
@@ -0,0 +1,178 @@
|
||||
package database_test
|
||||
|
||||
import (
|
||||
"context"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
"github.com/google/uuid"
|
||||
"github.com/stretchr/testify/assert"
|
||||
"github.com/stretchr/testify/require"
|
||||
"gorm.io/gorm"
|
||||
"sneak.berlin/go/webhooker/internal/database"
|
||||
)
|
||||
|
||||
// livePragma reads a pragma off a live handle. Reading the DSN back
|
||||
// would prove only that the string was built; these tests assert that
|
||||
// SQLite actually applied it.
|
||||
func livePragma(t *testing.T, db *gorm.DB, name string) string {
|
||||
t.Helper()
|
||||
|
||||
var v string
|
||||
|
||||
row := db.Raw("pragma " + name).Row()
|
||||
require.NoError(t, row.Scan(&v))
|
||||
|
||||
return v
|
||||
}
|
||||
|
||||
func TestSQLiteDSNCarriesTheDurabilitySettings(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
dsn := database.SQLiteDSN(
|
||||
"/var/lib/webhooker/webhooker.db",
|
||||
database.SQLiteModeCreate,
|
||||
)
|
||||
|
||||
assert.Contains(t, dsn, "journal_mode%28WAL%29")
|
||||
assert.Contains(t, dsn, "busy_timeout%2810000%29")
|
||||
assert.Contains(t, dsn, "_txlock=immediate")
|
||||
assert.Contains(t, dsn, "mode=rwc")
|
||||
|
||||
// busy_timeout must come first. The driver runs these in order on
|
||||
// every new connection, and PRAGMA journal_mode takes a lock — a
|
||||
// connection opened while the database is busy would fail on that
|
||||
// pragma, with no busy handler yet installed to wait it out.
|
||||
assert.Less(
|
||||
t,
|
||||
strings.Index(dsn, "busy_timeout"),
|
||||
strings.Index(dsn, "journal_mode"),
|
||||
"busy_timeout must be applied before journal_mode",
|
||||
)
|
||||
|
||||
// cache=shared turns an in-process conflict into SQLITE_LOCKED,
|
||||
// which the busy handler does not retry. It must never come back.
|
||||
// See https://git.eeqj.de/sneak/webhooker/issues/256.
|
||||
assert.NotContains(t, strings.ToLower(dsn), "cache=shared")
|
||||
}
|
||||
|
||||
// TestPerWebhookDBAppliesPragmasOnALiveHandle is the check the issue
|
||||
// asks for by name: the settings are confirmed by querying the running
|
||||
// database, not by inspecting the connection string.
|
||||
func TestPerWebhookDBAppliesPragmasOnALiveHandle(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
mgr, lc := setupTestWebhookDBManager(t)
|
||||
ctx := context.Background()
|
||||
require.NoError(t, lc.Start(ctx))
|
||||
|
||||
defer func() { require.NoError(t, lc.Stop(ctx)) }()
|
||||
|
||||
webhookID := uuid.New().String()
|
||||
|
||||
db, err := mgr.GetDB(webhookID)
|
||||
require.NoError(t, err)
|
||||
|
||||
assert.Equal(
|
||||
t, "wal",
|
||||
strings.ToLower(livePragma(t, db, "journal_mode")),
|
||||
)
|
||||
assert.Equal(
|
||||
t, "10000", livePragma(t, db, "busy_timeout"),
|
||||
)
|
||||
}
|
||||
|
||||
func TestMainDBAppliesPragmasOnALiveHandle(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
ctx := context.Background()
|
||||
dir := t.TempDir()
|
||||
|
||||
sqlDB, err := database.OpenSQLite(
|
||||
filepath.Join(dir, database.MainDBFileName),
|
||||
database.SQLiteModeCreate,
|
||||
)
|
||||
require.NoError(t, err)
|
||||
|
||||
defer func() { require.NoError(t, sqlDB.Close()) }()
|
||||
|
||||
var journal string
|
||||
|
||||
require.NoError(t, sqlDB.
|
||||
QueryRowContext(ctx, "pragma journal_mode").
|
||||
Scan(&journal))
|
||||
assert.Equal(t, "wal", strings.ToLower(journal))
|
||||
|
||||
var busy string
|
||||
|
||||
require.NoError(t, sqlDB.
|
||||
QueryRowContext(ctx, "pragma busy_timeout").
|
||||
Scan(&busy))
|
||||
assert.Equal(t, "10000", busy)
|
||||
}
|
||||
|
||||
// TestConcurrentReaderDoesNotBlockWrites is the unit-scale form of the
|
||||
// reproduction in
|
||||
// https://git.eeqj.de/sneak/webhooker/issues/256: an operator's
|
||||
// long-held read of their own data used to make every concurrent write
|
||||
// fail. Under WAL the reader takes a snapshot and the writes proceed.
|
||||
func TestConcurrentReaderDoesNotBlockWrites(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
mgr, lc := setupTestWebhookDBManager(t)
|
||||
ctx := context.Background()
|
||||
require.NoError(t, lc.Start(ctx))
|
||||
|
||||
defer func() { require.NoError(t, lc.Stop(ctx)) }()
|
||||
|
||||
webhookID := uuid.New().String()
|
||||
|
||||
db, err := mgr.GetDB(webhookID)
|
||||
require.NoError(t, err)
|
||||
|
||||
// A second handle on the same file, holding a read transaction
|
||||
// open across every write below — what `sqlite3 <db> .dump` is.
|
||||
readerSQL, err := database.OpenSQLite(
|
||||
mgr.DBPath(webhookID), database.SQLiteModeExisting,
|
||||
)
|
||||
require.NoError(t, err)
|
||||
|
||||
defer func() { require.NoError(t, readerSQL.Close()) }()
|
||||
|
||||
readerConn, err := readerSQL.Conn(ctx)
|
||||
require.NoError(t, err)
|
||||
|
||||
defer func() { require.NoError(t, readerConn.Close()) }()
|
||||
|
||||
_, err = readerConn.ExecContext(ctx, "begin deferred")
|
||||
require.NoError(t, err)
|
||||
|
||||
_, err = readerConn.ExecContext(
|
||||
ctx, "select count(*) from events",
|
||||
)
|
||||
require.NoError(t, err)
|
||||
|
||||
for range 25 {
|
||||
err = db.Transaction(func(tx *gorm.DB) error {
|
||||
return tx.Create(&database.Event{
|
||||
WebhookID: webhookID,
|
||||
EntrypointID: uuid.New().String(),
|
||||
Method: "POST",
|
||||
Body: "{}",
|
||||
}).Error
|
||||
})
|
||||
require.NoError(t, err)
|
||||
}
|
||||
|
||||
_, err = readerConn.ExecContext(ctx, "commit")
|
||||
require.NoError(t, err)
|
||||
|
||||
var count int64
|
||||
|
||||
require.NoError(
|
||||
t,
|
||||
db.Model(&database.Event{}).Count(&count).Error,
|
||||
)
|
||||
assert.Equal(t, int64(25), count)
|
||||
}
|
||||
@@ -2,7 +2,6 @@ package database
|
||||
|
||||
import (
|
||||
"context"
|
||||
"database/sql"
|
||||
"errors"
|
||||
"fmt"
|
||||
"log/slog"
|
||||
@@ -14,6 +13,7 @@ import (
|
||||
"gorm.io/driver/sqlite"
|
||||
"gorm.io/gorm"
|
||||
"sneak.berlin/go/webhooker/internal/config"
|
||||
"sneak.berlin/go/webhooker/internal/datadir"
|
||||
"sneak.berlin/go/webhooker/internal/gormlog"
|
||||
"sneak.berlin/go/webhooker/internal/logger"
|
||||
)
|
||||
@@ -41,6 +41,11 @@ type WebhookDBManager struct {
|
||||
dataDir string
|
||||
dbs sync.Map // map[webhookID]*gorm.DB
|
||||
log *slog.Logger
|
||||
|
||||
// mu is held while a database is opened, deleted, or closed, so
|
||||
// each file has at most one open handle. Reading an already cached
|
||||
// handle does not take it.
|
||||
mu sync.Mutex
|
||||
}
|
||||
|
||||
// NewWebhookDBManager creates a new WebhookDBManager and
|
||||
@@ -54,8 +59,9 @@ func NewWebhookDBManager(
|
||||
log: params.Logger.Get(),
|
||||
}
|
||||
|
||||
// Create data directory if it doesn't exist
|
||||
err := os.MkdirAll(m.dataDir, dataDirPerm)
|
||||
// Create data directory if it doesn't exist. datadir.DirPerm is the
|
||||
// single source of the directory mode; either package may run first.
|
||||
err := os.MkdirAll(m.dataDir, datadir.DirPerm)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf(
|
||||
"creating data directory %s: %w",
|
||||
@@ -85,43 +91,39 @@ func (m *WebhookDBManager) GetDB(
|
||||
) (*gorm.DB, error) {
|
||||
// Fast path: already open
|
||||
if val, ok := m.dbs.Load(webhookID); ok {
|
||||
cachedDB, castOK := val.(*gorm.DB)
|
||||
if !castOK {
|
||||
return nil, fmt.Errorf(
|
||||
"%w for webhook %s",
|
||||
errInvalidCachedDBType,
|
||||
webhookID,
|
||||
)
|
||||
}
|
||||
return asGormDB(val, webhookID)
|
||||
}
|
||||
|
||||
return cachedDB, nil
|
||||
// Slow path: open the database under the lock, looking in the
|
||||
// cache again first. A caller that raced another one here then
|
||||
// waits for its handle instead of opening a second one.
|
||||
m.mu.Lock()
|
||||
defer m.mu.Unlock()
|
||||
|
||||
if val, ok := m.dbs.Load(webhookID); ok {
|
||||
return asGormDB(val, webhookID)
|
||||
}
|
||||
|
||||
// Slow path: open/create the database
|
||||
db, err := m.openDB(webhookID)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
|
||||
// Store it; if another goroutine beat us, close ours
|
||||
actual, loaded := m.dbs.LoadOrStore(webhookID, db)
|
||||
if loaded {
|
||||
// Another goroutine created it first; close our duplicate
|
||||
sqlDB, closeErr := db.DB()
|
||||
if closeErr == nil {
|
||||
_ = sqlDB.Close()
|
||||
}
|
||||
m.dbs.Store(webhookID, db)
|
||||
|
||||
existingDB, castOK := actual.(*gorm.DB)
|
||||
if !castOK {
|
||||
return nil, fmt.Errorf(
|
||||
"%w for webhook %s",
|
||||
errInvalidCachedDBType,
|
||||
webhookID,
|
||||
)
|
||||
}
|
||||
return db, nil
|
||||
}
|
||||
|
||||
return existingDB, nil
|
||||
// asGormDB returns a value read from the cache as the database
|
||||
// handle it is.
|
||||
func asGormDB(val any, webhookID string) (*gorm.DB, error) {
|
||||
db, ok := val.(*gorm.DB)
|
||||
if !ok {
|
||||
return nil, fmt.Errorf(
|
||||
"%w for webhook %s",
|
||||
errInvalidCachedDBType,
|
||||
webhookID,
|
||||
)
|
||||
}
|
||||
|
||||
return db, nil
|
||||
@@ -152,6 +154,11 @@ func (m *WebhookDBManager) DBExists(
|
||||
func (m *WebhookDBManager) DeleteDB(
|
||||
webhookID string,
|
||||
) error {
|
||||
// Held until the files are gone, so GetDB cannot open the file
|
||||
// again between the close and the removal.
|
||||
m.mu.Lock()
|
||||
defer m.mu.Unlock()
|
||||
|
||||
// Close and remove from cache
|
||||
if val, ok := m.dbs.LoadAndDelete(webhookID); ok {
|
||||
if gormDB, castOK := val.(*gorm.DB); castOK {
|
||||
@@ -185,6 +192,11 @@ func (m *WebhookDBManager) DeleteDB(
|
||||
// CloseAll closes all open per-webhook database connections.
|
||||
// Called during application shutdown.
|
||||
func (m *WebhookDBManager) CloseAll() error {
|
||||
// An open already under way finishes and is cached first, so it
|
||||
// is closed here rather than cached after this loop has passed.
|
||||
m.mu.Lock()
|
||||
defer m.mu.Unlock()
|
||||
|
||||
var lastErr error
|
||||
|
||||
m.dbs.Range(func(key, value any) bool {
|
||||
@@ -234,12 +246,11 @@ func (m *WebhookDBManager) openDB(
|
||||
webhookID string,
|
||||
) (*gorm.DB, error) {
|
||||
path := m.dbPath(webhookID)
|
||||
dbURL := fmt.Sprintf(
|
||||
"file:%s?cache=shared&mode=rwc",
|
||||
path,
|
||||
)
|
||||
|
||||
sqlDB, err := sql.Open("sqlite", dbURL)
|
||||
// See sqlite_open.go: WAL, a busy timeout, immediate-transaction
|
||||
// locking, and a bounded pool, all of which this file needs most —
|
||||
// it is the one every delivery worker writes to concurrently.
|
||||
sqlDB, err := OpenSQLite(path, SQLiteModeCreate)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf(
|
||||
"opening webhook database %s: %w",
|
||||
|
||||
@@ -1,10 +1,14 @@
|
||||
package database_test
|
||||
|
||||
import (
|
||||
"bytes"
|
||||
"context"
|
||||
"log/slog"
|
||||
"net/http"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"sync"
|
||||
"testing"
|
||||
|
||||
"github.com/google/uuid"
|
||||
@@ -104,6 +108,54 @@ func TestWebhookDBManager_CreateAndGetDB(t *testing.T) {
|
||||
assert.Equal(t, `{"test": true}`, readEvent.Body)
|
||||
}
|
||||
|
||||
// Many callers ask for one webhook's database at the same moment,
|
||||
// before it is cached. Only one of them may open the file; the others
|
||||
// must wait for its handle. openDB logs one "opened per-webhook
|
||||
// database" line per open, and those lines are what is counted.
|
||||
func TestWebhookDBManager_ConcurrentFirstTouchOpensOnce(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
var logs bytes.Buffer
|
||||
|
||||
mgr := database.NewTestWebhookDBManagerWithLogger(
|
||||
t.TempDir(),
|
||||
slog.New(slog.NewTextHandler(&logs, nil)),
|
||||
)
|
||||
|
||||
t.Cleanup(func() { assert.NoError(t, mgr.CloseAll()) })
|
||||
|
||||
webhookID := uuid.New().String()
|
||||
|
||||
const callers = 16
|
||||
|
||||
start := make(chan struct{})
|
||||
handles := make([]*gorm.DB, callers)
|
||||
errs := make([]error, callers)
|
||||
|
||||
var wg sync.WaitGroup
|
||||
|
||||
for i := range callers {
|
||||
wg.Go(func() {
|
||||
<-start
|
||||
|
||||
handles[i], errs[i] = mgr.GetDB(webhookID)
|
||||
})
|
||||
}
|
||||
|
||||
close(start)
|
||||
wg.Wait()
|
||||
|
||||
for i := range callers {
|
||||
require.NoError(t, errs[i])
|
||||
assert.Same(t, handles[0], handles[i])
|
||||
}
|
||||
|
||||
assert.Equal(
|
||||
t, 1,
|
||||
strings.Count(logs.String(), "opened per-webhook database"),
|
||||
)
|
||||
}
|
||||
|
||||
func TestWebhookDBManager_DeleteDB(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
|
||||
@@ -29,9 +29,11 @@ import (
|
||||
// process that was killed with SIGKILL blocks nothing.
|
||||
const LockFileName = "webhooker.lock"
|
||||
|
||||
// dirPerm is the mode Acquire creates DATA_DIR with. It matches what
|
||||
// internal/database uses, since whichever runs first creates it.
|
||||
const dirPerm = 0o750
|
||||
// DirPerm is the mode DATA_DIR is created with. It is the single
|
||||
// source of that mode: internal/database consumes it rather than
|
||||
// keeping its own copy, so the two packages that both create the
|
||||
// directory cannot drift into disagreeing about its permissions.
|
||||
const DirPerm = 0o750
|
||||
|
||||
// ErrLocked reports that another live process holds the data
|
||||
// directory. Callers that need to know whether a deployment is running
|
||||
@@ -64,7 +66,7 @@ func Acquire(dir string) (*Lock, error) {
|
||||
return nil, ErrNoDir
|
||||
}
|
||||
|
||||
err := os.MkdirAll(dir, dirPerm)
|
||||
err := os.MkdirAll(dir, DirPerm)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf(
|
||||
"creating data directory %s: %w", dir, err,
|
||||
|
||||
@@ -76,12 +76,20 @@ func (cb *CircuitBreaker) Allow() bool {
|
||||
}
|
||||
}
|
||||
|
||||
// CooldownRemaining returns how much time is left before
|
||||
// an open circuit transitions to half-open.
|
||||
// CooldownRemaining returns how long a delivery that Allow refused
|
||||
// should wait before it is tried again. Closed, it returns zero.
|
||||
// Open, it returns what is left of the cooldown, or zero once that
|
||||
// has passed. Half-open, it returns the whole cooldown: the one
|
||||
// probe delivery is still in flight, and if it fails the circuit
|
||||
// reopens for that long.
|
||||
func (cb *CircuitBreaker) CooldownRemaining() time.Duration {
|
||||
cb.mu.Lock()
|
||||
defer cb.mu.Unlock()
|
||||
|
||||
if cb.state == CircuitHalfOpen {
|
||||
return cb.cooldown
|
||||
}
|
||||
|
||||
if cb.state != CircuitOpen {
|
||||
return 0
|
||||
}
|
||||
|
||||
@@ -267,7 +267,7 @@ func TestCircuitBreaker_CooldownRemaining_ClosedReturnsZero(
|
||||
)
|
||||
}
|
||||
|
||||
func TestCircuitBreaker_CooldownRemaining_HalfOpenReturnsZero(
|
||||
func TestCircuitBreaker_CooldownRemaining_HalfOpenReturnsCooldown(
|
||||
t *testing.T,
|
||||
) {
|
||||
t.Parallel()
|
||||
@@ -282,9 +282,11 @@ func TestCircuitBreaker_CooldownRemaining_HalfOpenReturnsZero(
|
||||
|
||||
require.True(t, cb.Allow())
|
||||
|
||||
assert.Equal(t, time.Duration(0),
|
||||
// The cooldown newShortCooldownCB gives the breaker.
|
||||
assert.Equal(t, 50*time.Millisecond,
|
||||
cb.CooldownRemaining(),
|
||||
"half-open circuit should have zero cooldown remaining",
|
||||
"a delivery refused while half-open should wait "+
|
||||
"a whole cooldown",
|
||||
)
|
||||
}
|
||||
|
||||
|
||||
+977
-74
File diff suppressed because it is too large
Load Diff
@@ -2,7 +2,6 @@ package delivery_test
|
||||
|
||||
import (
|
||||
"context"
|
||||
"database/sql"
|
||||
"encoding/json"
|
||||
"fmt"
|
||||
"io"
|
||||
@@ -70,11 +69,12 @@ func iMainDB(t *testing.T) *gorm.DB {
|
||||
t.TempDir(), "main-test.db",
|
||||
)
|
||||
|
||||
dsn := fmt.Sprintf(
|
||||
"file:%s?cache=shared&mode=rwc", dbPath,
|
||||
// Opened the way the service opens the main database, so these
|
||||
// tests cannot pass against journal and locking settings
|
||||
// production does not use.
|
||||
sqlDB, err := database.OpenSQLite(
|
||||
dbPath, database.SQLiteModeCreate,
|
||||
)
|
||||
|
||||
sqlDB, err := sql.Open("sqlite", dsn)
|
||||
require.NoError(t, err)
|
||||
|
||||
t.Cleanup(func() { _ = sqlDB.Close() })
|
||||
@@ -377,6 +377,17 @@ func TestProcessRetryTask_SuccessfulRetry(t *testing.T) {
|
||||
|
||||
bodyStr := event.Body
|
||||
cfg := iHTTPConfig(ts.URL)
|
||||
|
||||
// The target row exists because the engine confirms a scheduled
|
||||
// retry's target has not been deleted before it runs it. A retry
|
||||
// task whose target id names no row at all is a state the service
|
||||
// does not produce: the handler read that target to build the
|
||||
// task. See https://git.eeqj.de/sneak/webhooker/issues/107.
|
||||
iCreateTarget(
|
||||
t, s.MainDB, targetID, s.WebhookID, "retry-target",
|
||||
database.TargetTypeHTTP, cfg, 5,
|
||||
)
|
||||
|
||||
task := iTask(
|
||||
d, event, s.WebhookID, targetID,
|
||||
"retry-target", cfg, 5, 2, &bodyStr,
|
||||
@@ -456,6 +467,12 @@ func TestProcessRetryTask_LargeBody_FetchFromDB(
|
||||
)
|
||||
|
||||
cfg := iHTTPConfig(ts.URL)
|
||||
|
||||
iCreateTarget(
|
||||
t, s.MainDB, targetID, s.WebhookID, "retry-large",
|
||||
database.TargetTypeHTTP, cfg, 5,
|
||||
)
|
||||
|
||||
task := iTask(
|
||||
d, event, s.WebhookID, targetID,
|
||||
"retry-large", cfg, 5, 2, nil,
|
||||
@@ -558,6 +575,12 @@ func TestWorkerLifecycle_ProcessesRetryChannel(
|
||||
|
||||
bodyStr := event.Body
|
||||
cfg := iHTTPConfig(ts.URL)
|
||||
|
||||
iCreateTarget(
|
||||
t, s.MainDB, targetID, s.WebhookID, "retry-chan-test",
|
||||
database.TargetTypeHTTP, cfg, 5,
|
||||
)
|
||||
|
||||
task := iTask(
|
||||
d, event, s.WebhookID, targetID,
|
||||
"retry-chan-test", cfg, 5, 2, &bodyStr,
|
||||
|
||||
@@ -2,6 +2,8 @@ package delivery_test
|
||||
|
||||
import (
|
||||
"context"
|
||||
"fmt"
|
||||
"path/filepath"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
@@ -269,3 +271,88 @@ func TestEngine_StopHookHonoursStopTimeout(t *testing.T) {
|
||||
|
||||
requireStopHookExpires(t, lc.hooks[0], "delivery engine")
|
||||
}
|
||||
|
||||
// deliverToArchive runs one delivery to a database target through
|
||||
// the running engine and returns the webhook's archive file path.
|
||||
// The archive writer holds the file open afterwards.
|
||||
func deliverToArchive(t *testing.T, s iSetup) string {
|
||||
t.Helper()
|
||||
|
||||
deliveryID, task := seedLogTask(t, s)
|
||||
task.TargetType = database.TargetTypeDatabase
|
||||
|
||||
s.Engine.Notify([]delivery.Task{task})
|
||||
|
||||
iWaitForDelivered(t, s.WebhookDB, deliveryID)
|
||||
|
||||
return filepath.Join(
|
||||
filepath.Dir(s.DBMgr.DBPath(s.WebhookID)),
|
||||
fmt.Sprintf("archive-%s.db", s.WebhookID),
|
||||
)
|
||||
}
|
||||
|
||||
// TestEngine_StopHookClosesArchives is the regression test for an
|
||||
// archive split across two files by a clean stop. The engine never
|
||||
// closed its archive writers, so after a stop the archived rows
|
||||
// could sit in archive-{id}.db-wal while archive-{id}.db held no
|
||||
// table at all, and copying the .db on its own gave an empty
|
||||
// database.
|
||||
func TestEngine_StopHookClosesArchives(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
s := newISetup(t)
|
||||
|
||||
lc := startEngineViaHook(t, s.Engine)
|
||||
|
||||
path := deliverToArchive(t, s)
|
||||
require.FileExists(
|
||||
t, path+"-wal",
|
||||
"an open archive should have a -wal for the stop to remove",
|
||||
)
|
||||
|
||||
require.NoError(t, lc.hooks[0].OnStop(context.Background()))
|
||||
|
||||
wals, err := filepath.Glob(
|
||||
filepath.Join(filepath.Dir(path), "archive-*.db-wal"),
|
||||
)
|
||||
require.NoError(t, err)
|
||||
require.Empty(
|
||||
t, wals, "a clean stop must leave no archive -wal behind",
|
||||
)
|
||||
|
||||
// With no -wal beside it, the row can only be in the .db.
|
||||
count, err := countArchivedRows(path)
|
||||
require.NoError(t, err)
|
||||
require.Equal(t, int64(1), count)
|
||||
}
|
||||
|
||||
// TestEngine_StopHookTimeoutLeavesArchivesOpen covers a stop whose
|
||||
// budget runs out while a worker is still running. The archive
|
||||
// writers are left open, as a kill would leave them: closing them
|
||||
// would wait for any write in progress, and that worker would then
|
||||
// open new writers that nothing closes.
|
||||
func TestEngine_StopHookTimeoutLeavesArchivesOpen(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
s := newISetup(t)
|
||||
|
||||
lc := startEngineViaHook(t, s.Engine)
|
||||
|
||||
deliverToArchive(t, s)
|
||||
|
||||
release := make(chan struct{})
|
||||
|
||||
t.Cleanup(func() {
|
||||
close(release)
|
||||
s.Engine.EvictWebhook(s.WebhookID)
|
||||
})
|
||||
|
||||
s.Engine.ExportWedgeWorker(release)
|
||||
|
||||
requireStopHookExpires(t, lc.hooks[0], "delivery engine")
|
||||
|
||||
require.True(
|
||||
t, s.Engine.ExportArchiveHandleOpen(s.WebhookID),
|
||||
"a stop that timed out must not close archive writers",
|
||||
)
|
||||
}
|
||||
|
||||
@@ -3,7 +3,6 @@ package delivery_test
|
||||
import (
|
||||
"bytes"
|
||||
"context"
|
||||
"database/sql"
|
||||
"encoding/json"
|
||||
"fmt"
|
||||
"log/slog"
|
||||
@@ -18,6 +17,7 @@ import (
|
||||
"time"
|
||||
|
||||
"github.com/google/uuid"
|
||||
"github.com/prometheus/client_golang/prometheus"
|
||||
"github.com/stretchr/testify/assert"
|
||||
"github.com/stretchr/testify/require"
|
||||
"gorm.io/driver/sqlite"
|
||||
@@ -25,6 +25,7 @@ import (
|
||||
_ "modernc.org/sqlite"
|
||||
"sneak.berlin/go/webhooker/internal/database"
|
||||
"sneak.berlin/go/webhooker/internal/delivery"
|
||||
"sneak.berlin/go/webhooker/internal/metrics"
|
||||
)
|
||||
|
||||
// testContentType is the event content type used in tests.
|
||||
@@ -37,11 +38,12 @@ func testWebhookDB(t *testing.T) *gorm.DB {
|
||||
t.TempDir(), "events-test.db",
|
||||
)
|
||||
|
||||
dsn := fmt.Sprintf(
|
||||
"file:%s?cache=shared&mode=rwc", dbPath,
|
||||
// Opened the way the service opens a per-webhook database, so
|
||||
// these tests cannot pass against journal and locking settings
|
||||
// production does not use.
|
||||
sqlDB, err := database.OpenSQLite(
|
||||
dbPath, database.SQLiteModeCreate,
|
||||
)
|
||||
|
||||
sqlDB, err := sql.Open("sqlite", dsn)
|
||||
require.NoError(t, err)
|
||||
|
||||
t.Cleanup(func() { _ = sqlDB.Close() })
|
||||
@@ -894,6 +896,100 @@ func TestDeliverHTTP_CircuitBreakerBlocks(t *testing.T) {
|
||||
)
|
||||
}
|
||||
|
||||
// recordingScheduler keeps the delay of every retry it is asked to
|
||||
// schedule, and schedules nothing.
|
||||
type recordingScheduler struct {
|
||||
delays []time.Duration
|
||||
}
|
||||
|
||||
func (s *recordingScheduler) ScheduleRetry(
|
||||
_ delivery.Task, delay time.Duration,
|
||||
) {
|
||||
s.delays = append(s.delays, delay)
|
||||
}
|
||||
|
||||
// TestDeliverHTTP_HalfOpenBreakerDelaysQueuedTasks proves that while a
|
||||
// half-open breaker's one probe delivery is in flight, every other task
|
||||
// for the target is put back with a whole cooldown as its delay rather
|
||||
// than none, and that its status is written the first time the breaker
|
||||
// turns it away and not on each pass after that.
|
||||
func TestDeliverHTTP_HalfOpenBreakerDelaysQueuedTasks(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
db := testWebhookDB(t)
|
||||
e := testEngine(t, 1)
|
||||
|
||||
// Every write of retrying moves the retry counter, so on a registry
|
||||
// this test owns the counter is the number of those writes.
|
||||
reg := prometheus.NewRegistry()
|
||||
e.ExportSetMetrics(metrics.New(reg))
|
||||
|
||||
targetID := uuid.New().String()
|
||||
cb := newShortCooldownCB(t)
|
||||
e.ExportSetCircuitBreaker(targetID, cb)
|
||||
|
||||
for range delivery.ExportDefaultFailureThreshold {
|
||||
cb.RecordFailure()
|
||||
}
|
||||
|
||||
time.Sleep(60 * time.Millisecond)
|
||||
|
||||
require.True(t, cb.Allow(), "the probe delivery should go through")
|
||||
require.Equal(t, delivery.CircuitHalfOpen, cb.State())
|
||||
|
||||
cfg := newHTTPTargetConfig(
|
||||
"http://will-not-be-called.invalid",
|
||||
)
|
||||
sched := &recordingScheduler{}
|
||||
|
||||
const queued, passes = 3, 4
|
||||
|
||||
for range queued {
|
||||
event := seedEvent(t, db, `{"cb":"half-open"}`)
|
||||
dlv := seedDelivery(
|
||||
t, db, event.ID, targetID,
|
||||
database.DeliveryStatusPending,
|
||||
)
|
||||
|
||||
for range passes {
|
||||
// Each pass starts from the stored row, as a retry does.
|
||||
var row database.Delivery
|
||||
|
||||
require.NoError(t, db.First(
|
||||
&row, "id = ?", dlv.ID,
|
||||
).Error)
|
||||
|
||||
fix := buildHTTPFixture(
|
||||
row, event, targetID,
|
||||
"test-cb-half-open", cfg, 5, 1,
|
||||
)
|
||||
|
||||
e.ExportDeliverHTTPWithScheduler(
|
||||
context.TODO(), db, fix.Delivery, fix.Task, sched,
|
||||
)
|
||||
}
|
||||
|
||||
assertDeliveryStatus(t, db, dlv.ID,
|
||||
database.DeliveryStatusRetrying,
|
||||
)
|
||||
}
|
||||
|
||||
require.Len(t, sched.delays, queued*passes)
|
||||
|
||||
for _, delay := range sched.delays {
|
||||
// The cooldown newShortCooldownCB gives the breaker.
|
||||
assert.Equal(t, 50*time.Millisecond, delay,
|
||||
"a task turned away while half-open should wait "+
|
||||
"a whole cooldown",
|
||||
)
|
||||
}
|
||||
|
||||
assert.InDelta(t, float64(queued),
|
||||
mCounter(t, reg, mRetries, mTypeHTTP), 0,
|
||||
"status should be written once per task, not once per pass",
|
||||
)
|
||||
}
|
||||
|
||||
func TestGetCircuitBreaker_CreatesOnDemand(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
@@ -1070,6 +1166,10 @@ func TestIsForwardableHeader(t *testing.T) {
|
||||
assert.False(t,
|
||||
delivery.ExportIsForwardableHeader("Content-Length"),
|
||||
)
|
||||
|
||||
assert.False(t,
|
||||
delivery.ExportIsForwardableHeader("Content-Type"),
|
||||
)
|
||||
}
|
||||
|
||||
func TestTruncate(t *testing.T) {
|
||||
@@ -1151,6 +1251,81 @@ func TestDoHTTPRequest_ForwardsHeaders(t *testing.T) {
|
||||
)
|
||||
}
|
||||
|
||||
// The event's stored inbound headers carry the same Content-Type the
|
||||
// receiver saved as the event's ContentType, so a delivery could send
|
||||
// it twice. It must go out exactly once, with a Content-Type configured
|
||||
// on the target winning, then the event's ContentType.
|
||||
func TestApplyRequestHeaders_SendsOneContentType(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
cases := map[string]struct {
|
||||
inbound string
|
||||
event string
|
||||
configured string
|
||||
want []string
|
||||
}{
|
||||
"inbound and event agree": {
|
||||
inbound: testContentType,
|
||||
event: testContentType,
|
||||
want: []string{testContentType},
|
||||
},
|
||||
"inbound and event disagree": {
|
||||
inbound: "text/plain",
|
||||
event: testContentType,
|
||||
want: []string{testContentType},
|
||||
},
|
||||
"event has none": {
|
||||
inbound: testContentType,
|
||||
want: nil,
|
||||
},
|
||||
"target configures its own": {
|
||||
inbound: testContentType,
|
||||
event: testContentType,
|
||||
configured: "application/xml",
|
||||
want: []string{"application/xml"},
|
||||
},
|
||||
}
|
||||
|
||||
for name, tc := range cases {
|
||||
t.Run(name, func(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
inbound, err := json.Marshal(map[string][]string{
|
||||
headerContentType: {tc.inbound},
|
||||
})
|
||||
require.NoError(t, err)
|
||||
|
||||
cfg := &delivery.HTTPTargetConfig{}
|
||||
if tc.configured != "" {
|
||||
cfg.Headers = map[string]string{
|
||||
headerContentType: tc.configured,
|
||||
}
|
||||
}
|
||||
|
||||
req, err := http.NewRequestWithContext(
|
||||
context.Background(),
|
||||
http.MethodPost,
|
||||
"https://target.example.com/hook",
|
||||
http.NoBody,
|
||||
)
|
||||
require.NoError(t, err)
|
||||
|
||||
delivery.ExportApplyRequestHeaders(
|
||||
req,
|
||||
&database.Event{
|
||||
Headers: string(inbound),
|
||||
ContentType: tc.event,
|
||||
},
|
||||
cfg,
|
||||
)
|
||||
|
||||
assert.Equal(t,
|
||||
tc.want, req.Header.Values(headerContentType),
|
||||
)
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
func TestProcessDelivery_RoutesToCorrectHandler(
|
||||
t *testing.T,
|
||||
) {
|
||||
|
||||
@@ -96,7 +96,15 @@ func TestEventDBHoldsNoTargetRows(t *testing.T) {
|
||||
)
|
||||
assertNoTargetRows(t, dbPath)
|
||||
|
||||
// A retry.
|
||||
// A retry. Its target exists in the main database, because the
|
||||
// engine confirms a scheduled retry's target has not been
|
||||
// deleted before running it; see
|
||||
// https://git.eeqj.de/sneak/webhooker/issues/107.
|
||||
iCreateTarget(
|
||||
t, s.MainDB, targetID, s.WebhookID, "leaky-target",
|
||||
database.TargetTypeHTTP, cfg, 5,
|
||||
)
|
||||
|
||||
rd := iSeedDelivery(
|
||||
t, s.WebhookDB, event.ID, targetID,
|
||||
database.DeliveryStatusRetrying,
|
||||
|
||||
@@ -0,0 +1,442 @@
|
||||
package delivery_test
|
||||
|
||||
import (
|
||||
"context"
|
||||
"encoding/json"
|
||||
"io"
|
||||
"net/http"
|
||||
"net/http/httptest"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"github.com/google/uuid"
|
||||
"github.com/stretchr/testify/assert"
|
||||
"github.com/stretchr/testify/require"
|
||||
"gorm.io/gorm"
|
||||
"sneak.berlin/go/webhooker/internal/database"
|
||||
"sneak.berlin/go/webhooker/internal/delivery"
|
||||
)
|
||||
|
||||
// tsEventCreatedAt is the receipt time seeded on the events these
|
||||
// tests deliver. It is far enough from both the zero time and from
|
||||
// now that neither can be mistaken for it.
|
||||
func tsEventCreatedAt() time.Time {
|
||||
return time.Date(
|
||||
2026, time.March, 4, 5, 6, 7, 0, time.UTC,
|
||||
)
|
||||
}
|
||||
|
||||
// tsZeroStamp is what a Slack message renders when the event handed
|
||||
// to FormatSlackMessage carries no CreatedAt.
|
||||
const tsZeroStamp = "*Timestamp:* `0001-01-01T00:00:00Z`"
|
||||
|
||||
// tsEventBody is the body seeded on every event in this file. It is
|
||||
// small enough that a Task can inline it.
|
||||
const tsEventBody = `{"hello":"world"}`
|
||||
|
||||
// tsUndeliverableHook stands in for a Slack incoming webhook on the
|
||||
// tests that never send: the config parser requires a URL, but no
|
||||
// request is made.
|
||||
const tsUndeliverableHook = "https://hooks.slack.com/services/T/B/x"
|
||||
|
||||
// tsSink is a stand-in Slack incoming webhook that records the raw
|
||||
// body posted to it.
|
||||
type tsSink struct {
|
||||
*httptest.Server
|
||||
|
||||
bodies chan []byte
|
||||
}
|
||||
|
||||
func newTSSink(t *testing.T) *tsSink {
|
||||
t.Helper()
|
||||
|
||||
s := &tsSink{bodies: make(chan []byte, 8)}
|
||||
|
||||
s.Server = httptest.NewServer(http.HandlerFunc(
|
||||
func(w http.ResponseWriter, r *http.Request) {
|
||||
body, _ := io.ReadAll(r.Body)
|
||||
|
||||
select {
|
||||
case s.bodies <- body:
|
||||
default:
|
||||
}
|
||||
|
||||
w.WriteHeader(http.StatusOK)
|
||||
},
|
||||
))
|
||||
|
||||
t.Cleanup(s.Close)
|
||||
|
||||
return s
|
||||
}
|
||||
|
||||
// text returns the Slack message text from the single payload the
|
||||
// sink received.
|
||||
func (s *tsSink) text(t *testing.T) string {
|
||||
t.Helper()
|
||||
|
||||
select {
|
||||
case raw := <-s.bodies:
|
||||
t.Logf("raw slack payload: %s", raw)
|
||||
|
||||
var payload struct {
|
||||
Text string `json:"text"`
|
||||
}
|
||||
|
||||
require.NoError(t, json.Unmarshal(raw, &payload))
|
||||
|
||||
return payload.Text
|
||||
case <-time.After(5 * time.Second):
|
||||
t.Fatal("slack sink received no payload")
|
||||
|
||||
return ""
|
||||
}
|
||||
}
|
||||
|
||||
func tsSlackConfig(t *testing.T, url string) string {
|
||||
t.Helper()
|
||||
|
||||
data, err := json.Marshal(
|
||||
delivery.SlackTargetConfig{WebhookURL: url},
|
||||
)
|
||||
require.NoError(t, err)
|
||||
|
||||
return string(data)
|
||||
}
|
||||
|
||||
// tsSeedEvent writes an event whose CreatedAt is tsEventCreatedAt
|
||||
// rather than the write time, so an assertion on the rendered
|
||||
// timestamp cannot pass by accident against "roughly now".
|
||||
func tsSeedEvent(
|
||||
t *testing.T, db *gorm.DB, webhookID string,
|
||||
) database.Event {
|
||||
t.Helper()
|
||||
|
||||
event := database.Event{
|
||||
WebhookID: webhookID,
|
||||
EntrypointID: uuid.New().String(),
|
||||
Method: http.MethodPost,
|
||||
Headers: `{}`,
|
||||
Body: tsEventBody,
|
||||
ContentType: "application/json",
|
||||
}
|
||||
event.ID = uuid.New().String()
|
||||
event.CreatedAt = tsEventCreatedAt()
|
||||
event.UpdatedAt = tsEventCreatedAt()
|
||||
|
||||
require.NoError(t, db.Create(&event).Error)
|
||||
|
||||
var stored database.Event
|
||||
|
||||
require.NoError(t,
|
||||
db.First(&stored, "id = ?", event.ID).Error,
|
||||
)
|
||||
require.Equal(t,
|
||||
tsEventCreatedAt().UTC(), stored.CreatedAt.UTC(),
|
||||
"seeded created_at did not round-trip",
|
||||
)
|
||||
|
||||
return event
|
||||
}
|
||||
|
||||
// tsSeedTarget writes the slack target row into the main database.
|
||||
// The retry path confirms the target still exists before sending.
|
||||
func tsSeedTarget(
|
||||
t *testing.T, mainDB *gorm.DB, webhookID, config string,
|
||||
) database.Target {
|
||||
t.Helper()
|
||||
|
||||
target := database.Target{
|
||||
WebhookID: webhookID,
|
||||
Name: "slack-sink",
|
||||
Type: database.TargetTypeSlack,
|
||||
Config: config,
|
||||
Active: true,
|
||||
}
|
||||
|
||||
require.NoError(t, mainDB.Create(&target).Error)
|
||||
|
||||
return target
|
||||
}
|
||||
|
||||
func tsTask(
|
||||
d database.Delivery,
|
||||
event database.Event,
|
||||
webhookID string,
|
||||
target database.Target,
|
||||
attemptNum int,
|
||||
body *string,
|
||||
) delivery.Task {
|
||||
return delivery.Task{
|
||||
DeliveryID: d.ID,
|
||||
EventID: event.ID,
|
||||
WebhookID: webhookID,
|
||||
EntrypointID: event.EntrypointID,
|
||||
TargetID: target.ID,
|
||||
TargetName: target.Name,
|
||||
TargetType: database.TargetTypeSlack,
|
||||
TargetConfig: target.Config,
|
||||
MaxRetries: 0,
|
||||
Method: event.Method,
|
||||
Headers: event.Headers,
|
||||
ContentType: event.ContentType,
|
||||
Body: body,
|
||||
AttemptNum: attemptNum,
|
||||
}
|
||||
}
|
||||
|
||||
func tsAssertRealTimestamp(t *testing.T, text string) {
|
||||
t.Helper()
|
||||
|
||||
assert.NotContains(t, text, tsZeroStamp,
|
||||
"slack message carries the zero timestamp",
|
||||
)
|
||||
assert.Contains(t, text,
|
||||
"*Timestamp:* `"+
|
||||
tsEventCreatedAt().UTC().Format(time.RFC3339)+"`",
|
||||
"slack message does not carry the event's receipt time",
|
||||
)
|
||||
}
|
||||
|
||||
// tsCase is one end-to-end delivery of a seeded event to a slack
|
||||
// sink, over whichever engine path `process` names.
|
||||
type tsCase struct {
|
||||
// status is the delivery row's status before the engine runs.
|
||||
// The retry path refuses a delivery that is not retrying.
|
||||
status database.DeliveryStatus
|
||||
|
||||
// inlineBody mirrors a Task built for a body under
|
||||
// MaxInlineBodySize. When false the engine reads the body back
|
||||
// from the stored row.
|
||||
inlineBody bool
|
||||
|
||||
attemptNum int
|
||||
|
||||
process func(
|
||||
ctx context.Context, e *delivery.Engine, task *delivery.Task,
|
||||
)
|
||||
}
|
||||
|
||||
// run delivers one event through the named path and returns the
|
||||
// Slack message text the sink received.
|
||||
func (c tsCase) run(t *testing.T) (iSetup, database.Delivery, string) {
|
||||
t.Helper()
|
||||
|
||||
s := newISetup(t)
|
||||
sink := newTSSink(t)
|
||||
|
||||
cfg := tsSlackConfig(t, sink.URL)
|
||||
target := tsSeedTarget(t, s.MainDB, s.WebhookID, cfg)
|
||||
event := tsSeedEvent(t, s.WebhookDB, s.WebhookID)
|
||||
|
||||
d := iSeedDelivery(
|
||||
t, s.WebhookDB, event.ID, target.ID, c.status,
|
||||
)
|
||||
|
||||
var body *string
|
||||
|
||||
if c.inlineBody {
|
||||
bodyStr := event.Body
|
||||
body = &bodyStr
|
||||
}
|
||||
|
||||
task := tsTask(
|
||||
d, event, s.WebhookID, target, c.attemptNum, body,
|
||||
)
|
||||
|
||||
c.process(context.TODO(), s.Engine, &task)
|
||||
|
||||
return s, d, sink.text(t)
|
||||
}
|
||||
|
||||
// TestSlackFirstAttemptCarriesEventTimestamp covers the path an
|
||||
// event takes on its first delivery: the task comes from the
|
||||
// receiver and the engine reconstructs the event from it.
|
||||
func TestSlackFirstAttemptCarriesEventTimestamp(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
s, d, text := tsCase{
|
||||
status: database.DeliveryStatusPending,
|
||||
inlineBody: true,
|
||||
attemptNum: 1,
|
||||
process: func(
|
||||
ctx context.Context,
|
||||
e *delivery.Engine,
|
||||
task *delivery.Task,
|
||||
) {
|
||||
e.ExportProcessNewTask(ctx, task)
|
||||
},
|
||||
}.run(t)
|
||||
|
||||
tsAssertRealTimestamp(t, text)
|
||||
|
||||
iAssertStatus(t, s.WebhookDB, d.ID,
|
||||
database.DeliveryStatusDelivered,
|
||||
)
|
||||
}
|
||||
|
||||
// TestSlackFirstAttemptLargeBodyCarriesEventTimestamp covers the
|
||||
// first-attempt path for an event whose body exceeded
|
||||
// MaxInlineBodySize, so the task carries no body and the engine
|
||||
// reads it back from the stored row.
|
||||
func TestSlackFirstAttemptLargeBodyCarriesEventTimestamp(
|
||||
t *testing.T,
|
||||
) {
|
||||
t.Parallel()
|
||||
|
||||
_, _, text := tsCase{
|
||||
status: database.DeliveryStatusPending,
|
||||
inlineBody: false,
|
||||
attemptNum: 1,
|
||||
process: func(
|
||||
ctx context.Context,
|
||||
e *delivery.Engine,
|
||||
task *delivery.Task,
|
||||
) {
|
||||
e.ExportProcessNewTask(ctx, task)
|
||||
},
|
||||
}.run(t)
|
||||
|
||||
tsAssertRealTimestamp(t, text)
|
||||
}
|
||||
|
||||
// TestSlackRetryCarriesEventTimestamp covers the retry path, which
|
||||
// reconstructs the event from the same task the first attempt used.
|
||||
func TestSlackRetryCarriesEventTimestamp(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
s, d, text := tsCase{
|
||||
status: database.DeliveryStatusRetrying,
|
||||
inlineBody: true,
|
||||
attemptNum: 2,
|
||||
process: func(
|
||||
ctx context.Context,
|
||||
e *delivery.Engine,
|
||||
task *delivery.Task,
|
||||
) {
|
||||
e.ExportProcessRetryTask(ctx, task)
|
||||
},
|
||||
}.run(t)
|
||||
|
||||
tsAssertRealTimestamp(t, text)
|
||||
|
||||
iAssertStatus(t, s.WebhookDB, d.ID,
|
||||
database.DeliveryStatusDelivered,
|
||||
)
|
||||
}
|
||||
|
||||
// TestFormatSlackMessageOverTaskReconstructedEvent asserts on the
|
||||
// formatted message directly, over the event the delivery paths
|
||||
// reconstruct from a Task. It is the unit-level guard under the
|
||||
// end-to-end tests: revert the CreatedAt population in hydrateEvent
|
||||
// and this fails on the zero timestamp.
|
||||
func TestFormatSlackMessageOverTaskReconstructedEvent(
|
||||
t *testing.T,
|
||||
) {
|
||||
t.Parallel()
|
||||
|
||||
s := newISetup(t)
|
||||
|
||||
cfg := tsSlackConfig(t, tsUndeliverableHook)
|
||||
target := tsSeedTarget(t, s.MainDB, s.WebhookID, cfg)
|
||||
event := tsSeedEvent(t, s.WebhookDB, s.WebhookID)
|
||||
|
||||
d := iSeedDelivery(
|
||||
t, s.WebhookDB, event.ID, target.ID,
|
||||
database.DeliveryStatusPending,
|
||||
)
|
||||
|
||||
bodyStr := event.Body
|
||||
task := tsTask(d, event, s.WebhookID, target, 1, &bodyStr)
|
||||
|
||||
rebuilt, err := s.Engine.ExportEventForTask(
|
||||
s.WebhookDB, &task,
|
||||
)
|
||||
require.NoError(t, err)
|
||||
assert.False(t, rebuilt.CreatedAt.IsZero(),
|
||||
"reconstructed event carries the zero time",
|
||||
)
|
||||
assert.Equal(t,
|
||||
tsEventCreatedAt().UTC(), rebuilt.CreatedAt.UTC(),
|
||||
)
|
||||
|
||||
tsAssertRealTimestamp(
|
||||
t, delivery.FormatSlackMessage(&rebuilt),
|
||||
)
|
||||
}
|
||||
|
||||
// TestFormatSlackMessageZeroTimestamp asserts the rendering choice
|
||||
// directly, without going through the engine: a zero CreatedAt (the
|
||||
// shape a reaped-row fallback produces) renders as "unknown" rather
|
||||
// than the year-1 zero time, while a real CreatedAt still renders as
|
||||
// RFC3339.
|
||||
func TestFormatSlackMessageZeroTimestamp(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
zeroEvent := database.Event{
|
||||
Method: http.MethodPost,
|
||||
ContentType: testContentType,
|
||||
Body: tsEventBody,
|
||||
}
|
||||
|
||||
zeroText := delivery.FormatSlackMessage(&zeroEvent)
|
||||
|
||||
assert.NotContains(t, zeroText, "0001-01-01",
|
||||
"slack message carries the zero-time year",
|
||||
)
|
||||
assert.Contains(t, zeroText, "*Timestamp:* `unknown`",
|
||||
"slack message does not mark an unset receipt time as unknown",
|
||||
)
|
||||
|
||||
nonZeroEvent := zeroEvent
|
||||
nonZeroEvent.CreatedAt = tsEventCreatedAt()
|
||||
|
||||
nonZeroText := delivery.FormatSlackMessage(&nonZeroEvent)
|
||||
|
||||
assert.Contains(t, nonZeroText,
|
||||
"*Timestamp:* `"+
|
||||
tsEventCreatedAt().UTC().Format(time.RFC3339)+"`",
|
||||
"slack message does not render a real receipt time as RFC3339",
|
||||
)
|
||||
}
|
||||
|
||||
// TestEventReconstructionSurvivesAReapedRow pins the fallback: an
|
||||
// event row reaped by retention while its delivery still holds the
|
||||
// body inline is still delivered, with the receipt time unset,
|
||||
// rather than dropped.
|
||||
func TestEventReconstructionSurvivesAReapedRow(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
s := newISetup(t)
|
||||
|
||||
cfg := tsSlackConfig(t, tsUndeliverableHook)
|
||||
target := tsSeedTarget(t, s.MainDB, s.WebhookID, cfg)
|
||||
event := tsSeedEvent(t, s.WebhookDB, s.WebhookID)
|
||||
|
||||
d := iSeedDelivery(
|
||||
t, s.WebhookDB, event.ID, target.ID,
|
||||
database.DeliveryStatusPending,
|
||||
)
|
||||
|
||||
bodyStr := event.Body
|
||||
task := tsTask(d, event, s.WebhookID, target, 1, &bodyStr)
|
||||
|
||||
require.NoError(t, s.WebhookDB.Unscoped().Delete(
|
||||
&database.Event{}, "id = ?", event.ID,
|
||||
).Error)
|
||||
|
||||
rebuilt, err := s.Engine.ExportEventForTask(
|
||||
s.WebhookDB, &task,
|
||||
)
|
||||
require.NoError(t, err)
|
||||
assert.Equal(t, bodyStr, rebuilt.Body)
|
||||
assert.True(t, rebuilt.CreatedAt.IsZero())
|
||||
|
||||
// A task with no inlined body has nothing left to deliver, so
|
||||
// the same reaped row is an error there.
|
||||
noBody := task
|
||||
noBody.Body = nil
|
||||
|
||||
_, err = s.Engine.ExportEventForTask(s.WebhookDB, &noBody)
|
||||
require.Error(t, err)
|
||||
}
|
||||
@@ -33,6 +33,11 @@ const (
|
||||
// response is written against this number, so a test has to
|
||||
// be able to name it.
|
||||
ExportMaxBodyLog = maxBodyLog
|
||||
|
||||
// ExportPendingSweepMinAge is how long a delivery must sit at
|
||||
// pending before the sweep treats it as stranded. A test has to
|
||||
// name it to age a row past the bound.
|
||||
ExportPendingSweepMinAge = pendingSweepMinAge
|
||||
)
|
||||
|
||||
// ExportIsBlockedIP exposes isBlockedIP for testing.
|
||||
@@ -96,6 +101,19 @@ func (e *Engine) ExportDeliverHTTP(
|
||||
e.httpTarget.Deliver(ctx, webhookDB, d, task, e)
|
||||
}
|
||||
|
||||
// ExportDeliverHTTPWithScheduler delivers via the http target, handing
|
||||
// any retry to sched instead of the engine, so a test can see the
|
||||
// delay each retry is given.
|
||||
func (e *Engine) ExportDeliverHTTPWithScheduler(
|
||||
ctx context.Context,
|
||||
webhookDB *gorm.DB,
|
||||
d *database.Delivery,
|
||||
task *Task,
|
||||
sched Scheduler,
|
||||
) {
|
||||
e.httpTarget.Deliver(ctx, webhookDB, d, task, sched)
|
||||
}
|
||||
|
||||
// ExportDeliverDatabase delivers via the database target.
|
||||
func (e *Engine) ExportDeliverDatabase(
|
||||
webhookDB *gorm.DB, d *database.Delivery,
|
||||
@@ -146,6 +164,16 @@ func (e *Engine) ExportProcessRetryTask(
|
||||
e.processRetryTask(ctx, task)
|
||||
}
|
||||
|
||||
// ExportEventForTask exposes the event reconstruction the delivery
|
||||
// paths run: buildEventFromTask followed by hydrateEvent.
|
||||
func (e *Engine) ExportEventForTask(
|
||||
webhookDB *gorm.DB, task *Task,
|
||||
) (database.Event, error) {
|
||||
return e.hydrateEvent(
|
||||
webhookDB, buildEventFromTask(task), task,
|
||||
)
|
||||
}
|
||||
|
||||
// ExportProcessDelivery exposes processDelivery.
|
||||
func (e *Engine) ExportProcessDelivery(
|
||||
ctx context.Context,
|
||||
@@ -164,6 +192,14 @@ func (e *Engine) ExportGetCircuitBreaker(
|
||||
return e.httpTarget.getCircuitBreaker(targetID)
|
||||
}
|
||||
|
||||
// ExportSetCircuitBreaker makes cb the http target's circuit breaker
|
||||
// for targetID, so a test can use one with a short cooldown.
|
||||
func (e *Engine) ExportSetCircuitBreaker(
|
||||
targetID string, cb *CircuitBreaker,
|
||||
) {
|
||||
e.httpTarget.circuitBreakers.Store(targetID, cb)
|
||||
}
|
||||
|
||||
// ExportParseHTTPConfig exposes parseHTTPConfig.
|
||||
func (e *Engine) ExportParseHTTPConfig(
|
||||
configJSON string,
|
||||
@@ -286,6 +322,51 @@ func (e *Engine) ExportWedgeWorker(release <-chan struct{}) {
|
||||
})
|
||||
}
|
||||
|
||||
// ExportInflightHeld reports how many deliveries the engine currently
|
||||
// owns, so a test can prove ownership is released rather than leaked.
|
||||
func (e *Engine) ExportInflightHeld() int {
|
||||
return e.inflight.held()
|
||||
}
|
||||
|
||||
// ExportRetainDelivery takes the first reference on a delivery, as the
|
||||
// queueing side does. It lets a test put a delivery into the state a
|
||||
// worker or a full channel would, without running the pool.
|
||||
func (e *Engine) ExportRetainDelivery(deliveryID string) bool {
|
||||
return e.inflight.retainIdle(deliveryID)
|
||||
}
|
||||
|
||||
// ExportRecoverRetryingDeliveries exposes recoverRetryingDeliveries.
|
||||
func (e *Engine) ExportRecoverRetryingDeliveries(
|
||||
webhookDB *gorm.DB, webhookID string,
|
||||
) {
|
||||
e.recoverRetryingDeliveries(webhookDB, webhookID)
|
||||
}
|
||||
|
||||
// ExportFailMissingTarget exposes failMissingTarget, so a test can hand
|
||||
// it a delivery as a batch read it earlier.
|
||||
func (e *Engine) ExportFailMissingTarget(
|
||||
webhookDB *gorm.DB,
|
||||
webhookID string,
|
||||
d *database.Delivery,
|
||||
) {
|
||||
e.failMissingTarget(webhookDB, webhookID, d)
|
||||
}
|
||||
|
||||
// ExportSendRecoveredDeliveries exposes sendRecoveredDeliveries, so a
|
||||
// test can hand it a target map that lacks a delivery's target.
|
||||
func (e *Engine) ExportSendRecoveredDeliveries(
|
||||
ctx context.Context,
|
||||
webhookDB *gorm.DB,
|
||||
deliveries []database.Delivery,
|
||||
webhookID string,
|
||||
targetMap map[string]database.Target,
|
||||
settled map[string]struct{},
|
||||
) {
|
||||
e.sendRecoveredDeliveries(
|
||||
ctx, webhookDB, deliveries, webhookID, targetMap, settled,
|
||||
)
|
||||
}
|
||||
|
||||
// ExportDeliveryCh returns the delivery channel.
|
||||
func (e *Engine) ExportDeliveryCh() chan Task {
|
||||
return e.deliveryCh
|
||||
|
||||
@@ -0,0 +1,110 @@
|
||||
package delivery
|
||||
|
||||
import "sync"
|
||||
|
||||
// inflightSet records which deliveries the engine currently owns.
|
||||
//
|
||||
// A delivery is owned from the moment a task for it is handed to a
|
||||
// channel or to a retry timer until the engine has no further plan for
|
||||
// it in memory. Restart recovery and both arms of the periodic sweep
|
||||
// re-dispatch only deliveries the set does not hold, which is what
|
||||
// makes them exact rather than a guess about how long a row has sat at
|
||||
// pending.
|
||||
//
|
||||
// This replaces reasoning from timestamps. A delivery's row says
|
||||
// pending from creation until its outcome is written, which covers
|
||||
// four different situations — never dispatched, waiting in a channel,
|
||||
// being attempted right now, and genuinely stranded — and no column
|
||||
// distinguishes them. Only the engine knows which, and it knows
|
||||
// exactly. `deliveryChannelSize` is 10000 against 10 workers, so a
|
||||
// perfectly healthy delivery can wait far longer than any age bound
|
||||
// worth setting before its attempt even begins; an age bound alone
|
||||
// re-sends it. See
|
||||
// https://git.eeqj.de/sneak/webhooker/issues/256.
|
||||
//
|
||||
// In-memory state is sufficient because a data directory admits one
|
||||
// process: internal/datadir takes an flock on it at startup and a
|
||||
// second instance refuses to run. Deliveries owned by a process that
|
||||
// died are not in any successor's set, and restart recovery is what
|
||||
// picks those up.
|
||||
//
|
||||
// References are counted rather than held as a plain set because
|
||||
// ownership outlives the worker that took it. A target that schedules
|
||||
// a retry from inside Deliver adds a reference while the worker still
|
||||
// holds one, so the delivery stays owned across the gap between the
|
||||
// worker returning and the timer firing — the window in which a sweep
|
||||
// would otherwise find the row at retrying and send it again.
|
||||
//
|
||||
// The zero value is ready to use, and the Engine holds one by value.
|
||||
// That is deliberate: an engine built by a constructor that forgot to
|
||||
// initialise this would not refuse to re-dispatch anything, and the
|
||||
// symptom would be duplicate deliveries rather than a failure anybody
|
||||
// notices.
|
||||
type inflightSet struct {
|
||||
mu sync.Mutex
|
||||
ids map[string]int
|
||||
}
|
||||
|
||||
// retain adds a reference to a delivery the caller already knows the
|
||||
// engine owns, so that ownership survives the current holder letting
|
||||
// go. It cannot fail.
|
||||
func (s *inflightSet) retain(deliveryID string) {
|
||||
s.mu.Lock()
|
||||
defer s.mu.Unlock()
|
||||
|
||||
if s.ids == nil {
|
||||
s.ids = make(map[string]int)
|
||||
}
|
||||
|
||||
s.ids[deliveryID]++
|
||||
}
|
||||
|
||||
// retainIdle takes the first reference on a delivery, and reports
|
||||
// whether it got it. It fails when the engine already owns the
|
||||
// delivery, which is what makes two claimants — restart recovery and
|
||||
// the sweep run concurrently, or two sweep arms — mutually exclusive
|
||||
// rather than merely atomic.
|
||||
func (s *inflightSet) retainIdle(deliveryID string) bool {
|
||||
s.mu.Lock()
|
||||
defer s.mu.Unlock()
|
||||
|
||||
if s.ids[deliveryID] > 0 {
|
||||
return false
|
||||
}
|
||||
|
||||
if s.ids == nil {
|
||||
s.ids = make(map[string]int)
|
||||
}
|
||||
|
||||
s.ids[deliveryID] = 1
|
||||
|
||||
return true
|
||||
}
|
||||
|
||||
// release drops one reference. The delivery becomes eligible for
|
||||
// re-dispatch again once the last one goes.
|
||||
func (s *inflightSet) release(deliveryID string) {
|
||||
s.mu.Lock()
|
||||
defer s.mu.Unlock()
|
||||
|
||||
n := s.ids[deliveryID] - 1
|
||||
if n <= 0 {
|
||||
delete(s.ids, deliveryID)
|
||||
|
||||
return
|
||||
}
|
||||
|
||||
s.ids[deliveryID] = n
|
||||
}
|
||||
|
||||
// held reports how many deliveries the engine currently owns. It
|
||||
// exists so a test can assert that ownership is released rather than
|
||||
// leaked: a reference that is never dropped hides its delivery from
|
||||
// every sweep for the life of the process, which is the one way this
|
||||
// mechanism can fail silently.
|
||||
func (s *inflightSet) held() int {
|
||||
s.mu.Lock()
|
||||
defer s.mu.Unlock()
|
||||
|
||||
return len(s.ids)
|
||||
}
|
||||
@@ -0,0 +1,495 @@
|
||||
package delivery_test
|
||||
|
||||
import (
|
||||
"context"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"sync"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"github.com/google/uuid"
|
||||
"github.com/stretchr/testify/assert"
|
||||
"github.com/stretchr/testify/require"
|
||||
"sneak.berlin/go/webhooker/internal/database"
|
||||
"sneak.berlin/go/webhooker/internal/delivery"
|
||||
)
|
||||
|
||||
// These tests pin the rule that decides whether a delivery may be
|
||||
// handed back to a worker: the engine re-dispatches only what it does
|
||||
// not already own. Age alone is not that rule — a healthy delivery
|
||||
// waiting in a 10000-deep channel is old and must not be re-sent. See
|
||||
// https://git.eeqj.de/sneak/webhooker/issues/256.
|
||||
|
||||
// fSweepSetup seeds the main database with the webhook row the sweep
|
||||
// enumerates, and returns the setup.
|
||||
func fSweepSetup(
|
||||
t *testing.T, targetID, name string,
|
||||
) iSetup {
|
||||
t.Helper()
|
||||
|
||||
s := newISetup(t)
|
||||
|
||||
iCreateTarget(t, s.MainDB, targetID,
|
||||
s.WebhookID, name,
|
||||
database.TargetTypeLog, "", 0,
|
||||
)
|
||||
|
||||
require.NoError(t, s.MainDB.Create(&database.Webhook{
|
||||
BaseModel: database.BaseModel{ID: s.WebhookID},
|
||||
UserID: uuid.New().String(),
|
||||
Name: name,
|
||||
}).Error)
|
||||
|
||||
return s
|
||||
}
|
||||
|
||||
// fDrain collects every task the engine has queued.
|
||||
//
|
||||
// Every caller drives the dispatch paths synchronously and has already
|
||||
// waited for them to return, so anything they queued is in the channel
|
||||
// by now. The short grace covers nothing but scheduler jitter, and is
|
||||
// kept small because one of these tests runs the drain forty times.
|
||||
func fDrain(e *delivery.Engine) []delivery.Task {
|
||||
var out []delivery.Task
|
||||
|
||||
for {
|
||||
select {
|
||||
case task := <-e.ExportDeliveryCh():
|
||||
out = append(out, task)
|
||||
case task := <-e.ExportRetryCh():
|
||||
out = append(out, task)
|
||||
case <-time.After(25 * time.Millisecond):
|
||||
return out
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestArchiveHandleIsWAL closes the last gap in the durability
|
||||
// evidence: the main and per-webhook tiers each assert their journal
|
||||
// mode on a live handle, and the archive tier gets its settings from
|
||||
// the same code path but nothing checked the running file.
|
||||
func TestArchiveHandleIsWAL(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
w := delivery.NewExportArchiveWriter(
|
||||
filepath.Join(t.TempDir(), "archive-wal.db"),
|
||||
archiveTestLogger(), 0,
|
||||
)
|
||||
|
||||
require.NoError(t, w.Open(0))
|
||||
|
||||
var mode string
|
||||
|
||||
row := w.DB().Raw("pragma journal_mode").Row()
|
||||
require.NoError(t, row.Scan(&mode))
|
||||
assert.Equal(t, "wal", strings.ToLower(mode))
|
||||
|
||||
var busy string
|
||||
|
||||
row = w.DB().Raw("pragma busy_timeout").Row()
|
||||
require.NoError(t, row.Scan(&busy))
|
||||
assert.Equal(t, "10000", busy)
|
||||
}
|
||||
|
||||
// TestSweepLeavesAQueuedDeliveryAlone is the case the age bound cannot
|
||||
// see. The delivery is queued and untouched, so its row is arbitrarily
|
||||
// old and still perfectly healthy; only ownership distinguishes it
|
||||
// from a stranded one.
|
||||
func TestSweepLeavesAQueuedDeliveryAlone(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
targetID := uuid.New().String()
|
||||
s := fSweepSetup(t, targetID, "queued")
|
||||
|
||||
event := iSeedEvent(
|
||||
t, s.WebhookDB, s.WebhookID, `{"queued":true}`,
|
||||
)
|
||||
|
||||
d := iSeedDelivery(
|
||||
t, s.WebhookDB, event.ID, targetID,
|
||||
database.DeliveryStatusPending,
|
||||
)
|
||||
rAgePending(t, s.WebhookDB, d.ID)
|
||||
|
||||
// Queued exactly as the receiver queues it, and never dequeued:
|
||||
// no workers are running in this engine.
|
||||
s.Engine.Notify([]delivery.Task{{
|
||||
DeliveryID: d.ID,
|
||||
EventID: event.ID,
|
||||
WebhookID: s.WebhookID,
|
||||
TargetID: targetID,
|
||||
}})
|
||||
|
||||
require.Equal(t, 1, s.Engine.ExportInflightHeld())
|
||||
|
||||
s.Engine.ExportSweepWebhookRetries(
|
||||
context.Background(), s.WebhookID,
|
||||
)
|
||||
|
||||
tasks := fDrain(s.Engine)
|
||||
assert.Len(
|
||||
t, tasks, 1,
|
||||
"the sweep must not queue a delivery that is "+
|
||||
"already waiting for a worker",
|
||||
)
|
||||
}
|
||||
|
||||
// TestRecoveryAndSweepDoNotDoubleDispatch drives the two entry points
|
||||
// the engine starts concurrently against one aged pending row. Before
|
||||
// ownership they both dispatched it.
|
||||
func TestRecoveryAndSweepDoNotDoubleDispatch(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
targetID := uuid.New().String()
|
||||
s := fSweepSetup(t, targetID, "racing")
|
||||
|
||||
event := iSeedEvent(
|
||||
t, s.WebhookDB, s.WebhookID, `{"racing":true}`,
|
||||
)
|
||||
|
||||
ctx := context.Background()
|
||||
|
||||
for range 40 {
|
||||
d := iSeedDelivery(
|
||||
t, s.WebhookDB, event.ID, targetID,
|
||||
database.DeliveryStatusPending,
|
||||
)
|
||||
rAgePending(t, s.WebhookDB, d.ID)
|
||||
|
||||
var wg sync.WaitGroup
|
||||
|
||||
wg.Go(func() {
|
||||
s.Engine.ExportRecoverPendingDeliveries(
|
||||
ctx, s.WebhookDB, s.WebhookID,
|
||||
)
|
||||
})
|
||||
wg.Go(func() {
|
||||
s.Engine.ExportSweepWebhookRetries(
|
||||
ctx, s.WebhookID,
|
||||
)
|
||||
})
|
||||
wg.Wait()
|
||||
|
||||
tasks := fDrain(s.Engine)
|
||||
require.Len(
|
||||
t, tasks, 1,
|
||||
"delivery %s dispatched %d times",
|
||||
d.ID, len(tasks),
|
||||
)
|
||||
|
||||
// No worker runs in this engine, so the reference the winner
|
||||
// took is never released and earlier iterations' deliveries
|
||||
// stay owned — which is itself the property under test, since
|
||||
// both paths see them on every subsequent pass.
|
||||
}
|
||||
}
|
||||
|
||||
// TestConcurrentClaimsOfOneDeliveryYieldOneOwner exercises the
|
||||
// exclusion directly, rather than arguing it from a SQL predicate.
|
||||
func TestConcurrentClaimsOfOneDeliveryYieldOneOwner(
|
||||
t *testing.T,
|
||||
) {
|
||||
t.Parallel()
|
||||
|
||||
eng := newISetup(t).Engine
|
||||
deliveryID := uuid.New().String()
|
||||
|
||||
var (
|
||||
wg sync.WaitGroup
|
||||
mu sync.Mutex
|
||||
won int
|
||||
)
|
||||
|
||||
for range 64 {
|
||||
wg.Go(func() {
|
||||
if eng.ExportRetainDelivery(deliveryID) {
|
||||
mu.Lock()
|
||||
won++
|
||||
mu.Unlock()
|
||||
}
|
||||
})
|
||||
}
|
||||
|
||||
wg.Wait()
|
||||
|
||||
assert.Equal(t, 1, won)
|
||||
assert.Equal(t, 1, eng.ExportInflightHeld())
|
||||
}
|
||||
|
||||
// TestOwnershipIsReleasedAfterDelivery guards the other direction: a
|
||||
// leaked reference hides a delivery from every sweep for the life of
|
||||
// the process.
|
||||
func TestOwnershipIsReleasedAfterDelivery(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
s := newISetup(t)
|
||||
targetID := uuid.New().String()
|
||||
|
||||
iCreateTarget(t, s.MainDB, targetID,
|
||||
s.WebhookID, "released",
|
||||
database.TargetTypeLog, "", 0,
|
||||
)
|
||||
|
||||
event := iSeedEvent(
|
||||
t, s.WebhookDB, s.WebhookID, `{"released":true}`,
|
||||
)
|
||||
|
||||
d := iSeedDelivery(
|
||||
t, s.WebhookDB, event.ID, targetID,
|
||||
database.DeliveryStatusPending,
|
||||
)
|
||||
|
||||
s.Engine.ExportStart()
|
||||
|
||||
defer func() {
|
||||
require.NoError(
|
||||
t, s.Engine.ExportStop(context.Background()),
|
||||
)
|
||||
}()
|
||||
|
||||
body := `{"released":true}`
|
||||
|
||||
s.Engine.Notify([]delivery.Task{{
|
||||
DeliveryID: d.ID,
|
||||
EventID: event.ID,
|
||||
WebhookID: s.WebhookID,
|
||||
TargetID: targetID,
|
||||
TargetName: "released",
|
||||
TargetType: database.TargetTypeLog,
|
||||
Body: &body,
|
||||
EntrypointID: event.EntrypointID,
|
||||
}})
|
||||
|
||||
iWaitForDelivered(t, s.WebhookDB, d.ID)
|
||||
|
||||
assert.Eventually(
|
||||
t,
|
||||
func() bool {
|
||||
return s.Engine.ExportInflightHeld() == 0
|
||||
},
|
||||
2*time.Second, 20*time.Millisecond,
|
||||
"the delivery stayed owned after it was delivered",
|
||||
)
|
||||
}
|
||||
|
||||
// TestNotifyAfterRecoveryDoesNotSendAgain is the startup race of
|
||||
// https://git.eeqj.de/sneak/webhooker/issues/299. The receiver has
|
||||
// written a delivery, restart recovery finds it pending, sends it and
|
||||
// releases it, and only then does the receiver's Notify for it arrive.
|
||||
// Nothing owns the delivery by then, so Notify takes it.
|
||||
func TestNotifyAfterRecoveryDoesNotSendAgain(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
targetID := uuid.New().String()
|
||||
s := fSweepSetup(t, targetID, "recovered")
|
||||
|
||||
event := iSeedEvent(
|
||||
t, s.WebhookDB, s.WebhookID, `{"recovered":true}`,
|
||||
)
|
||||
|
||||
d := iSeedDelivery(
|
||||
t, s.WebhookDB, event.ID, targetID,
|
||||
database.DeliveryStatusPending,
|
||||
)
|
||||
|
||||
s.Engine.ExportStart()
|
||||
|
||||
defer func() {
|
||||
require.NoError(
|
||||
t, s.Engine.ExportStop(context.Background()),
|
||||
)
|
||||
}()
|
||||
|
||||
// Restart recovery sends the delivery and lets it go.
|
||||
iWaitForDelivered(t, s.WebhookDB, d.ID)
|
||||
require.Eventually(
|
||||
t,
|
||||
func() bool {
|
||||
return s.Engine.ExportInflightHeld() == 0
|
||||
},
|
||||
5*time.Second, 20*time.Millisecond,
|
||||
)
|
||||
|
||||
body := event.Body
|
||||
|
||||
s.Engine.Notify([]delivery.Task{{
|
||||
DeliveryID: d.ID,
|
||||
EventID: event.ID,
|
||||
WebhookID: s.WebhookID,
|
||||
TargetID: targetID,
|
||||
TargetName: "recovered",
|
||||
TargetType: database.TargetTypeLog,
|
||||
Body: &body,
|
||||
EntrypointID: event.EntrypointID,
|
||||
}})
|
||||
|
||||
// Notify took the delivery, and a worker releases it once it has
|
||||
// run the task.
|
||||
require.Eventually(
|
||||
t,
|
||||
func() bool {
|
||||
return s.Engine.ExportInflightHeld() == 0
|
||||
},
|
||||
5*time.Second, 20*time.Millisecond,
|
||||
)
|
||||
|
||||
assert.Len(
|
||||
t, iResults(t, s.WebhookDB, d.ID), 1,
|
||||
"the delivery was sent a second time",
|
||||
)
|
||||
}
|
||||
|
||||
// TestRetryingRecoverySkipsASuccessfulResult is the retrying-side twin
|
||||
// of the pending reconcile. A second attempt that reached the receiver
|
||||
// and whose status write then failed sits at retrying holding a
|
||||
// successful result, and re-sending it is the same duplicate.
|
||||
func TestRetryingRecoverySkipsASuccessfulResult(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
targetID := uuid.New().String()
|
||||
s := fSweepSetup(t, targetID, "retry-settled")
|
||||
|
||||
event := iSeedEvent(
|
||||
t, s.WebhookDB, s.WebhookID, `{"retry":true}`,
|
||||
)
|
||||
|
||||
d := iSeedDelivery(
|
||||
t, s.WebhookDB, event.ID, targetID,
|
||||
database.DeliveryStatusRetrying,
|
||||
)
|
||||
rSeedResult(t, s.WebhookDB, d.ID, 1, false)
|
||||
rSeedResult(t, s.WebhookDB, d.ID, 2, true)
|
||||
|
||||
s.Engine.ExportRecoverRetryingDeliveries(
|
||||
s.WebhookDB, s.WebhookID,
|
||||
)
|
||||
|
||||
assert.Empty(
|
||||
t, fDrain(s.Engine),
|
||||
"a retrying delivery holding a successful result "+
|
||||
"must not be sent again",
|
||||
)
|
||||
|
||||
iAssertStatus(
|
||||
t, s.WebhookDB, d.ID,
|
||||
database.DeliveryStatusDelivered,
|
||||
)
|
||||
}
|
||||
|
||||
// TestRetryingSweepSkipsASuccessfulResult is the same rule on the
|
||||
// periodic sweep's retrying arm.
|
||||
func TestRetryingSweepSkipsASuccessfulResult(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
targetID := uuid.New().String()
|
||||
s := fSweepSetup(t, targetID, "retry-swept")
|
||||
|
||||
event := iSeedEvent(
|
||||
t, s.WebhookDB, s.WebhookID, `{"swept":true}`,
|
||||
)
|
||||
|
||||
d := iSeedDelivery(
|
||||
t, s.WebhookDB, event.ID, targetID,
|
||||
database.DeliveryStatusRetrying,
|
||||
)
|
||||
rSeedResult(t, s.WebhookDB, d.ID, 1, false)
|
||||
rSeedResult(t, s.WebhookDB, d.ID, 2, true)
|
||||
|
||||
s.Engine.ExportSweepWebhookRetries(
|
||||
context.Background(), s.WebhookID,
|
||||
)
|
||||
|
||||
assert.Empty(t, fDrain(s.Engine))
|
||||
|
||||
iAssertStatus(
|
||||
t, s.WebhookDB, d.ID,
|
||||
database.DeliveryStatusDelivered,
|
||||
)
|
||||
|
||||
var attempts int64
|
||||
|
||||
require.NoError(t, s.WebhookDB.
|
||||
Model(&database.DeliveryResult{}).
|
||||
Where("delivery_id = ?", d.ID).
|
||||
Count(&attempts).Error)
|
||||
assert.Equal(
|
||||
t, int64(2), attempts,
|
||||
"settling must not invent an attempt",
|
||||
)
|
||||
}
|
||||
|
||||
// TestScheduledRetryIsNotSweptDuringBackoff closes the window between
|
||||
// a target scheduling a retry and the timer firing. The row says
|
||||
// retrying and nothing is running, which is exactly what an orphaned
|
||||
// retry looks like from the database.
|
||||
func TestScheduledRetryIsNotSweptDuringBackoff(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
targetID := uuid.New().String()
|
||||
s := fSweepSetup(t, targetID, "backoff")
|
||||
|
||||
event := iSeedEvent(
|
||||
t, s.WebhookDB, s.WebhookID, `{"backoff":true}`,
|
||||
)
|
||||
|
||||
d := iSeedDelivery(
|
||||
t, s.WebhookDB, event.ID, targetID,
|
||||
database.DeliveryStatusRetrying,
|
||||
)
|
||||
|
||||
s.Engine.ExportScheduleRetry(delivery.Task{
|
||||
DeliveryID: d.ID,
|
||||
EventID: event.ID,
|
||||
WebhookID: s.WebhookID,
|
||||
TargetID: targetID,
|
||||
AttemptNum: 2,
|
||||
}, time.Hour)
|
||||
|
||||
require.Equal(t, 1, s.Engine.ExportInflightHeld())
|
||||
|
||||
s.Engine.ExportSweepWebhookRetries(
|
||||
context.Background(), s.WebhookID,
|
||||
)
|
||||
|
||||
assert.Empty(
|
||||
t, fDrain(s.Engine),
|
||||
"the sweep must not duplicate a retry that is "+
|
||||
"already scheduled",
|
||||
)
|
||||
}
|
||||
|
||||
// TestRedispatchStampsTheRow pins the cadence control: a stranded
|
||||
// delivery that has just been handed out is not selected again by the
|
||||
// next tick a minute later.
|
||||
func TestRedispatchStampsTheRow(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
targetID := uuid.New().String()
|
||||
s := fSweepSetup(t, targetID, "stamped")
|
||||
|
||||
event := iSeedEvent(
|
||||
t, s.WebhookDB, s.WebhookID, `{"stamped":true}`,
|
||||
)
|
||||
|
||||
d := iSeedDelivery(
|
||||
t, s.WebhookDB, event.ID, targetID,
|
||||
database.DeliveryStatusPending,
|
||||
)
|
||||
rAgePending(t, s.WebhookDB, d.ID)
|
||||
|
||||
ctx := context.Background()
|
||||
|
||||
s.Engine.ExportSweepWebhookRetries(ctx, s.WebhookID)
|
||||
require.Len(t, fDrain(s.Engine), 1)
|
||||
|
||||
var row database.Delivery
|
||||
|
||||
require.NoError(t, s.WebhookDB.
|
||||
First(&row, "id = ?", d.ID).Error)
|
||||
assert.WithinDuration(
|
||||
t, time.Now(), row.UpdatedAt, time.Minute,
|
||||
"a re-dispatched delivery must be stamped so the "+
|
||||
"next tick does not select it again",
|
||||
)
|
||||
}
|
||||
@@ -229,6 +229,13 @@ func mExhaustRetries(t *testing.T, s iSetup) {
|
||||
body := event.Body
|
||||
cfg := iHTTPConfig(ts.URL)
|
||||
|
||||
// The retry below is only run if its target still exists; see
|
||||
// https://git.eeqj.de/sneak/webhooker/issues/107.
|
||||
iCreateTarget(
|
||||
t, s.MainDB, targetID, s.WebhookID, "metrics-fail",
|
||||
database.TargetTypeHTTP, cfg, 2,
|
||||
)
|
||||
|
||||
first := iTask(
|
||||
d, event, s.WebhookID, targetID,
|
||||
"metrics-fail", cfg, 2, 1, &body,
|
||||
@@ -289,6 +296,13 @@ func TestDeliveryMetrics_CircuitBreakerGauge(t *testing.T) {
|
||||
// rather than the budget is what stops the delivery.
|
||||
maxRetries := delivery.ExportDefaultFailureThreshold + 5
|
||||
|
||||
// The retries below are only run if their target still exists;
|
||||
// see https://git.eeqj.de/sneak/webhooker/issues/107.
|
||||
iCreateTarget(
|
||||
t, s.MainDB, targetID, s.WebhookID, "metrics-trip",
|
||||
database.TargetTypeHTTP, cfg, maxRetries,
|
||||
)
|
||||
|
||||
first := iTask(
|
||||
d, event, s.WebhookID, targetID,
|
||||
"metrics-trip", cfg, maxRetries, 1, &body,
|
||||
@@ -353,6 +367,11 @@ func TestDeliveryMetrics_BreakerBlockedIsNotAnAttempt(
|
||||
cfg := iHTTPConfig(ts.URL)
|
||||
maxRetries := delivery.ExportDefaultFailureThreshold + 5
|
||||
|
||||
iCreateTarget(
|
||||
t, s.MainDB, targetID, s.WebhookID, "metrics-blocked",
|
||||
database.TargetTypeHTTP, cfg, maxRetries,
|
||||
)
|
||||
|
||||
first := iTask(
|
||||
d, event, s.WebhookID, targetID,
|
||||
"metrics-blocked", cfg, maxRetries, 1, &body,
|
||||
@@ -393,9 +412,10 @@ func TestDeliveryMetrics_BreakerBlockedIsNotAnAttempt(
|
||||
|
||||
s.Engine.ExportProcessRetryTask(context.TODO(), &blocked)
|
||||
|
||||
// The breaker refused it: rescheduled, so the retry counter
|
||||
// moved, but nothing was attempted or timed.
|
||||
assert.InDelta(t, retriesBefore+1,
|
||||
// The breaker refused it: rescheduled without rewriting the
|
||||
// retrying status it already had, so the retry counter did not
|
||||
// move, and nothing was attempted or timed.
|
||||
assert.InDelta(t, retriesBefore,
|
||||
mCounter(t, reg, mRetries, mTypeHTTP), 0)
|
||||
assert.InDelta(t, threshold,
|
||||
mCounter(t, reg, mAttempts, mTypeHTTP), 0)
|
||||
|
||||
@@ -3,8 +3,6 @@ package delivery_test
|
||||
import (
|
||||
"bytes"
|
||||
"context"
|
||||
"database/sql"
|
||||
"fmt"
|
||||
"log/slog"
|
||||
"net/http"
|
||||
"path/filepath"
|
||||
@@ -54,12 +52,10 @@ func (q *qdSyncBuf) String() string {
|
||||
func qdMainDB(t *testing.T, log *slog.Logger) *gorm.DB {
|
||||
t.Helper()
|
||||
|
||||
dsn := fmt.Sprintf(
|
||||
"file:%s?cache=shared&mode=rwc",
|
||||
sqlDB, err := database.OpenSQLite(
|
||||
filepath.Join(t.TempDir(), "main-gormlog.db"),
|
||||
database.SQLiteModeCreate,
|
||||
)
|
||||
|
||||
sqlDB, err := sql.Open("sqlite", dsn)
|
||||
require.NoError(t, err)
|
||||
|
||||
t.Cleanup(func() { _ = sqlDB.Close() })
|
||||
|
||||
@@ -0,0 +1,378 @@
|
||||
package delivery_test
|
||||
|
||||
import (
|
||||
"context"
|
||||
"net/http"
|
||||
"net/http/httptest"
|
||||
"sync/atomic"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"github.com/google/uuid"
|
||||
"github.com/stretchr/testify/assert"
|
||||
"github.com/stretchr/testify/require"
|
||||
"gorm.io/gorm"
|
||||
"sneak.berlin/go/webhooker/internal/database"
|
||||
"sneak.berlin/go/webhooker/internal/delivery"
|
||||
)
|
||||
|
||||
// These tests cover the delivery half of
|
||||
// https://git.eeqj.de/sneak/webhooker/issues/256: a delivery that
|
||||
// reached its receiver but whose bookkeeping write failed used to be
|
||||
// left at pending and re-sent on the next restart, giving the receiver
|
||||
// a second copy while the event log recorded one attempt.
|
||||
|
||||
// rSeedResult records a DeliveryResult against a delivery, standing in
|
||||
// for the attempt row the send path writes before the status.
|
||||
func rSeedResult(
|
||||
t *testing.T,
|
||||
db *gorm.DB,
|
||||
deliveryID string,
|
||||
attemptNum int,
|
||||
success bool,
|
||||
) {
|
||||
t.Helper()
|
||||
|
||||
require.NoError(t, db.Create(&database.DeliveryResult{
|
||||
DeliveryID: deliveryID,
|
||||
AttemptNum: attemptNum,
|
||||
Success: success,
|
||||
}).Error)
|
||||
}
|
||||
|
||||
// rAgePending backdates a delivery past the sweep's age bound, which is
|
||||
// what separates a stranded delivery from one a worker still holds.
|
||||
func rAgePending(
|
||||
t *testing.T, db *gorm.DB, deliveryID string,
|
||||
) {
|
||||
t.Helper()
|
||||
|
||||
old := time.Now().Add(
|
||||
-2 * delivery.ExportPendingSweepMinAge,
|
||||
)
|
||||
|
||||
require.NoError(t, db.Model(&database.Delivery{}).
|
||||
Where("id = ?", deliveryID).
|
||||
UpdateColumn("updated_at", old).Error)
|
||||
}
|
||||
|
||||
func TestRecoverySkipsPendingWithSuccessfulResult(
|
||||
t *testing.T,
|
||||
) {
|
||||
t.Parallel()
|
||||
|
||||
s := newISetup(t)
|
||||
targetID := uuid.New().String()
|
||||
|
||||
iCreateTarget(t, s.MainDB, targetID,
|
||||
s.WebhookID, "already-delivered",
|
||||
database.TargetTypeLog, "", 0,
|
||||
)
|
||||
|
||||
event := iSeedEvent(
|
||||
t, s.WebhookDB, s.WebhookID, `{"delivered":true}`,
|
||||
)
|
||||
|
||||
// The delivery whose send succeeded and whose result row landed:
|
||||
// only the status write failed, so it sits at pending.
|
||||
done := iSeedDelivery(
|
||||
t, s.WebhookDB, event.ID, targetID,
|
||||
database.DeliveryStatusPending,
|
||||
)
|
||||
rSeedResult(t, s.WebhookDB, done.ID, 1, true)
|
||||
|
||||
// A delivery that was genuinely never attempted.
|
||||
fresh := iSeedDelivery(
|
||||
t, s.WebhookDB, event.ID, targetID,
|
||||
database.DeliveryStatusPending,
|
||||
)
|
||||
|
||||
s.Engine.ExportRecoverPendingDeliveries(
|
||||
context.Background(), s.WebhookDB, s.WebhookID,
|
||||
)
|
||||
|
||||
select {
|
||||
case task := <-s.Engine.ExportDeliveryCh():
|
||||
assert.Equal(
|
||||
t, fresh.ID, task.DeliveryID,
|
||||
"only the unattempted delivery may be re-sent",
|
||||
)
|
||||
case <-time.After(2 * time.Second):
|
||||
t.Fatal("expected the unattempted delivery")
|
||||
}
|
||||
|
||||
select {
|
||||
case task := <-s.Engine.ExportDeliveryCh():
|
||||
t.Fatalf(
|
||||
"re-sent an already delivered delivery: %s",
|
||||
task.DeliveryID,
|
||||
)
|
||||
case <-time.After(200 * time.Millisecond):
|
||||
}
|
||||
|
||||
// It is settled rather than merely skipped: leaving it pending
|
||||
// would strand it again on the next sweep.
|
||||
iAssertStatus(
|
||||
t, s.WebhookDB, done.ID,
|
||||
database.DeliveryStatusDelivered,
|
||||
)
|
||||
}
|
||||
|
||||
// TestRecoveryContinuesTheAttemptNumbering pins the audit trail: a
|
||||
// recovered delivery that already recorded two attempts is re-sent as
|
||||
// attempt three, not as attempt one again.
|
||||
func TestRecoveryContinuesTheAttemptNumbering(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
s := newISetup(t)
|
||||
targetID := uuid.New().String()
|
||||
|
||||
iCreateTarget(t, s.MainDB, targetID,
|
||||
s.WebhookID, "numbering",
|
||||
database.TargetTypeLog, "", 0,
|
||||
)
|
||||
|
||||
event := iSeedEvent(
|
||||
t, s.WebhookDB, s.WebhookID, `{"numbering":true}`,
|
||||
)
|
||||
|
||||
d := iSeedDelivery(
|
||||
t, s.WebhookDB, event.ID, targetID,
|
||||
database.DeliveryStatusPending,
|
||||
)
|
||||
|
||||
rSeedResult(t, s.WebhookDB, d.ID, 1, false)
|
||||
rSeedResult(t, s.WebhookDB, d.ID, 2, false)
|
||||
|
||||
s.Engine.ExportRecoverPendingDeliveries(
|
||||
context.Background(), s.WebhookDB, s.WebhookID,
|
||||
)
|
||||
|
||||
select {
|
||||
case task := <-s.Engine.ExportDeliveryCh():
|
||||
assert.Equal(t, d.ID, task.DeliveryID)
|
||||
assert.Equal(t, 3, task.AttemptNum)
|
||||
case <-time.After(2 * time.Second):
|
||||
t.Fatal("expected the delivery to be recovered")
|
||||
}
|
||||
}
|
||||
|
||||
// TestSweepRecoversStrandedPending is the half that removes the
|
||||
// restart requirement: a delivery left at pending is picked up by the
|
||||
// periodic sweep.
|
||||
func TestSweepRecoversStrandedPending(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
targetID := uuid.New().String()
|
||||
s := fSweepSetup(t, targetID, "stranded")
|
||||
|
||||
event := iSeedEvent(
|
||||
t, s.WebhookDB, s.WebhookID, `{"stranded":true}`,
|
||||
)
|
||||
|
||||
stranded := iSeedDelivery(
|
||||
t, s.WebhookDB, event.ID, targetID,
|
||||
database.DeliveryStatusPending,
|
||||
)
|
||||
rAgePending(t, s.WebhookDB, stranded.ID)
|
||||
|
||||
// A delivery a worker may still be holding: young, and therefore
|
||||
// none of the sweep's business.
|
||||
inFlight := iSeedDelivery(
|
||||
t, s.WebhookDB, event.ID, targetID,
|
||||
database.DeliveryStatusPending,
|
||||
)
|
||||
|
||||
s.Engine.ExportSweepWebhookRetries(
|
||||
context.Background(), s.WebhookID,
|
||||
)
|
||||
|
||||
select {
|
||||
case task := <-s.Engine.ExportDeliveryCh():
|
||||
assert.Equal(t, stranded.ID, task.DeliveryID)
|
||||
case <-time.After(2 * time.Second):
|
||||
t.Fatal("expected the stranded delivery")
|
||||
}
|
||||
|
||||
select {
|
||||
case task := <-s.Engine.ExportDeliveryCh():
|
||||
t.Fatalf(
|
||||
"swept an in-flight delivery: %s",
|
||||
task.DeliveryID,
|
||||
)
|
||||
case <-time.After(200 * time.Millisecond):
|
||||
}
|
||||
|
||||
iAssertStatus(
|
||||
t, s.WebhookDB, inFlight.ID,
|
||||
database.DeliveryStatusPending,
|
||||
)
|
||||
}
|
||||
|
||||
// TestSweepClaimsAStrandedDeliveryOnlyOnce guards the repeat the sweep
|
||||
// would otherwise be: the row stays pending for as long as the attempt
|
||||
// runs, and a sweep a minute later must not send it a second time.
|
||||
func TestSweepClaimsAStrandedDeliveryOnlyOnce(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
targetID := uuid.New().String()
|
||||
s := fSweepSetup(t, targetID, "claimed")
|
||||
|
||||
event := iSeedEvent(
|
||||
t, s.WebhookDB, s.WebhookID, `{"claimed":true}`,
|
||||
)
|
||||
|
||||
d := iSeedDelivery(
|
||||
t, s.WebhookDB, event.ID, targetID,
|
||||
database.DeliveryStatusPending,
|
||||
)
|
||||
rAgePending(t, s.WebhookDB, d.ID)
|
||||
|
||||
ctx := context.Background()
|
||||
|
||||
s.Engine.ExportSweepWebhookRetries(ctx, s.WebhookID)
|
||||
|
||||
select {
|
||||
case task := <-s.Engine.ExportDeliveryCh():
|
||||
assert.Equal(t, d.ID, task.DeliveryID)
|
||||
case <-time.After(2 * time.Second):
|
||||
t.Fatal("expected the stranded delivery")
|
||||
}
|
||||
|
||||
// The delivery is still pending — nothing has run it yet — but
|
||||
// the claim must keep the next sweep off it.
|
||||
iAssertStatus(
|
||||
t, s.WebhookDB, d.ID,
|
||||
database.DeliveryStatusPending,
|
||||
)
|
||||
|
||||
s.Engine.ExportSweepWebhookRetries(ctx, s.WebhookID)
|
||||
|
||||
select {
|
||||
case task := <-s.Engine.ExportDeliveryCh():
|
||||
t.Fatalf(
|
||||
"sent a claimed delivery again: %s",
|
||||
task.DeliveryID,
|
||||
)
|
||||
case <-time.After(200 * time.Millisecond):
|
||||
}
|
||||
}
|
||||
|
||||
// TestSweepSettlesStrandedPendingWithoutResending is the sweep's own
|
||||
// version of the reconcile: a stranded delivery holding a successful
|
||||
// result is settled where it stands, and the receiver hears nothing.
|
||||
func TestSweepSettlesStrandedPendingWithoutResending(
|
||||
t *testing.T,
|
||||
) {
|
||||
t.Parallel()
|
||||
|
||||
targetID := uuid.New().String()
|
||||
s := fSweepSetup(t, targetID, "settled")
|
||||
|
||||
event := iSeedEvent(
|
||||
t, s.WebhookDB, s.WebhookID, `{"settled":true}`,
|
||||
)
|
||||
|
||||
d := iSeedDelivery(
|
||||
t, s.WebhookDB, event.ID, targetID,
|
||||
database.DeliveryStatusPending,
|
||||
)
|
||||
rSeedResult(t, s.WebhookDB, d.ID, 1, true)
|
||||
rAgePending(t, s.WebhookDB, d.ID)
|
||||
|
||||
s.Engine.ExportSweepWebhookRetries(
|
||||
context.Background(), s.WebhookID,
|
||||
)
|
||||
|
||||
select {
|
||||
case task := <-s.Engine.ExportDeliveryCh():
|
||||
t.Fatalf(
|
||||
"re-sent a delivery that already succeeded: %s",
|
||||
task.DeliveryID,
|
||||
)
|
||||
case <-time.After(200 * time.Millisecond):
|
||||
}
|
||||
|
||||
iAssertStatus(
|
||||
t, s.WebhookDB, d.ID,
|
||||
database.DeliveryStatusDelivered,
|
||||
)
|
||||
|
||||
var attempts int64
|
||||
|
||||
require.NoError(t, s.WebhookDB.
|
||||
Model(&database.DeliveryResult{}).
|
||||
Where("delivery_id = ?", d.ID).
|
||||
Count(&attempts).Error)
|
||||
assert.Equal(
|
||||
t, int64(1), attempts,
|
||||
"settling must not invent an attempt",
|
||||
)
|
||||
}
|
||||
|
||||
// TestFailedResultWriteLeavesDeliveryRecoverable is the rule the
|
||||
// targets now follow: a bookkeeping write that fails must not advance
|
||||
// the status, because pending and retrying are the states the sweeps
|
||||
// recover and delivered is a claim the database refused to record.
|
||||
func TestFailedResultWriteLeavesDeliveryRecoverable(
|
||||
t *testing.T,
|
||||
) {
|
||||
t.Parallel()
|
||||
|
||||
s := newISetup(t)
|
||||
targetID := uuid.New().String()
|
||||
|
||||
var hits atomic.Int64
|
||||
|
||||
ts := httptest.NewServer(http.HandlerFunc(
|
||||
func(w http.ResponseWriter, _ *http.Request) {
|
||||
hits.Add(1)
|
||||
w.WriteHeader(http.StatusOK)
|
||||
},
|
||||
))
|
||||
defer ts.Close()
|
||||
|
||||
event := iSeedEvent(
|
||||
t, s.WebhookDB, s.WebhookID, `{"unwritable":true}`,
|
||||
)
|
||||
|
||||
d := iSeedDelivery(
|
||||
t, s.WebhookDB, event.ID, targetID,
|
||||
database.DeliveryStatusPending,
|
||||
)
|
||||
|
||||
// Drop the table the attempt row goes in, so the send succeeds
|
||||
// and only the bookkeeping write fails.
|
||||
require.NoError(
|
||||
t,
|
||||
s.WebhookDB.Exec("drop table delivery_results").Error,
|
||||
)
|
||||
|
||||
full := &database.Delivery{
|
||||
EventID: event.ID,
|
||||
TargetID: targetID,
|
||||
Status: database.DeliveryStatusPending,
|
||||
Event: event,
|
||||
Target: database.Target{
|
||||
Name: "unwritable",
|
||||
Type: database.TargetTypeHTTP,
|
||||
Config: iHTTPConfig(ts.URL),
|
||||
},
|
||||
}
|
||||
full.ID = d.ID
|
||||
|
||||
s.Engine.ExportDeliverHTTP(
|
||||
context.Background(), s.WebhookDB, full,
|
||||
&delivery.Task{DeliveryID: d.ID, AttemptNum: 1},
|
||||
)
|
||||
|
||||
assert.Equal(
|
||||
t, int64(1), hits.Load(),
|
||||
"the send itself must still happen",
|
||||
)
|
||||
|
||||
iAssertStatus(
|
||||
t, s.WebhookDB, d.ID,
|
||||
database.DeliveryStatusPending,
|
||||
)
|
||||
}
|
||||
@@ -170,7 +170,8 @@ func TestDelivery_CrossOriginRedirectDropsOriginScopedHeaders(
|
||||
// Stripping must not fire within the configured origin, or every
|
||||
// destination that redirects its own path would lose its
|
||||
// credential and start answering 401 — and would lose the inbound
|
||||
// signature the receiver verifies.
|
||||
// signature header the target endpoint verifies. webhooker's own
|
||||
// receiver verifies no signature; it only forwards the header.
|
||||
func TestDelivery_SameOriginRedirectKeepsOriginScopedHeaders(
|
||||
t *testing.T,
|
||||
) {
|
||||
@@ -338,10 +339,11 @@ func TestRedirectPolicy_StopsAtHopCap(t *testing.T) {
|
||||
// The set the redirect policy strips is whatever the delivery path
|
||||
// actually put on the wire, so a header added to the forward set is
|
||||
// covered without a second edit. A header the event never carried
|
||||
// is not in the set, and the delivery path's own two are deliberately
|
||||
// excluded: Content-Type describes the body, which a 307 carries
|
||||
// across hosts, and the inbound User-Agent every real sender supplies
|
||||
// is overwritten before the request goes out.
|
||||
// is not in the set, and neither is the inbound Content-Type, because
|
||||
// it is not forwarded. Two more are deliberately excluded: a
|
||||
// Content-Type configured on the target describes the body, which a
|
||||
// 307 carries across hosts, and the inbound User-Agent every real
|
||||
// sender supplies is overwritten before the request goes out.
|
||||
func TestApplyRequestHeaders_ReportsOriginScopedNames(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
@@ -370,6 +372,7 @@ func TestApplyRequestHeaders_ReportsOriginScopedNames(t *testing.T) {
|
||||
&delivery.HTTPTargetConfig{
|
||||
Headers: map[string]string{
|
||||
probeHeaderName: probeHeaderValue,
|
||||
"Content-Type": testContentType,
|
||||
},
|
||||
},
|
||||
)
|
||||
@@ -377,7 +380,11 @@ func TestApplyRequestHeaders_ReportsOriginScopedNames(t *testing.T) {
|
||||
assert.Equal(t,
|
||||
[]string{probeHeaderName, inboundHeaderName}, names,
|
||||
"both header classes are reported, and only those: "+
|
||||
"Host is never forwarded, Content-Type and "+
|
||||
"User-Agent are the delivery path's own",
|
||||
"Host and the inbound Content-Type are never "+
|
||||
"forwarded, User-Agent is the delivery path's own",
|
||||
)
|
||||
assert.NotContains(t, names, "Content-Type",
|
||||
"a Content-Type configured on the target must survive "+
|
||||
"a cross-origin 307/308 with the body it describes",
|
||||
)
|
||||
}
|
||||
|
||||
@@ -26,7 +26,7 @@ var (
|
||||
"hostname resolved to no IP addresses",
|
||||
)
|
||||
errBlockedIP = errors.New(
|
||||
"blocked private/reserved IP range",
|
||||
"blocked private, reserved or cloud metadata address",
|
||||
)
|
||||
errBlockedMetadata = errors.New(
|
||||
"blocked link-local or cloud instance metadata " +
|
||||
@@ -37,9 +37,10 @@ var (
|
||||
)
|
||||
)
|
||||
|
||||
// blockedNetworks contains all private/reserved IP ranges
|
||||
// that should be blocked to prevent SSRF attacks. An operator
|
||||
// can permit specific blocks out of this set with
|
||||
// blockedNetworks is the default blocklist: the private and
|
||||
// reserved IP ranges, plus the public cloud metadata addresses,
|
||||
// that are blocked to prevent SSRF attacks. An operator can
|
||||
// permit specific blocks out of this set with
|
||||
// ALLOWED_EGRESS_CIDRS; see Guard.
|
||||
//
|
||||
//nolint:gochecknoglobals // package-level network list is appropriate here
|
||||
@@ -122,6 +123,8 @@ func init() {
|
||||
"::1/128",
|
||||
"fc00::/7",
|
||||
"fe80::/10",
|
||||
// Azure WireServer, a public address that serves VM credentials.
|
||||
"168.63.129.16/32",
|
||||
})
|
||||
|
||||
// Every entry is named. The set must not grow or shrink
|
||||
@@ -216,8 +219,8 @@ func matchesAny(networks []*net.IPNet, ip net.IP) bool {
|
||||
}
|
||||
|
||||
// isBlockedIP checks whether an IP address falls within
|
||||
// any blocked private/reserved network range, before any
|
||||
// operator allowlist is considered.
|
||||
// the default blocklist, before any operator allowlist is
|
||||
// considered.
|
||||
func isBlockedIP(ip net.IP) bool {
|
||||
return matchesAny(blockedNetworks, ip)
|
||||
}
|
||||
@@ -320,7 +323,7 @@ func (g *Guard) allows(ip net.IP) bool {
|
||||
//
|
||||
// 1. alwaysBlockedNetworks is refused before the allowlist is
|
||||
// consulted, so no configured CIDR reaches link-local or a
|
||||
// cloud instance metadata endpoint.
|
||||
// cloud metadata endpoint at a non-public address.
|
||||
// 2. The allowlist is consulted next, so a listed private
|
||||
// network becomes reachable.
|
||||
// 3. Everything else keeps the default blocklist's answer.
|
||||
|
||||
@@ -390,6 +390,41 @@ func TestGuardAllowlist_PublicUnaffected(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
// TestGuardAllowlist_AzureWireServerReopenable covers Azure's
|
||||
// WireServer, a public address that serves VM credentials. The
|
||||
// default guard refuses it, but because it is public it sits in
|
||||
// the default blocklist rather than the unconditional set, so an
|
||||
// operator who lists it can reach it.
|
||||
func TestGuardAllowlist_AzureWireServerReopenable(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
const wireServerIP = "168.63.129.16"
|
||||
|
||||
target := "http://" + wireServerIP + "/?comp=versions"
|
||||
|
||||
defaultGuard := delivery.NewTestGuard()
|
||||
|
||||
err := defaultGuard.ValidateTargetURL(context.Background(), target)
|
||||
require.Error(t, err,
|
||||
"WireServer must be refused with no allowlist set",
|
||||
)
|
||||
assert.NotContains(t, err.Error(), metadataRefusalClause,
|
||||
"WireServer must be refused by the default blocklist, "+
|
||||
"which an allowlist can override",
|
||||
)
|
||||
|
||||
assertDialRefused(t, defaultGuard, target)
|
||||
|
||||
listed := delivery.NewTestGuard(
|
||||
netip.MustParsePrefix(wireServerIP + "/32"),
|
||||
)
|
||||
|
||||
assert.NoError(t,
|
||||
listed.ValidateTargetURL(context.Background(), target),
|
||||
"an operator who lists WireServer must be able to reach it",
|
||||
)
|
||||
}
|
||||
|
||||
// TestGuardCheckIP_BothPathsShareOneDecision asserts that the
|
||||
// validator and the dialer are not two policies that happen to
|
||||
// agree: both are defined in terms of checkIP, so the exported
|
||||
|
||||
@@ -58,12 +58,17 @@ func (t *databaseTarget) Deliver(
|
||||
"error", err,
|
||||
)
|
||||
|
||||
t.eng.recordResult(
|
||||
recErr := t.eng.recordResult(
|
||||
webhookDB, d, 1, false, 0, "",
|
||||
err.Error(), elapsed.Milliseconds(),
|
||||
)
|
||||
if recErr != nil {
|
||||
t.eng.bookkeepingFailed(d, recErr)
|
||||
|
||||
t.eng.updateDeliveryStatus(
|
||||
return
|
||||
}
|
||||
|
||||
t.eng.settleStatus(
|
||||
webhookDB, d, d.Target.Type,
|
||||
database.DeliveryStatusFailed,
|
||||
)
|
||||
@@ -71,12 +76,17 @@ func (t *databaseTarget) Deliver(
|
||||
return
|
||||
}
|
||||
|
||||
t.eng.recordResult(
|
||||
recErr := t.eng.recordResult(
|
||||
webhookDB, d, 1, true, 0, "", "",
|
||||
elapsed.Milliseconds(),
|
||||
)
|
||||
if recErr != nil {
|
||||
t.eng.bookkeepingFailed(d, recErr)
|
||||
|
||||
t.eng.updateDeliveryStatus(
|
||||
return
|
||||
}
|
||||
|
||||
t.eng.settleStatus(
|
||||
webhookDB, d, d.Target.Type,
|
||||
database.DeliveryStatusDelivered,
|
||||
)
|
||||
@@ -267,6 +277,24 @@ func (t *databaseTarget) evict(webhookID string) {
|
||||
)
|
||||
}
|
||||
|
||||
// evictAll evicts every cached archive writer, exactly as evict
|
||||
// does for one webhook. The engine calls it at shutdown, once its
|
||||
// workers have returned. Closing the last handle on an archive
|
||||
// moves the contents of its -wal into the .db and removes the
|
||||
// -wal, so a clean stop leaves each archive as a single file.
|
||||
func (t *databaseTarget) evictAll() {
|
||||
t.mu.Lock()
|
||||
|
||||
writers := t.writers
|
||||
t.writers = nil
|
||||
|
||||
t.mu.Unlock()
|
||||
|
||||
for _, w := range writers {
|
||||
w.evict()
|
||||
}
|
||||
}
|
||||
|
||||
// sweepWebhook prunes one webhook's archive of rows older than
|
||||
// expiry, without requiring a write. It returns nil (nothing to
|
||||
// do) when the archive file does not exist, so a sweep never
|
||||
|
||||
@@ -1,7 +1,6 @@
|
||||
package delivery
|
||||
|
||||
import (
|
||||
"database/sql"
|
||||
"encoding/json"
|
||||
"errors"
|
||||
"fmt"
|
||||
@@ -12,6 +11,7 @@ import (
|
||||
|
||||
"gorm.io/driver/sqlite"
|
||||
"gorm.io/gorm"
|
||||
"sneak.berlin/go/webhooker/internal/database"
|
||||
"sneak.berlin/go/webhooker/internal/gormlog"
|
||||
)
|
||||
|
||||
@@ -30,13 +30,13 @@ const (
|
||||
// path: open the archive file, creating it if missing, so a
|
||||
// first write (or a write after the operator moved the file
|
||||
// away) recreates it.
|
||||
archiveModeCreate = "rwc"
|
||||
archiveModeCreate = database.SQLiteModeCreate
|
||||
|
||||
// archiveModeExisting is the SQLite URI mode used by the idle
|
||||
// sweep: open read-write but never create. A sweep must never
|
||||
// conjure an empty archive file for a webhook that has a
|
||||
// database target but has never received an event.
|
||||
archiveModeExisting = "rw"
|
||||
archiveModeExisting = database.SQLiteModeExisting
|
||||
)
|
||||
|
||||
var (
|
||||
@@ -273,9 +273,11 @@ func (w *archiveWriter) open(expiry time.Duration) error {
|
||||
func (w *archiveWriter) openMode(
|
||||
mode string, expiry time.Duration,
|
||||
) error {
|
||||
dbURL := fmt.Sprintf("file:%s?mode=%s", w.path, mode)
|
||||
|
||||
sqlDB, err := sql.Open("sqlite", dbURL)
|
||||
// Opened through database.OpenSQLite so an archive file carries
|
||||
// the same WAL journaling, busy timeout, immediate-transaction
|
||||
// locking, and pool bounds as every other database file. See
|
||||
// internal/database/sqlite_open.go.
|
||||
sqlDB, err := database.OpenSQLite(w.path, mode)
|
||||
if err != nil {
|
||||
return fmt.Errorf(
|
||||
"opening archive database %s: %w", w.path, err,
|
||||
|
||||
@@ -1,6 +1,7 @@
|
||||
package delivery_test
|
||||
|
||||
import (
|
||||
"context"
|
||||
"errors"
|
||||
"fmt"
|
||||
"net/http"
|
||||
@@ -361,3 +362,46 @@ func TestEvictWebhook_LaterDeliveryRecreatesWriter(t *testing.T) {
|
||||
"a later delivery should recreate the writer",
|
||||
)
|
||||
}
|
||||
|
||||
// TestEngineStop_WriteAfterStopIsRefused proves the engine's stop
|
||||
// closes each archive writer the way deleting its webhook does: a
|
||||
// write that reaches a writer after the stop is refused, reopens
|
||||
// nothing and adds no row.
|
||||
func TestEngineStop_WriteAfterStopIsRefused(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
eng, _ := evictTestEngine(t)
|
||||
|
||||
webhookDB := testWebhookDB(t)
|
||||
event := seedEvent(t, webhookDB, `{"archived":true}`)
|
||||
d := seedDatabaseTargetDelivery(t, webhookDB, event, "")
|
||||
|
||||
eng.ExportDeliverDatabase(webhookDB, d)
|
||||
|
||||
w := eng.ExportArchiveWriterFor(event.WebhookID)
|
||||
require.NotNil(t, w)
|
||||
require.True(t, w.HandleOpen())
|
||||
|
||||
require.NoError(t, eng.ExportStop(context.Background()))
|
||||
|
||||
err := w.Write(evictTestRow("ev-after-stop"), 0)
|
||||
|
||||
require.ErrorIs(
|
||||
t, err, delivery.ErrExportArchiveWriterEvicted,
|
||||
"a write after the stop must be refused",
|
||||
)
|
||||
assert.False(
|
||||
t, w.HandleOpen(),
|
||||
"a refused write must not reopen the archive",
|
||||
)
|
||||
assert.False(
|
||||
t, eng.ExportHasArchiveWriter(event.WebhookID),
|
||||
"the stop should empty the registry",
|
||||
)
|
||||
|
||||
count, err := countArchivedRows(w.Path())
|
||||
require.NoError(t, err)
|
||||
assert.Equal(
|
||||
t, int64(1), count, "the refused row must not be written",
|
||||
)
|
||||
}
|
||||
|
||||
@@ -11,10 +11,11 @@ import (
|
||||
"sneak.berlin/go/webhooker/internal/delivery"
|
||||
)
|
||||
|
||||
// Literals these tests repeat, named so that the header name and the
|
||||
// Literals these tests repeat, named so that the header names and the
|
||||
// keep-forever archive config each have one definition.
|
||||
const (
|
||||
headerAuthorization = "Authorization"
|
||||
headerContentType = "Content-Type"
|
||||
bearerValue = "Bearer abc"
|
||||
archiveConfigNever = "{\"expiry\":\"never\"}"
|
||||
)
|
||||
|
||||
@@ -77,14 +77,19 @@ func (c *httpCore) fireAndForget(
|
||||
) {
|
||||
c.eng.observeAttempt(d.Target.Type, res.elapsed())
|
||||
|
||||
c.eng.recordResult(
|
||||
err := c.eng.recordResult(
|
||||
webhookDB, d, 1, res.success,
|
||||
res.statusCode, res.respBody, res.errMsg,
|
||||
res.duration,
|
||||
)
|
||||
if err != nil {
|
||||
c.eng.bookkeepingFailed(d, err)
|
||||
|
||||
return
|
||||
}
|
||||
|
||||
if res.success {
|
||||
c.eng.updateDeliveryStatus(
|
||||
c.eng.settleStatus(
|
||||
webhookDB, d, d.Target.Type,
|
||||
database.DeliveryStatusDelivered,
|
||||
)
|
||||
@@ -92,7 +97,7 @@ func (c *httpCore) fireAndForget(
|
||||
return
|
||||
}
|
||||
|
||||
c.eng.updateDeliveryStatus(
|
||||
c.eng.settleStatus(
|
||||
webhookDB, d, d.Target.Type,
|
||||
database.DeliveryStatusFailed,
|
||||
)
|
||||
@@ -122,16 +127,25 @@ func (c *httpCore) withRetry(
|
||||
|
||||
c.eng.observeAttempt(d.Target.Type, res.elapsed())
|
||||
|
||||
c.eng.recordResult(
|
||||
err := c.eng.recordResult(
|
||||
webhookDB, d, attemptNum, res.success,
|
||||
res.statusCode, res.respBody, res.errMsg,
|
||||
res.duration,
|
||||
)
|
||||
if err != nil {
|
||||
// The breaker still learns the outcome: it describes the
|
||||
// target's health, which is unaffected by this database's.
|
||||
c.recordCircuitOutcome(cb, res.success)
|
||||
|
||||
c.eng.bookkeepingFailed(d, err)
|
||||
|
||||
return
|
||||
}
|
||||
|
||||
if res.success {
|
||||
cb.RecordSuccess()
|
||||
|
||||
c.eng.updateDeliveryStatus(
|
||||
c.eng.settleStatus(
|
||||
webhookDB, d, d.Target.Type,
|
||||
database.DeliveryStatusDelivered,
|
||||
)
|
||||
@@ -146,6 +160,20 @@ func (c *httpCore) withRetry(
|
||||
)
|
||||
}
|
||||
|
||||
// recordCircuitOutcome feeds one attempt's outcome to the target's
|
||||
// circuit breaker.
|
||||
func (c *httpCore) recordCircuitOutcome(
|
||||
cb *CircuitBreaker, success bool,
|
||||
) {
|
||||
if success {
|
||||
cb.RecordSuccess()
|
||||
|
||||
return
|
||||
}
|
||||
|
||||
cb.RecordFailure()
|
||||
}
|
||||
|
||||
func (c *httpCore) circuitBreakerBlock(
|
||||
webhookDB *gorm.DB,
|
||||
d *database.Delivery,
|
||||
@@ -169,10 +197,14 @@ func (c *httpCore) circuitBreakerBlock(
|
||||
"cooldown_remaining", remaining,
|
||||
)
|
||||
|
||||
c.eng.updateDeliveryStatus(
|
||||
webhookDB, d, d.Target.Type,
|
||||
database.DeliveryStatusRetrying,
|
||||
)
|
||||
// A delivery already at retrying is left as it is, so a task
|
||||
// the breaker keeps turning away writes nothing each time.
|
||||
if d.Status != database.DeliveryStatusRetrying {
|
||||
c.eng.settleStatus(
|
||||
webhookDB, d, d.Target.Type,
|
||||
database.DeliveryStatusRetrying,
|
||||
)
|
||||
}
|
||||
|
||||
retryTask := *task
|
||||
sched.ScheduleRetry(retryTask, remaining)
|
||||
@@ -189,7 +221,7 @@ func (c *httpCore) handleRetry(
|
||||
attemptNum int,
|
||||
) {
|
||||
if attemptNum >= maxRetries {
|
||||
c.eng.updateDeliveryStatus(
|
||||
c.eng.settleStatus(
|
||||
webhookDB, d, d.Target.Type,
|
||||
database.DeliveryStatusFailed,
|
||||
)
|
||||
@@ -197,7 +229,7 @@ func (c *httpCore) handleRetry(
|
||||
return
|
||||
}
|
||||
|
||||
c.eng.updateDeliveryStatus(
|
||||
c.eng.settleStatus(
|
||||
webhookDB, d, d.Target.Type,
|
||||
database.DeliveryStatusRetrying,
|
||||
)
|
||||
@@ -332,12 +364,17 @@ func (t *httpTarget) Deliver(
|
||||
"error", err,
|
||||
)
|
||||
|
||||
t.eng.recordResult(
|
||||
recErr := t.eng.recordResult(
|
||||
webhookDB, d, task.AttemptNum,
|
||||
false, 0, "", err.Error(), 0,
|
||||
)
|
||||
if recErr != nil {
|
||||
t.eng.bookkeepingFailed(d, recErr)
|
||||
|
||||
t.eng.updateDeliveryStatus(
|
||||
return
|
||||
}
|
||||
|
||||
t.eng.settleStatus(
|
||||
webhookDB, d, d.Target.Type,
|
||||
database.DeliveryStatusFailed,
|
||||
)
|
||||
@@ -504,6 +541,11 @@ func isForwardableHeader(name string) bool {
|
||||
"Upgrade", "Proxy-Authorization",
|
||||
"Proxy-Connection", "Content-Length":
|
||||
return false
|
||||
case "Content-Type":
|
||||
// applyRequestHeaders sets Content-Type itself. The receiver
|
||||
// already stored this inbound value as the event's
|
||||
// ContentType, so forwarding it too would send it twice.
|
||||
return false
|
||||
default:
|
||||
return true
|
||||
}
|
||||
@@ -516,6 +558,10 @@ func isForwardableHeader(name string) bool {
|
||||
// policy strips exactly that set on a hop that leaves the origin,
|
||||
// so the forward set is decided here and only here — a header added
|
||||
// to it is covered off-origin without a second edit elsewhere.
|
||||
//
|
||||
// Content-Type goes out once: a Content-Type configured on the target
|
||||
// wins, otherwise the event's ContentType, otherwise none. The inbound
|
||||
// Content-Type in the event's headers is never forwarded.
|
||||
func applyRequestHeaders(
|
||||
req *http.Request,
|
||||
event *database.Event,
|
||||
@@ -536,10 +582,10 @@ func applyRequestHeaders(
|
||||
|
||||
req.Header.Set("User-Agent", "webhooker/1.0")
|
||||
|
||||
// Content-Type describes the body being sent rather than the
|
||||
// sender, and the delivery path sets it from the event itself.
|
||||
// A 307/308 preserves the body across hosts, so stripping it
|
||||
// would send that body untyped.
|
||||
// A Content-Type configured on the target describes the body
|
||||
// being sent rather than the sender. A 307/308 preserves the
|
||||
// body across hosts, so stripping it would send that body
|
||||
// untyped.
|
||||
delete(originScoped, "Content-Type")
|
||||
|
||||
// User-Agent is overwritten just above, so an inbound one never
|
||||
|
||||
@@ -55,12 +55,17 @@ func (t *logTarget) Deliver(
|
||||
|
||||
t.eng.observeAttempt(d.Target.Type, elapsed)
|
||||
|
||||
t.eng.recordResult(
|
||||
err := t.eng.recordResult(
|
||||
webhookDB, d, 1, true, 0, "", "",
|
||||
elapsed.Milliseconds(),
|
||||
)
|
||||
if err != nil {
|
||||
t.eng.bookkeepingFailed(d, err)
|
||||
|
||||
t.eng.updateDeliveryStatus(
|
||||
return
|
||||
}
|
||||
|
||||
t.eng.settleStatus(
|
||||
webhookDB, d, d.Target.Type,
|
||||
database.DeliveryStatusDelivered,
|
||||
)
|
||||
|
||||
@@ -95,12 +95,17 @@ func (t *slackTarget) failConfig(
|
||||
d *database.Delivery,
|
||||
err error,
|
||||
) {
|
||||
t.eng.recordResult(
|
||||
recErr := t.eng.recordResult(
|
||||
webhookDB, d, 1,
|
||||
false, 0, "", err.Error(), 0,
|
||||
)
|
||||
if recErr != nil {
|
||||
t.eng.bookkeepingFailed(d, recErr)
|
||||
|
||||
t.eng.updateDeliveryStatus(
|
||||
return
|
||||
}
|
||||
|
||||
t.eng.settleStatus(
|
||||
webhookDB, d, d.Target.Type,
|
||||
database.DeliveryStatusFailed,
|
||||
)
|
||||
@@ -226,10 +231,15 @@ func FormatSlackMessage(
|
||||
event.ContentType,
|
||||
)
|
||||
|
||||
timestamp := "unknown"
|
||||
if !event.CreatedAt.IsZero() {
|
||||
timestamp = event.CreatedAt.UTC().Format(time.RFC3339)
|
||||
}
|
||||
|
||||
fmt.Fprintf(
|
||||
&b,
|
||||
"*Timestamp:* `%s`\n",
|
||||
event.CreatedAt.UTC().Format(time.RFC3339),
|
||||
timestamp,
|
||||
)
|
||||
|
||||
fmt.Fprintf(
|
||||
|
||||
@@ -0,0 +1,792 @@
|
||||
package delivery_test
|
||||
|
||||
import (
|
||||
"context"
|
||||
"net/http"
|
||||
"net/http/httptest"
|
||||
"sync/atomic"
|
||||
"testing"
|
||||
|
||||
"github.com/google/uuid"
|
||||
"github.com/stretchr/testify/assert"
|
||||
"github.com/stretchr/testify/require"
|
||||
"sneak.berlin/go/webhooker/internal/database"
|
||||
"sneak.berlin/go/webhooker/internal/delivery"
|
||||
)
|
||||
|
||||
// The two terminal-state gaps of
|
||||
// https://git.eeqj.de/sneak/webhooker/issues/107: a delivery failed
|
||||
// with nothing in its event log to say why, and a retrying delivery
|
||||
// whose target was deleted, which used to keep sending and then never
|
||||
// terminalise. Section 4 is the same deleted-target gap for a pending
|
||||
// delivery: https://git.eeqj.de/sneak/webhooker/issues/293.
|
||||
|
||||
// tUnknownType is a target type no build implements. It stands in for
|
||||
// a target whose type was written by a build that knew a type this one
|
||||
// does not.
|
||||
const tUnknownType = database.TargetType("pubsub")
|
||||
|
||||
// tSeedDeletedTarget creates a target, a delivery against it at the
|
||||
// given status with one recorded failed attempt, and then deletes the
|
||||
// target the way the source page does.
|
||||
//
|
||||
// It asserts the delete is soft, because that is the whole reason the
|
||||
// engine could not tell a deleted target from a target id that never
|
||||
// named a row: the surviving row is invisible to a scoped read.
|
||||
func tSeedDeletedTarget(
|
||||
t *testing.T,
|
||||
s iSetup,
|
||||
name, url string,
|
||||
status database.DeliveryStatus,
|
||||
) string {
|
||||
t.Helper()
|
||||
|
||||
targetID := uuid.New().String()
|
||||
|
||||
iCreateTarget(
|
||||
t, s.MainDB, targetID, s.WebhookID, name,
|
||||
database.TargetTypeHTTP, iHTTPConfig(url), 5,
|
||||
)
|
||||
|
||||
event := iSeedEvent(
|
||||
t, s.WebhookDB, s.WebhookID, `{"target":"deleted"}`,
|
||||
)
|
||||
|
||||
d := iSeedDelivery(
|
||||
t, s.WebhookDB, event.ID, targetID, status,
|
||||
)
|
||||
|
||||
iSeedFailedResult(t, s.WebhookDB, d.ID)
|
||||
|
||||
require.NoError(t, s.MainDB.Delete(
|
||||
&database.Target{}, "id = ?", targetID,
|
||||
).Error)
|
||||
|
||||
var scoped, unscoped int64
|
||||
|
||||
require.NoError(t, s.MainDB.
|
||||
Model(&database.Target{}).
|
||||
Where("id = ?", targetID).
|
||||
Count(&scoped).Error)
|
||||
|
||||
require.NoError(t, s.MainDB.Unscoped().
|
||||
Model(&database.Target{}).
|
||||
Where("id = ?", targetID).
|
||||
Count(&unscoped).Error)
|
||||
|
||||
require.Zero(t, scoped,
|
||||
"the deleted target is still visible to a scoped read",
|
||||
)
|
||||
require.Equal(t, int64(1), unscoped,
|
||||
"the delete was hard, so this test proves nothing about "+
|
||||
"the soft-delete case it exists for",
|
||||
)
|
||||
|
||||
return d.ID
|
||||
}
|
||||
|
||||
// tLastResult returns a delivery's final recorded attempt, asserting
|
||||
// the expected number of them.
|
||||
func tLastResult(
|
||||
t *testing.T,
|
||||
s iSetup,
|
||||
deliveryID string,
|
||||
want int,
|
||||
) database.DeliveryResult {
|
||||
t.Helper()
|
||||
|
||||
results := iResults(t, s.WebhookDB, deliveryID)
|
||||
require.Len(t, results, want)
|
||||
|
||||
return results[want-1]
|
||||
}
|
||||
|
||||
// --- 1. A failure with nothing recorded ---
|
||||
|
||||
func TestProcessDelivery_UnknownTargetType_RecordsWhy(
|
||||
t *testing.T,
|
||||
) {
|
||||
t.Parallel()
|
||||
|
||||
s := newISetup(t)
|
||||
|
||||
targetID := uuid.New().String()
|
||||
|
||||
event := iSeedEvent(
|
||||
t, s.WebhookDB, s.WebhookID, `{"unknown":"type"}`,
|
||||
)
|
||||
|
||||
seeded := iSeedDelivery(
|
||||
t, s.WebhookDB, event.ID, targetID,
|
||||
database.DeliveryStatusPending,
|
||||
)
|
||||
|
||||
target := database.Target{
|
||||
Name: "mystery",
|
||||
Type: tUnknownType,
|
||||
Config: iHTTPConfig("http://example.com/hook"),
|
||||
}
|
||||
target.ID = targetID
|
||||
|
||||
d := database.Delivery{
|
||||
EventID: event.ID,
|
||||
TargetID: targetID,
|
||||
Status: database.DeliveryStatusPending,
|
||||
Event: event,
|
||||
Target: target,
|
||||
}
|
||||
d.ID = seeded.ID
|
||||
|
||||
body := event.Body
|
||||
task := iTask(
|
||||
seeded, event, s.WebhookID, targetID, "mystery",
|
||||
target.Config, 0, 1, &body,
|
||||
)
|
||||
task.TargetType = tUnknownType
|
||||
|
||||
s.Engine.ExportProcessDelivery(
|
||||
context.Background(), s.WebhookDB, &d, &task,
|
||||
)
|
||||
|
||||
iAssertStatus(
|
||||
t, s.WebhookDB, d.ID, database.DeliveryStatusFailed,
|
||||
)
|
||||
|
||||
last := tLastResult(t, s, d.ID, 1)
|
||||
|
||||
assert.False(t, last.Success)
|
||||
assert.Equal(t, 1, last.AttemptNum)
|
||||
assert.Contains(t, last.Error, string(tUnknownType),
|
||||
"the recorded reason does not name the offending type",
|
||||
)
|
||||
}
|
||||
|
||||
// --- 2. A retrying delivery whose target is gone ---
|
||||
|
||||
func TestRecoverSingleRetry_TargetDeleted(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
s := newISetup(t)
|
||||
|
||||
iCreateWebhook(
|
||||
t, s.MainDB, s.WebhookID, "deleted-target-recovery",
|
||||
)
|
||||
|
||||
deliveryID := tSeedDeletedTarget(
|
||||
t, s, "gone-on-recovery", "http://example.com/hook",
|
||||
database.DeliveryStatusRetrying,
|
||||
)
|
||||
|
||||
s.Engine.ExportRecoverWebhookDeliveries(
|
||||
context.Background(), s.WebhookID,
|
||||
)
|
||||
|
||||
iAssertStatus(
|
||||
t, s.WebhookDB, deliveryID,
|
||||
database.DeliveryStatusFailed,
|
||||
)
|
||||
|
||||
last := tLastResult(t, s, deliveryID, 2)
|
||||
|
||||
assert.False(t, last.Success)
|
||||
assert.Equal(t, 2, last.AttemptNum)
|
||||
assert.Contains(t, last.Error, "gone-on-recovery")
|
||||
assert.Contains(t, last.Error, "was deleted")
|
||||
|
||||
assert.Empty(t, s.Engine.ExportRetryCh(),
|
||||
"a delivery whose target is gone was rescheduled",
|
||||
)
|
||||
assert.Zero(t, s.Engine.ExportInflightHeld(),
|
||||
"the terminal path leaked its ownership reference",
|
||||
)
|
||||
}
|
||||
|
||||
func TestSweepSingleRetry_TargetDeleted(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
s := newISetup(t)
|
||||
|
||||
iCreateWebhook(
|
||||
t, s.MainDB, s.WebhookID, "deleted-target-sweep",
|
||||
)
|
||||
|
||||
deliveryID := tSeedDeletedTarget(
|
||||
t, s, "gone-on-sweep", "http://example.com/hook",
|
||||
database.DeliveryStatusRetrying,
|
||||
)
|
||||
|
||||
// Twice, because the bug was an error the sweep repeated every
|
||||
// minute for the life of the database: the second sweep must
|
||||
// find nothing left to do.
|
||||
s.Engine.ExportSweepWebhookRetries(
|
||||
context.Background(), s.WebhookID,
|
||||
)
|
||||
s.Engine.ExportSweepWebhookRetries(
|
||||
context.Background(), s.WebhookID,
|
||||
)
|
||||
|
||||
iAssertStatus(
|
||||
t, s.WebhookDB, deliveryID,
|
||||
database.DeliveryStatusFailed,
|
||||
)
|
||||
|
||||
last := tLastResult(t, s, deliveryID, 2)
|
||||
|
||||
assert.Contains(t, last.Error, "gone-on-sweep")
|
||||
assert.Contains(t, last.Error, "was deleted")
|
||||
|
||||
assert.Empty(t, s.Engine.ExportRetryCh())
|
||||
assert.Zero(t, s.Engine.ExportInflightHeld())
|
||||
}
|
||||
|
||||
// TestSweepSingleRetry_TargetNeverExisted covers the other half of the
|
||||
// soft-delete distinction: an id with no row at all, deleted or
|
||||
// otherwise, must not be reported as something the operator deleted.
|
||||
func TestSweepSingleRetry_TargetNeverExisted(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
s := newISetup(t)
|
||||
|
||||
iCreateWebhook(
|
||||
t, s.MainDB, s.WebhookID, "target-never-existed",
|
||||
)
|
||||
|
||||
targetID := uuid.New().String()
|
||||
|
||||
event := iSeedEvent(
|
||||
t, s.WebhookDB, s.WebhookID, `{"target":"absent"}`,
|
||||
)
|
||||
|
||||
d := iSeedDelivery(
|
||||
t, s.WebhookDB, event.ID, targetID,
|
||||
database.DeliveryStatusRetrying,
|
||||
)
|
||||
|
||||
iSeedFailedResult(t, s.WebhookDB, d.ID)
|
||||
|
||||
s.Engine.ExportSweepWebhookRetries(
|
||||
context.Background(), s.WebhookID,
|
||||
)
|
||||
|
||||
iAssertStatus(
|
||||
t, s.WebhookDB, d.ID, database.DeliveryStatusFailed,
|
||||
)
|
||||
|
||||
last := tLastResult(t, s, d.ID, 2)
|
||||
|
||||
assert.Contains(t, last.Error, targetID)
|
||||
assert.Contains(t, last.Error, "no longer exists")
|
||||
assert.NotContains(t, last.Error, "was deleted",
|
||||
"an id that never named a row was reported as a deletion",
|
||||
)
|
||||
}
|
||||
|
||||
// TestFailMissingTarget_WritesNoTargetRow holds the new terminal path
|
||||
// to the same rule as the existing one: no target row, and so no
|
||||
// plaintext target config, may be written into the per-webhook event
|
||||
// database. See https://git.eeqj.de/sneak/webhooker/issues/206.
|
||||
func TestFailMissingTarget_WritesNoTargetRow(
|
||||
t *testing.T,
|
||||
) {
|
||||
t.Parallel()
|
||||
|
||||
s := newISetup(t)
|
||||
|
||||
iCreateWebhook(
|
||||
t, s.MainDB, s.WebhookID, "no-target-row-deleted",
|
||||
)
|
||||
|
||||
hookURL := "https://hooks.slack.com/services/T00/B00/x"
|
||||
|
||||
deliveryID := tSeedDeletedTarget(
|
||||
t, s, "credential-bearing", hookURL,
|
||||
database.DeliveryStatusRetrying,
|
||||
)
|
||||
|
||||
s.Engine.ExportSweepWebhookRetries(
|
||||
context.Background(), s.WebhookID,
|
||||
)
|
||||
|
||||
iAssertStatus(
|
||||
t, s.WebhookDB, deliveryID,
|
||||
database.DeliveryStatusFailed,
|
||||
)
|
||||
|
||||
var configs []string
|
||||
|
||||
require.NoError(t, s.WebhookDB.
|
||||
Table("targets").
|
||||
Pluck("config", &configs).Error)
|
||||
|
||||
assert.Empty(t, configs,
|
||||
"the deleted-target terminal path wrote a target row "+
|
||||
"into the per-webhook event database",
|
||||
)
|
||||
}
|
||||
|
||||
// --- 3. The scheduled retry chain ---
|
||||
|
||||
// tRetryChainSetup wires a counting sink and a retrying delivery
|
||||
// against a live target pointing at it, and returns the task a
|
||||
// scheduled retry would carry — config and all, snapshotted as
|
||||
// ScheduleRetry snapshots it.
|
||||
func tRetryChainSetup(
|
||||
t *testing.T,
|
||||
s iSetup,
|
||||
name string,
|
||||
hits *atomic.Int64,
|
||||
) (delivery.Task, string) {
|
||||
t.Helper()
|
||||
|
||||
ts := httptest.NewServer(http.HandlerFunc(
|
||||
func(w http.ResponseWriter, _ *http.Request) {
|
||||
hits.Add(1)
|
||||
w.WriteHeader(http.StatusOK)
|
||||
},
|
||||
))
|
||||
t.Cleanup(ts.Close)
|
||||
|
||||
iCreateWebhook(t, s.MainDB, s.WebhookID, name)
|
||||
|
||||
targetID := uuid.New().String()
|
||||
cfg := iHTTPConfig(ts.URL)
|
||||
|
||||
iCreateTarget(
|
||||
t, s.MainDB, targetID, s.WebhookID, name,
|
||||
database.TargetTypeHTTP, cfg, 5,
|
||||
)
|
||||
|
||||
event := iSeedEvent(
|
||||
t, s.WebhookDB, s.WebhookID, `{"chain":"retry"}`,
|
||||
)
|
||||
|
||||
d := iSeedDelivery(
|
||||
t, s.WebhookDB, event.ID, targetID,
|
||||
database.DeliveryStatusRetrying,
|
||||
)
|
||||
|
||||
iSeedFailedResult(t, s.WebhookDB, d.ID)
|
||||
|
||||
body := event.Body
|
||||
|
||||
return iTask(
|
||||
d, event, s.WebhookID, targetID, name, cfg, 5, 2, &body,
|
||||
), targetID
|
||||
}
|
||||
|
||||
// TestProcessRetryTask_TargetDeleted_MakesNoAttempt is the half the
|
||||
// deployability audit found worse than filed: terminalising on
|
||||
// recovery and sweep alone leaves the already-scheduled timer chain
|
||||
// running, and it holds the target's configuration from before the
|
||||
// deletion, so it goes on sending to a destination that was removed.
|
||||
func TestProcessRetryTask_TargetDeleted_MakesNoAttempt(
|
||||
t *testing.T,
|
||||
) {
|
||||
t.Parallel()
|
||||
|
||||
s := newISetup(t)
|
||||
|
||||
var hits atomic.Int64
|
||||
|
||||
task, targetID := tRetryChainSetup(
|
||||
t, s, "gone-mid-chain", &hits,
|
||||
)
|
||||
|
||||
require.NoError(t, s.MainDB.Delete(
|
||||
&database.Target{}, "id = ?", targetID,
|
||||
).Error)
|
||||
|
||||
s.Engine.ExportProcessRetryTask(
|
||||
context.Background(), &task,
|
||||
)
|
||||
|
||||
assert.Zero(t, hits.Load(),
|
||||
"a scheduled retry fired at a target the operator "+
|
||||
"had already deleted",
|
||||
)
|
||||
|
||||
iAssertStatus(
|
||||
t, s.WebhookDB, task.DeliveryID,
|
||||
database.DeliveryStatusFailed,
|
||||
)
|
||||
|
||||
last := tLastResult(t, s, task.DeliveryID, 2)
|
||||
|
||||
assert.False(t, last.Success)
|
||||
assert.Contains(t, last.Error, "was deleted")
|
||||
|
||||
assert.Zero(t, s.Engine.ExportInflightHeld())
|
||||
}
|
||||
|
||||
// TestProcessRetryTask_TargetPresent_StillDelivers is the guard's
|
||||
// mutation check: a liveness check that refused every retry would pass
|
||||
// the test above and break every retry there is.
|
||||
func TestProcessRetryTask_TargetPresent_StillDelivers(
|
||||
t *testing.T,
|
||||
) {
|
||||
t.Parallel()
|
||||
|
||||
s := newISetup(t)
|
||||
|
||||
var hits atomic.Int64
|
||||
|
||||
task, _ := tRetryChainSetup(t, s, "still-there", &hits)
|
||||
|
||||
s.Engine.ExportProcessRetryTask(
|
||||
context.Background(), &task,
|
||||
)
|
||||
|
||||
assert.Equal(t, int64(1), hits.Load())
|
||||
|
||||
iAssertStatus(
|
||||
t, s.WebhookDB, task.DeliveryID,
|
||||
database.DeliveryStatusDelivered,
|
||||
)
|
||||
}
|
||||
|
||||
// TestProcessRetryTask_TargetUnreadable_StillDelivers pins the other
|
||||
// half of the guard: only a target that is confirmed gone stops a
|
||||
// retry. A main database that cannot be read is a transient fault, and
|
||||
// a guard that abandoned deliveries on one would be a worse bug than
|
||||
// the one it fixes.
|
||||
func TestProcessRetryTask_TargetUnreadable_StillDelivers(
|
||||
t *testing.T,
|
||||
) {
|
||||
t.Parallel()
|
||||
|
||||
s := newISetup(t)
|
||||
|
||||
var hits atomic.Int64
|
||||
|
||||
task, _ := tRetryChainSetup(t, s, "unreadable-main", &hits)
|
||||
|
||||
sqlDB, err := s.MainDB.DB()
|
||||
require.NoError(t, err)
|
||||
require.NoError(t, sqlDB.Close())
|
||||
|
||||
s.Engine.ExportProcessRetryTask(
|
||||
context.Background(), &task,
|
||||
)
|
||||
|
||||
assert.Equal(t, int64(1), hits.Load(),
|
||||
"a retry was abandoned because the main database "+
|
||||
"could not be read, not because its target was gone",
|
||||
)
|
||||
|
||||
iAssertStatus(
|
||||
t, s.WebhookDB, task.DeliveryID,
|
||||
database.DeliveryStatusDelivered,
|
||||
)
|
||||
}
|
||||
|
||||
// TestRecoverSingleRetry_TargetUnreadable_LeavesDeliveryAlone is the
|
||||
// same rule on the recovery path. A read failure that is not
|
||||
// "record not found" must leave every retrying delivery of every
|
||||
// webhook exactly as it was.
|
||||
func TestRecoverSingleRetry_TargetUnreadable_LeavesDeliveryAlone(
|
||||
t *testing.T,
|
||||
) {
|
||||
t.Parallel()
|
||||
|
||||
s := newISetup(t)
|
||||
|
||||
iCreateWebhook(
|
||||
t, s.MainDB, s.WebhookID, "unreadable-on-recovery",
|
||||
)
|
||||
|
||||
targetID := uuid.New().String()
|
||||
|
||||
iCreateTarget(
|
||||
t, s.MainDB, targetID, s.WebhookID, "healthy",
|
||||
database.TargetTypeHTTP,
|
||||
iHTTPConfig("http://example.com/hook"), 5,
|
||||
)
|
||||
|
||||
event := iSeedEvent(
|
||||
t, s.WebhookDB, s.WebhookID, `{"still":"retrying"}`,
|
||||
)
|
||||
|
||||
d := iSeedDelivery(
|
||||
t, s.WebhookDB, event.ID, targetID,
|
||||
database.DeliveryStatusRetrying,
|
||||
)
|
||||
|
||||
iSeedFailedResult(t, s.WebhookDB, d.ID)
|
||||
|
||||
sqlDB, err := s.MainDB.DB()
|
||||
require.NoError(t, err)
|
||||
require.NoError(t, sqlDB.Close())
|
||||
|
||||
s.Engine.ExportRecoverRetryingDeliveries(
|
||||
s.WebhookDB, s.WebhookID,
|
||||
)
|
||||
|
||||
iAssertStatus(
|
||||
t, s.WebhookDB, d.ID,
|
||||
database.DeliveryStatusRetrying,
|
||||
)
|
||||
|
||||
assert.Len(t, iResults(t, s.WebhookDB, d.ID), 1,
|
||||
"an unreadable main database produced a terminal "+
|
||||
"failure row",
|
||||
)
|
||||
|
||||
assert.Zero(t, s.Engine.ExportInflightHeld())
|
||||
}
|
||||
|
||||
// --- 4. A pending delivery whose target is gone ---
|
||||
|
||||
func TestRecoverPending_TargetDeleted(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
s := newISetup(t)
|
||||
|
||||
deliveryID := tSeedDeletedTarget(
|
||||
t, s, "gone-while-pending", "http://example.com/hook",
|
||||
database.DeliveryStatusPending,
|
||||
)
|
||||
|
||||
s.Engine.ExportRecoverWebhookDeliveries(
|
||||
context.Background(), s.WebhookID,
|
||||
)
|
||||
|
||||
iAssertStatus(
|
||||
t, s.WebhookDB, deliveryID,
|
||||
database.DeliveryStatusFailed,
|
||||
)
|
||||
|
||||
last := tLastResult(t, s, deliveryID, 2)
|
||||
|
||||
assert.False(t, last.Success)
|
||||
assert.Equal(t, 2, last.AttemptNum)
|
||||
assert.Contains(t, last.Error, "gone-while-pending")
|
||||
assert.Contains(t, last.Error, "was deleted")
|
||||
|
||||
assert.Empty(t, fDrain(s.Engine),
|
||||
"a delivery whose target is gone was sent",
|
||||
)
|
||||
assert.Zero(t, s.Engine.ExportInflightHeld(),
|
||||
"the terminal path leaked its ownership reference",
|
||||
)
|
||||
}
|
||||
|
||||
// TestRecoverPending_TargetDeleted_LeavesAnOwnedDeliveryAlone: the
|
||||
// terminal write takes ownership like every other recovery write, so a
|
||||
// delivery the engine still holds is not failed underneath its worker.
|
||||
func TestRecoverPending_TargetDeleted_LeavesAnOwnedDeliveryAlone(
|
||||
t *testing.T,
|
||||
) {
|
||||
t.Parallel()
|
||||
|
||||
s := newISetup(t)
|
||||
|
||||
deliveryID := tSeedDeletedTarget(
|
||||
t, s, "gone-but-owned", "http://example.com/hook",
|
||||
database.DeliveryStatusPending,
|
||||
)
|
||||
|
||||
require.True(t, s.Engine.ExportRetainDelivery(deliveryID))
|
||||
|
||||
s.Engine.ExportRecoverWebhookDeliveries(
|
||||
context.Background(), s.WebhookID,
|
||||
)
|
||||
|
||||
iAssertStatus(
|
||||
t, s.WebhookDB, deliveryID,
|
||||
database.DeliveryStatusPending,
|
||||
)
|
||||
|
||||
assert.Len(t, iResults(t, s.WebhookDB, deliveryID), 1,
|
||||
"a delivery the engine owns was failed underneath it",
|
||||
)
|
||||
}
|
||||
|
||||
// TestFailMissingTarget_LeavesASettledDeliveryAlone: the recovery paths
|
||||
// read their batch before taking ownership, and a worker may send a
|
||||
// delivery and let it go in between. The terminal write goes by the row
|
||||
// as it is now, not as the batch read it.
|
||||
func TestFailMissingTarget_LeavesASettledDeliveryAlone(
|
||||
t *testing.T,
|
||||
) {
|
||||
t.Parallel()
|
||||
|
||||
s := newISetup(t)
|
||||
|
||||
deliveryID := tSeedDeletedTarget(
|
||||
t, s, "gone-after-sending", "http://example.com/hook",
|
||||
database.DeliveryStatusPending,
|
||||
)
|
||||
|
||||
var batch database.Delivery
|
||||
|
||||
require.NoError(t, s.WebhookDB.First(
|
||||
&batch, "id = ?", deliveryID,
|
||||
).Error)
|
||||
|
||||
// A worker settles the delivery after the batch was read.
|
||||
require.NoError(t, s.WebhookDB.Model(&database.Delivery{}).
|
||||
Where("id = ?", deliveryID).
|
||||
Update("status", database.DeliveryStatusDelivered).Error)
|
||||
|
||||
s.Engine.ExportFailMissingTarget(
|
||||
s.WebhookDB, s.WebhookID, &batch,
|
||||
)
|
||||
|
||||
iAssertStatus(
|
||||
t, s.WebhookDB, deliveryID,
|
||||
database.DeliveryStatusDelivered,
|
||||
)
|
||||
|
||||
assert.Len(t, iResults(t, s.WebhookDB, deliveryID), 1,
|
||||
"a delivery settled after the batch read was then failed",
|
||||
)
|
||||
assert.Zero(t, s.Engine.ExportInflightHeld())
|
||||
}
|
||||
|
||||
// TestSweepPending_TargetDeleted sweeps twice over a batch that also
|
||||
// holds a healthy stranded delivery. The one whose target is gone is
|
||||
// failed once and then left alone; the healthy one is queued by the
|
||||
// first sweep and not again by the second.
|
||||
func TestSweepPending_TargetDeleted(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
liveTargetID := uuid.New().String()
|
||||
s := fSweepSetup(t, liveTargetID, "still-there")
|
||||
|
||||
deliveryID := tSeedDeletedTarget(
|
||||
t, s, "gone-on-pending-sweep", "http://example.com/hook",
|
||||
database.DeliveryStatusPending,
|
||||
)
|
||||
rAgePending(t, s.WebhookDB, deliveryID)
|
||||
|
||||
event := iSeedEvent(
|
||||
t, s.WebhookDB, s.WebhookID, `{"target":"live"}`,
|
||||
)
|
||||
|
||||
healthy := iSeedDelivery(
|
||||
t, s.WebhookDB, event.ID, liveTargetID,
|
||||
database.DeliveryStatusPending,
|
||||
)
|
||||
rAgePending(t, s.WebhookDB, healthy.ID)
|
||||
|
||||
ctx := context.Background()
|
||||
|
||||
s.Engine.ExportSweepWebhookRetries(ctx, s.WebhookID)
|
||||
|
||||
tasks := fDrain(s.Engine)
|
||||
require.Len(t, tasks, 1,
|
||||
"the first sweep did not queue the healthy delivery",
|
||||
)
|
||||
assert.Equal(t, healthy.ID, tasks[0].DeliveryID)
|
||||
|
||||
s.Engine.ExportSweepWebhookRetries(ctx, s.WebhookID)
|
||||
|
||||
assert.Empty(t, fDrain(s.Engine),
|
||||
"the second sweep queued a delivery again",
|
||||
)
|
||||
|
||||
iAssertStatus(
|
||||
t, s.WebhookDB, deliveryID,
|
||||
database.DeliveryStatusFailed,
|
||||
)
|
||||
|
||||
last := tLastResult(t, s, deliveryID, 2)
|
||||
|
||||
assert.Contains(t, last.Error, "gone-on-pending-sweep")
|
||||
assert.Contains(t, last.Error, "was deleted")
|
||||
}
|
||||
|
||||
// TestSendRecoveredDeliveries_TargetMissingFromMap: the batch's target
|
||||
// map is empty when its query failed, so every delivery in the batch is
|
||||
// looked up on its own. A healthy one is sent to the target that lookup
|
||||
// finds.
|
||||
func TestSendRecoveredDeliveries_TargetMissingFromMap(
|
||||
t *testing.T,
|
||||
) {
|
||||
t.Parallel()
|
||||
|
||||
s := newISetup(t)
|
||||
|
||||
targetID := uuid.New().String()
|
||||
|
||||
iCreateTarget(
|
||||
t, s.MainDB, targetID, s.WebhookID, "found-on-lookup",
|
||||
database.TargetTypeLog, "", 0,
|
||||
)
|
||||
|
||||
event := iSeedEvent(
|
||||
t, s.WebhookDB, s.WebhookID, `{"map":"empty"}`,
|
||||
)
|
||||
|
||||
d := iSeedDelivery(
|
||||
t, s.WebhookDB, event.ID, targetID,
|
||||
database.DeliveryStatusPending,
|
||||
)
|
||||
|
||||
s.Engine.ExportSendRecoveredDeliveries(
|
||||
context.Background(), s.WebhookDB,
|
||||
[]database.Delivery{d}, s.WebhookID,
|
||||
map[string]database.Target{}, nil,
|
||||
)
|
||||
|
||||
tasks := fDrain(s.Engine)
|
||||
require.Len(t, tasks, 1,
|
||||
"the healthy delivery was not queued exactly once",
|
||||
)
|
||||
assert.Equal(t, d.ID, tasks[0].DeliveryID)
|
||||
assert.Equal(t, targetID, tasks[0].TargetID)
|
||||
assert.Equal(t, database.TargetTypeLog, tasks[0].TargetType)
|
||||
|
||||
iAssertStatus(
|
||||
t, s.WebhookDB, d.ID,
|
||||
database.DeliveryStatusPending,
|
||||
)
|
||||
}
|
||||
|
||||
// TestRecoverPending_TargetUnreadable_LeavesDeliveryAlone: a failed
|
||||
// read of the main database is not a deleted target. Restart recovery
|
||||
// holds every pending delivery of the webhook in one batch, so failing
|
||||
// on this would fail all of them.
|
||||
func TestRecoverPending_TargetUnreadable_LeavesDeliveryAlone(
|
||||
t *testing.T,
|
||||
) {
|
||||
t.Parallel()
|
||||
|
||||
s := newISetup(t)
|
||||
|
||||
targetID := uuid.New().String()
|
||||
|
||||
iCreateTarget(
|
||||
t, s.MainDB, targetID, s.WebhookID, "healthy",
|
||||
database.TargetTypeLog, "", 0,
|
||||
)
|
||||
|
||||
event := iSeedEvent(
|
||||
t, s.WebhookDB, s.WebhookID, `{"still":"pending"}`,
|
||||
)
|
||||
|
||||
d := iSeedDelivery(
|
||||
t, s.WebhookDB, event.ID, targetID,
|
||||
database.DeliveryStatusPending,
|
||||
)
|
||||
|
||||
sqlDB, err := s.MainDB.DB()
|
||||
require.NoError(t, err)
|
||||
require.NoError(t, sqlDB.Close())
|
||||
|
||||
s.Engine.ExportRecoverPendingDeliveries(
|
||||
context.Background(), s.WebhookDB, s.WebhookID,
|
||||
)
|
||||
|
||||
iAssertStatus(
|
||||
t, s.WebhookDB, d.ID,
|
||||
database.DeliveryStatusPending,
|
||||
)
|
||||
|
||||
assert.Empty(t, iResults(t, s.WebhookDB, d.ID),
|
||||
"an unreadable main database produced a terminal "+
|
||||
"failure row",
|
||||
)
|
||||
assert.Empty(t, fDrain(s.Engine))
|
||||
assert.Zero(t, s.Engine.ExportInflightHeld())
|
||||
}
|
||||
@@ -92,11 +92,10 @@ func (h *Handlers) HandleEventBodyDownload() http.HandlerFunc {
|
||||
// once per range.
|
||||
//
|
||||
// One consequence is worth keeping in view: the read finishes
|
||||
// before the client is written to, so no read lock is held for
|
||||
// the length of a slow download. These per-webhook databases
|
||||
// run in SQLite's default journal mode rather than WAL, so a
|
||||
// lock held that long would block the receiver from recording
|
||||
// new events.
|
||||
// before the client is written to, so nothing is held open for
|
||||
// the length of a slow download. Under WAL a read no longer
|
||||
// blocks the receiver, but it does pin the WAL against
|
||||
// checkpointing, and a download can last minutes.
|
||||
func (h *Handlers) serveEventBody(
|
||||
w http.ResponseWriter,
|
||||
r *http.Request,
|
||||
|
||||
@@ -139,10 +139,10 @@ func (h *Handlers) resubmitEvent(
|
||||
}
|
||||
|
||||
// Read before the write transaction is opened. The body can be up
|
||||
// to the 1 MB ingest cap, and holding a read of it inside the
|
||||
// transaction would extend how long the per-webhook database is
|
||||
// locked against the receiver, which runs these files in
|
||||
// SQLite's default journal mode rather than WAL.
|
||||
// to the 1 MB ingest cap, and every transaction on these files
|
||||
// takes the write lock at BEGIN (_txlock=immediate, see
|
||||
// internal/database/sqlite_open.go), so reading inside it would
|
||||
// hold that lock against the receiver for the length of the read.
|
||||
src, found, err := loadResubmitSource(
|
||||
webhookDB, webhook.ID, eventID.String(),
|
||||
)
|
||||
|
||||
@@ -0,0 +1,283 @@
|
||||
package handlers_test
|
||||
|
||||
import (
|
||||
"context"
|
||||
"crypto/tls"
|
||||
"net/http"
|
||||
"net/http/httptest"
|
||||
"regexp"
|
||||
"testing"
|
||||
|
||||
"github.com/go-chi/chi"
|
||||
"github.com/stretchr/testify/assert"
|
||||
"github.com/stretchr/testify/require"
|
||||
"sneak.berlin/go/webhooker/internal/database"
|
||||
"sneak.berlin/go/webhooker/internal/handlers"
|
||||
"sneak.berlin/go/webhooker/internal/session"
|
||||
)
|
||||
|
||||
// The only two schemes a rendered entrypoint URL may carry,
|
||||
// whatever the request claimed.
|
||||
const (
|
||||
schemeHTTPS = "https"
|
||||
schemeHTTP = "http"
|
||||
)
|
||||
|
||||
// entrypointURLPattern captures the entrypoint URL the source
|
||||
// detail page renders, which is the operator-visible product of
|
||||
// BaseURL. Asserting on the extracted string rather than on a
|
||||
// substring of the page proves the raw header value cannot reach
|
||||
// the scheme by any route.
|
||||
var entrypointURLPattern = regexp.MustCompile(
|
||||
`<code id="entrypoint-url-[^"]*"[^>]*>([^<]*)</code>`,
|
||||
)
|
||||
|
||||
// baseURLFixture is one started app plus the webhook whose
|
||||
// entrypoint URL the BaseURL cases read.
|
||||
type baseURLFixture struct {
|
||||
handlers *handlers.Handlers
|
||||
session *session.Session
|
||||
webhook string
|
||||
path string
|
||||
}
|
||||
|
||||
// newBaseURLFixture starts the app and seeds a webhook with one
|
||||
// entrypoint.
|
||||
func newBaseURLFixture(t *testing.T) *baseURLFixture {
|
||||
t.Helper()
|
||||
|
||||
var (
|
||||
h *handlers.Handlers
|
||||
sess *session.Session
|
||||
db *database.Database
|
||||
)
|
||||
|
||||
app := newTestApp(t, &h, &sess, &db)
|
||||
app.RequireStart()
|
||||
|
||||
t.Cleanup(app.RequireStop)
|
||||
|
||||
wh := seedWebhook(t, db)
|
||||
seedEntrypoint(t, db, wh.ID)
|
||||
|
||||
return &baseURLFixture{
|
||||
handlers: h,
|
||||
session: sess,
|
||||
webhook: wh.ID,
|
||||
path: "ep-" + wh.ID,
|
||||
}
|
||||
}
|
||||
|
||||
// entrypointURL renders the source detail page for the fixture's
|
||||
// webhook over a request the caller shapes, and returns the
|
||||
// entrypoint URL as an operator would copy it.
|
||||
func (f *baseURLFixture) entrypointURL(
|
||||
t *testing.T,
|
||||
host string,
|
||||
shape func(*http.Request),
|
||||
) string {
|
||||
t.Helper()
|
||||
|
||||
req := httptest.NewRequestWithContext(
|
||||
context.Background(),
|
||||
http.MethodGet,
|
||||
"/source/"+f.webhook,
|
||||
nil,
|
||||
)
|
||||
req.Host = host
|
||||
|
||||
shape(req)
|
||||
|
||||
for _, c := range authenticatedCookies(
|
||||
t, f.session, deleteTestUserID, deleteTestUsername,
|
||||
) {
|
||||
req.AddCookie(c)
|
||||
}
|
||||
|
||||
rctx := chi.NewRouteContext()
|
||||
rctx.URLParams.Add(paramSourceID, f.webhook)
|
||||
req = req.WithContext(
|
||||
context.WithValue(
|
||||
req.Context(), chi.RouteCtxKey, rctx,
|
||||
),
|
||||
)
|
||||
|
||||
w := httptest.NewRecorder()
|
||||
f.handlers.HandleSourceDetail().ServeHTTP(w, req)
|
||||
|
||||
require.Equal(t, http.StatusOK, w.Code)
|
||||
|
||||
match := entrypointURLPattern.FindStringSubmatch(w.Body.String())
|
||||
require.Len(
|
||||
t, match, 2,
|
||||
"the page must render exactly one entrypoint URL",
|
||||
)
|
||||
|
||||
return match[1]
|
||||
}
|
||||
|
||||
// forwardedProto returns a request shaper setting
|
||||
// X-Forwarded-Proto, or leaving the request alone for "".
|
||||
func forwardedProto(value string) func(*http.Request) {
|
||||
return func(r *http.Request) {
|
||||
if value == "" {
|
||||
return
|
||||
}
|
||||
|
||||
r.Header.Set("X-Forwarded-Proto", value)
|
||||
}
|
||||
}
|
||||
|
||||
// baseURLCase is one X-Forwarded-Proto spelling and the scheme
|
||||
// the rendered entrypoint URL owes it.
|
||||
type baseURLCase struct {
|
||||
name string
|
||||
header string
|
||||
scheme string
|
||||
why string
|
||||
}
|
||||
|
||||
// baseURLCases enumerate the spellings a proxy really emits. The
|
||||
// scheme is only ever http or https: the header value itself is
|
||||
// never a scheme, however it is spelled.
|
||||
func baseURLCases() []baseURLCase {
|
||||
return []baseURLCase{
|
||||
{
|
||||
name: "lowercase",
|
||||
header: schemeHTTPS,
|
||||
scheme: schemeHTTPS,
|
||||
why: "the ordinary spelling",
|
||||
},
|
||||
{
|
||||
name: "uppercase",
|
||||
header: "HTTPS",
|
||||
scheme: schemeHTTPS,
|
||||
why: "the token is case-insensitive; the scheme " +
|
||||
"in a copyable URL is not",
|
||||
},
|
||||
{
|
||||
name: "chain with plaintext inner hop",
|
||||
header: "https, http",
|
||||
scheme: schemeHTTPS,
|
||||
why: "a chained proxy appends its hop; the " +
|
||||
"leftmost element faces the client",
|
||||
},
|
||||
{
|
||||
name: "chain of two TLS hops",
|
||||
header: "https,https",
|
||||
scheme: schemeHTTPS,
|
||||
why: "appended chain with no space after the comma",
|
||||
},
|
||||
{
|
||||
name: "trailing space",
|
||||
header: "https ",
|
||||
scheme: schemeHTTPS,
|
||||
why: "whitespace is not part of the token",
|
||||
},
|
||||
{
|
||||
name: "plaintext",
|
||||
header: schemeHTTP,
|
||||
scheme: schemeHTTP,
|
||||
why: "the negative control: the proxy reports plaintext",
|
||||
},
|
||||
{
|
||||
name: "no header",
|
||||
header: "",
|
||||
scheme: schemeHTTP,
|
||||
why: "a plaintext request asserting nothing is http",
|
||||
},
|
||||
{
|
||||
name: "garbage token",
|
||||
header: "javascript:alert(1)//",
|
||||
scheme: schemeHTTP,
|
||||
why: "anything that is not https is not TLS, and " +
|
||||
"the token never becomes the scheme",
|
||||
},
|
||||
}
|
||||
}
|
||||
|
||||
// TestSourceDetailBaseURL_ForwardedProtoSpellings is the
|
||||
// regression test for the entrypoint URL an operator pastes into
|
||||
// the sending system: a header spelling that used to land in the
|
||||
// scheme verbatim produced a URL no sender could deliver to.
|
||||
func TestSourceDetailBaseURL_ForwardedProtoSpellings(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
const host = "hooks.example.com"
|
||||
|
||||
fixture := newBaseURLFixture(t)
|
||||
|
||||
for _, tc := range baseURLCases() {
|
||||
t.Run(tc.name, func(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
assert.Equal(
|
||||
t,
|
||||
tc.scheme+"://"+host+"/webhook/"+fixture.path,
|
||||
fixture.entrypointURL(
|
||||
t, host, forwardedProto(tc.header),
|
||||
),
|
||||
"X-Forwarded-Proto %q: %s", tc.header, tc.why,
|
||||
)
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// TestSourceDetailBaseURL_DirectTLSBeatsPlaintextHeader pins the
|
||||
// precedence the old code had backwards: it let any present
|
||||
// header overwrite what the connection itself proved, so a
|
||||
// direct-TLS request behind a proxy reporting http rendered an
|
||||
// http URL.
|
||||
func TestSourceDetailBaseURL_DirectTLSBeatsPlaintextHeader(
|
||||
t *testing.T,
|
||||
) {
|
||||
t.Parallel()
|
||||
|
||||
const host = "hooks.example.com"
|
||||
|
||||
fixture := newBaseURLFixture(t)
|
||||
|
||||
got := fixture.entrypointURL(t, host, func(r *http.Request) {
|
||||
r.TLS = &tls.ConnectionState{}
|
||||
r.Header.Set("X-Forwarded-Proto", "http")
|
||||
})
|
||||
|
||||
assert.Equal(
|
||||
t,
|
||||
"https://"+host+"/webhook/"+fixture.path,
|
||||
got,
|
||||
"a connection this process terminated with TLS "+
|
||||
"outranks a header claiming plaintext",
|
||||
)
|
||||
}
|
||||
|
||||
// TestSourceDetailBaseURL_KeepsHostAuthority pins the host half
|
||||
// of the URL: it is taken from the request unchanged, so the
|
||||
// deployments that do not sit on port 443 still get a URL that
|
||||
// works. Constraining the host would break exactly these.
|
||||
func TestSourceDetailBaseURL_KeepsHostAuthority(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
fixture := newBaseURLFixture(t)
|
||||
|
||||
hosts := []string{
|
||||
"hooks.example.com:8443",
|
||||
"[2001:db8::1]:8443",
|
||||
"internal-host",
|
||||
}
|
||||
|
||||
for _, host := range hosts {
|
||||
t.Run(host, func(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
assert.Equal(
|
||||
t,
|
||||
"https://"+host+"/webhook/"+fixture.path,
|
||||
fixture.entrypointURL(
|
||||
t, host, forwardedProto("HTTPS"),
|
||||
),
|
||||
"the authority must survive verbatim, port and all",
|
||||
)
|
||||
})
|
||||
}
|
||||
}
|
||||
@@ -13,6 +13,7 @@ import (
|
||||
"gorm.io/gorm"
|
||||
"sneak.berlin/go/webhooker/internal/database"
|
||||
"sneak.berlin/go/webhooker/internal/delivery"
|
||||
"sneak.berlin/go/webhooker/internal/reqtls"
|
||||
)
|
||||
|
||||
// WebhookListItem holds data for the webhook list view.
|
||||
@@ -427,16 +428,16 @@ func (h *Handlers) renderSourceDetail(
|
||||
}
|
||||
}
|
||||
|
||||
host := r.Host
|
||||
scheme := "https"
|
||||
|
||||
if r.TLS == nil {
|
||||
scheme = "http"
|
||||
scheme := "http"
|
||||
if reqtls.IsTLS(r) {
|
||||
scheme = "https"
|
||||
}
|
||||
|
||||
if fwdProto := r.Header.Get("X-Forwarded-Proto"); fwdProto != "" {
|
||||
scheme = fwdProto
|
||||
}
|
||||
// The host is the client's Host header, unvalidated. It is
|
||||
// inert only because source_detail.html renders BaseURL as
|
||||
// text inside a <code> element; putting it in an href or any
|
||||
// other URL context needs it constrained first.
|
||||
baseURL := scheme + "://" + r.Host
|
||||
|
||||
// The template calls Webhook methods, which take pointer
|
||||
// receivers; html/template cannot address a value stored in a map.
|
||||
@@ -448,7 +449,7 @@ func (h *Handlers) renderSourceDetail(
|
||||
"Entrypoints": NewEntrypointViews(entrypoints),
|
||||
"Targets": delivery.NewTargetViews(targets),
|
||||
"Events": events,
|
||||
"BaseURL": scheme + "://" + host,
|
||||
"BaseURL": baseURL,
|
||||
}
|
||||
|
||||
h.renderTemplate(w, r, "source_detail.html", data)
|
||||
|
||||
@@ -300,3 +300,80 @@ func TestEntrypointCopyButtonIsProgressiveEnhancement(t *testing.T) {
|
||||
"the page must render to completion, not abort partway",
|
||||
)
|
||||
}
|
||||
|
||||
// maxRetriesHelp is the wording both target forms must carry. The
|
||||
// delivery core makes max_retries attempts in total, not that many
|
||||
// retries on top of a first try (a fresh delivery starts at attempt 1
|
||||
// and target_http gives up once the attempt number reaches
|
||||
// max_retries), and 0 is special-cased to a single fire-and-forget
|
||||
// attempt with no circuit breaker.
|
||||
const maxRetriesHelp = "This is the total number of delivery attempts, " +
|
||||
"not retries on top of the first: a value of 3 makes three attempts " +
|
||||
"in all. 0 means a single attempt with no retries and no circuit " +
|
||||
"breaker."
|
||||
|
||||
// TestTargetFormMaxRetriesCopyMatchesBehaviour pins the max_retries
|
||||
// help text on both the create form (the add-target form on the webhook
|
||||
// detail page) and the edit form, so the copy cannot drift back to
|
||||
// calling the number a retry count.
|
||||
func TestTargetFormMaxRetriesCopyMatchesBehaviour(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
var h *handlers.Handlers
|
||||
|
||||
var sess *session.Session
|
||||
|
||||
app := newTestApp(t, &h, &sess)
|
||||
app.RequireStart()
|
||||
|
||||
t.Cleanup(app.RequireStop)
|
||||
|
||||
webhook := &database.Webhook{Name: "wh", RetentionDays: 14}
|
||||
webhook.ID = testWebhookID
|
||||
|
||||
entrypoint := database.Entrypoint{Path: "abc123"}
|
||||
entrypoint.ID = "ep-1"
|
||||
|
||||
createBody := renderPage(
|
||||
t, h, sess, "source_detail.html", map[string]any{
|
||||
dataKeyWebhook: webhook,
|
||||
"Entrypoints": handlers.NewEntrypointViews(
|
||||
[]database.Entrypoint{entrypoint},
|
||||
),
|
||||
"Targets": delivery.NewTargetViews(nil),
|
||||
"Events": []database.Event{},
|
||||
"BaseURL": "https://hooks.example.com",
|
||||
},
|
||||
)
|
||||
|
||||
assert.Contains(
|
||||
t, createBody, maxRetriesHelp,
|
||||
"the add-target form must explain max_retries as total attempts",
|
||||
)
|
||||
|
||||
// A slack target exercises the same max_retries field while needing
|
||||
// only Config.URL from the edit template, so the test data stays
|
||||
// minimal. The Target key mirrors the field names the template reads
|
||||
// off the handler's view value.
|
||||
editBody := renderPage(
|
||||
t, h, sess, "target_edit.html", map[string]any{
|
||||
dataKeyWebhook: webhook,
|
||||
"Target": map[string]any{
|
||||
"ID": "tg-1",
|
||||
"Name": "t",
|
||||
"Type": "slack",
|
||||
"Active": true,
|
||||
"MaxRetries": 3,
|
||||
"Config": map[string]any{
|
||||
"URL": "https://hooks.slack.com/services/x",
|
||||
},
|
||||
},
|
||||
dataKeyError: "",
|
||||
},
|
||||
)
|
||||
|
||||
assert.Contains(
|
||||
t, editBody, maxRetriesHelp,
|
||||
"the target edit form must explain max_retries as total attempts",
|
||||
)
|
||||
}
|
||||
|
||||
@@ -380,9 +380,8 @@ func csrfTookStrictPath(
|
||||
|
||||
// TestCSRF_ForwardedProtoSpellingsTakeStrictPath runs the header
|
||||
// spellings a real proxy emits through the middleware. The environment
|
||||
// is dev -- the DEFAULT when WEBHOOKER_ENVIRONMENT is unset -- to pin
|
||||
// that the routing is a per-request transport decision and owes
|
||||
// nothing to configuration.
|
||||
// is set to dev -- the permissive setting -- to pin that the routing is
|
||||
// a per-request transport decision and owes nothing to configuration.
|
||||
func TestCSRF_ForwardedProtoSpellingsTakeStrictPath(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
|
||||
@@ -133,6 +133,11 @@ func (w *recoverResponseWriter) Unwrap() http.ResponseWriter {
|
||||
// what the access log records and the metrics count, and outside the
|
||||
// sentryhttp handler, whose Repanic option depends on something
|
||||
// further out recovering what it re-raises.
|
||||
//
|
||||
// Unlike http.Error on its own, it deletes any Set-Cookie the handler
|
||||
// set before panicking, because a request that failed must not hand
|
||||
// the client a credential; every other header is left to http.Error.
|
||||
// See https://git.eeqj.de/sneak/webhooker/issues/193.
|
||||
func (s *Middleware) Recoverer() func(http.Handler) http.Handler {
|
||||
return func(next http.Handler) http.Handler {
|
||||
return http.HandlerFunc(func(
|
||||
@@ -164,6 +169,8 @@ func (s *Middleware) Recoverer() func(http.Handler) http.Handler {
|
||||
return
|
||||
}
|
||||
|
||||
rw.Header().Del("Set-Cookie")
|
||||
|
||||
http.Error(
|
||||
rw,
|
||||
http.StatusText(
|
||||
|
||||
@@ -304,16 +304,44 @@ func TestRecovererRepanicsErrAbortHandler(t *testing.T) {
|
||||
)
|
||||
}
|
||||
|
||||
// TestRecovererDropsSetCookieFromTheRecovered500 covers a handler that
|
||||
// sets a cookie and a redirect target and then panics before sending
|
||||
// anything. A request that failed must not hand the client a
|
||||
// credential, so the 500 carries no cookie; Location is left alone.
|
||||
func TestRecovererDropsSetCookieFromTheRecovered500(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
probe := newRecovererProbe(
|
||||
t, false,
|
||||
func(w http.ResponseWriter, _ *http.Request) {
|
||||
w.Header().Set("Set-Cookie", "session=x")
|
||||
w.Header().Set("Location", "/after")
|
||||
|
||||
panic(panicMarker)
|
||||
},
|
||||
)
|
||||
|
||||
resp, err := probe.get(t)
|
||||
require.NoError(t, err)
|
||||
require.NoError(t, resp.Body.Close())
|
||||
|
||||
assert.Equal(t, http.StatusInternalServerError, resp.StatusCode)
|
||||
assert.Empty(t, resp.Cookies())
|
||||
assert.Equal(t, "/after", resp.Header.Get("Location"))
|
||||
}
|
||||
|
||||
// TestRecovererKeepsAnAlreadyCommittedResponse covers a handler that
|
||||
// panics after sending its status. The bytes are already on the wire,
|
||||
// so a second WriteHeader would change nothing the client sees and
|
||||
// would draw net/http's "superfluous response.WriteHeader" report.
|
||||
// cookie included, so a second WriteHeader would change nothing the
|
||||
// client sees and would draw net/http's "superfluous
|
||||
// response.WriteHeader" report.
|
||||
func TestRecovererKeepsAnAlreadyCommittedResponse(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
probe := newRecovererProbe(
|
||||
t, false,
|
||||
func(w http.ResponseWriter, _ *http.Request) {
|
||||
w.Header().Set("Set-Cookie", "session=x")
|
||||
w.WriteHeader(committedStatus)
|
||||
_, _ = w.Write([]byte("partial"))
|
||||
|
||||
@@ -331,6 +359,7 @@ func TestRecovererKeepsAnAlreadyCommittedResponse(t *testing.T) {
|
||||
|
||||
assert.Equal(t, committedStatus, resp.StatusCode)
|
||||
assert.Equal(t, "partial", string(body))
|
||||
assert.Len(t, resp.Cookies(), 1)
|
||||
|
||||
record := probe.panicRecord(t)
|
||||
assert.Equal(t, panicMarker, record["panic"])
|
||||
|
||||
@@ -5,7 +5,7 @@
|
||||
// several packages, by hand, and the answers disagreed. The session
|
||||
// cookie's Secure attribute was decided at startup from the configured
|
||||
// environment while the CSRF cookie's was decided per-request, so a
|
||||
// deployment behind a TLS proxy in the default environment emitted one
|
||||
// deployment behind a TLS proxy in the dev environment emitted one
|
||||
// Secure cookie and one non-Secure cookie on the same response.
|
||||
// Everything kept working, which is exactly why nobody noticed.
|
||||
//
|
||||
|
||||
@@ -70,7 +70,7 @@ func (s *Server) serveUntilShutdown() {
|
||||
err := s.httpServer.ListenAndServe()
|
||||
if err != nil && !errors.Is(err, http.ErrServerClosed) {
|
||||
s.log.Error("listen error", "error", err)
|
||||
s.shutdownOnListenFailure()
|
||||
s.shutdownWithFailure()
|
||||
}
|
||||
}
|
||||
|
||||
|
||||
@@ -93,7 +93,7 @@ func requireListenFailureExit(t *testing.T, env *testEnv) {
|
||||
select {
|
||||
case sig := <-app.Wait():
|
||||
require.Equal(
|
||||
t, server.ListenFailureExitCode, sig.ExitCode,
|
||||
t, server.StartupFailureExitCode, sig.ExitCode,
|
||||
"listen failure must exit non-zero",
|
||||
)
|
||||
case <-time.After(listenFailureDeadline):
|
||||
|
||||
@@ -92,11 +92,25 @@ func (s *Server) setupGlobalMiddleware() {
|
||||
func (s *Server) setupRoutes() {
|
||||
s.router.Get("/", s.h.HandleIndex())
|
||||
|
||||
s.router.Mount(
|
||||
"/s",
|
||||
http.StripPrefix("/s", http.FileServer(http.FS(static.Static))),
|
||||
// Static assets answer GET and HEAD only. chi's default 405
|
||||
// carries no Allow header, so this group supplies its own.
|
||||
staticFiles := http.StripPrefix(
|
||||
"/s", http.FileServer(http.FS(static.Static)),
|
||||
)
|
||||
|
||||
s.router.Route("/s", func(r chi.Router) {
|
||||
r.MethodNotAllowed(func(w http.ResponseWriter, _ *http.Request) {
|
||||
w.Header().Set("Allow", "GET, HEAD")
|
||||
http.Error(
|
||||
w,
|
||||
"Method Not Allowed",
|
||||
http.StatusMethodNotAllowed,
|
||||
)
|
||||
})
|
||||
r.Method(http.MethodGet, "/*", staticFiles)
|
||||
r.Method(http.MethodHead, "/*", staticFiles)
|
||||
})
|
||||
|
||||
s.router.Route("/api/v1", func(_ chi.Router) {
|
||||
// API routes will be added here.
|
||||
})
|
||||
|
||||
+108
-19
@@ -7,6 +7,7 @@ import (
|
||||
"net/http/httptest"
|
||||
"net/url"
|
||||
"regexp"
|
||||
"slices"
|
||||
"strconv"
|
||||
"strings"
|
||||
"testing"
|
||||
@@ -220,9 +221,21 @@ func (e *testEnv) csrfFrom(
|
||||
// out of the markup has to be unescaped before it is submitted.
|
||||
token := html.UnescapeString(match[1])
|
||||
|
||||
combined := make([]*http.Cookie, 0, len(cookies))
|
||||
combined = append(combined, cookies...)
|
||||
combined = append(combined, w.Result().Cookies()...)
|
||||
// A cookie the page sets replaces the one of the same name, as in
|
||||
// a browser. Sent both, the server would read the first, older one.
|
||||
set := w.Result().Cookies()
|
||||
combined := make([]*http.Cookie, 0, len(cookies)+len(set))
|
||||
|
||||
for _, c := range cookies {
|
||||
replaced := slices.ContainsFunc(set, func(n *http.Cookie) bool {
|
||||
return n.Name == c.Name
|
||||
})
|
||||
if !replaced {
|
||||
combined = append(combined, c)
|
||||
}
|
||||
}
|
||||
|
||||
combined = append(combined, set...)
|
||||
|
||||
return token, combined
|
||||
}
|
||||
@@ -396,13 +409,15 @@ func (e *testEnv) storedHash(t *testing.T, username string) string {
|
||||
|
||||
// --- /s static group ---
|
||||
|
||||
// TestStaticServesEveryMethod pins what the static mount actually
|
||||
// answers. chi's Mount registers the handler for all methods and
|
||||
// http.FileServer only special-cases HEAD (by suppressing the body),
|
||||
// so a POST or a DELETE to an asset is served the file rather than
|
||||
// refused. The README documents this; the test is what keeps the two
|
||||
// from drifting.
|
||||
func TestStaticServesEveryMethod(t *testing.T) {
|
||||
// TestStaticServesOnlyGetAndHead pins the methods the static group
|
||||
// answers: GET and HEAD are served the asset, and the other methods
|
||||
// chi routes (POST, PUT, DELETE and the rest) are refused with 405
|
||||
// and an Allow header naming those two. A method chi does not route,
|
||||
// such as PROPFIND, is refused with 405 by the top-level router
|
||||
// before it reaches the static group, so it gets no Allow header.
|
||||
// The README documents this; the test is what keeps the two from
|
||||
// drifting.
|
||||
func TestStaticServesOnlyGetAndHead(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
env := newTestEnv(t)
|
||||
@@ -417,6 +432,7 @@ func TestStaticServesEveryMethod(t *testing.T) {
|
||||
http.MethodPost,
|
||||
http.MethodPut,
|
||||
http.MethodDelete,
|
||||
"PROPFIND",
|
||||
} {
|
||||
t.Run(method, func(t *testing.T) {
|
||||
t.Parallel()
|
||||
@@ -428,18 +444,38 @@ func TestStaticServesEveryMethod(t *testing.T) {
|
||||
w := httptest.NewRecorder()
|
||||
env.router.ServeHTTP(w, req)
|
||||
|
||||
assert.Equal(t, http.StatusOK, w.Code,
|
||||
"static mount answers every method")
|
||||
|
||||
if method == http.MethodHead {
|
||||
switch method {
|
||||
case http.MethodGet:
|
||||
assert.Equal(t, http.StatusOK, w.Code)
|
||||
assert.Equal(t, body, w.Body.Bytes(),
|
||||
"the asset itself is returned")
|
||||
case http.MethodHead:
|
||||
assert.Equal(t, http.StatusOK, w.Code)
|
||||
assert.Empty(t, w.Body.Bytes(),
|
||||
"HEAD must not carry a body")
|
||||
|
||||
return
|
||||
case "PROPFIND":
|
||||
assert.Equal(
|
||||
t, http.StatusMethodNotAllowed, w.Code,
|
||||
)
|
||||
assert.Empty(t, w.Header().Get("Allow"),
|
||||
"chi refuses a method it does not route "+
|
||||
"before the static group runs")
|
||||
assert.NotContains(
|
||||
t, w.Body.String(), string(body),
|
||||
"a refused method must not get the asset",
|
||||
)
|
||||
default:
|
||||
assert.Equal(
|
||||
t, http.StatusMethodNotAllowed, w.Code,
|
||||
)
|
||||
assert.Equal(
|
||||
t, "GET, HEAD", w.Header().Get("Allow"),
|
||||
)
|
||||
assert.NotContains(
|
||||
t, w.Body.String(), string(body),
|
||||
"a refused method must not get the asset",
|
||||
)
|
||||
}
|
||||
|
||||
assert.Equal(t, body, w.Body.Bytes(),
|
||||
"the asset itself is returned")
|
||||
})
|
||||
}
|
||||
}
|
||||
@@ -591,6 +627,59 @@ func TestPagesLogin_CorrectPasswordSurvivesASpentBudget(
|
||||
)
|
||||
}
|
||||
|
||||
// TestPagesLogin_CookiesFromAnEarlierDatabase is
|
||||
// https://git.eeqj.de/sneak/webhooker/issues/359. A new database
|
||||
// brings a new session key, and the operator's browser still holds
|
||||
// the session and CSRF cookies signed with the old one. Logging in
|
||||
// must work as from a fresh browser and leave cookies the new key
|
||||
// accepts.
|
||||
func TestPagesLogin_CookiesFromAnEarlierDatabase(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
const (
|
||||
username = "operator"
|
||||
password = "correct-horse-battery-staple"
|
||||
)
|
||||
|
||||
earlier := newTestEnv(t)
|
||||
earlierID, _ := earlier.seedUser(t, username, password)
|
||||
_, stale := earlier.csrfFrom(t, "/pages/login", nil)
|
||||
stale = append(stale, earlier.authCookies(t, earlierID, username)...)
|
||||
|
||||
env := newTestEnv(t)
|
||||
env.seedUser(t, username, password)
|
||||
|
||||
token, cookies := env.csrfFrom(t, "/pages/login", stale)
|
||||
|
||||
form := url.Values{}
|
||||
form.Set("csrf_token", token)
|
||||
form.Set("username", username)
|
||||
form.Set("password", password)
|
||||
|
||||
w := env.post("/pages/login", form, cookies)
|
||||
require.Equal(
|
||||
t, http.StatusSeeOther, w.Code,
|
||||
"a session cookie from another key must not fail the login",
|
||||
)
|
||||
|
||||
// The response deletes the old session cookie and then sets the
|
||||
// new one; a browser keeps the last.
|
||||
var fresh *http.Cookie
|
||||
|
||||
for _, c := range w.Result().Cookies() {
|
||||
if c.Name == session.SessionName {
|
||||
fresh = c
|
||||
}
|
||||
}
|
||||
|
||||
require.NotNil(t, fresh, "login must set a session cookie")
|
||||
assert.Equal(
|
||||
t, "/sources",
|
||||
env.get("/", []*http.Cookie{fresh}).Header().Get("Location"),
|
||||
"the new session cookie must authenticate",
|
||||
)
|
||||
}
|
||||
|
||||
// --- /user/{username} group ---
|
||||
|
||||
// TestPasswordChange_OversizeBody_RejectedAndPasswordUnchanged
|
||||
|
||||
@@ -0,0 +1,99 @@
|
||||
package server_test
|
||||
|
||||
import (
|
||||
"context"
|
||||
"net"
|
||||
"strconv"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"github.com/stretchr/testify/require"
|
||||
"go.uber.org/fx"
|
||||
"sneak.berlin/go/webhooker/internal/config"
|
||||
"sneak.berlin/go/webhooker/internal/globals"
|
||||
"sneak.berlin/go/webhooker/internal/server"
|
||||
)
|
||||
|
||||
// TestSentryInitFailure_ShutsDownTheApp pins that error reporting
|
||||
// which is configured and cannot be started ends the application
|
||||
// instead of serving without it.
|
||||
//
|
||||
// The measured defect logged `sentry init failure` and kept running,
|
||||
// so the deployment served traffic with reporting off while every
|
||||
// other signal — SENTRY_DSN still set, the startup summary's own
|
||||
// field — said it was on. Nothing later in the process can notice
|
||||
// that reports are going nowhere, which is why this exits rather than
|
||||
// degrades.
|
||||
//
|
||||
// The DSN is placed on a hand-built Config, which is the only way to
|
||||
// reach this branch at all: loadFromEnv now parses SENTRY_DSN with
|
||||
// sentry.NewDsn, the same call sentry.Init makes, so a DSN that
|
||||
// survives configuration cannot fail initialisation in the SDK
|
||||
// version this pins. The branch stays because that is a property of
|
||||
// the SDK's current implementation rather than of its contract.
|
||||
func TestSentryInitFailure_ShutsDownTheApp(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
port := freePort(t)
|
||||
|
||||
env := newTestEnvWithConfig(t, &config.Config{
|
||||
DataDir: t.TempDir(),
|
||||
Environment: config.EnvironmentDev,
|
||||
BindAddress: loopbackV4,
|
||||
Port: port,
|
||||
SentryDSN: "not-a-dsn",
|
||||
})
|
||||
|
||||
app := fx.New(
|
||||
fx.NopLogger,
|
||||
fx.Supply(env.log, env.cfg, env.mw, env.hnd),
|
||||
fx.Provide(globals.New, server.New),
|
||||
fx.Invoke(func(*server.Server) {}),
|
||||
)
|
||||
|
||||
startCtx, cancelStart := context.WithTimeout(
|
||||
context.Background(), lifecycleTimeout,
|
||||
)
|
||||
defer cancelStart()
|
||||
|
||||
require.NoError(t, app.Start(startCtx))
|
||||
|
||||
select {
|
||||
case sig := <-app.Wait():
|
||||
require.Equal(
|
||||
t, server.StartupFailureExitCode, sig.ExitCode,
|
||||
"a sentry failure must exit non-zero",
|
||||
)
|
||||
case <-time.After(listenFailureDeadline):
|
||||
t.Fatal("a sentry failure left the app running")
|
||||
}
|
||||
|
||||
// The stop sequence still has to complete: the failure must reach
|
||||
// shutdown through fx rather than around it.
|
||||
stopCtx, cancelStop := context.WithTimeout(
|
||||
context.Background(), lifecycleTimeout,
|
||||
)
|
||||
defer cancelStop()
|
||||
|
||||
require.NoError(t, app.Stop(stopCtx))
|
||||
|
||||
// And it must give up before it listens. A process that bound the
|
||||
// port and then exited would have accepted requests it could not
|
||||
// report on, which is the state under test in miniature.
|
||||
requireBindable(t, port)
|
||||
}
|
||||
|
||||
// requireBindable asserts that the port is free, which it is only if
|
||||
// the server under test never claimed it.
|
||||
func requireBindable(t *testing.T, port int) {
|
||||
t.Helper()
|
||||
|
||||
var listenCfg net.ListenConfig
|
||||
|
||||
listener, err := listenCfg.Listen(
|
||||
t.Context(), "tcp",
|
||||
net.JoinHostPort(loopbackV4, strconv.Itoa(port)),
|
||||
)
|
||||
require.NoError(t, err, "the server bound a port it then gave up")
|
||||
require.NoError(t, listener.Close())
|
||||
}
|
||||
+57
-26
@@ -51,12 +51,13 @@ const (
|
||||
minSentryFlush = 250 * time.Millisecond
|
||||
)
|
||||
|
||||
// ListenFailureExitCode is the status the process exits with when the
|
||||
// HTTP listener cannot be established, or dies for a reason other
|
||||
// than a requested shutdown. It must stay non-zero: systemd
|
||||
// `Restart=on-failure` and Docker's restart policies key off it, and a
|
||||
// zero exit would read as a deliberate stop.
|
||||
const ListenFailureExitCode = 1
|
||||
// StartupFailureExitCode is the status the process exits with when
|
||||
// the serving goroutine gives up: the HTTP listener cannot be
|
||||
// established or dies for a reason other than a requested shutdown, or
|
||||
// error reporting is configured and cannot be started. It must stay
|
||||
// non-zero: systemd `Restart=on-failure` and Docker's restart policies
|
||||
// key off it, and a zero exit would read as a deliberate stop.
|
||||
const StartupFailureExitCode = 1
|
||||
|
||||
// SentryFlushBudget reports how long the Sentry flush may run when
|
||||
// remaining is the time left on the fx stop context after the HTTP
|
||||
@@ -135,11 +136,25 @@ func New(lc fx.Lifecycle, params ServerParams) (*Server, error) {
|
||||
}
|
||||
|
||||
// Run configures Sentry and starts serving HTTP requests.
|
||||
//
|
||||
// A Sentry failure ends the application instead of listening. It runs
|
||||
// before the listener rather than after it so that the process never
|
||||
// binds a port it is about to give up.
|
||||
func (s *Server) Run() {
|
||||
s.configure()
|
||||
|
||||
// logging before sentry, because sentry logs
|
||||
s.enableSentry()
|
||||
err := s.enableSentry()
|
||||
if err != nil {
|
||||
s.log.Error(
|
||||
"SENTRY_DSN is set but error reporting could not be "+
|
||||
"started; refusing to serve with it off",
|
||||
"error", err,
|
||||
)
|
||||
s.shutdownWithFailure()
|
||||
|
||||
return
|
||||
}
|
||||
|
||||
s.serve()
|
||||
}
|
||||
@@ -150,11 +165,23 @@ func (s *Server) MaintenanceMode() bool {
|
||||
return s.params.Config.MaintenanceMode
|
||||
}
|
||||
|
||||
func (s *Server) enableSentry() {
|
||||
// enableSentry initialises the Sentry SDK when error reporting is
|
||||
// configured, and reports the failure when it is configured and cannot
|
||||
// be initialised. A DSN that is not set is not a failure: reporting
|
||||
// stays off and the server starts normally.
|
||||
//
|
||||
// There is no fallback to running with reporting off. An operator who
|
||||
// set SENTRY_DSN asked for failures to be visible, and serving traffic
|
||||
// with reporting quietly off is the one state nothing can ever tell
|
||||
// them about — the DSN is still set, so every later signal says it is
|
||||
// on. Config already refused a DSN the SDK cannot parse, which is what
|
||||
// a typo produces, so reaching this branch means the SDK refused
|
||||
// something that parsed: not a condition to guess at either.
|
||||
func (s *Server) enableSentry() error {
|
||||
s.sentryEnabled.Store(false)
|
||||
|
||||
if s.params.Config.SentryDSN == "" {
|
||||
return
|
||||
if !s.params.Config.SentryEnabled() {
|
||||
return nil
|
||||
}
|
||||
|
||||
err := sentry.Init(sentryClientOptions(
|
||||
@@ -166,19 +193,19 @@ func (s *Server) enableSentry() {
|
||||
),
|
||||
))
|
||||
if err != nil {
|
||||
s.log.Error("sentry init failure", "error", err)
|
||||
// Don't use fatal since we still want the service to run
|
||||
return
|
||||
return fmt.Errorf("initialising sentry: %w", err)
|
||||
}
|
||||
|
||||
s.log.Info("sentry error reporting activated")
|
||||
s.sentryEnabled.Store(true)
|
||||
|
||||
return nil
|
||||
}
|
||||
|
||||
// serve installs the signal watcher, starts the listener and blocks
|
||||
// until the server's context is cancelled. The process exit status is
|
||||
// fx's to decide — from a signal, or from the code
|
||||
// shutdownOnListenFailure hands the Shutdowner — so this reports
|
||||
// shutdownWithFailure hands the Shutdowner — so this reports
|
||||
// nothing back to its caller.
|
||||
func (s *Server) serve() {
|
||||
ctx, cancelFunc := context.WithCancel(context.Background())
|
||||
@@ -208,20 +235,24 @@ func (s *Server) serve() {
|
||||
// Do not call cleanShutdown() here to avoid double invocation.
|
||||
}
|
||||
|
||||
// shutdownOnListenFailure ends the application after the HTTP
|
||||
// listener failed. The fx OnStart hook returns as soon as the serving
|
||||
// goroutine is spawned, so nothing downstream of it ever learns that
|
||||
// the listen failed: fx reports RUNNING and the process sits alive
|
||||
// with nothing bound, which is invisible to systemd and Docker
|
||||
// restart policies. Asking the Shutdowner to stop the app with a
|
||||
// non-zero code is what turns that into a visible failure.
|
||||
// shutdownWithFailure ends the application non-zero from the serving
|
||||
// goroutine. It is how anything on that goroutine fails fatally: the
|
||||
// fx OnStart hook returns as soon as the goroutine is spawned, so
|
||||
// nothing downstream of it ever learns that the goroutine gave up. fx
|
||||
// reports RUNNING and the process sits alive having done neither what
|
||||
// it was asked nor anything visible instead, which systemd and
|
||||
// Docker restart policies cannot see. Asking the Shutdowner to stop
|
||||
// the app with a non-zero code is what turns that into a visible
|
||||
// failure, and it is the whole of "fatal" here — no panic, no
|
||||
// os.Exit, and every stop hook still runs.
|
||||
//
|
||||
// The context cancel that follows only unwinds serve()'s own wait.
|
||||
// The shutdown itself runs through fx's normal stop sequence, so the
|
||||
// clean-shutdown drain in cleanShutdown is reached unchanged.
|
||||
func (s *Server) shutdownOnListenFailure() {
|
||||
// The context cancel that follows only unwinds serve()'s own wait,
|
||||
// and is skipped before serve has installed one. The shutdown itself
|
||||
// runs through fx's normal stop sequence, so the clean-shutdown drain
|
||||
// in cleanShutdown is reached unchanged.
|
||||
func (s *Server) shutdownWithFailure() {
|
||||
err := s.params.Shutdowner.Shutdown(
|
||||
fx.ExitCode(ListenFailureExitCode),
|
||||
fx.ExitCode(StartupFailureExitCode),
|
||||
)
|
||||
if err != nil {
|
||||
s.log.Error("shutdown request failed", "error", err)
|
||||
|
||||
@@ -19,8 +19,8 @@ import (
|
||||
)
|
||||
|
||||
// The tests below exercise the securecookie codecs underneath the
|
||||
// store and nothing else: Session.Get only decodes, so no server-side
|
||||
// expiry check takes part in the result. They exist because
|
||||
// store and nothing else: they decode through the store itself, so no
|
||||
// server-side expiry check takes part in the result. They exist because
|
||||
// NewCookieStore gives its codecs a 30-day max age that assigning
|
||||
// store.Options does not override, which would let the codec accept a
|
||||
// cookie weeks past the cap the cookie attribute advertises.
|
||||
@@ -75,10 +75,11 @@ func restamp(
|
||||
return base64.URLEncoding.EncodeToString(payload)
|
||||
}
|
||||
|
||||
// decodeCookie feeds value back through the store's decode path.
|
||||
// decodeCookie feeds value back through the store's decode path. It
|
||||
// asks the store rather than Session.Get, which treats a cookie that
|
||||
// does not decode as absent and so hides the codec's reason.
|
||||
func decodeCookie(
|
||||
t *testing.T,
|
||||
s *session.Session,
|
||||
value string,
|
||||
) (*sessions.Session, error) {
|
||||
t.Helper()
|
||||
@@ -94,7 +95,7 @@ func decodeCookie(
|
||||
SameSite: http.SameSiteLaxMode,
|
||||
})
|
||||
|
||||
sess, err := s.Get(req)
|
||||
sess, err := session.NewStore(testKey()).Get(req, session.SessionName)
|
||||
require.NotNil(t, sess)
|
||||
|
||||
return sess, err
|
||||
@@ -105,7 +106,7 @@ func TestCodec_AcceptsCookieInsideAbsoluteCap(t *testing.T) {
|
||||
|
||||
s := testSession(t)
|
||||
|
||||
sess, err := decodeCookie(t, s, restamp(
|
||||
sess, err := decodeCookie(t, restamp(
|
||||
t,
|
||||
issuedCookie(t, s),
|
||||
time.Now().Add(-(testAbsoluteMaxAge-time.Hour)),
|
||||
@@ -126,7 +127,7 @@ func TestCodec_RejectsCookiePastAbsoluteCap(t *testing.T) {
|
||||
|
||||
s := testSession(t)
|
||||
|
||||
sess, err := decodeCookie(t, s, restamp(
|
||||
sess, err := decodeCookie(t, restamp(
|
||||
t,
|
||||
issuedCookie(t, s),
|
||||
time.Now().Add(-(testAbsoluteMaxAge+time.Hour)),
|
||||
|
||||
@@ -146,8 +146,8 @@ func newStore(key []byte) *sessions.CookieStore {
|
||||
//
|
||||
// This is decided per-request, not once at startup. Deciding it at
|
||||
// startup from the configured environment is what this replaces, and
|
||||
// it got the DEFAULT posture wrong: "dev" is the environment when
|
||||
// WEBHOOKER_ENVIRONMENT is unset, so a deployment terminating TLS at a
|
||||
// it got the DEFAULT posture wrong: "dev" was then the environment when
|
||||
// WEBHOOKER_ENVIRONMENT was unset, so a deployment terminating TLS at a
|
||||
// proxy without also setting the environment emitted the
|
||||
// authentication cookie with no Secure attribute -- silently, and on
|
||||
// the same response as a CSRF cookie that did have one.
|
||||
@@ -224,10 +224,22 @@ func New(
|
||||
}
|
||||
|
||||
// Get retrieves a session for the request.
|
||||
//
|
||||
// A session cookie that does not decode -- one signed with an earlier
|
||||
// session key, say, because the database was made anew -- is treated
|
||||
// as absent: the caller gets a new, empty session and no error, and
|
||||
// the next save replaces the cookie.
|
||||
func (s *Session) Get(
|
||||
r *http.Request,
|
||||
) (*sessions.Session, error) {
|
||||
return s.store.Get(r, SessionName)
|
||||
sess, err := s.store.Get(r, SessionName)
|
||||
if sess == nil {
|
||||
return nil, err
|
||||
}
|
||||
|
||||
// For a cookie that does not decode, gorilla/sessions returns a
|
||||
// new, empty session alongside the error that is dropped here.
|
||||
return sess, nil
|
||||
}
|
||||
|
||||
// GetKey returns the raw 32-byte authentication key used for
|
||||
|
||||
@@ -990,8 +990,8 @@ func sessionCookieFrom(
|
||||
|
||||
// TestSave_SecureFollowsRequestTransport is the regression test for
|
||||
// the defect this replaces: Secure was fixed at startup from the
|
||||
// configured environment, and "dev" is the environment when
|
||||
// WEBHOOKER_ENVIRONMENT is unset. A deployment behind a TLS proxy in
|
||||
// configured environment, and "dev" was then the environment when
|
||||
// WEBHOOKER_ENVIRONMENT was unset. A deployment behind a TLS proxy in
|
||||
// that DEFAULT posture shipped the authentication cookie with no
|
||||
// Secure attribute and said nothing about it.
|
||||
//
|
||||
|
||||
@@ -120,9 +120,12 @@
|
||||
<label class="text-sm text-gray-700">Timeout (seconds, blank = default):</label>
|
||||
<input type="number" name="timeout" min="0" max="300" :disabled="targetType !== 'http'" class="input text-sm w-24">
|
||||
</div>
|
||||
<div x-show="targetType === 'http'" class="flex gap-2 items-center">
|
||||
<label class="text-sm text-gray-700">Max retries (0 = fire-and-forget):</label>
|
||||
<input type="number" name="max_retries" value="0" min="0" max="20" class="input text-sm w-24">
|
||||
<div x-show="targetType === 'http'">
|
||||
<div class="flex gap-2 items-center">
|
||||
<label class="text-sm text-gray-700">Max retries:</label>
|
||||
<input type="number" name="max_retries" value="0" min="0" max="20" class="input text-sm w-24">
|
||||
</div>
|
||||
<p class="text-xs text-gray-500 mt-1">This is the total number of delivery attempts, not retries on top of the first: a value of 3 makes three attempts in all. 0 means a single attempt with no retries and no circuit breaker.</p>
|
||||
</div>
|
||||
<div x-show="targetType === 'slack'">
|
||||
<input type="url" name="url" placeholder="https://hooks.slack.com/services/..." :disabled="targetType !== 'slack'" class="input text-sm">
|
||||
|
||||
@@ -69,7 +69,7 @@
|
||||
<div class="form-group">
|
||||
<label for="max_retries" class="label">Max retries</label>
|
||||
<input type="number" id="max_retries" name="max_retries" value="{{.Target.MaxRetries}}" min="0" max="20" class="input">
|
||||
<p class="text-xs text-gray-500 mt-1">0 is fire-and-forget: one attempt, no circuit breaker.</p>
|
||||
<p class="text-xs text-gray-500 mt-1">This is the total number of delivery attempts, not retries on top of the first: a value of 3 makes three attempts in all. 0 means a single attempt with no retries and no circuit breaker.</p>
|
||||
</div>
|
||||
{{end}}
|
||||
|
||||
|
||||
Reference in New Issue
Block a user