1 Commits

Author SHA1 Message Date
55edebaea1 Set fx.StopTimeout inside the container stop grace (closes #134)
Some checks failed
check / check (push) Has been cancelled
fx defaults the stop timeout to 15s and the Dockerfile sets no
STOPSIGNAL or grace override, so Docker's 10s default SIGKILLs the
process five seconds before the bound can fire. Everything gated on
it — including the "shutdown timed out, goroutines still running"
error log that tells an operator a component is wedged — was
unreachable in the image this repo produces.

Set fx.StopTimeout to 5s: inside the grace with headroom for signal
delivery and process exit. The option set moves into newApp() so a
test can read (*fx.App).StopTimeout() back and pin it against
drift; dropping the option makes that test report fx's 15s default.

Lower the HTTP drain budget (server.ShutdownTimeout) from 5s to 3s.
fx bounds the whole stop sequence and returns without running its
remaining hooks once the stop context expires, so two equal values
meant a drain that used its full budget exhausted the sequence
budget at that instant and skipped every later hook — the delivery
engine, the healthcheck, the webhook DB manager and the database
close — in exactly the case where the drain mattered. The tail
hooks are microsecond-scale in normal operation, so 2s of remaining
budget is ample, and holding the total at 5s keeps a wide margin
under Docker's 10s grace. This does not make the database close
unconditional: the ArchiveSweeper and RetentionReaper hooks run
before the server and can still consume the whole budget.

Bound the Sentry flush by the remaining stop budget. The server's
stop hook is not only the drain: cleanShutdown calls sentry.Flush
after it, in the same hook, and sentry.Flush takes a bare duration
and honours no context. With SENTRY_DSN set to an unreachable
endpoint, a full-length drain plus a stalled 2s flush spent the
whole 5s sequence budget by itself and the tail hooks — database
close included — were skipped again, on a configuration the README
documents. SentryFlushBudget now clamps the flush to what is left
on the stop context less server.TailHookReserve, skipping it below
250ms rather than making a useless attempt, so a full-length drain
drops Sentry events instead of the database close.

TestStopTimeout_LeavesHeadroomForTailHooks now walks every drain
length the hook can produce and asserts drain plus flush still
leaves the 2s tail margin, so it covers the hook's real worst case
rather than the drain alone; unbounding the flush fails it at a
1.01s drain. TestSentryFlushBudget covers the clamp directly.

Also fix a latent coin flip in the shared stop-hook waiter. It
selected on the drained channel against ctx.Done() with no
preamble, and select picks uniformly among ready cases, so a
component that drained against an already-expired context reported
a timeout about half the time. Not reachable through fx, which
re-checks ctx.Err() before each hook, but the helper is shared and
a direct caller can reach it. waitDone now settles the drained case
in a non-blocking preamble first; the test drives it over 1000
passes, so a restored coin flip cannot pass by luck.

README records the real stop-hook order (ArchiveSweeper,
RetentionReaper, server, delivery.Engine, healthcheck,
WebhookDBManager, database close), the two timeouts and their
relationship, why the Sentry flush is clamped rather than allowed
its own fixed budget, and the container stop grace: that lowering
the grace below the bound puts SIGKILL back in front of it, and
that an expired stop context makes fx skip its remaining hooks, so
a wedge in the first-stopped component means the database close
never runs. The Package Layout tree gets its internal/lifecycle/
entry in the sorted position, dropping the out-of-order duplicate
this branch rebased onto.
2026-08-17 21:48:25 +00:00
68 changed files with 374 additions and 11079 deletions

View File

@@ -13,19 +13,39 @@ jobs:
uses: actions/checkout@11bd71901bbe5b1630ceea73d27597364c9af683 # v4.2.2 2024-10-23
with:
# The fingerprint step below needs history to find the last commit
# that touched the Docker build context, and the superseded-status
# step needs it to walk ancestors (it aborts on a shallow clone).
# that touched the Docker build context.
fetch-depth: 0
- name: Mark superseded run statuses
- name: Neutralize superseded run statuses
# Gitea cancels the in-flight run when another commit is pushed to the
# same branch and records the cancellation as `failure`, so a commit
# that was never tested reads as a test result. The script rewrites
# those statuses to say what happened. See its header for why the
# state stays `failure` and not `skipped`.
# that was never tested reads red. The cancellation is unconditional
# server-side for push events and cannot be disabled from a workflow
# file, so the superseding run rewrites those statuses to `skipped`.
# Only the exact cancellation status is touched; a real failure is
# left alone.
env:
GITEA_TOKEN: ${{ secrets.GITEA_TOKEN }}
run: script/ci-mark-superseded
run: |
set -eu
api="${GITHUB_API_URL}/repos/${GITHUB_REPOSITORY}"
ctx='check / check (push)'
for sha in $(git rev-list --max-count=20 "${GITHUB_SHA}^" || true); do
latest="$(curl -sf "${api}/commits/${sha}/status" | jq -r \
--arg c "$ctx" \
'[.statuses[] | select(.context == $c)][0] // empty
| "\(.status)|\(.description)"')" || continue
[ "$latest" = 'failure|Has been cancelled' ] || continue
curl -sf -X POST "${api}/statuses/${sha}" \
-H "Authorization: token ${GITEA_TOKEN}" \
-H 'Content-Type: application/json' \
-d "$(jq -nc --arg c "$ctx" '{
context: $c,
state: "skipped",
description: "Superseded by a newer commit; never tested"
}')" >/dev/null
echo "neutralized superseded status on ${sha}"
done
- name: Fingerprint the build context
# `.dockerignore` keeps docs out of the build context, so a docs-only

View File

@@ -19,14 +19,9 @@ RUN go mod download
# .dockerignore.
COPY . .
# Run formatting check and linter. golangci-lint is invoked directly rather
# than through `make lint`: this stage is already the pinned linter image, and
# script/lint is a wrapper that builds Dockerfile.lint, so calling it here
# would need a docker daemon inside the build. Keep these steps in step with
# Dockerfile.lint, including --network=none (see its header for why).
# Run formatting check and linter
RUN make fmt-check
RUN --network=none golangci-lint config verify --config .golangci.yml
RUN --network=none golangci-lint run --config .golangci.yml ./...
RUN make lint
# Build stage
# golang:1.26.1-bookworm (Debian-based), 2026-03-17
@@ -37,9 +32,7 @@ FROM golang:1.26.1-bookworm@sha256:4465644228bc2857a954b092167e12aa59c006a349228
# Depend on lint stage passing
COPY --from=lint /src/go.sum /dev/null
# jq is a runtime dependency of script/ci-mark-superseded, which the test
# suite executes.
RUN apt-get update && apt-get install -y --no-install-recommends make curl ca-certificates jq && rm -rf /var/lib/apt/lists/*
RUN apt-get update && apt-get install -y --no-install-recommends make curl ca-certificates && rm -rf /var/lib/apt/lists/*
WORKDIR /build

View File

@@ -1,37 +0,0 @@
# Lint-only image, built by script/lint. golangci-lint is never installed on
# the host: the repo is COPYed into the pinned image and linted as a build
# step, so a successful build IS a clean lint. This works even when the docker
# daemon is remote and bind mounts are impossible.
#
# script/lint passes --no-cache-filter=lint. Without it an unchanged tree
# replays the lint stage from cache and the build succeeds in under a second
# having run no linter at all. Do not drop that flag.
#
# The lint steps run with --network=none. `golangci-lint config verify` is
# documented as fetching its JSON schema over HTTPS, which would make linting
# depend on an unpinned remote artifact; this pinned image resolves the schema
# without any network, and --network=none enforces that rather than trusting
# it. It also proves no linter reaches out at analysis time. If a future image
# bump makes either step need the network, this build fails loudly instead of
# quietly acquiring an unpinned dependency.
# golangci/golangci-lint:v2.12.2 (Debian-based), 2026-08-07
# Using Debian-based image because mattn/go-sqlite3 (CGO) does not
# compile on Alpine musl (off64_t is a glibc type).
FROM golangci/golangci-lint:v2.12.2@sha256:5cceeef04e53efe1470638d4b4b4f5ceefd574955ab3941b2d9a68a8c9ad5240 AS deps
WORKDIR /src
# Copy go mod files first for better layer caching. This stage is cacheable;
# only the lint stage below is forced to re-execute.
COPY go.mod go.sum ./
RUN go mod download
FROM deps AS lint
COPY . .
# `run` silently ignores config keys it does not recognize, so a typo would
# disable a setting without a word. `config verify` is what catches that.
RUN --network=none golangci-lint config verify --config .golangci.yml
RUN --network=none golangci-lint run --config .golangci.yml ./...

683
README.md
View File

@@ -12,16 +12,14 @@ with retry support, logging, and observability. Category: infrastructure
### Prerequisites
- Go 1.26.1+ (the version in `go.mod`)
- Docker (for linting, for the test stage of the CI gate, and for
containerized deployment)
- golangci-lint v2.12.2 (the version pinned in `script/bootstrap` and
in the `Dockerfile`'s lint stage; `make bootstrap` installs it)
- Docker (for containerized deployment, and for the lint and test
stages of the CI gate)
- `curl`, used by `script/fetch-assets` to download the third-party
browser assets, which are not committed (`make bootstrap` installs
it if missing)
golangci-lint is not a prerequisite and must not be installed on the
host: `script/bootstrap` does not install it, and `make lint` runs the
digest-pinned linter image via `Dockerfile.lint`.
### Quick Start
```bash
@@ -29,9 +27,9 @@ digest-pinned linter image via `Dockerfile.lint`.
git clone https://git.eeqj.de/sneak/webhooker.git
cd webhooker
# Install Go dependencies and the third-party browser assets.
# `make deps` alone is not enough: it only runs go mod download/tidy,
# and the checks below need the fetched assets.
# Install Go dependencies, the pinned linter, and the third-party
# browser assets. `make deps` alone is not enough: it only runs
# go mod download/tidy, and the checks below need the fetched assets.
make bootstrap
# Run all checks (test, lint, format check)
@@ -54,7 +52,7 @@ make setup # Bootstrap + install git pre-commit hook
make assets # Fetch + verify third-party browser assets
make fmt # Format code (gofmt + goimports)
make fmt-check # Fail if gofmt would change anything (writes nothing)
make lint # Run golangci-lint in Docker (Dockerfile.lint)
make lint # Run golangci-lint
make test # Run tests with race detection
make check # test + lint + fmt-check (CI gate)
make build # Build binary to bin/webhooker
@@ -113,7 +111,7 @@ TTY detection, and security headers are always applied.
| `RETENTION_SWEEP_INTERVAL` | How often the retention reaper and archive sweeper run (Go duration, must be positive) | `1h` |
| `SESSION_IDLE_TIMEOUT` | Idle session timeout (Go duration) | `24h` |
| `RECEIVER_RATE_LIMIT` | Receiver requests/minute per IP per entrypoint (10x that per IP across the route) | `120` |
| `TRUSTED_PROXIES` | CIDRs whose forwarded headers are trusted (unset: all clients behind a proxy share one rate-limit bucket; a correct login password is never throttled either way) | `""` (none) |
| `TRUSTED_PROXIES` | CIDRs whose forwarded headers are trusted (unset: all clients behind a proxy share one rate-limit bucket) | `""` (none) |
#### Trusted proxies
@@ -136,15 +134,13 @@ That default is safe against forged headers, but leaving it unset in
production has a cost you must know about. Production runs behind a
TLS-terminating reverse proxy, so with `TRUSTED_PROXIES` unset every
request keys on the proxy's own address and all clients share a single
bucket per limit. The receiver limits become service-wide ceilings,
and the login endpoint's failure counting collapses onto one key, so a
stranger's wrong passwords throttle every other client's wrong
passwords.
What it cannot do is lock the operator out. The login endpoint
verifies credentials **before** it consults any limit and charges only
failures, so a correct password is never throttled no matter how full
the bucket is. See [Rate Limiting](#rate-limiting).
bucket per limit. For the login and password-change limits that is a
denial of service anyone can perform: a steady five POSTs per minute
from any address on the internet keeps the shared login bucket full,
and the operator's own login then returns HTTP 429 for as long as the
trickle continues. There is no second administrative path and no
bypass. Restarting the service clears the in-memory buckets, but a
sustained trickle re-locks them immediately.
The remedy is to set `TRUSTED_PROXIES` to your reverse proxy's
address, which restores per-client buckets. webhooker logs a warning
@@ -279,15 +275,13 @@ are inline commands with no script behind them. We provide:
- `script/fetch-assets` — download the third-party browser assets into
`static/`, verifying each against its pinned sha256
- `script/test` — run the test suite
- `script/lint` — run golangci-lint in Docker (see Linting below)
- `script/lint` — run golangci-lint
- `script/fmt` — format all code (writes)
- `script/fmt-check` — check formatting (read-only)
- `script/check` — run test, lint, and fmt-check
- `script/docker` — build the Docker image tagged via `script/projectname`
- `script/cibuild` — CI entrypoint: `docker build .` (the Dockerfile
runs the checks, so a green build implies a green repo)
- `script/ci-mark-superseded` — CI helper: mark the commits whose run a
newer push cancelled (see [CI gate honesty](#ci-gate-honesty))
- `script/precommit` — pre-commit checks (`go mod tidy` guard, then
`script/check`)
- `script/install-precommit` — install the git pre-commit hook that
@@ -384,12 +378,11 @@ It uses:
- **[gorilla/csrf](https://github.com/gorilla/csrf)** for CSRF
protection (cookie-based double-submit tokens)
- **[go-chi/httprate](https://github.com/go-chi/httprate)** for
sliding-window rate limiting of the password-change and webhook
receiver endpoints. The bucket is per client IP only when
sliding-window rate limiting of the login, password-change and
webhook receiver endpoints. The bucket is per client IP only when
`TRUSTED_PROXIES` names the reverse proxy; unset, every client
behind that proxy shares one bucket per limit. The login endpoint
counts failed attempts itself instead, so that a correct password is
never throttled (see [Rate Limiting](#rate-limiting))
behind that proxy shares one bucket per limit (see
[Rate Limiting](#rate-limiting))
- **[Prometheus](https://prometheus.io)** for metrics, served at
`/metrics` behind basic auth
- **[Sentry](https://sentry.io)** for optional error reporting
@@ -994,356 +987,21 @@ costs; log volume it caps rather than eliminates. A path that names no
entrypoint is recorded by the handler at `DEBUG`, and the aggregate
limiter logs its own rejections at `DEBUG` and without the path, so
neither appears at all under the default level. The per-entrypoint
limiter is the loud one: it logs every rejection at `WARN` with the
request path, which on this route is attacker-controlled text. A client
hammering a single invented path is served `RECEIVER_RATE_LIMIT`
limiter is the loud one: it still logs every rejection at `WARN` with
the request path, which on this route is attacker-controlled text. A
client hammering a single invented path is served `RECEIVER_RATE_LIMIT`
requests and has the rest of its aggregate budget rejected there, so
the aggregate limit is what bounds the _number_ of those `WARN` lines —
to under ten times `RECEIVER_RATE_LIMIT` per minute per client IP, 1080
at the defaults, where before it there was no bound at all. Their
_width_ is bounded by the field budgets below, the same ones the access
log spends. The access log is bounded by neither limit: every request
is recorded once at `INFO`, served or rejected alike.
What the access log does bound is the _content_ of those lines. A 3xx
or 4xx response logs the chi route pattern — `/webhook/{uuid}`,
`/user/{username}//`, or the literal `(unmatched)` when the request hit
no route at all — in place of the concrete URL. Those are the outcomes
an unauthenticated client can drive for free: 404 and 429 on any
invented receiver path, a login redirect on any invented profile path.
Logging the URL there would let a flood write text of its own choosing,
at a length of its own choosing, into the log. 2xx and 5xx responses
keep the concrete path — a success resolved against a static route or
against the operator's own data (on the receiver, against a stored
entrypoint UUID), and a 5xx is a bug in this service, where the exact
path is the evidence and no client can provoke one at will.
The query string is never logged; it is replaced by the fixed marker
`?(redacted)`. It is client-chosen on every route, and
`/.well-known/healthcheck` and `/s/*` answer 200 to anyone with no rate
limiter in front of them, so a query on a fixed 200 URL would otherwise
buy the same amplification as an invented path. Nothing debuggable is
lost: `page`, on the authenticated pagination links, is the only query
parameter this service reads.
Client-supplied request content does not leave the host by the other
route either. The Sentry SDK attaches the request to every event it
captures, independently of the access log, and `SendDefaultPII=false`
does not cover all of what it copies: the raw query string and the
first 10 KiB of the request body are both taken unconditionally, the
body precisely because these handlers call `ParseForm`. A `BeforeSend`
hook therefore replaces the query string and the body with
`(redacted)`, drops cookies and the remote-address environment, and
reduces the headers to a fixed allowlist — `Accept`, `Content-Length`,
`Content-Type`, `Host`, `Origin`, `Referer`, `User-Agent` and
`X-Request-Id`.
The same hook rewrites the request URL. The SDK builds it as
`scheme://host/path` from the concrete path, which on the receiver
route is `/webhook/<uuid>` in full — and that UUID is a write
capability, not an identifier: anyone holding it can post events this
service accepts and its targets then deliver. A tracker has its own
retention, access control and deletion policy, so the rule the access
log follows above does not carry across that boundary. What is sent is
the chi route pattern instead: `http://host/webhook/{uuid}`.
The scheme and the host are kept, and everything else in the URL is
discarded rather than edited, so a future SDK version that starts
appending a query string cannot widen this. The scheme has to survive
for the reason given below. The host is whatever the request's `Host`
header carried — this service validates no hostname, so on a directly
exposed deployment a client sets it — and that same header is on the
allowlist above, so scrubbing the host out of the URL would withhold
nothing that is not sent anyway.
The body, the query string and the URL are all handled on every route
rather than filtered by route. For the URL that is also what keeps the
event locatable: an error event is grouped by its exception and stack
trace, not by its URL, so replacing the path with the pattern costs no
grouping and the pattern still names the route in the UI. And an
unconditional rule cannot leak on a route somebody forgets to add to
it, which a route-conditional one can. For the body there is a second
reason: nothing debuggable is lost, because every handler reads its
fields with `PostFormValue`, so the body is exactly where the
credentials are — the target destination URL, the login password, both
password-change fields — and the one route whose body is genuine
signal is the receiver, whose body is already stored on the event and
served from the UI, so a tracker is not where anyone reads it.
The route is reachable from the hook only on the error dispatch.
`sentryhttp`'s recover path puts the request on the context it hands
to `RecoverWithContext`, and the SDK carries that context through to
`BeforeSend` as `hint.Context`, so
`hint.Context.Value(sentry.RequestContextKey)` yields the live request
and chi's `RoutePattern()` yields the matched pattern off it. The
transaction dispatch has no such request: a finished span captures
with a nil hint, which the client replaces with an empty one, so
`BeforeSendTransaction` sees no context at all. Tracing is off in this
service, so no transaction event is produced today, but the hook is
installed on both dispatches as a floor.
Where the pattern is out of reach — the transaction dispatch, an event
captured outside the router, or a request that matched no route — the
fallback is never the concrete path. The path becomes the literal
`/(redacted)`, so the URL reads `http://host/(redacted)`; a URL the
rewrite cannot parse into a scheme is withheld whole. A transaction
event additionally carries the SDK's own `METHOD /path` name, built
from the concrete path as well; it is rewritten on the same terms, to
`POST /webhook/{uuid}` where the pattern is known and `POST
/(redacted)` where it is not.
The headers are an allowlist for the same reason the rules above are
unconditional: the SDK's own filter removes four names and passes
everything else, which would ship `X-CSRF-Token` and the shared
secrets senders put on the receiver route. What survives still names
the failing route — scheme, host, route pattern, method — and
`X-Request-Id` ties the event to the local access log line that holds
the rest. Nothing dropped is needed for the likeliest use, debugging a
CSRF rejection. Its three inputs are the TLS decision, `Origin` and
`Referer`; the latter two are kept, and the first is the scheme of the
retained URL, because the SDK derives that scheme from
`r.TLS != nil || r.Header.Get("X-Forwarded-Proto") == "https"` — byte
for byte the predicate `internal/middleware/csrf.go` uses to choose
between the `csrf.Secure(true)` and `csrf.Secure(false)` handlers.
That is what the rewrite above preserves it for, and it is why
dropping `X-Forwarded-Proto` costs nothing. The dropped provider
headers (`X-GitHub-Event`, `X-Gitlab-Event` and the like) are real
signal but are recorded locally on the event, and
`Sentry-Trace`/`Baggage` are already reflected in the event's trace
context.
The remaining client-supplied fields are truncated rather than dropped,
each to a fixed budget: 512 bytes for `url`, `useragent` and `referer`,
128 for `request_id` (chi passes an inbound `X-Request-Id` header
through), and 32 for `method`. A truncated `User-Agent` is still worth
reading; an absent one is not. A cut value ends in `[truncated]`, which
is charged on top of the budget rather than inside it.
Each budget is spent in _encoded_ bytes, not in the bytes the client
sent. Every rune is charged what the wider of the two log handlers
emits for it: two bytes for a quotation mark, a backslash or a tab; six
for a non-printable rune below U+10000; ten for one at or above it,
which the text handler spells `\UXXXXXXXX`. Go's header parser accepts
all of them in a header value, so a budget counted raw would buy a
field several times its nominal size — and the line, not the header, is
what an operator has to store. Plain ASCII encodes one byte for one, so
a real browser's `User-Agent` still fits whole; a value built out of
escapes keeps a proportionally shorter prefix, which is the right
trade.
Net: **one `INFO` line per request, of at most 2,560 bytes.** That
ceiling is arithmetic, not an observation: 3 × (512 + 11) for `url`,
`useragent` and `referer`, plus 128 + 11 for `request_id`, plus 32 + 11
for `method`, plus a 336-byte fixed portion (the field names, the
punctuation, both timestamps at their longest, an IPv6 `remoteIP` with
a zone, the status and the latency) — 2,087 bytes, stated at 2,560 so
the figure has headroom. `internal/middleware/accesslog_test.go`
asserts it against 8 KB of client-chosen text in the path, in the
query, and in each of `User-Agent`, `Referer` and `X-Request-Id`,
including cases built from the characters the handlers escape, and
against the widest access log line the service can be made to write: a
5xx that keeps its concrete path while all three header fields are also
at their budget. Every case runs through both handlers
`internal/logger` can select — the JSON one and the text one it installs
on a tty — since the two do not escape alike and the ceiling is quoted
unqualified. Measured over a real connection, the widest access log line
is 1,972 bytes.
Multiply that ceiling by the request rate to size log storage. Note
that the rate is not bounded by the limits above on every route:
`/.well-known/healthcheck` and `/s/*` sit behind no limiter, so there
the multiplier is whatever the deployment will serve.
**The same ceiling covers every other line the service writes through
`slog` that carries text an unauthenticated client supplies**, with one
exception stated below it: the recovered-panic record, which carries a
whole goroutine stack alongside its client-supplied fields and so has
its own wider ceiling. The access log is not the only line a client can
put its own text into, and a budget that held for one line and not the
others would be worse than no stated budget at all. Every `slog` call an
unauthenticated request can reach spends the same per-field budget
through `internal/logfield`, and each carries strictly fewer
client-supplied fields than the access log does, so none of them can be
wider than it:
| Log line | Level | Client-chosen value | Reachable unauthenticated |
| ------------------------------------------ | ------- | ------------------- | ----------------------------------------- |
| `request body exceeds limit` (413) | `WARN` | path, method | yes — `MaxBodySize` precedes `RequireAuth` |
| `csrf: token validation failed` (403) | `WARN` | path, method | yes — `CSRF` precedes `RequireAuth` |
| `... rate limit exceeded` (429) | `WARN` | path | yes, on the receiver |
| `auth middleware: unauthenticated request` | `DEBUG` | path, method | yes, by definition |
| `entrypoint not found` | `DEBUG` | entrypoint UUID | yes, on the receiver |
| `user not found` / `invalid password` | `DEBUG` | username | yes, on the login form |
| `login failure limit exceeded` (429) | `WARN` | path | yes, on the login form |
| `password verification capacity exhausted` | `WARN` | path | yes, on the login form |
`DEBUG` being off by default is not a bound. An operator turning it on
to diagnose a flood must not thereby hand the flood an unbounded write,
so those lines are capped too.
The last two rows are capped defensively rather than against a
demonstrated width: chi routes `POST /pages/login` on a static pattern,
so `r.URL.Path` there is the 12-byte constant `/pages/login` and each
line lands near 120 bytes. `RecordLoginFailure` is nonetheless an
exported method taking any `*http.Request`, and a future caller on a
route with a URL parameter would widen the line. Since no request
through the mux can, both caps are pinned by tests that call those two
entry points directly with the path such a caller would supply.
Removing either cap fails 14 subtests.
`internal/middleware/logbound_test.go` and
`internal/handlers/logbound_test.go` drive 8 KB of client-chosen text
at each of these — 1 KB at `invalid password`, whose accounts are
shared with the successful-login line, where a username past 4 KB
overflows the session cookie and answers 500 before that line is
written — through both handlers, and through seven fills: plain text
as the baseline, and then the quotation mark, backslash, tab, newline,
C0 control and astral non-printable, six characters the wider of the
two handlers spends more on than the client spent sending them. Every
case holds each line to the 2,560-byte ceiling. That per-line ceiling
is what the figure above states, and every row establishes it.
Three of the sites go further and bound the whole flood's output — the
total bytes a run of distinct invented values wrote, which is the
shape an operator sizing storage cares about. They are
`request body exceeds limit`
(`TestMaxBodySize_FloodOfOversizePathsDoesNotGrowTheLog`),
`entrypoint not found` and `user not found` (the last two through
`assertBoundedFlood`). The other rows carry no aggregate assertion;
the per-line ceiling is what they establish.
`internal/logfield/logfield_test.go` measures the per-rune charge
against what the handlers really emit, over roughly 3,000 code points on
each, so an undercharged rune fails a test rather than quietly
falsifying the ceiling.
**It covers GORM's statement logging as well.** GORM's own default
logger printed the fully interpolated SQL — parameters and all — to
standard output on every statement that returned an error, including a
plain record-not-found, at a level no operator setting reached. Two of
this service's lookups miss by design on unauthenticated routes: the
entrypoint lookup behind `/webhook/{uuid}` and the user lookup behind
the login form, whose path segment and submitted username the client
picks outright. Every
`gorm.Open` in the service now installs the adapter in
`internal/gormlog` instead. It writes through the same `slog` logger as
everything else, so its lines take the level the operator set and the
handler `internal/logger` selected, and every value it emits is spent
through the same `internal/logfield` budget. A record-not-found is not
logged as an error: it is the expected outcome on both of those paths,
and each handler already records its own miss at `DEBUG` — bounded, per
the table above — without the SQL. Slow statements are kept, at `WARN`,
above the same 200 ms threshold GORM used and with the statement
bounded, because that report is the one thing GORM's logger gave an
operator that nothing else here does. The adapter orders its cases
exactly as GORM's own `Trace` orders them, so a statement that both
missed and ran slow is still reported as slow, and dropping the miss
costs an operator no report `IgnoreRecordNotFoundError` would have
kept. A GORM line spends at most two of those budgets — the statement
and the driver error — against a smaller fixed portion than the access
log's, and `internal/gormlog/gormlog_test.go` asserts each line against
`MaxAccessLogLineBytes` directly rather than leaving it as arithmetic.
What that ceiling does **not** cover, stated here so the figure is not
read as more than it is:
- **Lines carrying an authenticated operator's own input**, which are
not truncated at all. `webhook created` logs the submitted `name`
verbatim and `target URL blocked by SSRF protection` logs the target
host (both `internal/handlers/source_management.go`), as do the
`target_name` lines in `internal/delivery/engine.go` and
`internal/delivery/target_http.go`. The only bound on any of them is
the 1 MB form body cap, so a 100 KB `name` writes a single line of
roughly 600 KB — measured. This is deliberate: every one of these
requires an authenticated operator on a service with no
self-registration, and truncating the operator's own configuration
echoed back would cost debuggability against no adversary. It does
mean the 2,560-byte figure sizes unauthenticated traffic, not the
operator's own administrative requests.
- The **`log` delivery target**, which writes the whole inbound event —
headers and body — to the log. This one is deliberate: capping it
would defeat the target, since emitting the payload is the delivery.
It costs nothing unless an authenticated operator creates a target of
that type on a specific webhook, and each line it writes is bounded
per event by the 1 MB receiver body cap. Adding one is a decision to
spend log volume on that webhook's payloads.
- **Two writers that do not go through `internal/logger` at all**, both
on standard error. `fx` prints the dependency graph and the lifecycle
hooks through its default console logger at startup and shutdown —
nothing calls `fx.WithLogger`, and `fx.New` builds that logger over
`os.Stderr`. The Go runtime writes a panic or a fatal error itself; a
panic in a background worker rather than in a request handler is the
case that reaches it, since nothing recovers those. Neither carries a
client-chosen value at a client-chosen length: the five `panic` calls
in this service are invariant guards over constants and over
`crypto/rand`.
- **`net/http`'s own faults**, which are _not_ a separate writer.
`internal/server/http.go` builds its server with a nil `ErrorLog`, so
`net/http` falls back to the `log` package's default logger — and
`internal/logger` calls `slog.SetDefault`, which redirects that logger
into whichever handler it installed. Those lines therefore arrive on
standard output, shaped like every other line, at `INFO`. They are not
truncated. A handler panic is no longer one of them: the recover
middleware below answers it and writes it as the bounded record
described there instead, and `internal/server/recoverer_test.go`
requires that `http: panic serving` appear in neither of the process's
two streams when a panic is driven through the production router. The
one panic still handed back to `net/http` is `http.ErrAbortHandler`,
which it special-cases and does not log at all. What is left on this
path is `net/http`'s own diagnostics, whose values are the runtime's,
not a client's.
Wider than that 2,560-byte ceiling, and stated separately rather than
carved out of it: the record a recovered panic produces. The recover
middleware in `internal/middleware` answers `500` and writes one `ERROR`
record through `internal/logger` carrying the panic value, the stack and
the request id — the same `request_id` the access log line for that
request carries, which is how the two are joined. It replaced chi's
`middleware.Recoverer`, which on a current Go release crashed inside its
own stack pretty-printer: the connection was dropped rather than
answered, and what reached the operator described that crash rather than
the fault behind it.
That record is bounded the same way, in the same encoded bytes and
through the same `internal/logfield` budget: 512 for the panic value,
because a handler is free to build one out of the request, 128 for the
request id, which a client supplies outright through `X-Request-Id`,
and 8,192 for the stack, cut at its far end so that the panic site
survives a cut and `net/http`'s accept frames are what is lost. Net:
**at most 10,240 bytes, once per recovered panic** — 9,121 by the
arithmetic (523 + 8,203 + 139 + a 256-byte fixed portion), stated at
10,240 for headroom.
Those two numbers are the claim; the measurements below only
illustrate it. `internal/middleware/recoverer_test.go` drives all
three growable fields past their budgets on one record, over both
handlers, and measured 9,009 bytes on the JSON handler and
8,9828,983 on the text one in one checkout. Neither is an invariant:
the stack's own content decides where its cut lands, so the figures
move by a byte or so between runs. The real case is far below both —
through the shipped middleware chain the whole record measures
roughly 3,960 bytes over a roughly 3,690-byte stack, taken by
`internal/server/recoverer_test.go` from the process's own file
descriptors while driving a panic through the production router over
a real server in a subprocess. That pair moves further still, since
`debug.Stack()` embeds absolute source paths and so depends on where
the tree is checked out: four checkouts have reported 3,959, 3,961,
3,984 and 4,026. What the tests assert is the ceiling, that every
client-supplied field was cut, and that the shipped chain's stack
arrived uncut — never the numbers.
the aggregate limit is what bounds those `WARN` lines — to under ten
times `RECEIVER_RATE_LIMIT` per minute per client IP, 1080 at the
defaults, where before it there was no bound at all. The access log is
bounded by neither limit: every request is recorded once at `INFO` with
its full URL, served or rejected alike.
Every limiter here — receiver, login, and password change — identifies
the client the same way, through one shared key function: the
connection's own address, unless the peer is listed in
`TRUSTED_PROXIES`, in which case the forwarded client address is used
instead. That address becomes a bucket by family: IPv4 keys on the full
address, IPv6 on its `/64` prefix. A routed `/64` is the normal
residential and mobile IPv6 allocation, so keying IPv6 per address would
let one subscriber rotate source addresses and mint a fresh bucket per
request, evading these limits at the network layer without spoofing
anything; the cost is that distinct clients inside one `/64` share a
bucket. IPv4-mapped addresses (`::ffff:1.2.3.4`) key as the IPv4 address
they carry. See [Trusted proxies](#trusted-proxies). Deployed without that
instead. See [Trusted proxies](#trusted-proxies). Deployed without that
variable set, a client behind a reverse proxy shares one bucket with
every other client behind the same proxy. Set `TRUSTED_PROXIES` to the
proxy's address to get per-client limits back. What the shared bucket
@@ -1360,117 +1018,14 @@ opposite directions:
and all entrypoints, where the per-entrypoint limit's capacity still
grows with the number of entrypoints. Any deployment with more than a
handful of busy entrypoints must set `TRUSTED_PROXIES`.
- For the **login and password-change** limits it costs precision, not
availability. Login failures from every client land in one counter,
so a stranger's wrong passwords make the operator's own wrong
passwords answer `429` sooner; the operator's _correct_ password is
never affected, because it is never counted. Production deployments
should still set `TRUSTED_PROXIES`; webhooker warns at startup
whenever it is empty, in any environment.
#### The login endpoint
The login `POST` is the one endpoint with no pre-emptive limiter in
front of it, and that is deliberate. A limiter that spends budget on
arrival is a lockout in this deployment shape: sharing one bucket, a
stranger sending five POSTs a minute — about 0.08 requests per second,
from anywhere — keeps it permanently full, and the operator has no
second administrative path. So the handler inverts the order:
1. **Credentials are verified first, and only a failed attempt spends
budget.** A correct password is never rate-limited, whatever the
counters hold. This is what guarantees the admin UI stays
reachable.
2. **Failures are counted per (client bucket, submitted username)**,
five per minute, after which further _failures_ from that pair are
answered `429` with a `Retry-After`. That `429` is a label on the
response, not a gate in front of the work: the credential check has
already run by the time the counter is consulted, so a throttled
client's guess is still evaluated. See the guessing rate below. A
successful login clears the counter, so mistyping a few times and
then getting it right leaves you unthrottled. Because the submitted
username is attacker-controlled, at most 1024 username counters and
1024 fallback address counters are tracked; past the first cap
failures fall back to the address counter, and past both they are
answered as throttled without being recorded. Total tracked state
is under half a megabyte and does not grow with the number of
usernames an attacker invents.
3. **Concurrent password verifications are capped at two, and the
queue for them at 16.** Verifying before counting means every login
request costs an Argon2id hash, and Argon2id here is 64 MB per
hash — two slots is a 128 MB ceiling on password hashing. Every
endpoint that hashes a password takes a slot, including the
password-change endpoint, which holds one across both the
verification and the new hash. A request that waits five seconds
without getting a slot is answered `503 Service Unavailable` and no
hash is computed for it. The wait alone does not bound memory, only
how long one request holds some, so the number of waiters is capped
as well. Size the queue from what a parked waiter actually retains,
not from the 1 MB body cap: that caps the raw body read, while the
body-cap, CSRF and form-parsing middleware all run before the
guard, so a waiter holds its parsed form plus its request header
block for the whole wait. Measured on the pinned Go 1.26.1
toolchain, as the heap delta with 64 waiters parked in the handler,
an ordinary two-field login form retains ~0 MB, a 1 MB urlencoded
body at Go's 10,000-parameter parse cap retains 2.82 MB (3.09 MB
with `%41` escapes), and the ~0.9 MB of headers the 1 MB header cap
allows takes it to **4.18 MB** — the retained parse and the headers
dominate, not the raw body. So the cap is 16 waiters: 16 x 4.18 MB
is about 67 MB of committed queue memory, and two slots drain a
full 16-deep queue in roughly 0.6 s, far inside the five-second
deadline. A request arriving past the cap is shed with `503`
immediately instead of joining the queue. **Peak commitment for the
endpoint is therefore about 203 MB**: 128 MB of Argon2id, plus the
18 requests holding a parsed form — 16 queued and the 2 being
hashed — at about 75 MB. That 203 MB is _live_ commitment, not
resident size: the Go collector lets the heap reach roughly twice
the live set before collecting, with transient parse garbage on top
of it. The independent review of this endpoint fired 18 adversarial
requests at an idle guard and measured a peak `HeapAlloc` of
392 MB. **Provision on the order of 400 MB**, not for the 203 MB
itemised here and not for the hashing budget alone.
An unknown username is verified against a dummy hash rather than
rejected early, so a nonexistent account costs the same time as a real
one and the response cannot be used to enumerate usernames.
**This raises online guessing throughput by about 300x, and that is
the trade.** Because the credential check always precedes the counter,
what bounds online brute force is the semaphore, not the failure
counter. Two slots at the cost of one Argon2id verification is on the
order of **27 guesses per second, about 2.3 million per day**, against
5 per minute under the pre-emptive limiter this replaced. Treat that
figure as a lower bound rather than a ceiling: it was measured with
Go's race detector enabled, so real hardware verifies faster and
guesses faster. Choose the admin password to survive millions of
online guesses per day — a long random passphrase, not a memorable
one. Rate-limiting `POST /pages/login` at the reverse proxy, where the
real client address is visible, is the way to put a cheaper bound back
on top.
The residual exposure is a bounded, self-clearing loss of login
**availability** — not merely of latency. A flood can keep both
verification slots busy, and a request that neither gets a slot within
five seconds nor finds room in the queue is answered `503`. Above
roughly 27 requests per second the operator is not served slowly, it
is shed: its chance per attempt is about the ratio of service rate to
flood rate, so at 400 requests per second it is roughly one attempt in
fourteen. A sufficiently determined flood still denies login for as
long as it runs.
What changed is the price and the aftermath. Denying login used to
cost an attacker 0.08 requests per second from anywhere; it now costs
30 or more sustained, about 400 times as much. Nothing accumulates
while the flood runs, nothing needs resetting when it stops, and the
operator's correct password succeeds on the first attempt afterwards.
Restarting the service is **not** a remedy: a restart clears the
failure counters, which are not what is saturated, and the flood
re-fills both verification slots on its first two requests. The
remedies are to block the source at the reverse proxy, or to
rate-limit `POST /pages/login` there — the one place a limit can be
applied without reintroducing the lockout, because the proxy sees the
real client address. Setting `TRUSTED_PROXIES` does not stop the
saturation, but it makes the source visible in the failure logs.
- For the **login and password-change** limits it costs availability of
the only administrative path, which is not safe at all. Five POSTs
per minute from any address on the internet keeps the single shared
login bucket full, and the operator's own login returns HTTP 429 for
as long as that trickle continues. A restart clears the in-memory
buckets and a resumed trickle re-locks them. Production deployments
must set `TRUSTED_PROXIES`; webhooker warns at startup whenever it is
empty, in any environment.
Finer-grained per-webhook rate limits (configured in the web UI and
enforced in the webhook handler) can layer on top of this env-level
@@ -1491,8 +1046,8 @@ abuse limit later; they are tracked as future work.
| Method | Path | Description |
| ------ | --------------- | ----------- |
| `GET` | `/pages/login` | Login page (not rate limited) |
| `POST` | `/pages/login` | Login form submission. Credentials are verified before any limit is consulted, so a correct password is never throttled; 5 FAILED attempts per minute per bucket per submitted username, then `429`. `503` if no verification slot frees up within 5s, or immediately if 16 requests are already queued for one (see [Rate Limiting](#rate-limiting)) |
| `GET` | `/pages/login` | Login page (not rate limited; the limiter applies to POST only) |
| `POST` | `/pages/login` | Login form submission (5 per minute per bucket, then 429) |
| `POST` | `/pages/logout` | Logout (destroys session) |
#### Authenticated Endpoints
@@ -1500,7 +1055,7 @@ abuse limit later; they are tracked as future work.
| Method | Path | Description |
| ------ | ------------------------ | ----------- |
| `GET` | `/user/{username}` | User profile page |
| `POST` | `/user/{username}/password` | Change the user's password (5 per minute per bucket, then `429`; `503` if no verification slot frees up within 5s, or immediately if 16 requests are already queued for one) |
| `POST` | `/user/{username}/password` | Change the user's password (5 per minute per bucket, then 429) |
| `GET` | `/sources` | List user's webhooks |
| `GET` | `/sources/new` | Create webhook form |
| `POST` | `/sources/new` | Create webhook submission |
@@ -1570,10 +1125,6 @@ webhooker/
│ │ └── webhook_db_manager.go # Per-webhook DB lifecycle manager
│ ├── globals/
│ │ └── globals.go # Build-time variables (appname, version, arch)
│ ├── gormlog/
│ │ └── gormlog.go # GORM's logger.Interface on top of slog, bounded
│ ├── logfield/
│ │ └── logfield.go # Encoded-byte budget for client-supplied log values
│ ├── delivery/
│ │ ├── engine.go # Event-driven delivery engine (channel + timer based)
│ │ ├── circuit_breaker.go # Per-target circuit breaker for http/slack targets with retries
@@ -1606,7 +1157,6 @@ webhooker/
│ │ ├── middleware.go # Logging, CORS, Auth, Metrics, MetricsAuth, SecurityHeaders, MaxBodySize
│ │ ├── csrf.go # CSRF protection middleware (gorilla/csrf)
│ │ ├── ratelimit.go # Per-IP rate limiting middleware (go-chi/httprate)
│ │ ├── loginguard.go # Login failure counters and the Argon2id verification semaphore
│ │ └── testing.go # NewForTest: Middleware without the fx lifecycle
│ ├── server/
│ │ ├── server.go # Server struct, fx lifecycle, signal handling
@@ -1626,7 +1176,6 @@ webhooker/
├── templates/ # Go HTML templates (base, login, sources, etc.)
├── script/ # Scripts to Rule Them All entrypoints
├── Dockerfile # Three stages: lint, test+build, Alpine runtime
├── Dockerfile.lint # Lint-only image built by script/lint
├── Makefile # 10 of 16 targets shim script/; 6 are inline
├── go.mod / go.sum
└── .golangci.yml # Linter configuration
@@ -1669,32 +1218,19 @@ to record results.
Applied to all routes in this order:
1. **RequestID**Generate unique request IDs (chi built-in)
2. **SecurityHeaders** — Production security headers on every response
1. **Recoverer**Panic recovery (chi built-in)
2. **RequestID** — Generate unique request IDs (chi built-in)
3. **SecurityHeaders** — Production security headers on every response
(HSTS, X-Content-Type-Options, X-Frame-Options, CSP, Referrer-Policy,
Permissions-Policy)
3. **Logging** — Structured request logging (method, URL, status,
4. **Logging** — Structured request logging (method, URL, status,
latency, remote IP, user agent, request ID)
4. **Metrics** — Prometheus HTTP metrics (if `METRICS_USERNAME` is set)
5. **CORS** — Cross-origin resource sharing headers
6. **Timeout** — 60-second request timeout
7. **Recoverer** — Panic recovery: one `ERROR` record through
`internal/logger` and a `500`
5. **Metrics** — Prometheus HTTP metrics (if `METRICS_USERNAME` is set)
6. **CORS** — Cross-origin resource sharing headers
7. **Timeout** — 60-second request timeout
8. **Sentry** — Error reporting to Sentry (if `SENTRY_DSN` is set;
configured with `Repanic: true` so panics still reach Recoverer)
Recoverer sits seventh rather than first, and both neighbours are the
reason. It runs **inside** everything that observes the response, so
the `500` it writes for a panicking handler is the status the access
log records and the metrics count; registered first, as chi's own
`middleware.Recoverer` was, the same request was logged as a `200` that
the client never received. It runs **outside** the Sentry handler, so
`Repanic: true` has something to re-raise into: an operator with
`SENTRY_DSN` set keeps the report, and one without it now gets the
local record instead of nothing. What that placement gives up is
recovery of a panic in the six entries above it, none of which does
more than set a header or start a timer.
Additionally, form endpoints (`/pages`, `/user/*`, `/sources`,
`/source/*`) apply a **MaxBodySize** middleware that limits
POST/PUT/PATCH request bodies to 1 MB. It is registered ahead of the
@@ -1717,11 +1253,9 @@ one that lies about its length, is hard-capped by
Those same four route groups then apply **CSRF** and **NoCache**
(`Cache-Control: no-store`, `Pragma: no-cache`), and every group except
`/pages` applies **RequireAuth**. The rate limiters are per-route
rather than global: **PasswordChangeRateLimit** on
`/user/{username}/password` and **ReceiverRateLimit** on
`/webhook/{uuid}`. There is deliberately none on `/pages/login` — that
endpoint counts failures inside the handler, after the credential
check, see [The login endpoint](#the-login-endpoint).
rather than global: **LoginRateLimit** on `/pages/login`,
**PasswordChangeRateLimit** on `/user/{username}/password`, and
**ReceiverRateLimit** on `/webhook/{uuid}`.
### Authentication
@@ -1759,27 +1293,14 @@ check, see [The login endpoint](#the-login-endpoint).
both at target creation time (URL validation) and at delivery time
(custom HTTP transport with SSRF-safe dialer that validates resolved
IPs before connecting, preventing DNS rebinding attacks)
- **Login limiting is inverted, deliberately.** The login `POST` has
no pre-emptive rate limiter in front of it. Credentials are
verified first and only a _failed_ attempt spends budget, so a
correct password is never throttled and no flood of wrong ones can
deny the operator the only administrative path. Failures are
counted per (bucket, submitted username), five per minute, after
which further failures are answered `429` with a `Retry-After`.
What bounds brute force is not that counter but the cap of two
concurrent Argon2id verifications: a throttled client's guess is
still evaluated, so roughly 27 guesses a second get through and the
admin password has to carry that load (see
[The login endpoint](#the-login-endpoint)). `GET` requests to the
login page are not limited
- **Password-change rate limiting** via [go-chi/httprate](https://github.com/go-chi/httprate):
sliding-window rate limiter, 5 POST attempts per minute per bucket.
It runs behind session auth, so only a client already holding a
valid session reaches it, and an operator throttled out of changing
a password can still log in. The bucket is per client IP only when
- **Login rate limiting** via [go-chi/httprate](https://github.com/go-chi/httprate):
sliding-window rate limiter on the login endpoint, 5 POST attempts
per minute per bucket, to slow brute-force attacks. GET requests to
the login page are not limited. The password-change endpoint carries
the same 5-per-minute limit. The bucket is per client IP only when
`TRUSTED_PROXIES` names the reverse proxy; unset, every client
shares one bucket, which costs precision rather than availability
(see [Rate Limiting](#rate-limiting)). webhooker warns at startup
shares one bucket and the login becomes remotely deniable (see
[Rate Limiting](#rate-limiting)). webhooker warns at startup
whenever `TRUSTED_PROXIES` is empty
- Prometheus metrics behind basic auth
- Static assets embedded in binary (no filesystem access needed at
@@ -1860,37 +1381,6 @@ Two operational consequences follow from bounding the sequence:
no shutdown diagnostics at all. Keep the deployment's grace above
the stop timeout.
### Linting
golangci-lint never runs on the host. `script/lint` builds
`Dockerfile.lint`, which copies the repo into the digest-pinned
golangci-lint image and lints as a build step, so a successful build is
a clean lint. A host binary would share one cache and one lock with
every other checkout on the machine, which has produced both invented
findings attributed to other worktrees and unearned passes.
Three properties are load-bearing:
- `script/lint` passes `--no-cache-filter=lint`. Without it an unchanged
tree replays the lint layer from cache and the build exits 0 in under
a second having linted nothing. The `deps` stage stays cacheable, so
module downloads are not repeated. Invalidation is scoped to the one
stage; never prune the shared build cache.
- `script/lint` does not trust that flag. Docker silently ignores
`--no-cache-filter` for a stage name that does not match, so a stage
rename or a one-character typo would restore the cached false green
with no warning and a fast exit 0. The script therefore tees the
build output and treats a run as a pass only if golangci-lint's own
summary line (`N issues.` / `N issues:`) appears in it: no summary,
no lint, whatever the exit code says.
- Both lint steps use `RUN --network=none`. `golangci-lint config
verify` is documented as fetching its JSON schema over HTTPS, which
would be an unpinned remote dependency; the pinned image resolves the
schema without network access, and `--network=none` enforces that
instead of trusting it. Verify is worth keeping because
`golangci-lint run` silently ignores config keys it does not
recognize, so a typo would disable a setting with no warning.
### Docker
The Dockerfile uses a three-stage build. Each stage is pinned by
@@ -1899,8 +1389,7 @@ version is fixed independently of the compiler's:
1. **Lint stage** (`golangci/golangci-lint:v2.12.2`, Debian-based) —
installs `make`, downloads dependencies, copies the source, and runs
`make fmt-check`, then `golangci-lint config verify` and
`golangci-lint run`, both with `--network=none`.
`make fmt-check` then `make lint`.
2. **Builder stage** (`golang:1.26.1-bookworm`) — depends on the lint
stage passing (it copies a file from it), runs `script/fetch-assets`
to download and verify the third-party browser assets, then runs
@@ -1911,21 +1400,20 @@ version is fixed independently of the compiler's:
runs as the non-root `webhooker` user (UID 1000), exposes port 8080,
and includes a health check against `/.well-known/healthcheck`.
The lint stage invokes `golangci-lint` directly rather than `make lint`:
it is already the pinned linter image, and `make lint` builds
`Dockerfile.lint`, which would need a docker daemon inside this build.
Both check stages use Debian rather than Alpine because
`gorm.io/driver/sqlite` pulls in `mattn/go-sqlite3`, which needs CGO
and does not compile against musl. Only the final binary is statically
linked, which is what lets it run on the Alpine runtime image.
`script/cibuild` — `docker build .` — is the CI gate: the checks run
inside the image, so a build that succeeds is a repo that is formatted,
linted, tested and compiled. `script/lint` also uses Docker
(`Dockerfile.lint`, see Linting above), so `make lint` and `make check`
run the same pinned linter version the gate does; only `script/test`
and `script/fmt-check` run on the host.
`script/cibuild``docker build .` — is the CI gate: the four check
targets run inside the image, so a build that succeeds is a repo that
is formatted, linted, tested and compiled. Only `script/cibuild` and
`script/docker` involve Docker. `script/lint`, and therefore
`make lint` and `make check`, run whatever `golangci-lint` is on the
host, which can be a different version from the pinned one — so the
container is the authoritative lint result
([issue #109](https://git.eeqj.de/sneak/webhooker/issues/109) tracks
routing local linting through it as well).
#### CI gate honesty
@@ -1938,8 +1426,8 @@ the hash of the last commit that touched the build context, so:
- Any commit that changes code (including a squash merge whose tree
matches an already-built branch) gets a new fingerprint, invalidates
the `COPY . .` layer of both check stages, and really runs
`make fmt-check`, `golangci-lint`, `make test`, and `make build`. A
run that reports success ran them.
`make fmt-check`, `make lint`, `make test`, and `make build`. A run
that reports success ran them.
- A docs-only commit leaves the fingerprint unchanged — `.dockerignore`
excludes `*.md`, `LICENSE` and `.editorconfig` from the context
anyway — so the image replays from cache and costs seconds.
@@ -1950,32 +1438,11 @@ way.
A separate workflow step, run before the fingerprint is written, covers
a second way the gate lied: Gitea cancels an in-flight run when a newer
commit lands on the same branch and records that cancellation as a
`failure` status, so a commit nothing ever tested reads as a test
result. Cancellation is unconditional server-side for push events, so
the superseding run calls `script/ci-mark-superseded`, which rewrites
that exact status to `failure` /
`Superseded by a newer commit; never tested`.
The state stays `failure` on purpose: Gitea's combined status folds
`skipped` into `success`, so marking a never-tested commit `skipped`
made the status API report green for it, indistinguishable from a commit
that passed. Reading a commit's status on this repo therefore goes:
- `success` / `Successful in ...` — the checks ran and passed.
- `failure` / `Failing after ...` — the checks ran and failed.
- `failure` / `Superseded by a newer commit; never tested` — the run was
cancelled, by a newer push or by hand, and nothing was verified about
this commit. Test the commit itself before concluding anything about
it.
Genuine failures and successes are never touched, and no status is left
`pending`, which would block the commit indefinitely. The step derives
its context string from the workflow name, the job **id** and the event.
That is deliberately not byte-identical to Gitea's own rule, which uses
the job's display `name:` where the runner exports the id, so giving the
job a `name:` — or renaming the workflow — makes the derived context
stop matching. The step fails loudly when no status on the commit
carries that context, so no rename can silently disable the rewrite.
`failure` status, marking a commit red that was never tested.
Cancellation is unconditional server-side for
push events, so the superseding run rewrites the exact
`Has been cancelled` status to `skipped`. Genuine failures are never
touched.
## TODO

View File

@@ -1,6 +1,6 @@
---
title: Repository Policies
last_modified: 2026-08-07
last_modified: 2026-07-06
---
This document covers repository structure, tooling, and workflow standards. Code
@@ -189,13 +189,8 @@ style conventions are in separate documents:
module under test to verify it compiles/parses. There is no excuse for
`make test` to be a no-op.
- `make test` must complete in under 60 seconds. That is the hard cap, and a
suite that exceeds it fails. Under 20 seconds is the target. A suite between
20 and 60 seconds is still green, but the overage must be filed as an
improvement bug against that repo. Add a 90-second timeout to the test
invocation in the Makefile (`go test -timeout 90s`). The backstop deliberately
sits above the hard cap so that it catches a genuinely hung test rather than a
merely slow one.
- `make test` must complete in under 20 seconds. Add a 30-second timeout in the
Makefile.
- **`make test` should use the conditional verbose rerun pattern.** Run tests
without `-v` (verbose) first. If tests fail, automatically rerun with `-v` to
@@ -214,9 +209,9 @@ style conventions are in separate documents:
```makefile
test:
@go test -timeout 90s -race -cover ./... || \
@go test -timeout 30s -race -cover ./... || \
{ echo "--- Rerunning with -v for details ---"; \
go test -timeout 90s -race -v ./...; exit 1; }
go test -timeout 30s -race -v ./...; exit 1; }
```
Python example:
@@ -265,10 +260,7 @@ style conventions are in separate documents:
- `.golangci.yml` is standardized and must _NEVER_ be modified by an agent, only
manually by the user. Fetch from
`https://git.eeqj.de/sneak/prompts/raw/branch/main/.golangci.yml`. The
canonical golangci-lint version is v2.12.2 (released 2026-05-06), installed
commit-pinned via
`go install github.com/golangci/golangci-lint/v2/cmd/golangci-lint@c0d3ddc9cf3faa61a4e378e879ece580256d76e5`.
`https://git.eeqj.de/sneak/prompts/raw/branch/main/.golangci.yml`.
- When pinning images or packages by hash, add a comment above the reference
with the version and date (YYYY-MM-DD).

140
TODO.md
View File

@@ -24,142 +24,30 @@ event retention (#63), the database archiving target (#43), the admin
password change flow (#65), policy compliance (#6), pinned lint tooling
(#55), and fail-loud configuration parsing (#80).
`next` holds the **complete 1.0.0 milestone**: every issue in it is
closed, and it is verified green both by CI and by cache-defeated
container runs (`docker build --no-cache-filter=lint
--no-cache-filter=builder`).
One caveat on reading a green check, narrower than it used to be. A
docs-only commit deliberately replays from the layer cache (#119), so a
green status on such a commit evidences a replay rather than an executed
run; a code commit invalidates the `COPY` layer and genuinely executes.
Superseded runs are no longer the hazard they were: before #152 they
were recorded as `skipped` and rolled up green, and before #119 a warm
layer cache let the gate report success without executing anything,
replaying the previous build's console log so the lie looked like a real
run. Both are fixed. Note: `TODO.md` was deliberately
`next` holds the completed 1.0.0 milestone: every issue in it is closed,
and it is verified green by cache-defeated container runs
(`docker build --no-cache-filter=lint --no-cache-filter=builder`). The
CI status is not independently claimed here: a superseded run is
recorded as `skipped` and still rolls up green, so a commit status on
`next` does not by itself evidence an executed check (#152). Before
#119, a warm layer cache also let the gate report success without
executing anything, and replayed the previous build's console log so
the lie looked like a real run. Note: `TODO.md` was deliberately
deleted from this repo in f9a9569 (2026-03-01, #6); its content was
folded into the README TODO section, which this draft reconstructs as
of 2026-07-06.
# Next Step
Merge the milestone PR (#111) to `main` and tag 1.0.0 from it. The
milestone is empty and `next` is green; nothing else blocks the tag.
Merge the milestone PR to `main` and tag 1.0.0 from it.
Three items belong to the owner, none of them blocking. #150 was decided
by the manager rather than left to stall the queue and is flagged on the
issue for reversal if that call was wrong. #112 (whether `Completed
Steps` should exist at all, given it once conflicted on every unit) is
unanswered; the provisional ruling in force is that issue branches do
not touch this file. #198 records that `make test` is past the org 20s
target — 46s of test execution inside a 62.8s CI layer — and turns on
which quantity the 60s hard cap governs; it is scoped as the improvement
bug the 20-60s band requires, and should be milestoned instead if the
cap is read as covering the whole invocation.
After the tag, the largest open cluster is the unmilestoned follow-up
backlog these units generated: #183, #184, #185, #190, #191, #193 and
#198.
Two decisions are open and belong to the owner, neither blocking the
tag: #115 (mask the `http` target's destination URL, implemented
speculatively and awaiting a yes or no) and #125 (whether IPv6
rate-limit keys should bucket by `/64`).
# Completed Steps
- 2026-08-18 Raise `script/test`'s per-package timeout from 30s to 90s,
matching the org-wide backstop. `go test` applies `-timeout` per
package, and `internal/handlers` had grown past the old budget: a
cache-defeated build failed outright at `GOMAXPROCS=4`, and every run
under deliberate host load breached 30s. The measurement table lives
in the script (#194)
- 2026-08-18 Re-sync `REPO_POLICIES.md` from `prompts`. The local copy
was stale and still mandated a 20s test target with a 30s timeout,
which the org replaced with a 60s cap and a 90s backstop. A synced
copy is not a source; reading it as one nearly produced a PR against
`prompts` proposing a change already merged there (#196)
- 2026-08-18 Report handler panics through the logger and answer 500.
chi v1.5.5's `Recoverer` scans for a `panic(0x` frame the runtime no
longer emits, then indexes `pkg[-1:]`, so it panicked inside its own
stack printer before writing a byte: the recovery never ran, the
client got a dropped connection instead of a 500, and the original
panic was lost. A local middleware replaces it, bounded by
`MaxPanicLogLineBytes` (#187)
- 2026-08-18 Route GORM's logger through `slog` and bound it. Every
`gorm.Open` left `logger.Default` in place at `Warn` with
`IgnoreRecordNotFoundError` false, so **every record-not-found
printed the fully interpolated SQL to stdout** — including the
client-chosen path on `/webhook/{uuid}` and the submitted username on
the login form, at no level the operator set and outside
`internal/logger` entirely. Three call sites, not the two the issue
named (#178)
- 2026-08-18 Bound every `slog` line against client-chosen text. Eight
sites reachable unauthenticated, found by reading every `slog` call in
the tree rather than only the one reported; the budget moved to a
shared `internal/logfield` so no second truncation exists. `DEBUG`
being off by default is not a bound and is not treated as one (#176)
- 2026-08-18 Stop a slow host turning a login-guard test into a
segfault. A non-fatal `assert` on an acquire result was dereferenced
on the next line, so one timing miss killed the whole
`internal/middleware` binary and reddened CI for unrelated PRs. The
fix also removed a real production race — `acquire` could shed a
request with a slot standing free, because Go picks uniformly among
ready `select` cases (#186)
- 2026-08-18 Send the chi route pattern to Sentry rather than the
concrete path. The receiver's path carries the entrypoint capability
token, so every Sentry event from `/webhook/{uuid}` shipped a live
credential to a third party. Request `Data`, `QueryString`, `Cookies`
and `Env` are dropped and headers reduced to an allowlist (#179)
- 2026-08-18 Read form fields from the POST body only. `r.FormValue`
merges the query string, so a login could be driven by URL parameters
— putting the password somewhere that lands in access logs, proxy
logs and browser history (#160)
- 2026-08-18 Verify login credentials before spending rate-limit
budget, so a flood of wrong passwords cannot lock out the account it
is guessing at. The manager took this decision rather than stall the
queue; it is flagged on the issue for reversal (#150)
- 2026-08-18 Run all linting in Docker via `Dockerfile.lint`. Host lint
was wrong in both directions from version skew and shared caches.
`script/lint` asserts the summary line, because `--no-cache-filter`
silently ignores a stage name it does not match — the flag that makes
the gate meaningful fails open (#109)
- 2026-08-18 Serve an event's full stored body over HTTP. The list
query truncates for rendering, and that truncated value was the only
way to read a body, so the full payload was unreachable (#157)
- 2026-08-18 Bound the access log line against client-chosen text.
`internal/logfield` budgets by *encoded* bytes, not runes, so a
handler's JSON escaping cannot multiply a field past its allowance
(#146)
- 2026-08-18 Mark superseded CI commits `failure` rather than
`skipped`. A skipped run rolls up green, so a commit that was never
tested reported success (#152)
- 2026-08-18 Set `fx.StopTimeout` inside the container stop grace, so
shutdown hooks are bounded by a deadline the orchestrator will
actually honour rather than being killed mid-flush (#134)
- 2026-08-17 Bucket IPv6 rate-limit keys by `/64`. A single allocation
hands out 2^64 addresses, so per-address keying let one client mint
unlimited buckets. Manager decision, recorded on the issue (#125)
- 2026-08-17 Correct release-blocking README and startup-warning
inaccuracies, including claims about behaviour the code does not have
(#151)
- 2026-08-17 Fetch and verify Alpine.js at build time against
`static/vendor.sha256` instead of committing the minified blob, so
the dependency is pinned by hash rather than by trust (#145)
- 2026-08-17 Bound the event log's rendered bodies in the query itself,
so a large stored payload cannot be read into memory just to be
truncated for display (#135)
- 2026-08-17 Mask the `http` target's destination URL in the UI: it can
carry a bearer credential in its path or query, and was rendered
verbatim. Manager decision to mask unconditionally (#115)
- 2026-08-14 Bound shutdown hooks by their stop context, so a hook that
hangs cannot hold the process past its grace period (#102)
- 2026-08-14 Render templates via a buffer rather than the
`ResponseWriter`, so a template error part-way through cannot commit
a 200 and then fail — the response is written only once it is whole
(#123)
- 2026-08-14 Align the session codec's max-age with the 7-day absolute
cap. The codec accepted cookies the session layer considered expired,
so the cap was enforced in one place and not the other (#108)
- 2026-08-12 Warn when `TRUSTED_PROXIES` is empty in production, where
the safe default silently discards forwarded headers and every client
rate-limits as the proxy's address (#149)
- 2026-08-12 Bound the receiver rate limit per client IP across the
whole `/webhook/*` route. The existing limiter keyed on the request
path and `/webhook/{uuid}` matches any single segment, so a client

2
go.mod
View File

@@ -17,7 +17,6 @@ require (
github.com/stretchr/testify v1.8.4
go.uber.org/fx v1.20.1
golang.org/x/crypto v0.38.0
gopkg.in/yaml.v3 v3.0.1
gorm.io/driver/sqlite v1.5.4
gorm.io/gorm v1.25.5
modernc.org/sqlite v1.28.0
@@ -53,6 +52,7 @@ require (
golang.org/x/text v0.25.0 // indirect
golang.org/x/tools v0.21.1-0.20240508182429-e35e4ccd0d2d // indirect
google.golang.org/protobuf v1.31.0 // indirect
gopkg.in/yaml.v3 v3.0.1 // indirect
lukechampine.com/uint128 v1.2.0 // indirect
modernc.org/cc/v3 v3.40.0 // indirect
modernc.org/ccgo/v3 v3.16.13 // indirect

View File

@@ -1,387 +0,0 @@
package ciscript_test
import (
"maps"
"os"
"os/exec"
"path/filepath"
"slices"
"strings"
"testing"
"github.com/stretchr/testify/require"
"gopkg.in/yaml.v3"
)
const (
// supersededDesc is the description script/ci-mark-superseded
// writes, and the one an earlier revision of it wrote alongside a
// `skipped` state.
supersededDesc = "Superseded by a newer commit; never tested"
// liveContext is the commit-status context Gitea uses for this
// repository's runs, as seen in its API. The script derives it from
// the workflow and job names rather than hardcoding it; the
// derivation is checked against this value below.
liveContext = "check / check (push)"
scriptPath = "../../script/ci-mark-superseded"
workflow = "../../.gitea/workflows/check.yml"
// failure is the only state that neither folds into a combined
// `success` (as `skipped` does) nor blocks the commit forever (as
// `pending` does).
failure = "failure"
)
// repo is a throwaway git history: parent is the commit a run would be
// cancelled on, head the commit that superseded it.
type repo struct {
dir string
head string
parent string
}
// scriptEnv is the run identity the Gitea runner exports and the script
// builds its context string from.
type scriptEnv struct {
workflow string
job string
event string
}
func defaultEnv() scriptEnv {
return scriptEnv{workflow: "check", job: "check", event: "push"}
}
func cancelled() commitStatus {
return commitStatus{
Context: liveContext,
Status: failure,
Description: "Has been cancelled",
}
}
func running() commitStatus {
return commitStatus{
Context: liveContext,
Status: "pending",
Description: "Has started running",
}
}
func TestMarkSuperseded(t *testing.T) {
t.Parallel()
cases := map[string]struct {
parent commitStatus
wantMark bool
}{
"a cancelled run is marked": {
parent: cancelled(),
wantMark: true,
},
"a laundered skipped status is marked": {
parent: commitStatus{
Context: liveContext,
Status: "skipped",
Description: supersededDesc,
},
wantMark: true,
},
"a genuine failure is left alone": {
parent: commitStatus{
Context: liveContext,
Status: failure,
Description: "Failing after 3m1s",
},
wantMark: false,
},
"a passing run is left alone": {
parent: commitStatus{
Context: liveContext,
Status: "success",
Description: "Successful in 2m52s",
},
wantMark: false,
},
"another context is left alone": {
parent: commitStatus{
Context: "other / other (push)",
Status: failure,
Description: "Has been cancelled",
},
wantMark: false,
},
}
for name, tc := range cases {
t.Run(name, func(t *testing.T) {
t.Parallel()
requireTools(t)
history := newRepo(t)
fake, api := newFakeGitea(t)
fake.setStatus(history.head, running())
fake.setStatus(history.parent, tc.parent)
out, err := runScript(t, history, api, defaultEnv())
require.NoError(t, err, out)
posted := fake.postedFor(history.parent)
if !tc.wantMark {
require.Empty(t, posted)
return
}
require.Equal(t, []postedStatus{{
Context: liveContext,
// Not `skipped`: Gitea's combined status folds
// that into `success`, which is what made a
// never-tested commit read green.
State: failure,
Description: supersededDesc,
}}, posted)
})
}
}
// A second run must not rewrite what the first one wrote, or every
// later push would post a duplicate status.
func TestMarkSupersededIsIdempotent(t *testing.T) {
t.Parallel()
requireTools(t)
history := newRepo(t)
fake, api := newFakeGitea(t)
fake.setStatus(history.head, running())
fake.setStatus(history.parent, cancelled())
for range 2 {
out, err := runScript(t, history, api, defaultEnv())
require.NoError(t, err, out)
}
require.Len(t, fake.postedFor(history.parent), 1)
}
// Renaming the workflow or the job changes the context string Gitea
// uses. The script must say so instead of quietly matching nothing.
func TestMarkSupersededRejectsAnUnknownContext(t *testing.T) {
t.Parallel()
requireTools(t)
history := newRepo(t)
fake, api := newFakeGitea(t)
fake.setStatus(history.head, running())
fake.setStatus(history.parent, cancelled())
env := defaultEnv()
env.job = "renamed"
out, err := runScript(t, history, api, env)
require.Error(t, err)
require.Contains(t, out, "renamed")
require.Contains(t, out, liveContext)
require.Empty(t, fake.postedFor(history.parent))
}
// ANCESTOR_LIMIT is a documented knob. A value that is set but unusable
// must abort: handing it to git and discarding the exit status left the
// walk empty and the step green, marking nothing.
func TestMarkSupersededRejectsAnUnparseableAncestorLimit(t *testing.T) {
t.Parallel()
requireTools(t)
history := newRepo(t)
fake, api := newFakeGitea(t)
fake.setStatus(history.head, running())
fake.setStatus(history.parent, cancelled())
out, err := runScript(
t, history, api, defaultEnv(), "ANCESTOR_LIMIT=twenty",
)
require.Error(t, err)
require.Contains(t, out, "ANCESTOR_LIMIT")
require.Contains(t, out, "twenty")
require.Empty(t, fake.postedFor(history.parent))
}
// A status read that fails is not the same as a commit with nothing to
// do. Losing curl's exit status through a pipe made the two identical
// and left a laundered commit laundered with no signal.
func TestMarkSupersededFailsOnAnUnreadableAncestorStatus(t *testing.T) {
t.Parallel()
requireTools(t)
history := newRepo(t)
fake, api := newFakeGitea(t)
fake.setStatus(history.head, running())
fake.setStatus(history.parent, cancelled())
fake.failStatusRead(history.parent)
out, err := runScript(t, history, api, defaultEnv())
require.Error(t, err)
require.Contains(t, out, history.parent)
require.Contains(t, out, "cannot read commit statuses")
require.Empty(t, fake.postedFor(history.parent))
}
// A shallow clone cannot resolve the parent, so it is indistinguishable
// from a root commit to rev-parse and the walk would exit 0 having
// marked nothing. It must abort instead: dropping `fetch-depth: 0` from
// the checkout step is one edit, and a silent no-op there restores the
// false-green bug this script exists to prevent.
func TestMarkSupersededRejectsAShallowRepository(t *testing.T) {
t.Parallel()
requireTools(t)
history := shallowClone(t, newRepo(t))
fake, api := newFakeGitea(t)
fake.setStatus(history.head, running())
fake.setStatus(history.parent, cancelled())
out, err := runScript(t, history, api, defaultEnv())
require.Error(t, err)
require.Contains(t, out, "shallow repository")
require.Empty(t, fake.postedFor(history.parent))
require.Empty(t, fake.postedFor(history.head))
}
// shallowClone returns the same history as a depth-1 clone. The `file://`
// URL is required: git ignores --depth for a plain local path.
func shallowClone(t *testing.T, history repo) repo {
t.Helper()
dir := t.TempDir()
//nolint:gosec // fixed argv, arguments are test-local paths
cmd := exec.CommandContext(t.Context(), "git", "clone", "-q",
"--depth=1", "file://"+history.dir, dir)
out, err := cmd.CombinedOutput()
require.NoError(t, err, string(out))
return repo{dir: dir, head: history.head, parent: history.parent}
}
// The derived context must equal the one Gitea actually uses, which is
// built from the same workflow and job names.
func TestDerivedContextMatchesGitea(t *testing.T) {
t.Parallel()
requireTools(t)
name, job := workflowIdentity(t)
history := newRepo(t)
fake, api := newFakeGitea(t)
fake.setStatus(history.head, running())
fake.setStatus(history.parent, cancelled())
out, err := runScript(t, history, api, scriptEnv{
workflow: name,
job: job,
event: "push",
})
require.NoError(t, err, out)
posted := fake.postedFor(history.parent)
require.Len(t, posted, 1)
require.Equal(t, liveContext, posted[0].Context)
}
// workflowIdentity reads the workflow name and its single job id out of
// the checked-in workflow file.
func workflowIdentity(t *testing.T) (string, string) {
t.Helper()
raw, err := os.ReadFile(workflow)
require.NoError(t, err)
var parsed struct {
Name string `yaml:"name"`
Jobs map[string]any `yaml:"jobs"`
}
require.NoError(t, yaml.Unmarshal(raw, &parsed))
jobs := slices.Collect(maps.Keys(parsed.Jobs))
require.Len(t, jobs, 1)
return parsed.Name, jobs[0]
}
func runScript(
t *testing.T, history repo, api string, env scriptEnv,
extra ...string,
) (string, error) {
t.Helper()
script, err := filepath.Abs(scriptPath)
require.NoError(t, err)
//nolint:gosec // fixed argv, repo-local script under test
cmd := exec.CommandContext(t.Context(), "sh", script)
cmd.Dir = history.dir
cmd.Env = append(os.Environ(),
"GITHUB_API_URL="+api,
"GITHUB_REPOSITORY=sneak/webhooker",
"GITHUB_SHA="+history.head,
"GITHUB_WORKFLOW="+env.workflow,
"GITHUB_JOB="+env.job,
"GITHUB_EVENT_NAME="+env.event,
"GITEA_TOKEN=test-token",
)
cmd.Env = append(cmd.Env, extra...)
out, err := cmd.CombinedOutput()
return string(out), err
}
func newRepo(t *testing.T) repo {
t.Helper()
dir := t.TempDir()
git := func(args ...string) string {
//nolint:gosec // fixed argv, arguments are test constants
cmd := exec.CommandContext(t.Context(), "git", args...)
cmd.Dir = dir
out, err := cmd.CombinedOutput()
require.NoError(t, err, string(out))
return strings.TrimSpace(string(out))
}
commit := func(message string) string {
git(
"-c", "user.email=ci@example.invalid",
"-c", "user.name=ci",
"-c", "commit.gpgsign=false",
"commit", "-q", "--allow-empty", "-m", message,
)
return git("rev-parse", "HEAD")
}
git("init", "-q", "-b", "main")
parent := commit("parent")
head := commit("head")
return repo{dir: dir, head: head, parent: parent}
}
func requireTools(t *testing.T) {
t.Helper()
for _, tool := range []string{"sh", "git", "curl", "jq"} {
_, err := exec.LookPath(tool)
if err != nil {
t.Skipf("%s is not installed: %v", tool, err)
}
}
}

View File

@@ -1,10 +0,0 @@
// Package ciscript holds the tests for the repository's CI shell
// scripts in script/. It carries no runtime code: the scripts run on
// the CI runner, not inside the binary, but their behaviour still has
// to be verified by the test suite.
//
// The scripts under test are outside the Go build graph, so `go test`'s
// result cache serves a stale PASS when only a script changed: run the
// container build, or GOFLAGS=-count=1, to trust a result here after
// editing script/.
package ciscript

View File

@@ -1,162 +0,0 @@
package ciscript_test
import (
"encoding/json"
"net/http"
"net/http/httptest"
"sync"
"testing"
)
// commitStatus is the part of an entry in Gitea's combined-status
// response that script/ci-mark-superseded reads.
type commitStatus struct {
Context string `json:"context"`
Status string `json:"status"`
Description string `json:"description"`
}
// postedStatus is the part of a create-status request body the script
// writes.
type postedStatus struct {
Context string `json:"context"`
State string `json:"state"`
Description string `json:"description"`
}
// fakeGitea serves the two endpoints the script talks to. Like Gitea,
// the newest status for a context replaces the previous one, so a
// second run of the script sees what the first one wrote.
type fakeGitea struct {
mu sync.Mutex
statuses map[string][]commitStatus
posted map[string][]postedStatus
// failRead is a commit whose combined-status read answers HTTP
// 500, standing in for a status API that is down.
failRead string
}
// newFakeGitea returns the fake and the base URL to hand the script as
// GITHUB_API_URL.
func newFakeGitea(t *testing.T) (*fakeGitea, string) {
t.Helper()
fake := &fakeGitea{
mu: sync.Mutex{},
statuses: map[string][]commitStatus{},
posted: map[string][]postedStatus{},
failRead: "",
}
srv := httptest.NewServer(fake.routes())
t.Cleanup(srv.Close)
return fake, srv.URL
}
func (f *fakeGitea) routes() http.Handler {
mux := http.NewServeMux()
mux.HandleFunc(
"GET /repos/{owner}/{repo}/commits/{sha}/status",
f.handleCombined,
)
mux.HandleFunc(
"POST /repos/{owner}/{repo}/statuses/{sha}",
f.handleCreate,
)
return mux
}
func (f *fakeGitea) handleCombined(
w http.ResponseWriter, r *http.Request,
) {
f.mu.Lock()
defer f.mu.Unlock()
sha := r.PathValue("sha")
if f.failRead != "" && f.failRead == sha {
http.Error(w, "boom", http.StatusInternalServerError)
return
}
body := struct {
Statuses []commitStatus `json:"statuses"`
}{Statuses: f.statuses[sha]}
payload, err := json.Marshal(body)
if err != nil {
http.Error(w, err.Error(), http.StatusInternalServerError)
return
}
w.Header().Set("Content-Type", "application/json")
_, _ = w.Write(payload)
}
func (f *fakeGitea) handleCreate(w http.ResponseWriter, r *http.Request) {
var got postedStatus
err := json.NewDecoder(r.Body).Decode(&got)
if err != nil {
http.Error(w, err.Error(), http.StatusBadRequest)
return
}
sha := r.PathValue("sha")
f.mu.Lock()
defer f.mu.Unlock()
f.posted[sha] = append(f.posted[sha], got)
f.replaceLocked(sha, commitStatus{
Context: got.Context,
Status: got.State,
Description: got.Description,
})
w.WriteHeader(http.StatusCreated)
}
// failStatusRead makes the combined-status read for one commit answer
// HTTP 500.
func (f *fakeGitea) failStatusRead(sha string) {
f.mu.Lock()
defer f.mu.Unlock()
f.failRead = sha
}
// setStatus gives a commit its latest status for a context.
func (f *fakeGitea) setStatus(sha string, status commitStatus) {
f.mu.Lock()
defer f.mu.Unlock()
f.replaceLocked(sha, status)
}
// postedFor returns the statuses the script created for a commit.
func (f *fakeGitea) postedFor(sha string) []postedStatus {
f.mu.Lock()
defer f.mu.Unlock()
return append([]postedStatus(nil), f.posted[sha]...)
}
// replaceLocked requires f.mu.
func (f *fakeGitea) replaceLocked(sha string, status commitStatus) {
for i, existing := range f.statuses[sha] {
if existing.Context == status.Context {
f.statuses[sha][i] = status
return
}
}
f.statuses[sha] = append(f.statuses[sha], status)
}

View File

@@ -430,14 +430,10 @@ func loadFromEnv() (*Config, error) {
// what is in front of the process, which this code cannot observe:
// with nothing in front, the peer is the client and the limits are
// per-client as intended; behind a reverse proxy the peer is the proxy
// for every request, so all clients share one bucket per limiter.
//
// The login endpoint no longer spends budget on arrival — it verifies
// credentials first and charges only failures — so a shared bucket
// cannot deny the operator a correct password. What it does collapse
// is the failure counting: one client's wrong passwords throttle
// everyone else's wrong passwords, and the receiver's limits become
// service-wide ceilings.
// for every request, so all clients share one bucket per limiter. The
// login limiter's bucket is the dangerous one: any remote client can
// keep it full, which denies the only administrative login to everyone
// until the process restarts.
//
// The warning is deliberately not gated on WEBHOOKER_ENVIRONMENT. That
// variable defaults to dev, so gating on it would silence the warning
@@ -458,11 +454,11 @@ func (c *Config) warnSharedRateLimitBucket(log *slog.Logger) {
"this process that is the client itself and the limits "+
"are per-client as intended. Behind a reverse proxy the "+
"peer is the proxy on every request, so all clients "+
"share one bucket per limit: the receiver limits become "+
"service-wide ceilings, and one client's failed logins "+
"throttle every other client's failed logins — a "+
"correct password still gets in. If anything proxies to "+
"this process, set TRUSTED_PROXIES to its address.",
"share one bucket per limit and any remote client can "+
"keep the login limit full, denying the admin login "+
"the only administrative path — until restart. If "+
"anything proxies to this process, set TRUSTED_PROXIES "+
"to its address.",
"environment", c.Environment,
"trustedProxies", len(c.TrustedProxies),
)

View File

@@ -629,9 +629,8 @@ func testTrustedProxiesSuccess(
// TestSharedRateLimitBucketWarning covers the startup warning that
// tells an operator a deployment behind a reverse proxy shares one
// rate-limit bucket between every client, which turns the receiver
// limits into service-wide ceilings and collapses login failure
// counting. It must fire whenever TRUSTED_PROXIES is empty,
// rate-limit bucket between every client, which makes the admin login
// remotely deniable. It must fire whenever TRUSTED_PROXIES is empty,
// in any environment: WEBHOOKER_ENVIRONMENT defaults to dev, so gating
// on it would silence the warning for exactly the operator who never
// configured the deployment. It stays quiet once proxies are named.
@@ -708,15 +707,7 @@ func TestSharedRateLimitBucketWarning(t *testing.T) {
assert.Contains(t, logged, `"level":"WARN"`)
assert.Contains(t, logged, "TRUSTED_PROXIES")
assert.Contains(t, logged, "share one bucket")
assert.Contains(
t, logged, "throttle every other client's failed logins",
)
// The warning must not claim a lockout the login
// endpoint no longer permits: credentials are verified
// before any budget is spent.
assert.Contains(
t, logged, "a correct password still gets in",
)
assert.Contains(t, logged, "denying the admin login")
// The text must stay accurate for a developer with
// nothing in front of the process, where an empty
// list costs nothing.

View File

@@ -17,7 +17,6 @@ import (
"gorm.io/gorm"
_ "modernc.org/sqlite" // Pure Go SQLite driver
"sneak.berlin/go/webhooker/internal/config"
"sneak.berlin/go/webhooker/internal/gormlog"
"sneak.berlin/go/webhooker/internal/logger"
)
@@ -156,10 +155,7 @@ func (d *Database) connect() error {
// Then use it with GORM
db, err := gorm.Open(sqlite.Dialector{
Conn: sqlDB,
}, &gorm.Config{
// Never leave this at GORM's default. See internal/gormlog.
Logger: gormlog.New(d.log),
})
}, &gorm.Config{})
if err != nil {
d.log.Error(
"failed to connect to database",

View File

@@ -65,9 +65,3 @@ func (r *RetentionReaper) ExportWedgeLoop(
func (r *RetentionReaper) ExportSetInterval(d time.Duration) {
r.interval = d
}
// DummyPasswordHashForTest exposes the encoded hash that unknown
// usernames are verified against.
func DummyPasswordHashForTest() string {
return dummyPasswordHash()
}

View File

@@ -2,16 +2,12 @@ package database
import "time"
// APIKey represents an API key for a user.
//
// Key is a bearer credential, so it is never marshalled with the
// model. A creation handler that has to show it once returns it in its
// own response type.
// APIKey represents an API key for a user
type APIKey struct {
BaseModel
UserID string `gorm:"type:uuid;not null" json:"userId"`
Key string `gorm:"uniqueIndex;not null" json:"-"`
Key string `gorm:"uniqueIndex;not null" json:"key"`
Description string `json:"description"`
LastUsedAt *time.Time `json:"lastUsedAt,omitempty"`

View File

@@ -1,107 +0,0 @@
package database_test
import (
"encoding/json"
"testing"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
"sneak.berlin/go/webhooker/internal/database"
)
// keptField is a non-secret value planted alongside each secret, so
// the assertions below cannot pass by the model marshalling to nothing.
const keptField = "keepme"
// marshalModel encodes a model the way a future JSON handler would.
func marshalModel(t *testing.T, v any) string {
t.Helper()
encoded, err := json.Marshal(v)
require.NoError(t, err)
return string(encoded)
}
// TestModelsDoNotMarshalTheirSecrets pins the barrier for the JSON
// path. The /api/v1 route group exists and is empty; delivery's
// TargetView masks the credential for the HTML path only, so without
// these tags the first handler that marshals a model serialises the
// secret with it. Each field below is a live credential:
//
// - Target.Config holds an incoming-webhook URL whose path segments
// are the bearer token.
// - APIKey.Key is a bearer token outright.
// - Setting.Value holds the session encryption key.
// - User.Password holds the Argon2 hash, and was already tagged.
func TestModelsDoNotMarshalTheirSecrets(t *testing.T) {
t.Parallel()
const marker = "QQMODELMARKERQQ"
cases := []struct {
name string
model any
}{
{
name: "target config",
model: database.Target{
Name: keptField,
Type: database.TargetTypeSlack,
Config: `{"webhookUrl":"https://h/s/` + marker + `"}`,
},
},
{
name: "api key",
model: database.APIKey{
Description: keptField,
Key: marker,
},
},
{
name: "setting value",
model: database.Setting{
Key: keptField,
Value: marker,
},
},
{
name: "user password hash",
model: database.User{
Username: keptField,
Password: marker,
},
},
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
t.Parallel()
encoded := marshalModel(t, tc.model)
assert.NotContains(t, encoded, marker)
assert.Contains(t, encoded, keptField)
})
}
}
// TestWebhookMarshalsNoTargetConfig covers the nested case: a webhook
// marshalled with its targets preloaded must not carry the credential
// through the association either.
func TestWebhookMarshalsNoTargetConfig(t *testing.T) {
t.Parallel()
const marker = "QQNESTEDMARKERQQ"
encoded := marshalModel(t, database.Webhook{
Name: keptField,
Targets: []database.Target{{
Name: "slack",
Config: `{"webhookUrl":"https://h/s/` + marker + `"}`,
}},
})
assert.NotContains(t, encoded, marker)
assert.Contains(t, encoded, keptField)
}

View File

@@ -3,9 +3,6 @@ package database
// Setting stores application-level key-value configuration.
// Used for auto-generated values like the session encryption key.
type Setting struct {
Key string `gorm:"primaryKey" json:"key"`
// Value holds the session encryption key, so it is never
// marshalled with the model.
Value string `gorm:"type:text;not null" json:"-"`
Key string `gorm:"primaryKey" json:"key"`
Value string `gorm:"type:text;not null" json:"value"`
}

View File

@@ -20,14 +20,8 @@ type Target struct {
Type TargetType `gorm:"not null" json:"type"`
Active bool `gorm:"default:true" json:"active"`
// Configuration fields (JSON stored based on type).
//
// json:"-" because the blob holds the target's credential — a
// Slack incoming-webhook URL, or an http destination whose path
// segments are the secret. delivery.TargetView is the masking
// barrier for the HTML path; this tag is the barrier for any
// handler that marshals the model itself.
Config string `gorm:"type:text" json:"-"` // JSON configuration
// Configuration fields (JSON stored based on type)
Config string `gorm:"type:text" json:"config"` // JSON configuration
// For HTTP targets (max_retries=0 means fire-and-forget,
// >0 enables retries with backoff)

View File

@@ -8,7 +8,6 @@ import (
"fmt"
"math/big"
"strings"
"sync"
"golang.org/x/crypto/argon2"
)
@@ -30,10 +29,6 @@ const hashParts = 6
// triggers per-character-class complexity enforcement.
const minPasswordComplexityLen = 4
// dummyPasswordLen is the length of the throwaway password behind
// dummyPasswordHash.
const dummyPasswordLen = 32
// Sentinel errors returned by decodeHash.
var (
errInvalidHashFormat = errors.New("invalid hash format")
@@ -127,38 +122,6 @@ func VerifyPassword(
return subtle.ConstantTimeCompare(hash, otherHash) == 1, nil
}
// dummyPasswordHash is an encoded Argon2id hash of a random
// password, computed once on first use. Nothing can match it: the
// password it encodes is discarded as soon as it is hashed. It is
// process-wide because building it per request would add a second
// 64 MB Argon2id pass to every login for an unknown username.
//
//nolint:gochecknoglobals // computed once, see above
var dummyPasswordHash = sync.OnceValue(func() string {
password, err := GenerateRandomPassword(dummyPasswordLen)
if err != nil {
panic(fmt.Sprintf("generating the dummy password: %v", err))
}
hash, err := HashPassword(password)
if err != nil {
panic(fmt.Sprintf("hashing the dummy password: %v", err))
}
return hash
})
// VerifyDummyPassword performs a credential verification that cannot
// succeed, at the same cost as a real one.
//
// Login must charge an unknown username the same work as a known
// one. Returning early for an account that does not exist answers in
// microseconds where a real account takes tens of milliseconds, which
// is a username oracle any client can read off the response time.
func VerifyDummyPassword(password string) {
_, _ = VerifyPassword(password, dummyPasswordHash())
}
// decodeHash extracts parameters, salt, and hash from an
// encoded hash string.
func decodeHash(

View File

@@ -191,41 +191,3 @@ func TestHashPasswordUniqueness(t *testing.T) {
)
}
}
// TestVerifyDummyPassword_DoesRealWork covers the anti-enumeration
// path. Login charges an unknown username a verification against a
// dummy hash so that a nonexistent account is not answered in
// microseconds where a real one takes tens of milliseconds. That only
// works if the dummy hash is a real, decodable Argon2id hash: a
// malformed one would make VerifyPassword fail on the decode and
// return before hashing anything.
func TestVerifyDummyPassword_DoesRealWork(t *testing.T) {
t.Parallel()
// Runs the OnceValue that builds the dummy hash, so a panic in
// it surfaces here rather than on a live login.
database.VerifyDummyPassword("whatever was submitted")
dummy := database.DummyPasswordHashForTest()
// A hash the verifier cannot decode would make VerifyPassword
// return on the decode error, before hashing anything — the
// timing oracle this path exists to close.
valid, err := database.VerifyPassword("whatever", dummy)
if err != nil {
t.Fatalf(
"the dummy hash must decode like a real one: %v", err,
)
}
if valid {
t.Error("nothing may authenticate against the dummy hash")
}
if !strings.HasPrefix(dummy, "$argon2id$") {
t.Errorf(
"the dummy hash must use the same algorithm as real "+
"hashes, got %q", dummy,
)
}
}

View File

@@ -14,7 +14,6 @@ import (
"gorm.io/driver/sqlite"
"gorm.io/gorm"
"sneak.berlin/go/webhooker/internal/config"
"sneak.berlin/go/webhooker/internal/gormlog"
"sneak.berlin/go/webhooker/internal/logger"
)
@@ -249,10 +248,7 @@ func (m *WebhookDBManager) openDB(
db, err := gorm.Open(sqlite.Dialector{
Conn: sqlDB,
}, &gorm.Config{
// Never leave this at GORM's default. See internal/gormlog.
Logger: gormlog.New(m.log),
})
}, &gorm.Config{})
if err != nil {
_ = sqlDB.Close()

View File

@@ -339,12 +339,7 @@ func TestWebhookDBManager_MultipleWebhooks(t *testing.T) {
var events []database.Event
require.NoError(t, db2.Find(&events).Error)
// require, not assert: this is exactly the regression the test
// guards, so the empty slice is the expected failure, and a
// non-fatal length check would index into it on the next line and
// panic the whole package test binary instead of failing here.
require.Len(t, events, 1)
assert.Len(t, events, 1)
assert.Equal(t, "PUT", events[0].Method)
}

View File

@@ -12,7 +12,6 @@ import (
"gorm.io/driver/sqlite"
"gorm.io/gorm"
"sneak.berlin/go/webhooker/internal/gormlog"
)
// archiveExpiryNever is the expiry sentinel (and default) that
@@ -283,11 +282,7 @@ func (w *archiveWriter) openMode(
}
gdb, err := gorm.Open(
sqlite.Dialector{Conn: sqlDB}, &gorm.Config{
// Never leave this at GORM's default. See
// internal/gormlog.
Logger: gormlog.New(w.log),
},
sqlite.Dialector{Conn: sqlDB}, &gorm.Config{},
)
if err != nil {
_ = sqlDB.Close()

View File

@@ -1,155 +0,0 @@
package delivery_test
import (
"bytes"
"log"
"log/slog"
"path/filepath"
"strings"
"sync"
"testing"
"time"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
"gorm.io/gorm"
gormlogger "gorm.io/gorm/logger"
"sneak.berlin/go/webhooker/internal/delivery"
"sneak.berlin/go/webhooker/internal/middleware"
)
// archiveGORMTailMarker sits at the far end of the value this file
// drives into an archive lookup. Its presence in a log line means the
// whole value reached the log, so nothing truncated it.
const archiveGORMTailMarker = "ENDOFCLIENTVALUE"
// archiveGORMFillBytes is how much text the lookup carries. It is far
// past every budget in play.
const archiveGORMFillBytes = 8 << 10
// gormDefaultBuf collects what GORM's package-level default logger
// writes, if anything reaches it.
type gormDefaultBuf struct {
mu sync.Mutex
b bytes.Buffer
}
func (g *gormDefaultBuf) Write(p []byte) (int, error) {
g.mu.Lock()
defer g.mu.Unlock()
return g.b.Write(p)
}
func (g *gormDefaultBuf) String() string {
g.mu.Lock()
defer g.mu.Unlock()
return g.b.String()
}
// captureArchiveGORMDefault replaces GORM's package-level default
// logger with one configured exactly as GORM configures its own,
// writing to a buffer.
//
// This duplicates the detector in internal/handlers rather than
// sharing it: a test helper cannot cross a package's test boundary
// without exporting production code to carry it, and a logging
// detector is not worth a production symbol. What it detects is the
// third gorm.Open in this service, at
// internal/delivery/target_database_archive.go — the archive writer,
// whose type is unexported, so nothing outside this package can drive
// it.
func captureArchiveGORMDefault(t *testing.T) *gormDefaultBuf {
t.Helper()
buf := &gormDefaultBuf{}
orig := gormlogger.Default
gormlogger.Default = gormlogger.New(
log.New(buf, "", log.LstdFlags),
gormlogger.Config{
SlowThreshold: 200 * time.Millisecond,
LogLevel: gormlogger.Warn,
IgnoreRecordNotFoundError: false,
Colorful: false,
},
)
t.Cleanup(func() { gormlogger.Default = orig })
return buf
}
// TestArchiveWriter_NeverUsesGORMsDefaultLogger pins the archive
// writer's gorm.Open to the adapter.
//
// Restore a bare &gorm.Config{} at
// internal/delivery/target_database_archive.go and this fails: the
// default logger prints the fully interpolated SELECT on every
// ErrRecordNotFound, so the client-chosen event id below arrives whole
// and unbounded on stdout, answering to no level the operator set.
//
// Not parallel: gormlogger.Default is process-global. Go runs every
// non-parallel top-level test to completion before it resumes the
// parallel ones.
//
//nolint:paralleltest // Deliberately sequential; see above.
func TestArchiveWriter_NeverUsesGORMsDefaultLogger(t *testing.T) {
var captured bytes.Buffer
gormDefault := captureArchiveGORMDefault(t)
w := delivery.NewExportArchiveWriter(
filepath.Join(t.TempDir(), "archive.db"),
slog.New(slog.NewTextHandler(
&captured, &slog.HandlerOptions{Level: slog.LevelDebug},
)),
0,
)
require.NoError(t, w.Open(0))
t.Cleanup(w.Evict)
// A lookup that misses, carrying a value the size of an inbound
// event id. Under the default logger this is the line that gets
// interpolated and printed.
value := strings.Repeat("\x01", archiveGORMFillBytes) +
archiveGORMTailMarker
var row delivery.ExportArchivedEvent
err := w.DB().Where("event_id = ?", value).First(&row).Error
require.ErrorIs(t, err, gorm.ErrRecordNotFound)
got := gormDefault.String()
assert.Empty(
t, got,
"GORM's default logger wrote %d bytes, so the archive "+
"writer's gorm.Open is back on a bare &gorm.Config{}; "+
"the first of them: %s",
len(got), got[:min(len(got), 300)],
)
// The adapter drops a miss, so this should be silent too — and
// whatever it does write stays inside the stated ceiling.
out := captured.String()
assert.NotContains(
t, out, archiveGORMTailMarker,
"the far end of the client-chosen value reached the log",
)
for line := range strings.SplitSeq(strings.TrimRight(out, "\n"), "\n") {
if line == "" {
continue
}
assert.LessOrEqual(
t, len(line), middleware.MaxAccessLogLineBytes,
"log line exceeded its bound: %s",
line[:min(len(line), 300)],
)
}
}

View File

@@ -11,17 +11,6 @@ import (
// inbound webhook — the full request body and headers, plus
// the method, content type, and the webhook and entrypoint
// ids — then records a single successful attempt.
//
// This is the one log call in the service that deliberately writes
// unbounded client-chosen bytes, so it is the one exception to the
// per-field budgets in internal/logfield and to the ceiling stated on
// middleware.MaxAccessLogLineBytes. Capping here would defeat the
// target: emitting the payload IS the delivery. It costs nothing by
// default — an authenticated operator has to create a target of this
// type on a specific webhook before a single line is written — and the
// bytes it writes are bounded per event by maxWebhookBodySize (1 MB).
// An operator who adds one is choosing to spend log volume on the
// payloads that webhook receives.
type logTarget struct {
eng *Engine
}

View File

@@ -1,17 +0,0 @@
package gormlog
import (
"log/slog"
"time"
)
// ExportNewWithSlowThreshold builds a Logger whose slow-statement
// threshold is d rather than DefaultSlowThreshold, so a test can pin
// which arm of Trace it is exercising instead of racing the clock on a
// loaded machine. The threshold is set at construction, like every
// other field, so the type's concurrency guarantee still holds.
func ExportNewWithSlowThreshold(
log *slog.Logger, d time.Duration,
) *Logger {
return &Logger{log: log, slowThreshold: d}
}

View File

@@ -1,168 +0,0 @@
// Package gormlog adapts GORM's logger onto the service's slog
// logger.
//
// GORM's own default logger is not usable here. It is built at package
// init with log.New(os.Stdout, ...) at LogLevel Warn with
// IgnoreRecordNotFoundError false, so it writes the fully interpolated
// SQL — parameters and all — for every statement that returns an
// error, including gorm.ErrRecordNotFound. Two of this service's
// lookups miss by design on unauthenticated routes: the entrypoint
// lookup on /webhook/{uuid}, whose path segment the client picks
// outright, and the user lookup behind the login form, whose username
// the client picks outright. Under the default logger each of those
// misses printed an unbounded, attacker-chosen string, at no level the
// operator can turn down, past every handler internal/logger installs.
//
// This adapter fixes all three properties at once: the lines get a
// level the operator controls, they are shaped by whichever handler
// internal/logger selected, and every value a client can influence is
// spent through logfield.Truncate.
package gormlog
import (
"context"
"errors"
"fmt"
"log/slog"
"time"
gormlogger "gorm.io/gorm/logger"
"sneak.berlin/go/webhooker/internal/logfield"
)
// DefaultSlowThreshold is the duration at or above which a statement
// is logged as slow. It is GORM's own default, kept deliberately: slow
// SQL is the one thing GORM's logger reports that nothing else in this
// service does, so silencing the logger outright would have cost real
// observability to fix a log-volume defect.
const DefaultSlowThreshold = 200 * time.Millisecond
// Logger implements gormlogger.Interface on top of an *slog.Logger.
//
// It is safe for concurrent use: every field is set at construction
// and never written again.
type Logger struct {
log *slog.Logger
slowThreshold time.Duration
}
// Interface compliance is asserted here rather than discovered at the
// gorm.Open call sites.
var _ gormlogger.Interface = (*Logger)(nil)
// New returns a GORM logger that writes through log.
func New(log *slog.Logger) *Logger {
return &Logger{
log: log,
slowThreshold: DefaultSlowThreshold,
}
}
// LogMode returns the logger unchanged.
//
// GORM's LogLevel is deliberately not honoured. Level is the operator's
// decision and it is expressed once, through LOG_LEVEL and the
// slog.LevelVar internal/logger holds; a second level knob inside the
// database layer could only disagree with it. The mapping from GORM's
// four categories onto slog levels is fixed in Trace below.
//
//nolint:ireturn // The interface return is GORM's signature, not a choice.
func (l *Logger) LogMode(gormlogger.LogLevel) gormlogger.Interface {
return l
}
// Info logs one of GORM's own informational messages.
func (l *Logger) Info(
ctx context.Context, msg string, data ...any,
) {
l.log.InfoContext(ctx, "gorm", "message", format(msg, data...))
}
// Warn logs one of GORM's own warnings.
func (l *Logger) Warn(
ctx context.Context, msg string, data ...any,
) {
l.log.WarnContext(ctx, "gorm", "message", format(msg, data...))
}
// Error logs one of GORM's own errors.
func (l *Logger) Error(
ctx context.Context, msg string, data ...any,
) {
l.log.ErrorContext(ctx, "gorm", "message", format(msg, data...))
}
// Trace reports the outcome of a single statement. GORM calls it for
// every statement it runs, so the cheap paths stay cheap: fc()
// renders the interpolated SQL and is called only on a branch that
// will actually emit.
//
// The arms are ordered exactly as GORM's own Trace orders them —
// non-record-not-found error, then slow, then the routine case — so
// that a statement which both misses and runs slow is still reported
// as slow. A miss is the likeliest statement to be slow, since it is
// the one that scans without finding a row, and ordering the drop
// ahead of the slow arm would have made this adapter less observant
// than the IgnoreRecordNotFoundError option it was chosen over.
func (l *Logger) Trace(
ctx context.Context,
begin time.Time,
fc func() (string, int64),
err error,
) {
elapsed := time.Since(begin)
switch {
case err != nil && !errors.Is(err, gormlogger.ErrRecordNotFound):
sql, rows := fc()
l.log.ErrorContext(ctx, "sql statement failed",
"error", logfield.Truncate(err.Error(), logfield.MaxBytes),
"sql", logfield.Truncate(sql, logfield.MaxBytes),
"rows", rows,
"elapsed_ms", elapsed.Milliseconds(),
)
case l.slowThreshold > 0 && elapsed >= l.slowThreshold:
sql, rows := fc()
l.log.WarnContext(ctx, "slow sql statement",
"sql", logfield.Truncate(sql, logfield.MaxBytes),
"rows", rows,
"elapsed_ms", elapsed.Milliseconds(),
"threshold_ms", l.slowThreshold.Milliseconds(),
)
case err != nil:
// gorm.ErrRecordNotFound is not an error on the paths that
// produce it here: an invented entrypoint UUID and an unknown
// username are the expected outcome of an unauthenticated
// request, not a fault. This is the IgnoreRecordNotFoundError
// behaviour, and it is unconditional rather than configurable
// because no caller in this service wants the other one — the
// two handlers that care already record the miss themselves,
// at DEBUG, without the SQL. A miss that ran slow has already
// been reported by the arm above.
return
case l.log.Enabled(ctx, slog.LevelDebug):
sql, rows := fc()
l.log.DebugContext(ctx, "sql statement",
"sql", logfield.Truncate(sql, logfield.MaxBytes),
"rows", rows,
"elapsed_ms", elapsed.Milliseconds(),
)
}
}
// format renders one of GORM's printf-style internal messages and
// bounds it. GORM builds these itself, but they can quote a value the
// statement carried, so they are spent through the same budget as
// everything else rather than trusted.
func format(msg string, data ...any) string {
if len(data) == 0 {
return logfield.Truncate(msg, logfield.MaxBytes)
}
return logfield.Truncate(
fmt.Sprintf(msg, data...), logfield.MaxBytes,
)
}

View File

@@ -1,436 +0,0 @@
package gormlog_test
import (
"bytes"
"context"
"database/sql"
"fmt"
"log/slog"
"path/filepath"
"strings"
"testing"
"time"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
"gorm.io/driver/sqlite"
"gorm.io/gorm"
_ "modernc.org/sqlite" // Pure Go SQLite driver.
"sneak.berlin/go/webhooker/internal/gormlog"
"sneak.berlin/go/webhooker/internal/middleware"
)
// fillBytes is how much client-chosen text each case drives into the
// statement. It is well past every budget in play, so a value that
// arrives short arrived short because something cut it.
const fillBytes = 8 << 10
// tailMarker sits at the far end of every generated value. A line that
// contains it carried the whole value, which means nothing cut it — so
// a value that merely happened to be short cannot pass for a truncated
// one.
const tailMarker = "ENDOFCLIENTVALUE"
// fills are the characters a client can drive into a SQL parameter,
// chosen for what the log handlers charge for them rather than for
// looking dangerous.
//
// The C0 control is the one that matters. Both handlers spell U+0001
// as a six-byte escape for the single byte it costs a client to send,
// which is the widest multiplier available in the basic multilingual
// plane and the case a raw-byte budget breaks on first. The astral
// non-printable costs ten under the text handler, four more than the
// JSON handler ever spends.
func fills() []struct {
name string
fill string
} {
return []struct {
name string
fill string
}{
{"plain", "x"},
{"quote", `"`},
{"backslash", `\`},
{"tab", "\t"},
{"newline", "\n"},
{"c0_control", "\x01"},
{"astral_nonprintable", "\U0001000C"},
}
}
// clientValue builds a value of at least fillBytes raw bytes out of
// fill, ending in tailMarker.
func clientValue(fill string) string {
var b strings.Builder
for b.Len() < fillBytes {
b.WriteString(fill)
}
b.WriteString(tailMarker)
return b.String()
}
// handlers are the two slog handlers internal/logger can install. The
// ceiling is quoted to operators unqualified, so every case is
// asserted under both.
func handlers() []struct {
name string
make func(*bytes.Buffer) slog.Handler
} {
opts := &slog.HandlerOptions{Level: slog.LevelDebug}
return []struct {
name string
make func(*bytes.Buffer) slog.Handler
}{
{"json", func(b *bytes.Buffer) slog.Handler {
return slog.NewJSONHandler(b, opts)
}},
{"text", func(b *bytes.Buffer) slog.Handler {
return slog.NewTextHandler(b, opts)
}},
}
}
type thing struct {
ID string `gorm:"primaryKey"`
Name string
}
// neverSlow is a slow-statement threshold no statement in this file
// can reach. Cases that are about a non-slow arm of Trace set it, so
// that a machine under load cannot turn a miss into a slow report and
// decide the outcome for them.
const neverSlow = time.Hour
// alwaysSlow makes every statement count as slow, so the slow arm is
// reached without the test waiting for it.
const alwaysSlow = time.Nanosecond
// openDB opens a real SQLite database behind the adapter under test,
// so every assertion below is made against SQL that GORM actually
// rendered rather than against a string a test wrote by hand. slow is
// the adapter's slow-statement threshold.
func openDB(
t *testing.T, buf *bytes.Buffer, h slog.Handler, slow time.Duration,
) *gorm.DB {
t.Helper()
sqlDB, err := sql.Open("sqlite", fmt.Sprintf(
"file:%s?mode=rwc",
filepath.Join(t.TempDir(), "gormlog.db"),
))
require.NoError(t, err)
t.Cleanup(func() { _ = sqlDB.Close() })
gl := gormlog.ExportNewWithSlowThreshold(slog.New(h), slow)
gdb, err := gorm.Open(
sqlite.Dialector{Conn: sqlDB},
&gorm.Config{Logger: gl},
)
require.NoError(t, err)
require.NoError(t, gdb.AutoMigrate(&thing{}))
// Migration chatter is not what any of these cases is about.
buf.Reset()
return gdb
}
// assertBounded holds every line the adapter wrote to the stated
// ceiling and proves each was cut rather than merely short.
func assertBounded(t *testing.T, out string) {
t.Helper()
assert.NotContains(
t, out, tailMarker,
"the far end of the client value reached the log, so "+
"nothing truncated it",
)
for line := range strings.SplitSeq(
strings.TrimRight(out, "\n"), "\n",
) {
if line == "" {
continue
}
assert.LessOrEqual(
t, len(line), middleware.MaxAccessLogLineBytes,
"log line exceeded its bound: %s",
line[:min(len(line), 300)],
)
}
}
// TestRecordNotFound_WritesNothing is the defect itself. GORM's own
// default logger prints the fully interpolated SELECT on every
// ErrRecordNotFound, and on this service's two unauthenticated
// lookups the interpolated parameter is whatever the client sent.
func TestRecordNotFound_WritesNothing(t *testing.T) {
t.Parallel()
for _, h := range handlers() {
for _, f := range fills() {
t.Run(h.name+"/"+f.name, func(t *testing.T) {
t.Parallel()
var buf bytes.Buffer
gdb := openDB(t, &buf, h.make(&buf), neverSlow)
var got thing
err := gdb.Where(
"id = ?", clientValue(f.fill),
).First(&got).Error
require.ErrorIs(t, err, gorm.ErrRecordNotFound)
assert.Empty(
t, buf.String(),
"a miss on a client-chosen key must not "+
"write a log line",
)
})
}
}
}
// TestSlowRecordNotFound_IsStillReportedSlow pins the arm ordering in
// Trace against the drop above.
//
// GORM's own Trace orders its cases error-that-is-not-a-miss, then
// slow, then routine, so IgnoreRecordNotFoundError: true — the cheap
// option this adapter was chosen over — still reports a miss that ran
// slow. An adapter that dropped the miss first would be strictly less
// observant than the option it replaced, on exactly the two lookups
// this package exists for. A miss is also the statement most likely to
// be slow, since it is the one that scans without finding a row.
func TestSlowRecordNotFound_IsStillReportedSlow(t *testing.T) {
t.Parallel()
for _, h := range handlers() {
for _, f := range fills() {
t.Run(h.name+"/"+f.name, func(t *testing.T) {
t.Parallel()
var buf bytes.Buffer
gdb := openDB(t, &buf, h.make(&buf), alwaysSlow)
var got thing
err := gdb.Where(
"id = ?", clientValue(f.fill),
).First(&got).Error
require.ErrorIs(t, err, gorm.ErrRecordNotFound)
assert.Contains(
t, buf.String(), "slow sql statement",
"a slow statement that missed was not "+
"reported as slow",
)
assertBounded(t, buf.String())
})
}
}
}
// TestRecordNotFoundFlood_DoesNotGrowWithInput states the definition
// of done directly: a flood of misses at two input sizes 64 times
// apart must cost the same number of bytes of log.
func TestRecordNotFoundFlood_DoesNotGrowWithInput(t *testing.T) {
t.Parallel()
const requests = 50
flood := func(t *testing.T, size int) int {
t.Helper()
var buf bytes.Buffer
gdb := openDB(
t, &buf,
slog.NewJSONHandler(&buf, &slog.HandlerOptions{
Level: slog.LevelDebug,
}),
neverSlow,
)
value := strings.Repeat("\x01", size)
for range requests {
var got thing
_ = gdb.Where("id = ?", value).First(&got).Error
}
return buf.Len()
}
small := flood(t, 128)
big := flood(t, 128*64)
assert.Equal(
t, small, big,
"log volume tracked the size of the client's input",
)
}
// TestStatementError_LineIsBounded covers the branch that does log.
// A driver error is not ErrRecordNotFound, so the interpolated
// statement is written — and on an insert the interpolated value is
// still whatever the client supplied.
func TestStatementError_LineIsBounded(t *testing.T) {
t.Parallel()
for _, h := range handlers() {
for _, f := range fills() {
t.Run(h.name+"/"+f.name, func(t *testing.T) {
t.Parallel()
var buf bytes.Buffer
gdb := openDB(t, &buf, h.make(&buf), neverSlow)
row := thing{ID: clientValue(f.fill), Name: "a"}
require.NoError(t, gdb.Create(&row).Error)
buf.Reset()
// The same primary key a second time: a UNIQUE
// constraint failure, which is an error GORM logs.
err := gdb.Create(&thing{
ID: row.ID, Name: "b",
}).Error
require.Error(t, err)
assert.Contains(
t, buf.String(), "sql statement failed",
)
assertBounded(t, buf.String())
})
}
}
}
// TestSucceedingStatement_LineIsBoundedOnEitherArm covers the two
// arms a statement that returns no error can take, over the same
// query, so neither can be bounded by accident of the other.
//
// - slow. Silencing GORM outright would have been the cheaper fix
// and would have cost this report, which is the one thing GORM's
// logger gave an operator that nothing else in this service does.
// - routine. The branch an operator reaches by turning the level
// down to DEBUG: every statement is reported, so every
// statement's interpolated parameters have to be bounded too.
func TestSucceedingStatement_LineIsBoundedOnEitherArm(t *testing.T) {
t.Parallel()
// "sql statement" is a substring of "slow sql statement", so the
// routine arm carries notWant as well: Contains alone cannot tell
// the two arms apart in that direction.
arms := []struct {
name string
slow time.Duration
want string
notWant string
}{
{"slow", alwaysSlow, "slow sql statement", ""},
{"routine", neverSlow, "sql statement", "slow sql statement"},
}
for _, a := range arms {
for _, h := range handlers() {
for _, f := range fills() {
name := a.name + "/" + h.name + "/" + f.name
t.Run(name, func(t *testing.T) {
t.Parallel()
var buf bytes.Buffer
gdb := openDB(t, &buf, h.make(&buf), a.slow)
var got []thing
require.NoError(t, gdb.Where(
"name = ?", clientValue(f.fill),
).Find(&got).Error)
assert.Contains(t, buf.String(), a.want)
if a.notWant != "" {
assert.NotContains(
t, buf.String(), a.notWant,
)
}
assertBounded(t, buf.String())
})
}
}
}
}
// TestGORMOwnMessages_AreBounded covers the three printf-style
// entry points. GORM builds these itself, but nothing stops one of
// them quoting a value the statement carried.
func TestGORMOwnMessages_AreBounded(t *testing.T) {
t.Parallel()
for _, h := range handlers() {
for _, f := range fills() {
t.Run(h.name+"/"+f.name, func(t *testing.T) {
t.Parallel()
var buf bytes.Buffer
gl := gormlog.New(slog.New(h.make(&buf)))
ctx := context.Background()
value := clientValue(f.fill)
gl.Info(ctx, "%s", value)
gl.Warn(ctx, "%s", value)
gl.Error(ctx, "%s", value)
// The no-argument form, which is how GORM reports
// most of its own conditions. Reached through a
// function value so the vet printf check does not
// read the message as a format string — which is
// also why the adapter does not.
noArgs := func(
f func(context.Context, string, ...any),
msg string,
) {
f(ctx, msg)
}
noArgs(gl.Info, value)
assertBounded(t, buf.String())
})
}
}
}
// TestLogMode_KeepsTheOperatorsLevel records that GORM's own level
// knob is deliberately inert: level belongs to LOG_LEVEL, and a
// second one inside the database layer could only disagree with it.
func TestLogMode_KeepsTheOperatorsLevel(t *testing.T) {
t.Parallel()
var buf bytes.Buffer
gl := gormlog.New(slog.New(slog.NewJSONHandler(
&buf, &slog.HandlerOptions{Level: slog.LevelDebug},
)))
assert.Same(t, gl, gl.LogMode(0))
}

View File

@@ -2,10 +2,8 @@ package handlers
import (
"net/http"
"strconv"
"sneak.berlin/go/webhooker/internal/database"
"sneak.berlin/go/webhooker/internal/logfield"
)
// HandleLoginPage returns a handler for the login page (GET)
@@ -41,10 +39,8 @@ func (h *Handlers) HandleLoginSubmit() http.HandlerFunc {
return
}
// PostFormValue, not FormValue: the credential must come
// from the body, never from the query string.
username := r.PostFormValue("username")
password := r.PostFormValue("password")
username := r.FormValue("username")
password := r.FormValue("password")
// Validate input
if username == "" || password == "" {
@@ -71,9 +67,7 @@ func (h *Handlers) HandleLoginSubmit() http.HandlerFunc {
h.log.Info(
"user logged in",
"username", logfield.Truncate(
username, logfield.MaxBytes,
),
"username", username,
"user_id", user.ID,
)
@@ -99,16 +93,6 @@ func (h *Handlers) renderLoginError(
// authenticateUser looks up and verifies a user's credentials.
// On failure it writes an HTTP response and returns an error.
//
// The credential check runs BEFORE any rate-limit budget is
// consulted, and only a failed check spends budget. That is what
// keeps the single administrative path reachable: behind the reverse
// proxy this deployment requires, with TRUSTED_PROXIES unset, every
// client shares one bucket, so a limiter spent on arrival lets any
// stranger deny the operator's own correct password indefinitely.
//
// Verifying first means every login POST costs an Argon2id hash, so
// the work is taken under a bounded number of verification slots.
func (h *Handlers) authenticateUser(
w http.ResponseWriter,
r *http.Request,
@@ -116,49 +100,16 @@ func (h *Handlers) authenticateUser(
) (database.User, error) {
var user database.User
release, ok := h.mw.BeginPasswordVerification(r.Context())
if !ok {
h.log.Warn(
"password verification capacity exhausted",
"path", logfield.Truncate(
r.URL.Path, logfield.MaxBytes,
),
)
h.renderLoginError(
w, r,
"The server is busy verifying credentials. "+
"Please try again.",
http.StatusServiceUnavailable,
)
return user, errVerificationBusy
}
defer release()
err := h.db.DB().Where(
"username = ?", username,
).First(&user).Error
if err != nil {
// A username that does not exist is charged the same work
// as one that does. Skipping the hash here would answer in
// microseconds where a real account takes tens of
// milliseconds, handing every client a username oracle.
h.dummyVerifications.Add(1)
database.VerifyDummyPassword(password)
// Login is unauthenticated, and the submitted username is
// a form field the client fills to any length the 1 MB
// body cap allows. On this branch it matched no row, so
// nothing else bounds it. The rate limiter caps how often
// the line is written, not how wide it is.
h.log.Debug(
"user not found",
"username", logfield.Truncate(
username, logfield.MaxBytes,
),
h.log.Debug("user not found", "username", username)
h.renderLoginError(
w, r,
"Invalid username or password",
http.StatusUnauthorized,
)
h.rejectLogin(w, r, username)
return user, err
}
@@ -175,60 +126,17 @@ func (h *Handlers) authenticateUser(
}
if !valid {
// Reached only once the username matched a stored row, so
// it is bounded by the operator's own data. Capped anyway,
// so that every username this unauthenticated endpoint
// logs is capped and no reader has to work out which
// branch narrowed it.
h.log.Debug(
"invalid password",
"username", logfield.Truncate(
username, logfield.MaxBytes,
),
)
h.rejectLogin(w, r, username)
return user, errInvalidPassword
}
// The password was correct, so forgive whatever failures this
// client accumulated: an operator who mistypes a few times and
// then gets it right must not stay throttled afterwards.
h.mw.ForgiveLoginFailures(r, username)
return user, nil
}
// rejectLogin counts one failed credential verification and answers
// it: 401 while this client still has failure budget against the
// submitted username, 429 with a Retry-After once it is spent.
//
// The 429 throttles wrong passwords only. A correct one never
// reaches here, so no amount of failure — from this client or any
// other sharing its bucket — can keep the operator out.
func (h *Handlers) rejectLogin(
w http.ResponseWriter,
r *http.Request,
username string,
) {
if !h.mw.RecordLoginFailure(r, username) {
h.log.Debug("invalid password", "username", username)
h.renderLoginError(
w, r,
"Invalid username or password",
http.StatusUnauthorized,
)
return
return user, errInvalidPassword
}
w.Header().Set("Retry-After", strconv.Itoa(int(
h.mw.LoginFailureInterval().Seconds(),
)))
h.renderLoginError(
w, r,
"Too many failed login attempts. Please try again later.",
http.StatusTooManyRequests,
)
return user, nil
}
// createAuthenticatedSession regenerates the session and stores

View File

@@ -1,455 +0,0 @@
package handlers_test
import (
"context"
"fmt"
"net/http"
"net/http/httptest"
"net/url"
"strings"
"sync"
"testing"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
"sneak.berlin/go/webhooker/internal/database"
"sneak.berlin/go/webhooker/internal/handlers"
"sneak.berlin/go/webhooker/internal/session"
)
const (
// operatorUser and operatorPassword are the single admin account
// these tests defend.
operatorUser = "admin"
operatorPassword = "correct horse battery staple"
// sharedProxyPeer is the whole point of this file. Production is
// required to run behind a TLS-terminating reverse proxy, and
// TRUSTED_PROXIES defaults to empty, so every client — attacker
// and operator alike — reaches the process from the proxy's
// address and shares one rate-limit bucket. Both parties in
// these tests therefore use the same RemoteAddr.
sharedProxyPeer = "10.0.0.1:44444"
// loginFailureLimit is the failure budget one client has against
// one submitted username. Restated here rather than imported
// from the middleware package, so that changing the production
// limit fails these tests instead of silently moving with them.
loginFailureLimit = 5
)
// seedOperator gives the bootstrapped admin account a password these
// tests know. The account itself is created at startup with a random
// password, which is exactly why its username is predictable to an
// attacker and why keying failures by username alone does not fix
// this issue.
func seedOperator(t *testing.T, db *database.Database) {
t.Helper()
hash, err := database.HashPassword(operatorPassword)
require.NoError(t, err)
result := db.DB().Model(&database.User{}).
Where("username = ?", operatorUser).
Update("password", hash)
require.NoError(t, result.Error)
require.EqualValues(
t, 1, result.RowsAffected,
"the bootstrap admin account must exist",
)
}
// loginPost builds a login form POST arriving from peer.
func loginPost(peer, username, password string) *http.Request {
form := url.Values{}
form.Set("username", username)
form.Set("password", password)
req := httptest.NewRequestWithContext(
context.Background(),
http.MethodPost,
"/pages/login",
strings.NewReader(form.Encode()),
)
req.Header.Set(
"Content-Type", "application/x-www-form-urlencoded",
)
req.RemoteAddr = peer
return req
}
// submitLogin drives one login POST through the handler.
func submitLogin(
h *handlers.Handlers, peer, username, password string,
) *httptest.ResponseRecorder {
w := httptest.NewRecorder()
h.HandleLoginSubmit().ServeHTTP(w, loginPost(
peer, username, password,
))
return w
}
// floodFailures sends attempts wrong-password logins for username
// from peer, which is what an attacker does.
func floodFailures(
t *testing.T,
h *handlers.Handlers,
peer, username string,
attempts int,
) {
t.Helper()
for i := range attempts {
w := submitLogin(h, peer, username, fmt.Sprintf("guess-%d", i))
require.NotEqual(
t, http.StatusSeeOther, w.Code,
"attempt %d must not authenticate", i,
)
}
}
// TestLogin_StrangersFloodCannotLockOutTheOperator is the
// done-criterion of https://git.eeqj.de/sneak/webhooker/issues/150.
//
// The attacker and the operator share one rate-limit bucket, because
// behind the mandated reverse proxy with TRUSTED_PROXIES unset every
// client keys on the proxy's address. The attacker floods the
// operator's own username — a single-admin product has a predictable
// one — far past the failure limit. The operator must still be able
// to log in with the correct password.
//
// This fails if credentials stop being verified ahead of the limiter.
func TestLogin_StrangersFloodCannotLockOutTheOperator(t *testing.T) {
t.Parallel()
var (
h *handlers.Handlers
db *database.Database
)
app := newTestApp(t, &h, &db)
app.RequireStart()
t.Cleanup(app.RequireStop)
seedOperator(t, db)
// Well past the limit, and from the same bucket the operator
// will arrive in.
floodFailures(
t, h, sharedProxyPeer, operatorUser,
loginFailureLimit*2,
)
w := submitLogin(
h, sharedProxyPeer, operatorUser, operatorPassword,
)
assert.Equal(
t, http.StatusSeeOther, w.Code,
"a correct password must never be throttled: the operator "+
"has no second administrative path",
)
assert.Equal(t, "/", w.Header().Get("Location"))
}
// TestLogin_StrangersFloodCannotDenyAnotherAccount is the
// cross-account half: flooding one username must not spend another
// account's budget, even from the same shared bucket.
func TestLogin_StrangersFloodCannotDenyAnotherAccount(t *testing.T) {
t.Parallel()
var (
h *handlers.Handlers
db *database.Database
)
app := newTestApp(t, &h, &db)
app.RequireStart()
t.Cleanup(app.RequireStop)
seedOperator(t, db)
floodFailures(
t, h, sharedProxyPeer, "someone-else",
loginFailureLimit*2,
)
w := submitLogin(h, sharedProxyPeer, operatorUser, "wrong")
assert.Equal(
t, http.StatusUnauthorized, w.Code,
"a flood against one username must not spend another "+
"account's failure budget",
)
}
// TestLogin_RepeatedWrongPasswordsAreThrottled is the brute-force
// half. Verifying before counting must not remove the throttle:
// repeated wrong passwords for one username from one client key run
// out of budget and are answered 429 with a Retry-After.
func TestLogin_RepeatedWrongPasswordsAreThrottled(t *testing.T) {
t.Parallel()
var (
h *handlers.Handlers
db *database.Database
)
app := newTestApp(t, &h, &db)
app.RequireStart()
t.Cleanup(app.RequireStop)
seedOperator(t, db)
for i := range loginFailureLimit - 1 {
w := submitLogin(
h, sharedProxyPeer, operatorUser,
fmt.Sprintf("guess-%d", i),
)
assert.Equal(
t, http.StatusUnauthorized, w.Code,
"attempt %d is still inside the budget", i,
)
}
w := submitLogin(h, sharedProxyPeer, operatorUser, "guess-last")
assert.Equal(
t, http.StatusTooManyRequests, w.Code,
"wrong passwords must still run out of budget",
)
assert.NotEmpty(
t, w.Header().Get("Retry-After"),
"a throttled login must say when to come back",
)
}
// TestLogin_SuccessForgivesEarlierMistakes covers the operator who
// mistypes several times and then gets it right: the successful
// attempt clears the counter, so the next mistake is answered 401
// rather than 429.
func TestLogin_SuccessForgivesEarlierMistakes(t *testing.T) {
t.Parallel()
var (
h *handlers.Handlers
db *database.Database
)
app := newTestApp(t, &h, &db)
app.RequireStart()
t.Cleanup(app.RequireStop)
seedOperator(t, db)
floodFailures(
t, h, sharedProxyPeer, operatorUser,
loginFailureLimit,
)
require.Equal(
t, http.StatusSeeOther,
submitLogin(
h, sharedProxyPeer, operatorUser, operatorPassword,
).Code,
)
w := submitLogin(h, sharedProxyPeer, operatorUser, "typo")
assert.Equal(
t, http.StatusUnauthorized, w.Code,
"a success must forgive the failures before it",
)
}
// TestLogin_UnknownUsernameCostsTheSameVerification is the
// username-enumeration guard. Verifying credentials before the
// limiter means response time is observable per attempt, so an
// unknown username must be charged an equivalent-cost verification
// against a dummy hash rather than returning early.
//
// The assertion is on the code path, not on wall-clock time: timing
// assertions are flaky, and what actually has to hold is that the
// hash is computed.
func TestLogin_UnknownUsernameCostsTheSameVerification(t *testing.T) {
t.Parallel()
var (
h *handlers.Handlers
db *database.Database
)
app := newTestApp(t, &h, &db)
app.RequireStart()
t.Cleanup(app.RequireStop)
seedOperator(t, db)
require.Zero(t, h.DummyVerificationsForTest())
// A username that exists, with the wrong password: a real
// Argon2id verification runs, and no dummy is needed.
require.Equal(
t, http.StatusUnauthorized,
submitLogin(h, sharedProxyPeer, operatorUser, "wrong").Code,
)
assert.Zero(
t, h.DummyVerificationsForTest(),
"a known username verifies against its own hash",
)
// A username that does not exist: indistinguishable response,
// and the equivalent-cost verification must have run.
require.Equal(
t, http.StatusUnauthorized,
submitLogin(h, sharedProxyPeer, "nosuchuser", "wrong").Code,
)
assert.Equal(
t, uint64(1), h.DummyVerificationsForTest(),
"an unknown username must still pay for a hash, or the "+
"response time says whether the account exists",
)
}
// TestLogin_ConcurrentLoginsAreAllAnswered covers the login path
// under the verification bound. The bound itself is pinned in the
// middleware package; what matters here is that funnelling every
// login through two slots does not lose or wedge a request — each one
// is answered, whether it got a slot or was shed with 503.
func TestLogin_ConcurrentLoginsAreAllAnswered(t *testing.T) {
t.Parallel()
const workers = 4
var (
h *handlers.Handlers
db *database.Database
)
app := newTestApp(t, &h, &db)
app.RequireStart()
t.Cleanup(app.RequireStop)
seedOperator(t, db)
var (
wg sync.WaitGroup
mu sync.Mutex
answers = map[int]int{}
)
for i := range workers {
wg.Go(func() {
w := submitLogin(
h, fmt.Sprintf("203.0.113.%d:5000", i),
operatorUser, fmt.Sprintf("guess-%d", i),
)
mu.Lock()
answers[w.Code]++
mu.Unlock()
})
}
wg.Wait()
mu.Lock()
defer mu.Unlock()
assert.Zero(
t, answers[http.StatusInternalServerError],
"concurrent logins must not error",
)
assert.Equal(
t, workers,
answers[http.StatusUnauthorized]+
answers[http.StatusTooManyRequests]+
answers[http.StatusServiceUnavailable],
"every concurrent login must be answered, whether it got "+
"a verification slot or was shed with 503",
)
}
// TestLogin_MissingCredentialsRejectedBeforeAnyHash pins that the
// empty-field check still runs ahead of the verification slot, so a
// client sending nothing cannot occupy one.
func TestLogin_MissingCredentialsRejectedBeforeAnyHash(t *testing.T) {
t.Parallel()
var (
h *handlers.Handlers
db *database.Database
)
app := newTestApp(t, &h, &db)
app.RequireStart()
t.Cleanup(app.RequireStop)
seedOperator(t, db)
w := submitLogin(h, sharedProxyPeer, "", "")
assert.Equal(t, http.StatusBadRequest, w.Code)
assert.Zero(
t, h.DummyVerificationsForTest(),
"an empty submission must not cost a hash",
)
}
// TestLogin_SuccessCreatesSession is the control for the tests above:
// the success path they assert on really does authenticate.
func TestLogin_SuccessCreatesSession(t *testing.T) {
t.Parallel()
var (
h *handlers.Handlers
db *database.Database
sess *session.Session
)
app := newTestApp(t, &h, &db, &sess)
app.RequireStart()
t.Cleanup(app.RequireStop)
seedOperator(t, db)
w := submitLogin(
h, sharedProxyPeer, operatorUser, operatorPassword,
)
require.Equal(t, http.StatusSeeOther, w.Code)
require.NotEmpty(
t, w.Result().Cookies(), "a session cookie must be issued",
)
next := httptest.NewRequestWithContext(
context.Background(), http.MethodGet, "/", nil,
)
// Login regenerates the session, so the response carries two
// Set-Cookie headers under the same name: one expiring the
// pre-login cookie and one issuing the new one. A browser keeps
// only the second, so replay only the one that is not an
// expiry.
for _, c := range w.Result().Cookies() {
if c.MaxAge >= 0 {
next.AddCookie(c)
}
}
s, err := sess.Get(next)
require.NoError(t, err)
assert.True(
t, sess.IsAuthenticated(s),
"the issued cookie must carry an authenticated session",
)
}

View File

@@ -1,199 +0,0 @@
package handlers
import (
"database/sql"
"errors"
"net/http"
"strconv"
"github.com/go-chi/chi"
"github.com/google/uuid"
"gorm.io/gorm"
"sneak.berlin/go/webhooker/internal/database"
)
// eventBodyQuery reads one event's stored body as bytes. The cast
// to blob is what makes the driver hand back the stored bytes
// rather than a string conversion, so Content-Length taken from
// the result matches what goes on the wire. The soft-delete
// predicate is spelled out because Raw bypasses GORM's default
// scope, and it is what stops a reaped event still being
// downloadable.
const eventBodyQuery = "SELECT cast(body as blob) " +
"FROM events WHERE id = ? AND webhook_id = ? AND deleted_at IS NULL"
// HandleEventBodyDownload serves one event's stored body in
// full, which the event log page cannot: it caps each rendered
// body at maxRenderedBodyBytes.
//
// The bytes are attacker-supplied — anyone who can reach the
// public receiver chooses them — and this route hands them back
// inside the operator's own authenticated origin, so the
// response is deliberately not renderable. Content-Disposition
// makes the browser download rather than display it, and the
// octet-stream type plus nosniff stop it being interpreted as
// HTML or script. Without those a stored payload would execute
// as the logged-in operator. The application's CSP does not
// help here: script-src allows 'unsafe-inline' from 'self', so
// a document served from this origin could run its own inline
// script.
func (h *Handlers) HandleEventBodyDownload() http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
webhook, ok := h.ownedWebhook(w, r)
if !ok {
return
}
// Parsing the id before use serves two purposes: a
// malformed id can never reach the SQL or the response
// header, and the canonical form below is drawn from
// uuid's own fixed alphabet rather than from the
// request, so the Content-Disposition value cannot be
// steered by a client.
eventID, err := uuid.Parse(chi.URLParam(r, "eventID"))
if err != nil {
http.NotFound(w, r)
return
}
h.serveEventBody(w, r, webhook, eventID.String())
}
}
// serveEventBody writes the named event's stored body to w.
//
// The event must belong to webhook, which is what keeps this
// route from reading any event in the system by id alone. Two
// things enforce that and they are not equally strong. The
// operative one is that events live in a per-webhook SQLite
// file, so a sibling webhook's event is not in the database
// being queried at all. The webhook_id predicate on the query
// below is the second guard, and it is currently redundant
// against that isolation; it is there so the scoping survives
// any future change that puts more than one webhook's events in
// one file.
//
// The body is read in one query and held whole in memory while
// it is written. That costs roughly two body-sized allocations
// per concurrent download, not one: the driver's column buffer
// and the copy database/sql makes in convertAssign when a
// []byte column is scanned into a *[]byte are live at the same
// time. Measured allocation is ~2x the body plus ~45 KB, so at
// the 1 MB ingest cap a download costs ~2 MB of Go heap. On
// top of that, SQLite's own materialisation of the column
// value sits in the driver's allocator outside the Go heap, so
// process peak is higher again: 2x is a floor, not a ceiling.
// There is no cheaper bound available — database/sql exposes
// no incremental handle on a SQLite BLOB, and reading byte
// ranges with substr does not avoid the cost either, because
// SQLite materialises the whole column value to evaluate each
// substr call. Range reads only pay for that materialisation
// once per range.
//
// One consequence is worth keeping in view: the read finishes
// before the client is written to, so no read lock is held for
// the length of a slow download. These per-webhook databases
// run in SQLite's default journal mode rather than WAL, so a
// lock held that long would block the receiver from recording
// new events.
func (h *Handlers) serveEventBody(
w http.ResponseWriter,
r *http.Request,
webhook database.Webhook,
eventID string,
) {
if !h.dbMgr.DBExists(webhook.ID) {
http.NotFound(w, r)
return
}
webhookDB, err := h.dbMgr.GetDB(webhook.ID)
if err != nil {
h.serverError(w, "failed to get webhook database", err)
return
}
body, found, err := eventBody(webhookDB, webhook.ID, eventID)
if err != nil {
h.serverError(w, "failed to read event body", err)
return
}
// A miss is a 404 whether the event belongs to another
// webhook or does not exist at all, so the response does
// not report which. Reading the body before any header is
// written is also what keeps an event reaped mid-request
// from producing a torn response: either the read finds the
// row and the whole body is served, or it does not and the
// response is a clean 404.
if !found {
http.NotFound(w, r)
return
}
setEventBodyHeaders(w, eventID, int64(len(body)))
_, err = w.Write(body)
if err != nil {
// The status and Content-Length are already committed,
// so the client sees a short download. There is no way
// to report a 500 from here; the log is the record.
h.log.Error(
"failed to write event body",
"webhook_id", webhook.ID,
"event_id", eventID,
"error", err,
)
}
}
// eventBody returns an event's stored body and whether the event
// exists within the webhook.
func eventBody(
webhookDB *gorm.DB,
webhookID, eventID string,
) ([]byte, bool, error) {
var body []byte
err := webhookDB.Raw(
eventBodyQuery, eventID, webhookID,
).Row().Scan(&body)
if errors.Is(err, sql.ErrNoRows) {
return nil, false, nil
}
if err != nil {
return nil, false, err
}
return body, true, nil
}
// setEventBodyHeaders applies the response headers that make
// this route safe to hand attacker-supplied bytes through. See
// HandleEventBodyDownload for why they are a security control
// and not a formatting choice.
//
// nosniff is also set by the global SecurityHeaders middleware.
// It is repeated here so the guarantee belongs to the route
// that needs it rather than to a middleware someone could
// reorder or scope away.
func setEventBodyHeaders(
w http.ResponseWriter,
eventID string,
size int64,
) {
w.Header().Set("Content-Type", "application/octet-stream")
w.Header().Set("X-Content-Type-Options", "nosniff")
w.Header().Set(
"Content-Disposition",
`attachment; filename="webhooker-event-`+eventID+`.bin"`,
)
w.Header().Set("Content-Length", strconv.FormatInt(size, 10))
}

View File

@@ -1,506 +0,0 @@
package handlers_test
import (
"context"
"net/http"
"net/http/httptest"
"net/url"
"strconv"
"strings"
"testing"
"github.com/go-chi/chi"
"github.com/google/uuid"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
"gorm.io/gorm/clause"
"sneak.berlin/go/webhooker/internal/database"
"sneak.berlin/go/webhooker/internal/handlers"
"sneak.berlin/go/webhooker/internal/session"
)
// paramEventID is the chi URL parameter the body download
// handler reads.
const paramEventID = "eventID"
// otherTestUserID owns webhooks the session user must not be
// able to read.
const otherTestUserID = "other-user-id"
// seedWebhookFor inserts a webhook owned by the given user.
func seedWebhookFor(
t *testing.T,
db *database.Database,
userID string,
) *database.Webhook {
t.Helper()
wh := &database.Webhook{
UserID: userID,
Name: "wh-" + userID,
}
require.NoError(
t,
db.DB().Omit(clause.Associations).Create(wh).Error,
)
return wh
}
// fetchEventBody runs the real download handler as the test user
// for the given source and event ids.
func fetchEventBody(
t *testing.T,
h *handlers.Handlers,
sess *session.Session,
sourceID, eventID string,
) *httptest.ResponseRecorder {
t.Helper()
// The path is escaped and the raw id goes in the route
// context, which is what chi hands a handler: the param is
// already percent-decoded by the time it is read.
req := httptest.NewRequestWithContext(
context.Background(),
http.MethodGet,
"/source/"+url.PathEscape(sourceID)+
"/logs/"+url.PathEscape(eventID)+"/body",
nil,
)
for _, c := range authenticatedCookies(
t, sess, deleteTestUserID, deleteTestUsername,
) {
req.AddCookie(c)
}
rctx := chi.NewRouteContext()
rctx.URLParams.Add(paramSourceID, sourceID)
rctx.URLParams.Add(paramEventID, eventID)
req = req.WithContext(
context.WithValue(
req.Context(), chi.RouteCtxKey, rctx,
),
)
w := httptest.NewRecorder()
h.HandleEventBodyDownload().ServeHTTP(w, req)
return w
}
// TestHandleEventBodyDownload_ServesOversizeBodyInFull is the
// capability the render cap took away: a body far above what the
// event log page will show comes back whole and byte-identical,
// with the headers that keep it from being rendered.
func TestHandleEventBodyDownload_ServesOversizeBodyInFull(
t *testing.T,
) {
t.Parallel()
var (
h *handlers.Handlers
sess *session.Session
db *database.Database
dbMgr *database.WebhookDBManager
)
app := newTestApp(t, &h, &sess, &db, &dbMgr)
app.RequireStart()
t.Cleanup(app.RequireStop)
// Far above the render cap, with multibyte runes and a
// distinctive tail, so a body that the log page can only
// show a slice of comes back whole and in order.
const sentinel = "TAIL-SENTINEL-1f4a9c"
stored := strings.Repeat("A", 200*1024) +
strings.Repeat(snowman, 1000) + sentinel
wh := seedWebhook(t, db)
evt := seedEventWithBody(t, dbMgr, wh.ID, stored)
w := fetchEventBody(t, h, sess, wh.ID, evt.ID)
require.Equal(t, http.StatusOK, w.Code)
assert.Greater(t, len(stored), bodyCap)
assert.Equal(t, stored, w.Body.String())
assert.Equal(
t, strconv.Itoa(len(stored)),
w.Header().Get("Content-Length"),
)
}
// TestHandleEventBodyDownload_BodiesRoundTripByteIdentical
// covers the sizes and byte values a stored body can actually
// take: empty, one byte, either side of the render cap, and
// bytes that are not text at all. Content-Length has to equal
// the bytes written in every case, since it is derived from the
// same read that produces them.
func TestHandleEventBodyDownload_BodiesRoundTripByteIdentical(
t *testing.T,
) {
t.Parallel()
// A NUL, invalid UTF-8 and a multibyte rune, so nothing on
// the path can be treating the body as text.
binary := "\x00\x01\xff\xfe" + snowman + "\x00tail"
cases := map[string]string{
"empty": "",
"single byte": "x",
"one below cap": strings.Repeat("b", bodyCap-1),
"exactly cap": strings.Repeat("c", bodyCap),
"one above cap": strings.Repeat("d", bodyCap+1),
"binary": binary,
}
for name, stored := range cases {
t.Run(name, func(t *testing.T) {
t.Parallel()
var (
h *handlers.Handlers
sess *session.Session
db *database.Database
dbMgr *database.WebhookDBManager
)
app := newTestApp(t, &h, &sess, &db, &dbMgr)
app.RequireStart()
t.Cleanup(app.RequireStop)
wh := seedWebhook(t, db)
evt := seedEventWithBody(t, dbMgr, wh.ID, stored)
w := fetchEventBody(t, h, sess, wh.ID, evt.ID)
require.Equal(t, http.StatusOK, w.Code)
assert.Equal(t, stored, w.Body.String())
assert.Equal(
t, strconv.Itoa(len(stored)),
w.Header().Get("Content-Length"),
)
assert.Equal(
t, len(stored), w.Body.Len(),
"Content-Length must equal bytes written",
)
})
}
}
// TestHandleEventBodyDownload_HeadersAreNotRenderable pins the
// response headers that stop attacker-supplied bytes executing
// in the operator's own origin. They are a security control, not
// presentation.
func TestHandleEventBodyDownload_HeadersAreNotRenderable(
t *testing.T,
) {
t.Parallel()
var (
h *handlers.Handlers
sess *session.Session
db *database.Database
dbMgr *database.WebhookDBManager
)
app := newTestApp(t, &h, &sess, &db, &dbMgr)
app.RequireStart()
t.Cleanup(app.RequireStop)
wh := seedWebhook(t, db)
evt := seedEventWithBody(t, dbMgr, wh.ID, `{"small":true}`)
w := fetchEventBody(t, h, sess, wh.ID, evt.ID)
require.Equal(t, http.StatusOK, w.Code)
assert.Equal(
t, "application/octet-stream",
w.Header().Get("Content-Type"),
)
assert.Equal(
t, "nosniff",
w.Header().Get("X-Content-Type-Options"),
)
disposition := w.Header().Get("Content-Disposition")
assert.Equal(
t,
`attachment; filename="webhooker-event-`+evt.ID+`.bin"`,
disposition,
)
}
// TestHandleEventBodyDownload_ScriptBodyStaysInert proves a
// stored HTML payload is handed back as an attachment of opaque
// bytes rather than as anything a browser will execute. The
// bytes themselves are unaltered: this route reports what was
// delivered.
func TestHandleEventBodyDownload_ScriptBodyStaysInert(
t *testing.T,
) {
t.Parallel()
var (
h *handlers.Handlers
sess *session.Session
db *database.Database
dbMgr *database.WebhookDBManager
)
app := newTestApp(t, &h, &sess, &db, &dbMgr)
app.RequireStart()
t.Cleanup(app.RequireStop)
const payload = `<html><script>alert(document.cookie)` +
`</script></html>`
wh := seedWebhook(t, db)
evt := seedEventWithBody(t, dbMgr, wh.ID, payload)
w := fetchEventBody(t, h, sess, wh.ID, evt.ID)
require.Equal(t, http.StatusOK, w.Code)
assert.Equal(t, payload, w.Body.String())
contentType := w.Header().Get("Content-Type")
assert.Equal(t, "application/octet-stream", contentType)
assert.NotContains(t, contentType, "html")
assert.NotContains(t, contentType, "xml")
assert.NotContains(t, contentType, "javascript")
assert.Contains(
t, w.Header().Get("Content-Disposition"), "attachment",
)
assert.Equal(
t, "nosniff",
w.Header().Get("X-Content-Type-Options"),
)
}
// TestHandleEventBodyDownload_OtherUsersEvent404s is the
// authorization test the definition of done asks for: an event
// stored under a webhook the session user does not own is not
// readable, and the miss does not distinguish itself from a
// nonexistent one.
func TestHandleEventBodyDownload_OtherUsersEvent404s(
t *testing.T,
) {
t.Parallel()
var (
h *handlers.Handlers
sess *session.Session
db *database.Database
dbMgr *database.WebhookDBManager
)
app := newTestApp(t, &h, &sess, &db, &dbMgr)
app.RequireStart()
t.Cleanup(app.RequireStop)
const theirPayload = "OTHER-USERS-PAYLOAD-8b1d"
theirs := seedWebhookFor(t, db, otherTestUserID)
evt := seedEventWithBody(t, dbMgr, theirs.ID, theirPayload)
w := fetchEventBody(t, h, sess, theirs.ID, evt.ID)
assert.Equal(t, http.StatusNotFound, w.Code)
assert.NotContains(t, w.Body.String(), theirPayload)
}
// TestHandleEventBodyDownload_EventOfAnotherWebhook404s pins
// that holding a valid event id is not enough: the event has to
// belong to the webhook in the path. Both webhooks here are the
// session user's and both have event databases, so the
// ownership check cannot be what produces the 404.
//
// What does produce it is the per-webhook database file rather
// than the webhook_id predicate on the query — removing that
// predicate leaves this test green, because the sibling's event
// is in a different file. The test is kept as the behavioural
// guard the route owes; see serveEventBody for which mechanism
// is load-bearing.
func TestHandleEventBodyDownload_EventOfAnotherWebhook404s(
t *testing.T,
) {
t.Parallel()
var (
h *handlers.Handlers
sess *session.Session
db *database.Database
dbMgr *database.WebhookDBManager
)
app := newTestApp(t, &h, &sess, &db, &dbMgr)
app.RequireStart()
t.Cleanup(app.RequireStop)
const other = "BELONGS-TO-THE-OTHER-WEBHOOK-3c7e"
mine := seedWebhook(t, db)
seedEventWithBody(t, dbMgr, mine.ID, `{"mine":true}`)
sibling := seedWebhook(t, db)
evt := seedEventWithBody(t, dbMgr, sibling.ID, other)
w := fetchEventBody(t, h, sess, mine.ID, evt.ID)
assert.Equal(t, http.StatusNotFound, w.Code)
assert.NotContains(t, w.Body.String(), other)
}
// TestHandleEventBodyDownload_UnknownEvent404s covers the plain
// miss, including an id that is not a uuid at all and so never
// reaches the query or the response header.
func TestHandleEventBodyDownload_UnknownEvent404s(t *testing.T) {
t.Parallel()
var (
h *handlers.Handlers
sess *session.Session
db *database.Database
dbMgr *database.WebhookDBManager
)
app := newTestApp(t, &h, &sess, &db, &dbMgr)
app.RequireStart()
t.Cleanup(app.RequireStop)
wh := seedWebhook(t, db)
seedEventWithBody(t, dbMgr, wh.ID, `{"mine":true}`)
for _, id := range []string{
uuid.New().String(),
`../../etc/passwd`,
"not-a-uuid",
`x"; rm -rf /`,
} {
w := fetchEventBody(t, h, sess, wh.ID, id)
assert.Equal(
t, http.StatusNotFound, w.Code,
"event id %q", id,
)
assert.Empty(
t, w.Header().Get("Content-Disposition"),
"event id %q must not reach a header", id,
)
}
}
// TestHandleEventBodyDownload_ReapedEvent404s pins what happens
// when the retention reaper takes an event out from under this
// route. The body is read in one query before any header is
// written, so a reaped event cannot produce a partial download:
// it is a clean 404 with no Content-Length and no
// Content-Disposition. Both removals the codebase performs are
// covered — the reaper hard-deletes, and a soft-deleted row is
// excluded by the query's own deleted_at predicate rather than
// by GORM's default scope, which Raw bypasses.
func TestHandleEventBodyDownload_ReapedEvent404s(t *testing.T) {
t.Parallel()
for name, hard := range map[string]bool{
"soft deleted": false,
"hard deleted": true,
} {
t.Run(name, func(t *testing.T) {
t.Parallel()
var (
h *handlers.Handlers
sess *session.Session
db *database.Database
dbMgr *database.WebhookDBManager
)
app := newTestApp(t, &h, &sess, &db, &dbMgr)
app.RequireStart()
t.Cleanup(app.RequireStop)
const payload = "REAPED-PAYLOAD-4d2a"
wh := seedWebhook(t, db)
evt := seedEventWithBody(t, dbMgr, wh.ID, payload)
webhookDB, err := dbMgr.GetDB(wh.ID)
require.NoError(t, err)
del := webhookDB
if hard {
del = del.Unscoped()
}
require.NoError(
t,
del.Delete(&database.Event{}, "id = ?", evt.ID).
Error,
)
w := fetchEventBody(t, h, sess, wh.ID, evt.ID)
assert.Equal(t, http.StatusNotFound, w.Code)
assert.NotContains(t, w.Body.String(), payload)
assert.Empty(t, w.Header().Get("Content-Length"))
assert.Empty(
t, w.Header().Get("Content-Disposition"),
)
})
}
}
// TestHandleSourceLogs_TruncationMarkerLinksToDownload proves
// the page tells the reader where the rest of the body is, and
// only when there is a rest to fetch.
func TestHandleSourceLogs_TruncationMarkerLinksToDownload(
t *testing.T,
) {
t.Parallel()
var (
h *handlers.Handlers
sess *session.Session
db *database.Database
dbMgr *database.WebhookDBManager
)
app := newTestApp(t, &h, &sess, &db, &dbMgr)
app.RequireStart()
t.Cleanup(app.RequireStop)
big := seedWebhook(t, db)
bigEvt := seedEventWithBody(
t, dbMgr, big.ID, strings.Repeat("A", 4*bodyCap),
)
page := renderSourceLogsPage(t, h, sess, big.ID)
assert.Contains(
t, page,
"/source/"+big.ID+"/logs/"+bigEvt.ID+"/body",
)
small := seedWebhook(t, db)
smallEvt := seedEventWithBody(
t, dbMgr, small.ID, `{"kept":"whole"}`,
)
page = renderSourceLogsPage(t, h, sess, small.ID)
assert.NotContains(
t, page,
"/source/"+small.ID+"/logs/"+smallEvt.ID+"/body",
)
}

View File

@@ -25,14 +25,13 @@ const bodyCap = handlers.MaxRenderedBodyBytesForTest
const snowman = "☃"
// seedEventWithBody records one event with the given body in the
// webhook's own database and returns it, so a caller that needs
// the generated event id can have it.
// webhook's own database.
func seedEventWithBody(
t *testing.T,
dbMgr *database.WebhookDBManager,
webhookID string,
body string,
) *database.Event {
) {
t.Helper()
webhookDB, err := dbMgr.GetDB(webhookID)
@@ -48,8 +47,6 @@ func seedEventWithBody(
require.NoError(t, webhookDB.Omit(
clause.Associations,
).Create(event).Error)
return event
}
// seedAndProject stores one body and returns the projection the

View File

@@ -2,31 +2,15 @@ package handlers
import (
"html/template"
"log/slog"
"net/http"
"sneak.berlin/go/webhooker/internal/database"
)
// SetLogForTest replaces the handler's logger, so the handlers_test
// package can assert on what a log line actually contains rather than
// on what it is meant to contain.
func (s *Handlers) SetLogForTest(log *slog.Logger) {
s.log = log
}
// MaxRenderedBodyBytesForTest exposes the event log's body cap
// to the handlers_test package.
const MaxRenderedBodyBytesForTest = maxRenderedBodyBytes
// DummyVerificationsForTest reports how many equivalent-cost
// verifications were charged for usernames that do not exist. It
// lets a test prove the anti-enumeration path ran without timing
// anything.
func (s *Handlers) DummyVerificationsForTest() uint64 {
return s.dummyVerifications.Load()
}
// TrimPartialRuneForTest exposes trimPartialRune for use in the
// handlers_test package.
func TrimPartialRuneForTest(b []byte) []byte {

View File

@@ -1,462 +0,0 @@
package handlers_test
import (
"bytes"
"context"
"io"
"log"
"net/http"
"net/http/httptest"
"os"
"strconv"
"strings"
"sync"
"testing"
"time"
"github.com/go-chi/chi"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
"gorm.io/gorm"
gormlogger "gorm.io/gorm/logger"
"sneak.berlin/go/webhooker/internal/database"
"sneak.berlin/go/webhooker/internal/handlers"
"sneak.berlin/go/webhooker/internal/middleware"
)
// gormBoundTailMarker sits at the far end of every client-chosen value
// this file sends. Its presence in the log means the whole value
// reached the log, so a value that merely happened to be short cannot
// pass for a truncated one.
const gormBoundTailMarker = "ENDOFCLIENTVALUE"
// gormBoundFills are the characters a client can drive through the
// receiver path segment and the login username, chosen for what a log
// handler charges for them.
//
// The bare C0 control is the one that matters: both handlers spell
// U+0001 as a six-byte escape for the one byte it costs to send, the
// widest multiplier available below U+10000 and the case a raw-byte
// budget breaks on first. GORM's default logger applies no budget at
// all, so under the mutation every one of these arrives whole.
func gormBoundFills() []struct {
name string
fill string
} {
return []struct {
name string
fill string
}{
{"plain", "x"},
{"quote", `"`},
{"backslash", `\`},
{"tab", "\t"},
{"newline", "\n"},
{"c0_control", "\x01"},
{"astral_nonprintable", "\U0001000C"},
}
}
// syncBuf collects captured output from the goroutine draining the
// pipe.
type syncBuf struct {
mu sync.Mutex
b bytes.Buffer
}
func (s *syncBuf) Write(p []byte) (int, error) {
s.mu.Lock()
defer s.mu.Unlock()
return s.b.Write(p)
}
func (s *syncBuf) String() string {
s.mu.Lock()
defer s.mu.Unlock()
return s.b.String()
}
func (s *syncBuf) reset() {
s.mu.Lock()
defer s.mu.Unlock()
s.b.Reset()
}
// stdoutCapture redirects os.Stdout for the duration of a test.
//
// internal/logger builds its handler over os.Stdout at construction
// time, so redirecting the variable before the application is built
// captures everything the service logger — and therefore the GORM
// adapter, which writes through it — emits.
type stdoutCapture struct {
buf *syncBuf
r *os.File
w *os.File
orig *os.File
done chan struct{}
seq int
}
func captureStdout(t *testing.T) *stdoutCapture {
t.Helper()
r, w, err := os.Pipe()
require.NoError(t, err)
c := &stdoutCapture{
buf: &syncBuf{},
r: r,
w: w,
orig: os.Stdout,
done: make(chan struct{}),
}
os.Stdout = w
go func() {
defer close(c.done)
_, _ = io.Copy(c.buf, r)
}()
t.Cleanup(func() {
os.Stdout = c.orig
_ = w.Close()
<-c.done
_ = r.Close()
})
return c
}
// drain returns everything written since the previous drain and
// clears the buffer.
//
// A sentinel is pushed through the same pipe and waited for, so the
// draining goroutine is known to have caught up before the buffer is
// read. Without it the comparison below would race the reader rather
// than measure the writers.
func (c *stdoutCapture) drain(t *testing.T) string {
t.Helper()
c.seq++
sentinel := "\n<<drain-" + strconv.Itoa(c.seq) + ">>\n"
_, err := c.w.WriteString(sentinel)
require.NoError(t, err)
deadline := time.Now().Add(10 * time.Second)
for !strings.Contains(c.buf.String(), sentinel) {
require.False(
t, time.Now().After(deadline),
"timed out waiting for captured output",
)
time.Sleep(time.Millisecond)
}
out := strings.Replace(c.buf.String(), sentinel, "", 1)
c.buf.reset()
return out
}
// teeStdout writes to a buffer and to whatever os.Stdout is at the
// moment of the write.
//
// The second half is the point. GORM's package-level default logger
// resolves os.Stdout once, at package init, so a logger built over the
// variable would keep writing to the real terminal no matter what a
// test redirects. Resolving it per write puts the bytes a defaulted
// gorm.Config would cost in production into the same capture as
// everything else internal/logger emits, which is what lets the volume
// assertions below measure the whole writer set rather than one member
// of it.
type teeStdout struct {
buf *syncBuf
}
func (w teeStdout) Write(p []byte) (int, error) {
_, _ = os.Stdout.Write(p)
return w.buf.Write(p)
}
// captureGORMDefault replaces GORM's package-level default logger with
// one configured exactly as GORM configures its own, writing to a
// buffer and to os.Stdout.
//
// This is the mutation detector. gormlogger.Default is what a bare
// &gorm.Config{} installs, and its config here is GORM's verbatim —
// Warn, IgnoreRecordNotFoundError false — so a reverted call site
// behaves as it would in production rather than as a test dialed it.
// With every gorm.Open in this service naming its own logger, nothing
// consults this value and the buffer stays empty; revert any one of
// the three and the interpolated SQL lands here.
func captureGORMDefault(t *testing.T) *syncBuf {
t.Helper()
buf := &syncBuf{}
orig := gormlogger.Default
gormlogger.Default = gormlogger.New(
log.New(teeStdout{buf: buf}, "", log.LstdFlags),
gormlogger.Config{
SlowThreshold: 200 * time.Millisecond,
LogLevel: gormlogger.Warn,
IgnoreRecordNotFoundError: false,
Colorful: false,
},
)
t.Cleanup(func() { gormlogger.Default = orig })
return buf
}
// floodUnauthenticated drives reps requests at each of the two
// unauthenticated lookups that miss by design, for every fill, with a
// client-chosen value of size raw bytes.
func floodUnauthenticated(
t *testing.T, h *handlers.Handlers, size, reps int,
) int {
t.Helper()
requests := 0
for _, f := range gormBoundFills() {
var b strings.Builder
for b.Len() < size {
b.WriteString(f.fill)
}
b.WriteString(gormBoundTailMarker)
value := b.String()
for range reps {
postWebhook(t, h, value)
postUnknownLogin(t, h, value)
requests += 2
}
}
return requests
}
// floodPerWebhook drives the same client-chosen values at the second
// gorm.Open site, the per-webhook database internal/database's
// WebhookDBManager opens.
//
// That site is behind authentication in production, so this is not
// part of the unauthenticated flood above and is counted separately.
// It is here because the ceiling the README states covers every
// writer, and the manager is one of them: with nothing driving it, a
// bare &gorm.Config{} could be restored at
// internal/database/webhook_db_manager.go and the whole suite would
// stay green.
func floodPerWebhook(
t *testing.T, mgr *database.WebhookDBManager, size, reps int,
) int {
t.Helper()
requests := 0
for _, f := range gormBoundFills() {
var b strings.Builder
for b.Len() < size {
b.WriteString(f.fill)
}
b.WriteString(gormBoundTailMarker)
value := b.String()
db, err := mgr.GetDB("pin-" + f.name)
require.NoError(t, err)
for range reps {
var got database.Event
err = db.Where("id = ?", value).First(&got).Error
require.ErrorIs(t, err, gorm.ErrRecordNotFound)
requests++
}
}
return requests
}
// postWebhook drives the receiver with an invented entrypoint path.
// The route pattern matches any single segment, so every byte of the
// value is the client's, and the lookup behind it misses by design.
func postWebhook(
t *testing.T, h *handlers.Handlers, entrypoint string,
) {
t.Helper()
req := httptest.NewRequestWithContext(
context.Background(), http.MethodPost, "/webhook/x",
strings.NewReader("{}"),
)
rctx := chi.NewRouteContext()
rctx.URLParams.Add("uuid", entrypoint)
req = req.WithContext(context.WithValue(
req.Context(), chi.RouteCtxKey, rctx,
))
w := httptest.NewRecorder()
h.HandleWebhook().ServeHTTP(w, req)
require.Equal(t, http.StatusNotFound, w.Code)
}
// postUnknownLogin submits the login form with an unknown username,
// through the postLogin helper in logbound_test.go. The field is
// bounded only by the 1 MB body cap, and the lookup behind it misses
// by design.
func postUnknownLogin(
t *testing.T, h *handlers.Handlers, username string,
) {
t.Helper()
// 401 while the client still has failure budget against this
// username, 429 once the login guard has taken it away. Both
// outcomes sit behind the user lookup, which is the query this
// test is here to drive.
require.Contains(
t,
[]int{http.StatusUnauthorized, http.StatusTooManyRequests},
postLogin(t, h, username),
)
}
// assertFloodBounded holds every captured line to the stated ceiling
// and proves nothing carried a whole client value.
func assertFloodBounded(t *testing.T, label, out string) {
t.Helper()
assert.NotContains(
t, out, gormBoundTailMarker,
"%s: the far end of a client-chosen value reached the "+
"log, so nothing truncated it", label,
)
for line := range strings.SplitSeq(
strings.TrimRight(out, "\n"), "\n",
) {
if line == "" {
continue
}
assert.LessOrEqual(
t, len(line), middleware.MaxAccessLogLineBytes,
"%s: log line exceeded its bound: %s",
label, line[:min(len(line), 300)],
)
}
}
// TestFlood_NoWriterGrowsWithTheInput is the definition of done for
// the GORM logger defect, stated over every writer at once, for two of
// this service's three gorm.Open sites: the main database behind the
// two unauthenticated lookups, and the per-webhook database the
// WebhookDBManager opens. The third, the archive writer, is pinned in
// internal/delivery, where its type lives.
//
// What each assertion is worth, since two of the three would pass
// against a service that had never been fixed if the capture were set
// up differently:
//
// - The gormDefault check is the sharp one. It fires the moment any
// gorm.Open in this service goes back to a bare &gorm.Config{}.
// - The volume and per-line checks bite only because the replaced
// default logger tees into os.Stdout, so a reverted call site
// shows up in the same capture as everything internal/logger
// writes — the way it would in production. Without that tee both
// were vacuous: at INFO the two handler misses log at DEBUG and
// the adapter drops the record-not-found, so the capture holds
// nothing but fixed-string warnings.
//
// The level is left where newTestApp leaves it, at INFO: the level an
// operator runs at by default, and the one the defect was visible at.
// The handlers' own miss lines sit at DEBUG and spend the same
// logfield budget as everything else, so they are not what makes
// either assertion above bite at any level.
//
// It is deliberately not parallel: it redirects os.Stdout and replaces
// gormlogger.Default, both of which are process-global. Go runs every
// non-parallel top-level test to completion before it resumes the
// parallel ones, so nothing else in this package is running while the
// capture is installed.
//
//nolint:paralleltest // Deliberately sequential; see above.
func TestFlood_NoWriterGrowsWithTheInput(t *testing.T) {
const (
smallBytes = 128
bigBytes = 8 << 10
reps = 5
)
gormDefault := captureGORMDefault(t)
capture := captureStdout(t)
var (
h *handlers.Handlers
mgr *database.WebhookDBManager
)
app := newTestApp(t, &h, &mgr)
app.RequireStart()
t.Cleanup(app.RequireStop)
// Startup chatter is not what this test measures.
capture.drain(t)
floodUnauthenticated(t, h, smallBytes, reps)
floodPerWebhook(t, mgr, smallBytes, reps)
small := capture.drain(t)
requests := floodUnauthenticated(t, h, bigBytes, reps)
requests += floodPerWebhook(t, mgr, bigBytes, reps)
big := capture.drain(t)
assertFloodBounded(t, "small flood", small)
assertFloodBounded(t, "big flood", big)
// GORM's default logger is what the defect was. Nothing in this
// service may reach it.
got := gormDefault.String()
assert.Empty(
t, got,
"GORM's default logger wrote %d bytes; the first of them: %s",
len(got), got[:min(len(got), 300)],
)
// The same flood, with 64 times the client-chosen input, must not
// buy 64 times the log. A few bytes of slack covers a latency
// field changing width; the input grew by roughly half a megabyte.
const slackPerRequest = 64
assert.LessOrEqual(
t, len(big), len(small)+slackPerRequest*requests,
"log volume tracked the size of the client's input: "+
"%d bytes at %d bytes of input per request, %d bytes "+
"at %d",
len(small), smallBytes, len(big), bigBytes,
)
}

View File

@@ -10,7 +10,6 @@ import (
"html/template"
"log/slog"
"net/http"
"sync/atomic"
"go.uber.org/fx"
"sneak.berlin/go/webhooker/internal/database"
@@ -40,12 +39,6 @@ const (
// errInvalidPassword is returned when a password does not match.
var errInvalidPassword = errors.New("invalid password")
// errVerificationBusy is returned when no password-verification slot
// became free before the wait elapsed, so no password was verified.
var errVerificationBusy = errors.New(
"password verification capacity exhausted",
)
//nolint:revive // HandlersParams is a standard fx naming convention.
type HandlersParams struct {
fx.In
@@ -56,7 +49,6 @@ type HandlersParams struct {
WebhookDBMgr *database.WebhookDBManager
Healthcheck *healthcheck.Healthcheck
Session *session.Session
Middleware *middleware.Middleware
Notifier delivery.Notifier
Evictor delivery.WebhookEvictor
}
@@ -70,15 +62,9 @@ type Handlers struct {
db *database.Database
dbMgr *database.WebhookDBManager
session *session.Session
mw *middleware.Middleware
notifier delivery.Notifier
evictor delivery.WebhookEvictor
templates map[string]*template.Template
// dummyVerifications counts the equivalent-cost verifications
// charged for usernames that do not exist. It exists so a test
// can prove that path runs without measuring wall-clock time.
dummyVerifications atomic.Uint64
}
// parsePageTemplate parses a page-specific template set from the
@@ -111,7 +97,6 @@ func New(
s.db = params.Database
s.dbMgr = params.WebhookDBMgr
s.session = params.Session
s.mw = params.Middleware
s.notifier = params.Notifier
s.evictor = params.Evictor

View File

@@ -20,7 +20,6 @@ import (
"sneak.berlin/go/webhooker/internal/handlers"
"sneak.berlin/go/webhooker/internal/healthcheck"
"sneak.berlin/go/webhooker/internal/logger"
"sneak.berlin/go/webhooker/internal/middleware"
"sneak.berlin/go/webhooker/internal/session"
)
@@ -83,7 +82,6 @@ func newTestApp(
func(r *recordingEvictor) delivery.WebhookEvictor {
return r
},
middleware.New,
handlers.New,
),
fx.Populate(targets...),

View File

@@ -1,543 +0,0 @@
package handlers_test
// The handler-side half of the log-field audit. Two slog calls in
// this package reach a value an UNAUTHENTICATED client picks outright
// and of a length it picks outright:
//
// - the unknown-entrypoint DEBUG line on /webhook/{uuid}, whose
// path segment matched no stored entrypoint and so is bounded by
// nothing;
// - the failed-login DEBUG lines, whose username is a form field.
//
// Both are at DEBUG, which is off in production by default. That is
// not a bound: an operator turning DEBUG on to diagnose a flood must
// not thereby hand the flood an unbounded write. Both spend the same
// internal/logfield budget as the access log, and both are held here
// to middleware.MaxAccessLogLineBytes.
//
// The two login lines past the username lookup — "invalid password"
// and "user logged in" — carry the same cap without needing it, since
// by then the value is a stored row rather than the client's. They are
// pinned here too, so the caps cannot be dropped silently.
//
// So is the "password verification capacity exhausted" WARN line,
// whose path chi pins to the constant "/pages/login" on the one route
// that reaches it. Its cap is defensive, and the test below drives the
// handler directly with the path a parameterised route would give it,
// because an unasserted cap is one a later edit removes for free.
import (
"bytes"
"context"
"io"
"log/slog"
"net/http"
"net/http/httptest"
"net/url"
"strings"
"testing"
"github.com/go-chi/chi"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
"sneak.berlin/go/webhooker/internal/database"
"sneak.berlin/go/webhooker/internal/handlers"
"sneak.berlin/go/webhooker/internal/middleware"
)
// floodRequests is the number of distinct invented values each flood
// drives through the call site under test.
const floodRequests = 32
// oversizedFillBytes is the length of the single client-chosen value
// used to show that line size does not track input size.
const oversizedFillBytes = 8192
// attackerMarker and tailMarker sit at the END of every oversized
// value, past every budget. Their absence from the log is what
// proves the value was cut rather than merely being short.
const (
attackerMarker = "QQATTACKERTEXTQQ"
tailMarker = "QQTRUNCATEDTAILQQ"
)
// escapeFills are the characters the log handlers escape, so a value
// built out of them costs more on the line than it did on the wire. A
// budget counted in raw bytes passes the plain case and fails these.
//
// U+1000C is unassigned, hence non-printable, and strconv.Quote
// spells it as a ten-byte \UXXXXXXXX while the JSON handler passes
// its four UTF-8 bytes through; only the text shape of these tests
// reaches that charge.
func escapeFills() map[string]string {
return map[string]string{
"plain": "x",
"quote": `"`,
"backslash": `\`,
"tab": "\t",
"newline": "\n",
// A C0 control neither handler has a short escape for, so
// each one costs six bytes on the line against the single
// byte it cost to send: the widest multiplier a client can
// drive, and the case a raw-byte budget breaks on first.
//
// This fill is load-bearing, not decoration. Budgeting raw
// bytes instead of encoded is caught by this fill alone,
// and only under the JSON handler, at 3,072 bytes against
// the 2,560 ceiling. Drop it and that mutation passes.
"control": "\x01",
"astral": "\U0001000C",
}
}
// logHandlers are the two handlers internal/logger can install.
func logHandlers() map[string]func(
io.Writer, *slog.HandlerOptions,
) slog.Handler {
return map[string]func(
io.Writer, *slog.HandlerOptions,
) slog.Handler{
"json": func(
w io.Writer, o *slog.HandlerOptions,
) slog.Handler {
return slog.NewJSONHandler(w, o)
},
"text": func(
w io.Writer, o *slog.HandlerOptions,
) slog.Handler {
return slog.NewTextHandler(w, o)
},
}
}
// oversizedFill builds an 8 KB client-chosen value out of
// repetitions of ch, with both markers at its far end.
func oversizedFill(ch string) string {
return "x" + strings.Repeat(ch, oversizedFillBytes) +
attackerMarker + tailMarker
}
// capturingHandlers builds a Handlers whose log is captured into the
// returned buffer at DEBUG through the named handler.
//
// extra is passed to fx.Populate alongside the Handlers, for the call
// sites that also need the database the client's value is looked up
// in, or the Middleware whose resource has to be exhausted before the
// branch under test is reached.
func capturingHandlers(
t *testing.T,
newHandler func(io.Writer, *slog.HandlerOptions) slog.Handler,
extra ...any,
) (*handlers.Handlers, *bytes.Buffer) {
t.Helper()
var h *handlers.Handlers
app := newTestApp(t, append([]any{&h}, extra...)...)
app.RequireStart()
t.Cleanup(app.RequireStop)
buf := new(bytes.Buffer)
h.SetLogForTest(slog.New(newHandler(
buf, &slog.HandlerOptions{Level: slog.LevelDebug},
)))
return h, buf
}
// logLines splits the captured buffer into non-empty lines, holding
// each to the stated per-line ceiling.
func logLines(t *testing.T, buf *bytes.Buffer) []string {
t.Helper()
var lines []string
for line := range strings.SplitSeq(
strings.TrimSpace(buf.String()), "\n",
) {
if line == "" {
continue
}
require.LessOrEqual(
t, len(line), middleware.MaxAccessLogLineBytes,
"log line exceeded its bound: %s", line,
)
lines = append(lines, line)
}
return lines
}
// assertNoClientText fails if the far end of the client-chosen input
// survived into the log.
func assertNoClientText(t *testing.T, buf *bytes.Buffer) {
t.Helper()
assert.NotContains(
t, buf.String(), attackerMarker,
"log carried attacker-chosen text",
)
assert.NotContains(
t, buf.String(), tailMarker,
"log carried the tail of the attacker-chosen text",
)
}
// receiverRouter mounts the real receiver handler at the production
// route pattern.
func receiverRouter(h *handlers.Handlers) *chi.Mux {
router := chi.NewRouter()
router.Post("/webhook/{uuid}", h.HandleWebhook())
return router
}
// postReceiver sends one POST at /webhook/<segment>.
//
// RawPath is cleared after parsing so chi routes on the decoded path
// and the handler sees the raw bytes rather than their percent-escaped
// spelling. That is the harder case for the budget: the escaped
// spelling is plain ASCII, which costs one byte per byte, while the
// decoded bytes are what the log handler has to escape.
func postReceiver(
t *testing.T, router *chi.Mux, segment string,
) int {
t.Helper()
req := httptest.NewRequestWithContext(
context.Background(),
http.MethodPost,
"/webhook/"+url.PathEscape(segment),
strings.NewReader(""),
)
req.URL.RawPath = ""
w := httptest.NewRecorder()
router.ServeHTTP(w, req)
return w.Code
}
// postLogin submits the login form with the given username and a
// non-empty password.
func postLogin(
t *testing.T, h *handlers.Handlers, username string,
) int {
t.Helper()
return postLoginWithPassword(t, h, username, "not-the-password")
}
// postLoginWithPassword submits the login form with both credentials
// chosen by the caller, so a test can reach the branches past the
// username lookup.
func postLoginWithPassword(
t *testing.T, h *handlers.Handlers, username, password string,
) int {
t.Helper()
form := url.Values{
"username": {username},
"password": {password},
}
req := httptest.NewRequestWithContext(
context.Background(),
http.MethodPost,
"/pages/login",
strings.NewReader(form.Encode()),
)
req.Header.Set(
"Content-Type", "application/x-www-form-urlencoded",
)
w := httptest.NewRecorder()
h.HandleLoginSubmit().ServeHTTP(w, req)
return w.Code
}
// TestUnknownEntrypoint_LogLineDoesNotTrackPathSize drives 8 KB of
// client-chosen path at the unauthenticated receiver's
// unknown-entrypoint DEBUG line and holds it to the same ceiling the
// access log states.
func TestUnknownEntrypoint_LogLineDoesNotTrackPathSize(t *testing.T) {
t.Parallel()
for handlerName, newHandler := range logHandlers() {
for fillName, fill := range escapeFills() {
t.Run(handlerName+"/"+fillName, func(t *testing.T) {
t.Parallel()
h, buf := capturingHandlers(t, newHandler)
router := receiverRouter(h)
for i := range floodRequests {
assert.Equal(
t,
http.StatusNotFound,
postReceiver(
t, router,
oversizedFill(fill)+
strings.Repeat("y", i),
),
)
}
lines := logLines(t, buf)
require.Len(t, lines, floodRequests)
assertNoClientText(t, buf)
assertBoundedFlood(t, buf.Len())
})
}
}
}
// TestFailedLogin_LogLineDoesNotTrackUsernameSize drives 8 KB of
// client-chosen username at the unauthenticated login endpoint's
// DEBUG line and holds it to the same ceiling.
func TestFailedLogin_LogLineDoesNotTrackUsernameSize(t *testing.T) {
t.Parallel()
for handlerName, newHandler := range logHandlers() {
for fillName, fill := range escapeFills() {
t.Run(handlerName+"/"+fillName, func(t *testing.T) {
t.Parallel()
h, buf := capturingHandlers(t, newHandler)
for i := range floodRequests {
assert.Equal(
t,
http.StatusUnauthorized,
postLogin(
t, h,
oversizedFill(fill)+
strings.Repeat("y", i),
),
)
}
lines := logLines(t, buf)
require.Len(t, lines, floodRequests)
assertNoClientText(t, buf)
assertBoundedFlood(t, buf.Len())
})
}
}
}
// storedUserPassword is the password held by the oversize accounts
// the test below creates.
const storedUserPassword = "correct-horse-battery-staple"
// storedFillBytes is the raw length of the client-chosen value in
// those accounts' usernames. It is well past the 512-byte field
// budget, so the line is still truncated, but short enough that the
// session cookie a successful login writes stays inside
// securecookie's 4 KB limit: the cookie is written BEFORE the
// "user logged in" line, so an 8 KB username answers 500 and never
// reaches it.
const storedFillBytes = 1024
// storedFill builds a username fill of storedFillBytes raw bytes out
// of repetitions of ch, with both markers at its far end.
func storedFill(ch string) string {
return "x" + strings.Repeat(ch, storedFillBytes/len(ch)) +
attackerMarker + tailMarker
}
// TestStoredUsername_LogLinesDoNotTrackUsernameSize pins the two
// login lines that are reached only AFTER the username matched a
// stored row: "invalid password" and "user logged in". Neither
// strictly needs its cap — the value is the operator's own data by
// then, not the client's — but both carry one so that every username
// this unauthenticated endpoint logs is capped, and an unasserted cap
// is one a later edit removes for free.
//
// One app per handler with the accounts created inside it, and no
// parallelism below that level: every account costs an Argon2id hash
// and every attempt costs a verification.
func TestStoredUsername_LogLinesDoNotTrackUsernameSize(t *testing.T) {
t.Parallel()
for handlerName, newHandler := range logHandlers() {
t.Run(handlerName, func(t *testing.T) {
t.Parallel()
var db *database.Database
h, buf := capturingHandlers(t, newHandler, &db)
hash, err := database.HashPassword(storedUserPassword)
require.NoError(t, err)
fills := escapeFills()
for fillName, fill := range fills {
username := storedFill(fill) + fillName
require.NoError(t, db.DB().Create(&database.User{
Username: username,
Password: hash,
}).Error)
// Matched the row, wrong secret: "invalid
// password".
assert.Equal(
t, http.StatusUnauthorized,
postLoginWithPassword(
t, h, username, "not-the-password",
),
)
// Matched the row, right secret: "user logged
// in".
assert.Equal(
t, http.StatusSeeOther,
postLoginWithPassword(
t, h, username, storedUserPassword,
),
)
}
lines := logLines(t, buf)
require.Len(t, lines, 2*len(fills))
assertNoClientText(t, buf)
})
}
}
// maxVerificationSlots bounds how many slots the loop below will
// take before it gives up, so a semaphore that never fills fails the
// test instead of hanging it. It is deliberately larger than the
// real concurrency bound, which is not exported to this package.
const maxVerificationSlots = 64
// canceledContext returns a context that is already done. A
// verification request carrying one takes the ctx.Done() branch of
// the semaphore's bounded wait immediately, so these cases turn on
// the semaphore being full rather than on a five-second timer firing.
// Nothing here is timing-dependent.
func canceledContext() context.Context {
ctx, cancel := context.WithCancel(context.Background())
cancel()
return ctx
}
// holdEveryVerificationSlot takes verification slots until one is
// refused, and releases them when the test ends. A free slot is
// handed out before any context is consulted, so a canceled context
// cannot make this loop stop early: it stops exactly when the slots
// are gone.
func holdEveryVerificationSlot(
t *testing.T, mw *middleware.Middleware,
) {
t.Helper()
for range maxVerificationSlots {
release, ok := mw.BeginPasswordVerification(canceledContext())
if !ok {
return
}
t.Cleanup(release)
}
require.Fail(t, "the verification semaphore never filled")
}
// postLoginAtPath submits the login form at a path of the caller's
// choosing, with a canceled context.
func postLoginAtPath(
t *testing.T, h *handlers.Handlers, path string,
) int {
t.Helper()
form := url.Values{
"username": {"someone"},
"password": {"not-the-password"},
}
req := httptest.NewRequestWithContext(
canceledContext(),
http.MethodPost,
path,
strings.NewReader(form.Encode()),
)
req.Header.Set(
"Content-Type", "application/x-www-form-urlencoded",
)
w := httptest.NewRecorder()
h.HandleLoginSubmit().ServeHTTP(w, req)
return w.Code
}
// TestVerificationCapacity_LogLineDoesNotTrackPathSize pins the cap
// on the "password verification capacity exhausted" WARN line.
//
// The one route that reaches it is chi's static "/pages/login", so no
// request through the mux can widen the line; the handler is driven
// directly here with the path a parameterised route would give it,
// which is what that cap exists for. Without this test, removing the
// logfield.Truncate there fails nothing.
func TestVerificationCapacity_LogLineDoesNotTrackPathSize(
t *testing.T,
) {
t.Parallel()
for handlerName, newHandler := range logHandlers() {
for fillName, fill := range escapeFills() {
t.Run(handlerName+"/"+fillName, func(t *testing.T) {
t.Parallel()
var mw *middleware.Middleware
h, buf := capturingHandlers(t, newHandler, &mw)
holdEveryVerificationSlot(t, mw)
assert.Equal(
t,
http.StatusServiceUnavailable,
postLoginAtPath(
t, h,
"/source/"+url.PathEscape(
oversizedFill(fill),
)+"/login",
),
)
lines := logLines(t, buf)
require.Len(t, lines, 1)
assertNoClientText(t, buf)
})
}
}
}
// assertBoundedFlood holds the whole flood's log output to what the
// stated per-line ceiling allows. The flood sent
// floodRequests * oversizedFillBytes bytes of client-chosen text;
// this is the assertion that the log did not grow with it.
func assertBoundedFlood(t *testing.T, got int) {
t.Helper()
sent := floodRequests * oversizedFillBytes
require.Less(
t, got, sent/2,
"log volume tracked the size of the flood's input",
)
require.LessOrEqual(
t, got,
floodRequests*middleware.MaxAccessLogLineBytes,
)
}

View File

@@ -1,7 +1,6 @@
package handlers
import (
"context"
"net/http"
"github.com/go-chi/chi"
@@ -43,14 +42,11 @@ func (h *Handlers) HandlePasswordChange() http.HandlerFunc {
}
successMessage, errorMessage, handled := h.applyPasswordChange(
r.Context(),
w,
sessionUsername,
// PostFormValue, not FormValue: the credential must
// come from the body, never from the query string.
r.PostFormValue("current_password"),
r.PostFormValue("new_password"),
r.PostFormValue("confirm_password"),
r.FormValue("current_password"),
r.FormValue("new_password"),
r.FormValue("confirm_password"),
)
if !handled {
return
@@ -70,30 +66,9 @@ func (h *Handlers) HandlePasswordChange() http.HandlerFunc {
// 500 response itself and returns handled=false, signalling the caller
// to stop without re-rendering the page.
func (h *Handlers) applyPasswordChange(
ctx context.Context,
w http.ResponseWriter,
username, currentPassword, newPassword, confirmPassword string,
) (string, string, bool) {
// This endpoint verifies one password and hashes another, at
// 64 MB each, so it takes a slot from the same bound the login
// endpoint uses. The bound is per hash, not per endpoint: leaving
// this path outside it would leave a hole in it. The slot is held
// across both hashes.
release, ok := h.mw.BeginPasswordVerification(ctx)
if !ok {
h.log.Warn("password verification capacity exhausted")
http.Error(
w,
"The server is busy verifying credentials. "+
"Please try again.",
http.StatusServiceUnavailable,
)
return "", "", false
}
defer release()
// Load the user row so we can verify the current password and
// persist the new hash.
var user database.User

View File

@@ -227,9 +227,9 @@ func (h *Handlers) HandleSourceCreateSubmit() http.HandlerFunc {
return
}
name := r.PostFormValue("name")
description := r.PostFormValue("description")
retentionStr := r.PostFormValue("retention_days")
name := r.FormValue("name")
description := r.FormValue("description")
retentionStr := r.FormValue("retention_days")
if name == "" {
w.WriteHeader(http.StatusBadRequest)
@@ -509,7 +509,7 @@ func (h *Handlers) applyWebhookEdit(
) {
// The body size cap is enforced by the MaxBodySize middleware,
// which runs before CSRF parses the form.
name := r.PostFormValue("name")
name := r.FormValue("name")
if name == "" {
data := map[string]any{
tmplKeyWebhook: webhook,
@@ -523,12 +523,12 @@ func (h *Handlers) applyWebhookEdit(
}
webhook.Name = name
webhook.Description = r.PostFormValue("description")
webhook.Description = r.FormValue("description")
// An empty field falls back to the stored value, so submitting the
// form without touching retention leaves the policy alone.
retentionDays, retErr := parseRetentionDays(
r.PostFormValue("retention_days"), webhook.RetentionDays,
r.FormValue("retention_days"), webhook.RetentionDays,
)
if retErr != nil {
data := map[string]any{
@@ -713,55 +713,29 @@ func (h *Handlers) evictArchiveWriterIfUnused(webhookID string) {
h.evictArchiveWriter(webhookID)
}
// ownedWebhook resolves the request's sourceID parameter to a
// webhook the session's user owns.
//
// Ownership and existence are decided by one query, so a
// webhook belonging to another user is indistinguishable from
// one that does not exist: both are a 404, and neither confirms
// the id. Callers that reach further into a webhook's data —
// the event log page and the event body download — share this
// one check rather than restating it, so the download cannot
// come to authorize differently from the page that links to it.
//
// It reports false once it has written the response, which is a
// redirect to the login page for an unauthenticated request and
// a 404 otherwise. The caller returns without writing more.
func (h *Handlers) ownedWebhook(
w http.ResponseWriter,
r *http.Request,
) (database.Webhook, bool) {
var webhook database.Webhook
userID, ok := h.getUserID(r)
if !ok {
http.Redirect(
w, r, "/pages/login", http.StatusSeeOther,
)
return database.Webhook{}, false
}
sourceID := chi.URLParam(r, "sourceID")
err := h.db.DB().Where(
"id = ? AND user_id = ?", sourceID, userID,
).First(&webhook).Error
if err != nil {
http.NotFound(w, r)
return database.Webhook{}, false
}
return webhook, true
}
// HandleSourceLogs shows the request/response logs for a
// webhook.
func (h *Handlers) HandleSourceLogs() http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
webhook, ok := h.ownedWebhook(w, r)
userID, ok := h.getUserID(r)
if !ok {
http.Redirect(
w, r, "/pages/login", http.StatusSeeOther,
)
return
}
sourceID := chi.URLParam(r, "sourceID")
var webhook database.Webhook
err := h.db.DB().Where(
"id = ? AND user_id = ?", sourceID, userID,
).First(&webhook).Error
if err != nil {
http.NotFound(w, r)
return
}
@@ -950,7 +924,7 @@ func (h *Handlers) HandleEntrypointCreate() http.HandlerFunc {
return
}
description := r.PostFormValue("description")
description := r.FormValue("description")
entrypoint := &database.Entrypoint{
WebhookID: webhook.ID,
@@ -1020,18 +994,11 @@ func (h *Handlers) processTargetCreate(
) {
// The body size cap is enforced by the MaxBodySize middleware,
// which runs before CSRF parses the form.
//
// Every field here is read with PostFormValue, not FormValue.
// FormValue falls back to the query string, which would let
// `POST /source/{id}/targets?url=https://hooks.slack.com/...`
// configure a target from a value the request line carries — and
// the request line, unlike the body, is what logs, proxies,
// Referer headers and error trackers record.
name := r.PostFormValue("name")
targetType := database.TargetType(r.PostFormValue("type"))
targetURL := r.PostFormValue("url")
maxRetriesStr := r.PostFormValue("max_retries")
expiry := r.PostFormValue("expiry")
name := r.FormValue("name")
targetType := database.TargetType(r.FormValue("type"))
targetURL := r.FormValue("url")
maxRetriesStr := r.FormValue("max_retries")
expiry := r.FormValue("expiry")
if name == "" {
http.Error(

View File

@@ -1,206 +0,0 @@
package handlers_test
import (
"bytes"
"context"
"log/slog"
"net/http"
"net/http/httptest"
"net/url"
"strings"
"testing"
"github.com/go-chi/chi"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
"sneak.berlin/go/webhooker/internal/config"
"sneak.berlin/go/webhooker/internal/database"
"sneak.berlin/go/webhooker/internal/middleware"
)
// targetSecretSegments are the path segments of an incoming-webhook
// URL. For Slack, Discord and Teams the path IS the bearer credential,
// so this string must not reach storage or the access log by way of
// the request line.
const targetSecretSegments = "T00000000/B00000000/QQTARGETSECRETQQ"
// targetSecretURL is a destination whose secret lives in its path. It
// uses a literal public address rather than a hostname so the SSRF
// check resolves nothing: with a hostname, a sandbox without DNS would
// reject the URL for the wrong reason and the test would pass even
// with the defect reintroduced.
const targetSecretURL = "https://93.184.216.34/services/" +
targetSecretSegments
// targetsForWebhook returns every target stored against a webhook.
func targetsForWebhook(
t *testing.T,
db *database.Database,
webhookID string,
) []database.Target {
t.Helper()
var targets []database.Target
require.NoError(
t,
db.DB().Where("webhook_id = ?", webhookID).
Find(&targets).Error,
)
return targets
}
// postTargetCreate drives HandleTargetCreate through the production
// access-log middleware and a chi route, so the logged url field is
// produced exactly as it ships, and returns the recorder plus the
// captured log.
func postTargetCreate(
t *testing.T,
env *sourceTestEnv,
webhookID string,
query string,
form url.Values,
) (*httptest.ResponseRecorder, string) {
t.Helper()
logBuf := new(bytes.Buffer)
mw := middleware.NewForTest(
slog.New(slog.NewJSONHandler(
logBuf, &slog.HandlerOptions{Level: slog.LevelInfo},
)),
&config.Config{Environment: config.EnvironmentDev},
nil,
)
router := chi.NewRouter()
router.Use(mw.Logging())
router.Post(
"/source/{sourceID}/targets",
env.handlers.HandleTargetCreate(),
)
target := "/source/" + webhookID + "/targets"
if query != "" {
target += "?" + query
}
body := ""
if form != nil {
body = form.Encode()
}
req := httptest.NewRequestWithContext(
context.Background(),
http.MethodPost,
target,
strings.NewReader(body),
)
req.Header.Set(
"Content-Type", "application/x-www-form-urlencoded",
)
for _, c := range env.cookies {
req.AddCookie(c)
}
w := httptest.NewRecorder()
router.ServeHTTP(w, req)
return w, logBuf.String()
}
// TestHandleTargetCreate_QueryStringURLDoesNotConfigureATarget is the
// regression test for the ingress leak. r.FormValue falls back to the
// query string when a field is absent from the POST body, so
//
// POST /source/{id}/targets?url=https://hooks.slack.com/services/...
//
// with an empty url field used to create a working target from a value
// carried on the request line — where logs, proxies, Referer headers
// and error trackers record it. The handler reads the body only, so
// the request is rejected for a missing URL and stores nothing.
//
// name and type are sent in the BODY on purpose: the request has to
// get past those two validations for the assertion to be about the url
// read specifically.
func TestHandleTargetCreate_QueryStringURLDoesNotConfigureATarget(
t *testing.T,
) {
t.Parallel()
env := setupSourceTest(t)
webhook := seedWebhookWithRetention(t, env.db, 30)
body := url.Values{}
body.Set("name", "leaky")
body.Set("type", string(database.TargetTypeSlack))
w, logged := postTargetCreate(
t, env, webhook.ID,
"url="+url.QueryEscape(targetSecretURL),
body,
)
assert.Equal(t, http.StatusBadRequest, w.Code)
targets := targetsForWebhook(t, env.db, webhook.ID)
assert.Empty(
t, targets,
"a query-string value must not populate a target config",
)
assert.NotContains(t, logged, targetSecretSegments)
assert.NotContains(t, logged, "93.184.216.34")
assert.NotEmpty(t, logged, "the access log line must still be written")
}
// TestHandleTargetCreate_BodyURLStillCreatesTheTarget is the positive
// control for the test above: the rejection has to come from where the
// value was read, not from the handler being broken.
func TestHandleTargetCreate_BodyURLStillCreatesTheTarget(t *testing.T) {
t.Parallel()
env := setupSourceTest(t)
webhook := seedWebhookWithRetention(t, env.db, 30)
body := url.Values{}
body.Set("name", "legit")
body.Set("type", string(database.TargetTypeSlack))
body.Set("url", targetSecretURL)
w, logged := postTargetCreate(t, env, webhook.ID, "", body)
assert.Equal(t, http.StatusSeeOther, w.Code)
targets := targetsForWebhook(t, env.db, webhook.ID)
require.Len(t, targets, 1)
assert.Contains(t, targets[0].Config, targetSecretSegments)
// The body carried the credential, so the access log must still
// not have it: the log records the request line only.
assert.NotContains(t, logged, targetSecretSegments)
}
// TestHandleTargetCreate_QueryStringCannotSupplyNameOrType covers the
// rest of the converted reads on this handler in one request: with an
// empty body, nothing the query carries is visible to it.
func TestHandleTargetCreate_QueryStringCannotSupplyNameOrType(
t *testing.T,
) {
t.Parallel()
env := setupSourceTest(t)
webhook := seedWebhookWithRetention(t, env.db, 30)
w, _ := postTargetCreate(
t, env, webhook.ID,
"name=leaky&type=slack&max_retries=9&expiry=30d&url="+
url.QueryEscape(targetSecretURL),
url.Values{},
)
assert.Equal(t, http.StatusBadRequest, w.Code)
assert.Contains(t, w.Body.String(), "Name is required")
assert.Empty(t, targetsForWebhook(t, env.db, webhook.ID))
}

View File

@@ -9,7 +9,6 @@ import (
"gorm.io/gorm"
"sneak.berlin/go/webhooker/internal/database"
"sneak.berlin/go/webhooker/internal/delivery"
"sneak.berlin/go/webhooker/internal/logfield"
)
const (
@@ -126,16 +125,9 @@ func (h *Handlers) lookupEntrypoint(
"path = ?", entrypointUUID,
).First(&entrypoint)
if result.Error != nil {
// The receiver is unauthenticated and /webhook/{uuid}
// matches any single segment, so this value is entirely
// client-chosen on exactly the branch where the lookup
// failed. DEBUG is off by default; the cap is what keeps
// turning it on from restoring an unbounded write.
h.log.Debug(
"entrypoint not found",
"path", logfield.Truncate(
entrypointUUID, logfield.MaxBytes,
),
"path", entrypointUUID,
)
http.NotFound(w, r)

View File

@@ -1,143 +0,0 @@
// Package logfield bounds the client-supplied values this service
// writes into its logs.
//
// Any log field whose content a client picks is spent against a budget
// here, in ENCODED bytes rather than in the bytes the client sent, so
// that escaping cannot multiply a field past its nominal size. One
// budget and one implementation serves the access log in
// internal/middleware and every other slog call that reaches a
// client-chosen path, header or form value; a second, ad-hoc
// truncation somewhere else in the tree is the thing this package
// exists to prevent.
package logfield
import (
"strings"
"unicode"
"unicode/utf8"
)
const (
// MaxBytes is the default budget for a log field whose value the
// client supplies outright: a URL, a path, a header, a form value.
// The budget is spent in ENCODED bytes (see Truncate), so 512 still
// holds a real browser's User-Agent whole — those are plain ASCII,
// which encodes one byte for one — while a value built from
// characters the encoder escapes keeps a shorter prefix. That is
// the intended trade: 500 quotation marks are not a debugging
// asset.
MaxBytes = 512
// TruncationMarker is appended to any field that was cut, so a
// short value and a truncated one cannot be confused. It is charged
// on top of the budget, not inside it.
TruncationMarker = "[truncated]"
)
// EncodedBytes is what r costs on the line once the log handler has
// escaped it, taking the worse of the two handlers internal/logger
// configures.
//
// slog's JSON handler escapes quote, backslash, newline, carriage
// return and tab to two bytes each, and every other C0 control plus
// LINE SEPARATOR and PARAGRAPH SEPARATOR to a six-byte \u escape; it
// passes every other rune through as its own UTF-8. Its text handler
// quotes with strconv.Quote, which spells a non-printable rune below
// U+10000 as \uXXXX but one at or above U+10000 as \UXXXXXXXX — ten
// bytes, not six. The text handler is therefore the worse of the two
// for every non-printable rune, and by four bytes apiece for the
// 955,086 unassigned, private-use and format code points on planes 1
// to 16.
//
// Charging ten there is what makes the stated line ceilings hold for
// the tty handler as well: U+1000C encodes as F0 90 80 8C, every byte
// >= 0x80, which httpguts.ValidHeaderFieldValue accepts and
// net/textproto does not strip, so a header can be filled with them.
//
// Both handlers pass printable runes through as their own UTF-8, so
// unicode.IsPrint separates the escaped cases from the plain ones for
// either handler.
func EncodedBytes(r rune) int {
const (
// A backslash and the character itself.
shortEscapeBytes = 2
// \uXXXX, which is also the width of \u00XX.
escapedRuneBytes = 6
// \UXXXXXXXX, strconv.Quote's spelling of a non-printable
// rune outside the basic multilingual plane.
escapedAstralRuneBytes = 10
// The first code point strconv.Quote spells with \U.
firstAstralRune = 0x10000
)
switch {
case r == '"' || r == '\\' || r == '\n' || r == '\r' || r == '\t':
return shortEscapeBytes
case !unicode.IsPrint(r) && r >= firstAstralRune:
return escapedAstralRuneBytes
case !unicode.IsPrint(r):
return escapedRuneBytes
default:
return utf8.RuneLen(r)
}
}
// Truncate caps s at maxBytes of ENCODED output, marking the value
// when it cuts.
//
// Budgeting raw bytes would not bound the line. Escaping only ever
// grows a value, so a raw budget spent on characters the encoder
// escapes buys a field several times its nominal size — and the line
// is the thing an operator is told to multiply by their request rate.
// Charging each rune what it will actually cost is what makes a stated
// ceiling true rather than merely larger. The visible consequence is
// that an escape-heavy value keeps a shorter prefix than a plain one,
// which is the correct trade.
//
// The result is always valid UTF-8. A cut on a byte boundary can split
// a multi-byte rune, and a header can carry bytes that were never
// valid UTF-8 to begin with; both are dropped rather than kept, since
// an encoder would otherwise spend six bytes replacing each one.
func Truncate(s string, maxBytes int) string {
// No rune encodes to fewer bytes than it occupies, so nothing past
// maxBytes raw can fit the budget. Slicing first bounds the scan
// below to the budget rather than to the size of the header the
// client sent.
window, cut := s, false
if len(window) > maxBytes {
window, cut = window[:maxBytes], true
}
var (
kept strings.Builder
spent int
)
for i := 0; i < len(window); {
r, size := utf8.DecodeRuneInString(window[i:])
if r == utf8.RuneError && size == 1 {
i += size
continue
}
cost := EncodedBytes(r)
if spent+cost > maxBytes {
cut = true
break
}
spent += cost
kept.WriteString(window[i : i+size])
i += size
}
if !cut {
return kept.String()
}
return kept.String() + TruncationMarker
}

View File

@@ -1,321 +0,0 @@
package logfield_test
import (
"bytes"
"io"
"log/slog"
"strings"
"testing"
"unicode/utf8"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
"sneak.berlin/go/webhooker/internal/logfield"
)
// budget is the field budget these tests spend. Small enough that a
// cut is unambiguous, large enough to hold several runes of every
// width.
const budget = 64
// sampleRunes is how many runes wide the values in the charge test
// are. The handlers add a constant per field — a pair of quotes when
// the value needs quoting — so the per-rune charge is only visible
// once it is amortised over a run of them.
const sampleRunes = 64
// quotingSlack is that constant: the pair of quotes a handler adds to
// a value that needs them and omits from one that does not.
const quotingSlack = 2
// newHandlers are the two handlers internal/logger can install. Time
// is dropped so a line's width is a function of its value alone —
// RFC3339Nano trims trailing zeros, so two consecutive timestamps do
// not render to the same number of bytes.
func newHandlers() map[string]func(io.Writer) slog.Handler {
opts := &slog.HandlerOptions{
Level: slog.LevelDebug,
ReplaceAttr: func(_ []string, a slog.Attr) slog.Attr {
if a.Key == slog.TimeKey {
return slog.Attr{}
}
return a
},
}
return map[string]func(io.Writer) slog.Handler{
"json": func(w io.Writer) slog.Handler {
return slog.NewJSONHandler(w, opts)
},
"text": func(w io.Writer) slog.Handler {
return slog.NewTextHandler(w, opts)
},
}
}
// renderedWidth is the number of bytes a handler writes for a line
// carrying value in a single attribute.
func renderedWidth(
newHandler func(io.Writer) slog.Handler,
value string,
) int {
buf := new(bytes.Buffer)
slog.New(newHandler(buf)).Info("m", "v", value)
return buf.Len()
}
// chargeTestRunes is the set of code points the charge test measures:
// every rune in the first two planes' worth of the BMP that the
// handlers are most likely to treat specially, the separators that
// only slog's JSON handler escapes, and a stratified sample across
// the rest of Unicode so the astral charge is exercised on more than
// one hand-picked rune.
func chargeTestRunes() []rune {
const (
denseCeiling = 0x800
stride = 1021
surrogateLo = 0xD800
surrogateHi = 0xDFFF
)
var runes []rune
keep := func(r rune) {
if r >= surrogateLo && r <= surrogateHi {
return
}
runes = append(runes, r)
}
for r := range rune(denseCeiling) {
keep(r)
}
for _, r := range []rune{
0x2028, 0x2029, 0x200B, 0x4E00, 0xE000, 0xFFFD,
0x1000C, 0x1F600, 0xE0001, 0x10FFFF,
} {
keep(r)
}
for r := rune(denseCeiling); r <= utf8.MaxRune; r += stride {
keep(r)
}
return runes
}
// TestEncodedBytes_ChargesAtLeastWhatTheHandlersEmit is the property
// the whole capping scheme rests on: a rune may not cost more on the
// line than the budget was charged for it. An undercharged rune is
// how a stated ceiling becomes false without any test noticing, so
// the charge is measured against what the handlers actually write
// rather than against the escaping rules as read.
func TestEncodedBytes_ChargesAtLeastWhatTheHandlersEmit(t *testing.T) {
t.Parallel()
for name, newHandler := range newHandlers() {
t.Run(name, func(t *testing.T) {
t.Parallel()
// 'a' is a printable ASCII rune, charged exactly one
// byte, so it is the zero point the other runes are
// measured against.
base := renderedWidth(
newHandler, strings.Repeat("a", sampleRunes),
)
for _, r := range chargeTestRunes() {
got := renderedWidth(
newHandler,
strings.Repeat(string(r), sampleRunes),
)
charged := sampleRunes *
(logfield.EncodedBytes(r) - 1)
require.LessOrEqual(
t, got-base, charged+quotingSlack,
"U+%04X costs more on the line than "+
"EncodedBytes charges for it",
r,
)
}
})
}
}
// TestTruncate_SpendsNoMoreThanTheBudget holds the result to the
// budget in ENCODED bytes, which is the unit the budget is stated in.
// A raw-byte cap passes the ASCII case here and fails every other
// one.
func TestTruncate_SpendsNoMoreThanTheBudget(t *testing.T) {
t.Parallel()
for name, fill := range map[string]string{
"plain": "x",
"quote": `"`,
"backslash": `\`,
"tab": "\t",
"newline": "\n",
"control": "\x01",
"astral": "\U0001000C",
// U+4E00, a printable multi-byte rune, charged its three
// UTF-8 bytes rather than an escape. Spelled numerically
// because gosmopolitan rejects Han in a string literal.
"cjk": string(rune(0x4E00)),
} {
t.Run(name, func(t *testing.T) {
t.Parallel()
got := logfield.Truncate(
strings.Repeat(fill, budget*8), budget,
)
require.True(
t, strings.HasSuffix(
got, logfield.TruncationMarker,
),
"an oversized value must be marked as cut",
)
spent := 0
for _, r := range strings.TrimSuffix(
got, logfield.TruncationMarker,
) {
spent += logfield.EncodedBytes(r)
}
assert.LessOrEqual(t, spent, budget)
assert.True(t, utf8.ValidString(got))
})
}
}
// TestTruncate_LeavesShortValuesAlone keeps the marker meaningful: a
// value that fits comes back byte for byte, so a marked value is
// always a cut one.
func TestTruncate_LeavesShortValuesAlone(t *testing.T) {
t.Parallel()
for _, s := range []string{
"", "GET", "/source/abc/edit", "Mozilla/5.0 (X11)",
} {
assert.Equal(t, s, logfield.Truncate(s, budget))
}
}
// TestTruncate_DropsInvalidUTF8 covers the bytes a header can carry
// that were never valid UTF-8. Keeping them would make the encoder
// spend six bytes apiece replacing them, which is exactly the
// amplification the budget exists to prevent.
func TestTruncate_DropsInvalidUTF8(t *testing.T) {
t.Parallel()
got := logfield.Truncate("a\xffb\xfe\xfec", budget)
assert.Equal(t, "abc", got)
assert.True(t, utf8.ValidString(got))
}
// encodedCost is what a whole string costs on a line, by the same
// accounting Truncate spends its budget with.
func encodedCost(s string) int {
total := 0
for _, r := range s {
total += logfield.EncodedBytes(r)
}
return total
}
// TestTruncate_SpendsEncodedBytesNotRawBytes is the zero-headroom
// version of TestTruncate_SpendsNoMoreThanTheBudget above, and of the
// line-length assertions elsewhere.
//
// A line ceiling has slack in it by construction, and a LessOrEqual
// against the budget cannot tell a budget spent exactly from one
// spent under. Here the budget is checked against exactly what it
// bought: a value built from a single rune must keep exactly
// MaxBytes/EncodedBytes(r) of them, with nothing spare. A raw-byte
// budget — cost := utf8.RuneLen(r) — fails this for every rune the
// handlers escape.
func TestTruncate_SpendsEncodedBytesNotRawBytes(t *testing.T) {
t.Parallel()
for name, r := range map[string]rune{
"plain": 'x',
"quote": '"',
"backslash": '\\',
"tab": '\t',
"newline": '\n',
"carriage_return": '\r',
"c0_control": '\x01',
"del": '\x7f',
// U+2028 LINE SEPARATOR, which only the JSON handler
// escapes.
"line_separator": '',
"astral_nonprintable": '\U0001000C',
"multibyte_printable": 'é',
// U+20AC, a three-byte printable rune, charged its own
// UTF-8 bytes rather than an escape.
"three_byte_printable": '€',
"emoji_printable": '\U0001F600',
} {
t.Run(name, func(t *testing.T) {
t.Parallel()
cost := logfield.EncodedBytes(r)
want := logfield.MaxBytes / cost
// Far past the budget under either accounting.
in := strings.Repeat(string(r), logfield.MaxBytes*2)
got := logfield.Truncate(in, logfield.MaxBytes)
require.True(
t, strings.HasSuffix(
got, logfield.TruncationMarker,
),
"a value past the budget must be marked",
)
kept := strings.TrimSuffix(
got, logfield.TruncationMarker,
)
assert.Equal(
t, want, utf8.RuneCountInString(kept),
"budget bought the wrong number of runes at "+
"%d encoded bytes each", cost,
)
assert.LessOrEqual(
t, encodedCost(kept), logfield.MaxBytes,
)
})
}
}
// TestTruncate_NeverSplitsARune covers a cut landing inside a
// multi-byte encoding rather than between two of them.
func TestTruncate_NeverSplitsARune(t *testing.T) {
t.Parallel()
// U+20AC, three bytes and printable, so a small budget lands
// inside an encoding rather than on a boundary.
in := strings.Repeat("€", logfield.MaxBytes)
for b := 1; b <= 16; b++ {
got := strings.TrimSuffix(
logfield.Truncate(in, b), logfield.TruncationMarker,
)
assert.True(
t, utf8.ValidString(got),
"budget %d produced invalid UTF-8", b,
)
assert.LessOrEqual(t, encodedCost(got), b)
}
}

View File

@@ -1,658 +0,0 @@
package middleware_test
import (
"bytes"
"context"
"encoding/json"
"log/slog"
"net/http"
"net/http/httptest"
"strings"
"testing"
"github.com/go-chi/chi"
chimw "github.com/go-chi/chi/middleware"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
"sneak.berlin/go/webhooker/internal/config"
"sneak.berlin/go/webhooker/internal/middleware"
)
// floodRequests is the number of distinct invented paths each flood
// test drives through the access log.
const floodRequests = 64
// attackerMarker is embedded in every invented path. No access log
// line for a redirected or rejected request may contain it.
const attackerMarker = "QQATTACKERTEXTQQ"
// maxLineBytes bounds a single access log line whose client-supplied
// fields are of ordinary size. Well above what the fixed fields need,
// well below the length of the oversized input the amplification tests
// send.
const maxLineBytes = 1024
// maxCappedLineBytes bounds a single access log line when every
// client-supplied field arrives oversized and is truncated to its
// budget. This is the number the README quotes as the per-line cost an
// operator sizes log storage against, and it is a bound on the
// ENCODED line, which is what the operator's disk holds.
const maxCappedLineBytes = 2560
// oversizedSegmentBytes is the length of the single attacker-chosen
// path segment, query string or header used to show line size does not
// track input size.
const oversizedSegmentBytes = 8192
// tailMarker is placed at the END of an oversized header value, so its
// absence from the log proves the value was truncated rather than
// merely being short.
const tailMarker = "QQTRUNCATEDTAILQQ"
// These mirror the middleware's own budgets, which are unexported.
// They are duplicated rather than exported so that widening a budget
// in the middleware has to be restated here deliberately.
const (
maxFieldBytes = 512
maxRequestIDBytes = 128
maxMethodBytes = 32
truncationSuffix = "[truncated]"
unmatchedRouteLiteral = "(unmatched)"
)
// capturingMiddleware returns a Middleware whose logger writes JSON
// lines into the returned buffer, so the access log can be asserted
// on directly.
func capturingMiddleware(t *testing.T) (*middleware.Middleware, *bytes.Buffer) {
t.Helper()
buf := new(bytes.Buffer)
log := slog.New(slog.NewJSONHandler(
buf,
&slog.HandlerOptions{Level: slog.LevelInfo},
))
cfg := &config.Config{Environment: config.EnvironmentDev}
return middleware.NewForTest(log, cfg, nil), buf
}
// capturingTextMiddleware is capturingMiddleware for the other handler
// internal/logger can select: slog's text handler, which
// internal/logger/logger.go installs when stderr is a tty. It escapes
// differently from the JSON one, so the line bound has to be asserted
// against both.
func capturingTextMiddleware(
t *testing.T,
) (*middleware.Middleware, *bytes.Buffer) {
t.Helper()
buf := new(bytes.Buffer)
log := slog.New(slog.NewTextHandler(
buf,
&slog.HandlerOptions{Level: slog.LevelInfo},
))
cfg := &config.Config{Environment: config.EnvironmentDev}
return middleware.NewForTest(log, cfg, nil), buf
}
// accessLogRouter mirrors the production route shapes that an
// unauthenticated client can reach: the public receiver, the
// authenticated profile route (which redirects to login rather than
// rejecting outright), the health check (which answers 200 to anyone,
// behind no rate limiter at all), and a plain static route.
func accessLogRouter(m *middleware.Middleware) *chi.Mux {
router := chi.NewRouter()
// Production registers RequestID ahead of Logging, and chi's
// RequestID passes an inbound X-Request-Id header straight
// through, so the request_id field is client-supplied too.
router.Use(chimw.RequestID)
router.Use(m.Logging())
router.Get(
"/.well-known/healthcheck",
func(w http.ResponseWriter, _ *http.Request) {
w.WriteHeader(http.StatusOK)
},
)
router.HandleFunc(
"/webhook/{uuid}",
func(w http.ResponseWriter, r *http.Request) {
// Stands in for the real handler: an unknown entrypoint
// UUID 404s, a known one succeeds.
if chi.URLParam(r, "uuid") != "known" {
http.Error(w, "not found", http.StatusNotFound)
return
}
w.WriteHeader(http.StatusOK)
},
)
router.Route("/user/{username}", func(r chi.Router) {
r.Get("/", func(w http.ResponseWriter, r *http.Request) {
http.Redirect(
w, r, "/pages/login", http.StatusSeeOther,
)
})
})
boom := func(w http.ResponseWriter, _ *http.Request) {
http.Error(w, "boom", http.StatusInternalServerError)
}
router.Get("/boom", boom)
// The 5xx branch keeps the concrete path, so it needs a route that
// answers 500 to a path of the client's choosing: that is where the
// url field and the header fields are both at their budget on the
// same line.
router.Get("/boom/*", boom)
return router
}
// accessLogEntries decodes the captured buffer into one map per
// logged line, holding every line to maxLineBytes.
func accessLogEntries(
t *testing.T,
buf *bytes.Buffer,
) []map[string]any {
t.Helper()
return accessLogEntriesWithin(t, buf, maxLineBytes)
}
// accessLogEntriesWithin decodes the captured buffer into one map per
// logged line, holding every line to bound bytes.
func accessLogEntriesWithin(
t *testing.T,
buf *bytes.Buffer,
bound int,
) []map[string]any {
t.Helper()
var entries []map[string]any
for line := range strings.SplitSeq(
strings.TrimSpace(buf.String()), "\n",
) {
if line == "" {
continue
}
require.LessOrEqual(
t, len(line), bound,
"access log line exceeded its bound",
)
var entry map[string]any
require.NoError(t, json.Unmarshal([]byte(line), &entry))
entries = append(entries, entry)
}
return entries
}
// get drives one GET through the router.
func get(t *testing.T, router *chi.Mux, target string) int {
t.Helper()
return getWithHeaders(t, router, target, nil)
}
// getWithHeaders drives one GET through the router with the supplied
// request headers set.
func getWithHeaders(
t *testing.T,
router *chi.Mux,
target string,
headers map[string]string,
) int {
t.Helper()
req := httptest.NewRequestWithContext(
context.Background(), http.MethodGet, target, nil,
)
for name, value := range headers {
req.Header.Set(name, value)
}
w := httptest.NewRecorder()
router.ServeHTTP(w, req)
return w.Code
}
// assertFloodIsBounded drives floodRequests distinct invented paths
// built by pathFor and asserts every logged line names wantURL, that
// none carries the invented text, and that the line count is exactly
// one per request.
func assertFloodIsBounded(
t *testing.T,
pathFor func(i int) string,
wantStatus int,
wantURL string,
) {
t.Helper()
m, buf := capturingMiddleware(t)
router := accessLogRouter(m)
for i := range floodRequests {
assert.Equal(t, wantStatus, get(t, router, pathFor(i)))
}
assert.NotContains(
t, buf.String(), attackerMarker,
"access log carried attacker-chosen path text",
)
entries := accessLogEntries(t, buf)
require.Len(t, entries, floodRequests)
for _, entry := range entries {
assert.Equal(t, wantURL, entry["url"])
assert.InDelta(
t, float64(wantStatus), entry["status"], 0,
)
}
}
func TestAccessLog_InventedReceiverPathsLogRoutePattern(t *testing.T) {
t.Parallel()
assertFloodIsBounded(
t,
func(i int) string {
return "/webhook/" + attackerMarker +
strings.Repeat("x", i) + "?q=" + attackerMarker
},
http.StatusNotFound,
"/webhook/{uuid}",
)
}
func TestAccessLog_InventedProfilePathsLogRoutePattern(t *testing.T) {
t.Parallel()
// The login redirect is a 3xx, not a 4xx, but it is just as free
// for an unauthenticated client to drive with invented input.
// The doubled slash is what chi's RoutePattern yields for a
// mounted subrouter's index route.
assertFloodIsBounded(
t,
func(i int) string {
return "/user/" + attackerMarker +
strings.Repeat("x", i) + "/"
},
http.StatusSeeOther,
"/user/{username}//",
)
}
func TestAccessLog_UnroutablePathsLogFixedLiteral(t *testing.T) {
t.Parallel()
assertFloodIsBounded(
t,
func(i int) string {
return "/" + attackerMarker + strings.Repeat("x", i)
},
http.StatusNotFound,
"(unmatched)",
)
}
// oversizedValue builds an 8 KB header value out of repetitions of ch,
// with the tail marker at its end.
//
// The leading 'x' is load-bearing for tab: net/textproto strips leading
// and trailing whitespace from a header value, so a value that were
// nothing but tabs would arrive empty over a real connection and the
// case would prove nothing.
func oversizedValue(ch string) string {
return "x" + strings.Repeat(ch, oversizedSegmentBytes) + tailMarker
}
// oversizedHeaders fills every client-supplied header the access log
// reads with the same value.
func oversizedHeaders(value string) map[string]string {
return map[string]string{
"User-Agent": value,
"Referer": value,
"X-Request-Id": value,
}
}
// sizeCase is one way of pointing 8 KB of client-chosen text at the
// access log.
type sizeCase struct {
target string
headers map[string]string
wantStatus int
wantURL string
bound int
}
// lineSizeCases enumerates every part of a request that reaches the
// access log, at 8 KB apiece.
func lineSizeCases() map[string]sizeCase {
cases := map[string]sizeCase{
"oversized path segment": {
target: "/webhook/" + attackerMarker +
strings.Repeat("x", oversizedSegmentBytes),
wantStatus: http.StatusNotFound,
wantURL: "/webhook/{uuid}",
bound: maxLineBytes,
},
// /.well-known/healthcheck answers 200 to anyone and has no
// rate limiter in front of it, so an oversized query appended
// to it would otherwise buy the same amplification as an
// invented 404 path, unauthenticated and unthrottled.
"oversized query on an unauthenticated 200": {
target: "/.well-known/healthcheck?q=" + attackerMarker +
strings.Repeat("x", oversizedSegmentBytes),
wantStatus: http.StatusOK,
wantURL: "/.well-known/healthcheck?(redacted)",
bound: maxLineBytes,
},
// These reach the line on every request, including one whose
// url field is correctly redacted.
"oversized headers": {
target: "/" + attackerMarker,
headers: oversizedHeaders(oversizedValue("h")),
wantStatus: http.StatusNotFound,
wantURL: unmatchedRouteLiteral,
bound: maxCappedLineBytes,
},
}
// The url field on a 5xx keeps the concrete path, so it reaches its
// own budget on the same line as the three header fields. That is
// the widest access log line the service can be made to write.
longPath := "/boom/" + strings.Repeat("x", oversizedSegmentBytes)
wantLongURL := longPath[:maxFieldBytes] + truncationSuffix
// escapeChars are the runes Go's header parser accepts in a header
// value and the log handler then escapes, coming out wider than
// they went in. A budget counted in raw bytes lets any of them buy
// a field several times its nominal size, so every one of them
// gets a case.
//
// The astral one is the case the JSON handler alone does not
// reach: U+1000C is unassigned, so it is non-printable, and
// strconv.Quote spells a non-printable rune at or above U+10000
// as a ten-byte \UXXXXXXXX. The JSON handler passes it through as
// its four UTF-8 bytes, so only the text-handler shape of this
// test holds the ten-byte charge honest.
escapeChars := map[string]string{
"quote": `"`,
"backslash": `\`,
"tab": "\t",
"astral": "\U0001000C",
}
for kind, char := range escapeChars {
fill := oversizedValue(char)
cases["oversized "+kind+" headers"] = sizeCase{
target: "/" + attackerMarker,
headers: oversizedHeaders(fill),
wantStatus: http.StatusNotFound,
wantURL: unmatchedRouteLiteral,
bound: maxCappedLineBytes,
}
cases["oversized "+kind+" headers with a 5xx concrete url"] =
sizeCase{
target: longPath,
headers: oversizedHeaders(fill),
wantStatus: http.StatusInternalServerError,
wantURL: wantLongURL,
bound: maxCappedLineBytes,
}
}
return cases
}
// TestAccessLog_LineSizeDoesNotTrackInputSize drives 8 KB of
// client-chosen text at the access log through each part of the
// request that reaches it, and holds the resulting line to a fixed
// bound in every case.
//
// The bound is on the ENCODED line, so the cases built out of
// characters the handler escapes are the ones that matter: a budget
// spent in raw bytes passes every plain-ASCII case here and still
// writes a line half again as long as the stated ceiling.
func TestAccessLog_LineSizeDoesNotTrackInputSize(t *testing.T) {
t.Parallel()
require.Equal(
t, middleware.MaxAccessLogLineBytes, maxCappedLineBytes,
"the README quotes this ceiling and the middleware derives "+
"it; they have to agree",
)
for name, tc := range lineSizeCases() {
t.Run(name, func(t *testing.T) {
t.Parallel()
m, buf := capturingMiddleware(t)
router := accessLogRouter(m)
assert.Equal(
t,
tc.wantStatus,
getWithHeaders(t, router, tc.target, tc.headers),
)
// accessLogEntriesWithin enforces the bound, which is
// orders of magnitude smaller than the input just sent.
entries := accessLogEntriesWithin(t, buf, tc.bound)
require.Len(t, entries, 1)
assert.Equal(t, tc.wantURL, entries[0]["url"])
// The markers sit at the far end of the client-chosen
// text, so their absence is what proves the redaction and
// the truncation actually ran.
assert.NotContains(
t, buf.String(), attackerMarker,
"access log carried attacker-chosen text",
)
assert.NotContains(
t, buf.String(), tailMarker,
"access log carried an untruncated client field",
)
})
}
}
// TestAccessLog_LineSizeDoesNotTrackInputSizeOnTheTextHandler runs the
// same cases through slog's text handler, which internal/logger
// selects on a tty.
//
// MaxAccessLogLineBytes is quoted to operators unqualified, so it has
// to hold for whichever handler is installed — and the two do not
// escape alike. The astral case is the one that separates them: the
// JSON handler emits U+1000C as its four UTF-8 bytes, while
// strconv.Quote spells it \U0001000C at ten. Charging six for it, as
// this code did, put a real 2,676-byte line on the wire here while
// every JSON case stayed comfortably inside the bound.
//
// Only the size bound is asserted; the url field's contents are the
// JSON shape's business above.
func TestAccessLog_LineSizeDoesNotTrackInputSizeOnTheTextHandler(
t *testing.T,
) {
t.Parallel()
for name, tc := range lineSizeCases() {
t.Run(name, func(t *testing.T) {
t.Parallel()
m, buf := capturingTextMiddleware(t)
router := accessLogRouter(m)
assert.Equal(
t,
tc.wantStatus,
getWithHeaders(t, router, tc.target, tc.headers),
)
line := strings.TrimSpace(buf.String())
require.NotEmpty(t, line)
assert.NotContains(
t, line, "\n", "expected exactly one log line",
)
require.LessOrEqual(
t, len(line), tc.bound,
"access log line exceeded its bound",
)
assert.Contains(t, line, "url=")
assert.NotContains(
t, line, attackerMarker,
"access log carried attacker-chosen text",
)
assert.NotContains(
t, line, tailMarker,
"access log carried an untruncated client field",
)
})
}
}
// TestAccessLog_OversizedMethodIsTruncated covers the last term in the
// MaxAccessLogLineBytes arithmetic that the size cases above cannot
// reach: Go accepts any RFC 7230 token as a method, and getWithHeaders
// only ever sends GET.
func TestAccessLog_OversizedMethodIsTruncated(t *testing.T) {
t.Parallel()
m, buf := capturingMiddleware(t)
router := accessLogRouter(m)
method := strings.Repeat("M", oversizedSegmentBytes) + attackerMarker
req := httptest.NewRequestWithContext(
context.Background(), method, "/"+attackerMarker, nil,
)
rec := httptest.NewRecorder()
router.ServeHTTP(rec, req)
entries := accessLogEntriesWithin(t, buf, maxLineBytes)
require.Len(t, entries, 1)
assert.Equal(
t,
strings.Repeat("M", maxMethodBytes)+truncationSuffix,
entries[0]["method"],
)
assert.NotContains(
t, buf.String(), attackerMarker,
"access log carried attacker-chosen text",
)
}
// TestAccessLog_OversizedHeadersKeepATruncatedPrefix checks the other
// half of the header cap: the fields are cut, not dropped, so a
// truncated User-Agent is still worth reading.
func TestAccessLog_OversizedHeadersKeepATruncatedPrefix(t *testing.T) {
t.Parallel()
m, buf := capturingMiddleware(t)
router := accessLogRouter(m)
assert.Equal(
t,
http.StatusNotFound,
getWithHeaders(
t, router, "/nope",
oversizedHeaders(oversizedValue("h")),
),
)
entries := accessLogEntriesWithin(t, buf, maxCappedLineBytes)
require.Len(t, entries, 1)
for key, budget := range map[string]int{
"useragent": maxFieldBytes,
"referer": maxFieldBytes,
"request_id": maxRequestIDBytes,
} {
value, ok := entries[0][key].(string)
require.True(t, ok, key)
assert.LessOrEqual(
t, len(value), budget+len(truncationSuffix), key,
)
assert.Contains(t, value, truncationSuffix, key)
assert.Contains(t, value, "hhhh", key)
}
}
func TestAccessLog_SuccessKeepsConcretePathAndRedactsQuery(
t *testing.T,
) {
t.Parallel()
m, buf := capturingMiddleware(t)
router := accessLogRouter(m)
assert.Equal(
t, http.StatusOK, get(t, router, "/webhook/known?src=ci"),
)
// The path resolved against a stored entrypoint, so it stays. The
// query never does: see TestAccessLog_UnauthenticatedSuccess...
entries := accessLogEntries(t, buf)
require.Len(t, entries, 1)
assert.Equal(t, "/webhook/known?(redacted)", entries[0]["url"])
assert.NotContains(t, buf.String(), "src=ci")
}
func TestAccessLog_ServerErrorKeepsConcreteURL(t *testing.T) {
t.Parallel()
m, buf := capturingMiddleware(t)
router := accessLogRouter(m)
assert.Equal(
t, http.StatusInternalServerError, get(t, router, "/boom"),
)
entries := accessLogEntries(t, buf)
require.Len(t, entries, 1)
assert.Equal(t, "/boom", entries[0]["url"])
}
func TestAccessLog_RetainsEveryOtherField(t *testing.T) {
t.Parallel()
m, buf := capturingMiddleware(t)
router := accessLogRouter(m)
assert.Equal(
t,
http.StatusNotFound,
get(t, router, "/webhook/"+attackerMarker),
)
entries := accessLogEntries(t, buf)
require.Len(t, entries, 1)
for _, key := range []string{
"request_start", "method", "url", "useragent", "request_id",
"referer", "proto", "remoteIP", "status", "latency_ms",
} {
assert.Contains(t, entries[0], key)
}
assert.Equal(t, http.MethodGet, entries[0]["method"])
assert.Equal(t, "HTTP/1.1", entries[0]["proto"])
}

View File

@@ -4,7 +4,6 @@ import (
"net/http"
"github.com/gorilla/csrf"
"sneak.berlin/go/webhooker/internal/logfield"
)
// CSRFToken retrieves the CSRF token from the request context.
@@ -43,22 +42,9 @@ func isClientTLS(r *http.Request) bool {
// csrf.Secure option is set at creation time, not per-request.
func (m *Middleware) CSRF() func(http.Handler) http.Handler {
csrfErrorHandler := http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
// CSRF is registered ahead of RequireAuth on every route
// group that uses it, so this WARN is reachable by an
// unauthenticated client: a POST with no token to
// /source/<any length of any text>/edit lands here. The
// method and path are capped against the same budgets as
// the access log. remote_addr is set by net/http from the
// accepted connection rather than by the client, and
// csrf.FailureReason returns one of gorilla/csrf's own
// fixed error values, so neither is client-sized.
m.log.Warn("csrf: token validation failed",
"method", logfield.Truncate(
r.Method, maxLogMethodBytes,
),
"path", logfield.Truncate(
r.URL.Path, logfield.MaxBytes,
),
"method", r.Method,
"path", r.URL.Path,
"remote_addr", r.RemoteAddr,
"reason", csrf.FailureReason(r),
)

View File

@@ -1,9 +1,7 @@
package middleware
import (
"context"
"net/http"
"time"
)
// NewLoggingResponseWriterForTest wraps newLoggingResponseWriter
@@ -37,79 +35,9 @@ func IsClientTLS(r *http.Request) bool {
return isClientTLS(r)
}
// LoginRateLimitConst exposes the loginRateLimit constant: the
// number of FAILED login attempts one client may make against one
// submitted username per interval.
// LoginRateLimitConst exposes the loginRateLimit constant.
const LoginRateLimitConst = loginRateLimit
// LoginFailureMaxKeysConst exposes the cap on each of the login
// guard's key sets.
const LoginFailureMaxKeysConst = loginFailureMaxKeys
// PasswordVerifyConcurrencyConst exposes the bound on concurrent
// Argon2id verifications.
const PasswordVerifyConcurrencyConst = passwordVerifyConcurrency
// PasswordVerifyMaxWaitersConst exposes the bound on how many
// requests may queue for a verification slot.
const PasswordVerifyMaxWaitersConst = passwordVerifyMaxWaiters
// LoginGuard is the login failure counter and verification
// semaphore, exposed for direct testing.
type LoginGuard = loginGuard
// NewLoginGuardForTest builds a guard with test-sized parameters.
func NewLoginGuardForTest(
limit int,
interval time.Duration,
maxKeys, concurrency, maxWaiters int,
wait time.Duration,
) *LoginGuard {
return newLoginGuard(
limit, interval, maxKeys, concurrency, maxWaiters, wait,
)
}
// QueuedWaitersForTest reports how many requests are currently
// queued for a verification slot.
func (g *LoginGuard) QueuedWaitersForTest() int {
return len(g.queue)
}
// SetNowForTest replaces the guard's clock.
func (g *LoginGuard) SetNowForTest(now func() time.Time) {
g.mu.Lock()
defer g.mu.Unlock()
g.now = now
}
// FailForTest exposes fail.
func (g *LoginGuard) FailForTest(clientKey, username string) bool {
return g.fail(clientKey, username)
}
// SucceedForTest exposes succeed.
func (g *LoginGuard) SucceedForTest(clientKey, username string) {
g.succeed(clientKey, username)
}
// AcquireForTest exposes acquire.
func (g *LoginGuard) AcquireForTest(
ctx context.Context,
) (func(), bool) {
return g.acquire(ctx)
}
// TrackedKeysForTest reports how many failure counters the guard
// holds, per-username and per-address respectively.
func (g *LoginGuard) TrackedKeysForTest() (int, int) {
g.mu.Lock()
defer g.mu.Unlock()
return len(g.byUser), len(g.byAddr)
}
// PasswordChangeRateLimitConst exposes the
// passwordChangeRateLimit constant.
const PasswordChangeRateLimitConst = passwordChangeRateLimit

View File

@@ -1,544 +0,0 @@
package middleware_test
// This file covers the log lines OUTSIDE the access log that carry a
// client-chosen value. accesslog_test.go bounds the one INFO line the
// Logging middleware writes; these are the separate slog calls that
// were never in that sweep and so never got the budget:
//
// - MaxBodySize's 413 rejection, at WARN, registered ahead of
// RequireAuth and therefore reachable unauthenticated at a URL of
// the client's choosing.
// - CSRF's 403 rejection, at WARN, also registered ahead of
// RequireAuth.
// - The rate limiters' 429 rejection, at WARN, on the
// unauthenticated receiver among others.
// - RequireAuth's own unauthenticated-request line, at DEBUG.
// - RecordLoginFailure's throttle rejection, at WARN. Its cap is
// defensive rather than load-bearing today: chi pins the one
// route that calls it to the constant path "/pages/login". The
// method is exported and takes any *http.Request, so the test
// below hands it the request a caller on a parameterised route
// would, which is what the cap exists for.
//
// Every case here holds the ENCODED line to
// middleware.MaxAccessLogLineBytes, under both handlers
// internal/logger can install, against 8 KB of client-chosen text
// built out of the characters those handlers escape. A budget spent
// in raw bytes passes the plain-ASCII cases and fails the rest.
import (
"bytes"
"context"
"io"
"log/slog"
"net/http"
"net/http/httptest"
"net/url"
"strings"
"testing"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
"sneak.berlin/go/webhooker/internal/config"
"sneak.berlin/go/webhooker/internal/middleware"
)
// bodyLimitBytes is the MaxBodySize cap these tests install. Any
// declared Content-Length above it takes the 413 branch.
const bodyLimitBytes = 1024
// declaredBodyBytes is the Content-Length an oversize request
// declares. Nothing is actually sent: the 413 branch fires off the
// declaration alone, which is what makes the attack free.
const declaredBodyBytes = bodyLimitBytes * 2
// receiverLimitPerMinute is the per-entrypoint receiver limit these
// tests install. The aggregate limiter sits at ten times this, so a
// flood stays under it and the rejections come from the
// per-entrypoint limiter, which is the one that logs the path.
const receiverLimitPerMinute = 8
// escapeFills are the characters a client can put in a request that
// the log handlers then escape, coming out wider than they went in.
// A budget counted in raw bytes lets any of them buy a field several
// times its nominal size.
//
// U+1000C is the case the JSON handler alone does not reach: it is
// unassigned, so it is non-printable, and strconv.Quote spells a
// non-printable rune at or above U+10000 as a ten-byte \UXXXXXXXX
// while the JSON handler passes its four UTF-8 bytes through. Only
// the text-handler shape of these tests holds that charge honest.
func escapeFills() map[string]string {
return map[string]string{
"plain": "x",
"quote": `"`,
"backslash": `\`,
"tab": "\t",
"newline": "\n",
// A C0 control neither handler has a short escape for, so
// each one costs six bytes on the line against the single
// byte it cost to send. This is the widest multiplier a
// client can drive, and the case a raw-byte budget breaks
// on first.
//
// This fill is load-bearing, not decoration. Budgeting raw
// bytes instead of encoded is caught by this fill alone,
// and only under the JSON handler, at 3,072 bytes against
// the 2,560 ceiling. Drop it and that mutation passes.
"control": "\x01",
"astral": "\U0001000C",
}
}
// logHandlers are the two handlers internal/logger can install: the
// JSON one, and the text one it selects when stderr is a tty. They do
// not escape alike, and MaxAccessLogLineBytes is quoted unqualified,
// so every case runs through both.
func logHandlers() map[string]func(
io.Writer, *slog.HandlerOptions,
) slog.Handler {
return map[string]func(
io.Writer, *slog.HandlerOptions,
) slog.Handler{
"json": func(
w io.Writer, o *slog.HandlerOptions,
) slog.Handler {
return slog.NewJSONHandler(w, o)
},
"text": func(
w io.Writer, o *slog.HandlerOptions,
) slog.Handler {
return slog.NewTextHandler(w, o)
},
}
}
// oversizedPathSegment builds an 8 KB client-chosen path segment out
// of repetitions of ch, percent-encoded so it survives URL parsing
// into r.URL.Path the way it would arriving off a socket.
//
// Both markers sit at the END, past every budget, so their absence
// from the log is what proves the value was cut rather than merely
// being short. The leading 'x' keeps the segment non-empty for fills
// that a parser might otherwise fold away.
func oversizedPathSegment(ch string) string {
return url.PathEscape(
"x" + strings.Repeat(ch, oversizedSegmentBytes) +
attackerMarker + tailMarker,
)
}
// capturingLogger returns a logger at DEBUG writing into the returned
// buffer through the named handler.
func capturingLogger(
newHandler func(io.Writer, *slog.HandlerOptions) slog.Handler,
) (*slog.Logger, *bytes.Buffer) {
buf := new(bytes.Buffer)
opts := &slog.HandlerOptions{Level: slog.LevelDebug}
return slog.New(newHandler(buf, opts)), buf
}
// capturingBoundMiddleware builds a Middleware with a real session
// manager (CSRF needs its key, RequireAuth needs its store) whose log
// is captured at DEBUG.
func capturingBoundMiddleware(
t *testing.T,
newHandler func(io.Writer, *slog.HandlerOptions) slog.Handler,
) (*middleware.Middleware, *bytes.Buffer) {
t.Helper()
log, buf := capturingLogger(newHandler)
cfg := &config.Config{
Environment: config.EnvironmentDev,
ReceiverRateLimit: receiverLimitPerMinute,
}
sess := newTestSessionManager(cfg, log, nil)
return middleware.NewForTest(log, cfg, sess), buf
}
// unreachable is a next-handler that fails the test if the middleware
// under test let the request through. Every site here rejects.
func unreachable(t *testing.T) http.Handler {
t.Helper()
return http.HandlerFunc(func(http.ResponseWriter, *http.Request) {
assert.Fail(t, "rejected request reached the next handler")
})
}
// logSite is one non-access-log call site that logs a client-chosen
// path. drive sends requests at it that all take the rejecting
// branch; linesPerRequest is how many log lines one such request
// produces there.
type logSite struct {
// build wraps the site's middleware around a handler that must
// not be reached.
build func(
t *testing.T, m *middleware.Middleware,
) http.Handler
// send issues one request for the given client-chosen path and
// returns the status. Some sites need a warm-up request before
// they reject, which send performs itself.
send func(h http.Handler, path string) int
// wantStatus is the status the rejecting branch answers with.
wantStatus int
}
// postOversize sends a POST whose declared Content-Length exceeds the
// body limit without sending a body, which is the whole cost of the
// attack on the MaxBodySize branch.
func postOversize(h http.Handler, path string) int {
req := httptest.NewRequestWithContext(
context.Background(), http.MethodPost, path, nil,
)
req.ContentLength = declaredBodyBytes
req.Header.Set(
"Content-Type", "application/x-www-form-urlencoded",
)
w := httptest.NewRecorder()
h.ServeHTTP(w, req)
return w.Code
}
// postNoToken sends a POST carrying no CSRF token and no session
// cookie, which is what an unauthenticated client sends.
func postNoToken(h http.Handler, path string) int {
req := httptest.NewRequestWithContext(
context.Background(), http.MethodPost, path,
strings.NewReader(""),
)
req.Header.Set(
"Content-Type", "application/x-www-form-urlencoded",
)
w := httptest.NewRecorder()
h.ServeHTTP(w, req)
return w.Code
}
// getNoSession sends a GET with no session cookie.
func getNoSession(h http.Handler, path string) int {
req := httptest.NewRequestWithContext(
context.Background(), http.MethodGet, path, nil,
)
w := httptest.NewRecorder()
h.ServeHTTP(w, req)
return w.Code
}
// logSites enumerates the call sites under test.
func logSites() map[string]logSite {
return map[string]logSite{
// The site this file exists for: WARN, on by default, and
// registered ahead of RequireAuth.
"maxbodysize 413": {
build: func(
t *testing.T, m *middleware.Middleware,
) http.Handler {
t.Helper()
return m.MaxBodySize(bodyLimitBytes)(
unreachable(t),
)
},
send: postOversize,
wantStatus: http.StatusRequestEntityTooLarge,
},
// Also ahead of RequireAuth, also WARN.
"csrf 403": {
build: func(
t *testing.T, m *middleware.Middleware,
) http.Handler {
t.Helper()
return m.CSRF()(unreachable(t))
},
send: postNoToken,
wantStatus: http.StatusForbidden,
},
// The per-entrypoint receiver limiter, unauthenticated. Its
// bucket is keyed on the path, so the first request through a
// fresh path is served and only the ones after it are
// rejected; sendUntilLimited absorbs that.
"receiver rate limit 429": {
build: func(
t *testing.T, m *middleware.Middleware,
) http.Handler {
t.Helper()
return m.ReceiverRateLimit()(okHandler())
},
send: sendUntilLimited,
wantStatus: http.StatusTooManyRequests,
},
// RequireAuth's own line. DEBUG is off in production by
// default, but turning it on to diagnose a flood must not
// restore an unbounded write.
"requireauth redirect": {
build: func(
t *testing.T, m *middleware.Middleware,
) http.Handler {
t.Helper()
return m.RequireAuth()(unreachable(t))
},
send: getNoSession,
wantStatus: http.StatusSeeOther,
},
}
}
// sendUntilLimited drives the per-entrypoint receiver limiter past
// its allowance on one path and returns the status of the rejected
// request. Every request before the last is served, and only the last
// one logs.
func sendUntilLimited(h http.Handler, path string) int {
code := http.StatusOK
for range receiverLimitPerMinute + 1 {
req := httptest.NewRequestWithContext(
context.Background(), http.MethodPost, path, nil,
)
req.RemoteAddr = "203.0.113.7:5555"
w := httptest.NewRecorder()
h.ServeHTTP(w, req)
code = w.Code
}
return code
}
// logLines splits the captured buffer into non-empty lines, holding
// each to bound bytes.
func logLines(t *testing.T, buf *bytes.Buffer, bound int) []string {
t.Helper()
var lines []string
for line := range strings.SplitSeq(
strings.TrimSpace(buf.String()), "\n",
) {
if line == "" {
continue
}
require.LessOrEqual(
t, len(line), bound,
"log line exceeded its bound: %s", line,
)
lines = append(lines, line)
}
return lines
}
// assertNoClientText fails if any marker from the far end of the
// client-chosen input survived into the log. Their absence is what
// distinguishes a real cut from a value that merely happened to be
// short.
func assertNoClientText(t *testing.T, buf *bytes.Buffer) {
t.Helper()
assert.NotContains(
t, buf.String(), attackerMarker,
"log carried attacker-chosen text",
)
assert.NotContains(
t, buf.String(), tailMarker,
"log carried the tail of the attacker-chosen text",
)
}
// TestLogLines_ClientChosenPathDoesNotSizeTheLine points 8 KB of
// client-chosen path at each non-access-log call site that logs one,
// through both handlers and through every character those handlers
// escape, and holds the resulting line to MaxAccessLogLineBytes.
//
// Removing any one of the logfield.Truncate calls at those sites
// fails this test: the line grows to roughly the size of the input,
// or to several times it on the escaping fills.
func TestLogLines_ClientChosenPathDoesNotSizeTheLine(t *testing.T) {
t.Parallel()
for siteName, site := range logSites() {
for handlerName, newHandler := range logHandlers() {
for fillName, fill := range escapeFills() {
name := siteName + "/" + handlerName + "/" + fillName
t.Run(name, func(t *testing.T) {
t.Parallel()
m, buf := capturingBoundMiddleware(
t, newHandler,
)
path := "/source/" +
oversizedPathSegment(fill) + "/edit"
assert.Equal(
t,
site.wantStatus,
site.send(site.build(t, m), path),
)
lines := logLines(
t, buf,
middleware.MaxAccessLogLineBytes,
)
require.NotEmpty(
t, lines,
"the site under test logged nothing, "+
"so the bound proves nothing",
)
assertNoClientText(t, buf)
})
}
}
}
}
// TestLoginThrottle_LogLineDoesNotTrackPathSize pins the cap on
// RecordLoginFailure's "login failure limit exceeded" WARN line.
//
// That site does not fit logSites above: it is not a middleware
// wrapping a handler but an exported method the login handler calls,
// and the only route that calls it today is chi's static
// "/pages/login", so no request through the mux can widen the line.
// Driving the method directly is therefore the whole point rather
// than a shortcut — it is exactly the call a second caller on a route
// with a URL parameter would make, and without this test removing the
// logfield.Truncate there fails nothing.
func TestLoginThrottle_LogLineDoesNotTrackPathSize(t *testing.T) {
t.Parallel()
for handlerName, newHandler := range logHandlers() {
for fillName, fill := range escapeFills() {
t.Run(handlerName+"/"+fillName, func(t *testing.T) {
t.Parallel()
m, buf := capturingBoundMiddleware(
t, newHandler,
)
req := httptest.NewRequestWithContext(
context.Background(),
http.MethodPost,
"/source/"+
oversizedPathSegment(fill)+"/login",
nil,
)
req.RemoteAddr = "203.0.113.9:5555"
// The budget is spent per client and username,
// so one more failure than the budget allows is
// what takes the throttled branch.
var throttled bool
for range middleware.LoginRateLimitConst + 1 {
throttled = m.RecordLoginFailure(
req, "someone",
)
}
require.True(
t, throttled,
"the throttled branch never ran, so the "+
"bound proves nothing",
)
lines := logLines(
t, buf, middleware.MaxAccessLogLineBytes,
)
require.NotEmpty(t, lines)
assertNoClientText(t, buf)
})
}
}
}
// TestMaxBodySize_FloodOfOversizePathsDoesNotGrowTheLog is the
// flood shape from the issue: an unauthenticated client posting
// oversize declarations at invented 8 KB paths, as fast as it likes.
//
// It asserts the property directly rather than by proxy — the bytes
// the flood writes to the operator's log do not track the bytes the
// flood sent. The same flood at a one-character path is the control:
// 8 KB of extra input per request buys at most the field budget, not
// 8 KB of log.
func TestMaxBodySize_FloodOfOversizePathsDoesNotGrowTheLog(
t *testing.T,
) {
t.Parallel()
for handlerName, newHandler := range logHandlers() {
for fillName, fill := range escapeFills() {
t.Run(handlerName+"/"+fillName, func(t *testing.T) {
t.Parallel()
flood := func(segment func(i int) string) int {
m, buf := capturingBoundMiddleware(
t, newHandler,
)
h := m.MaxBodySize(bodyLimitBytes)(
unreachable(t),
)
for i := range floodRequests {
assert.Equal(
t,
http.StatusRequestEntityTooLarge,
postOversize(
h,
"/source/"+segment(i)+"/edit",
),
)
}
lines := logLines(
t, buf,
middleware.MaxAccessLogLineBytes,
)
require.Len(t, lines, floodRequests)
assertNoClientText(t, buf)
return buf.Len()
}
sent := oversizedSegmentBytes * floodRequests
oversize := flood(func(i int) string {
return oversizedPathSegment(fill) +
strings.Repeat("y", i)
})
control := flood(func(i int) string {
return "a" + strings.Repeat("y", i)
})
// The whole point: 8 KB per request of extra
// client-chosen input bought a bounded amount of
// log, not a proportional amount.
assert.Less(
t, oversize-control, sent/2,
"log volume tracked the size of the flood's "+
"input",
)
assert.LessOrEqual(
t,
oversize,
floodRequests*
middleware.MaxAccessLogLineBytes,
)
})
}
}
}

View File

@@ -1,409 +0,0 @@
package middleware
import (
"context"
"crypto/sha256"
"encoding/hex"
"net/http"
"sync"
"time"
"sneak.berlin/go/webhooker/internal/logfield"
)
const (
// loginFailureMaxKeys bounds how many distinct failure counters
// each of the guard's two key sets holds. The submitted username
// is part of a key, so the key set is attacker-influenced and
// needs a hard cap or the limiter becomes the memory
// amplification surface it exists to protect.
//
// A single-admin deployment has a handful of legitimate (client,
// username) pairs, so 1024 is three orders of magnitude of
// headroom before a real operator can be pushed onto the
// fallback. It costs little: a counter is a ~64-byte key string,
// a 32-byte window and map overhead, call it 170 bytes, so both
// key sets full is 2 * 1024 * 170 bytes, under 0.4 MB.
loginFailureMaxKeys = 1024
// passwordVerifyConcurrency bounds how many Argon2id
// verifications may run at once across every password-verifying
// endpoint. Because credentials are now verified before any
// limiter budget is spent, an attacker can force one hash per
// request, and each hash allocates argon2Memory — 64 MB. Two
// slots commit at most 128 MB to password hashing, which fits
// inside the smallest container this service is realistically
// given alongside its own working set; four would commit 256 MB
// and crowd it. A single-admin product needs no concurrent
// logins at all, so the second slot exists only so that one
// stalled request does not serialise the endpoint.
passwordVerifyConcurrency = 2
// passwordVerifyWait is how long a request waits for a
// verification slot before it is answered 503. Slots are handed
// out in arrival order, so a legitimate request queues behind
// the requests already waiting rather than behind the flood as a
// whole. The wait is well inside the 60s request timeout.
passwordVerifyWait = 5 * time.Second
// passwordVerifyMaxWaiters bounds how many requests may be
// queued for a slot at once. Past it, acquire sheds immediately
// with 503 instead of joining the queue.
//
// The wait bounds how long one request occupies memory; this
// bounds how many do so at the same time, and without it the
// 128 MB hashing budget above is the smaller half of the real
// footprint. At the 400 req/s a saturation attack can offer, an
// unbounded queue would park ~2000 requests for the full five
// seconds.
//
// A waiter costs far more than maxFormBodySize suggests: that
// caps the raw body read, not what the parse retains. MaxBodySize,
// CSRF and ParseForm all run before acquire, so a parked waiter
// holds r.Form plus r.PostForm plus its header block for the
// whole wait. Measured on the pinned go1.26.1 toolchain, as the
// HeapAlloc delta across two GCs with 64 waiters parked in the
// handler: an ordinary two-field login form retains ~0 MB, but a
// 1 MB urlencoded body at Go's 10,000-parameter parse cap retains
// 2.82 MB (3.09 MB with %41 escapes), and adding the ~0.9 MB of
// headers httpMaxHeaderBytes allows takes it to 4.18 MB. The
// retained parse and the header block dominate; the raw body does
// not.
//
// Arithmetic, from the measured 4.18 MB worst case: 16 waiters
// commit ~67 MB of queue memory, and peak commitment for the
// endpoint is 128 MB of Argon2id plus the 18 requests that retain
// a parsed form — 16 queued and the 2 being hashed — at
// 18 * 4.18 MB, so ~75 MB: about 203 MB in all. Cross-check
// against the deadline: two slots at the ~27 verifications/s
// measured on a review host (with the race detector on, so the
// real rate is higher) drain a full 16-deep queue in about 0.6 s,
// far inside passwordVerifyWait.
//
// Those 203 MB are live bytes, not resident bytes: the Go
// collector lets the heap reach roughly twice the live set before
// collecting, with transient parse garbage on top. The review
// measured a peak HeapAlloc of 392 MB against this guard under 18
// adversarial requests, so provision on the order of 400 MB rather
// than 203 MB.
passwordVerifyMaxWaiters = 16
// failureKeyHashBytes is how much of the username digest goes
// into a failure key. 64 bits over at most loginFailureMaxKeys
// live keys makes a collision negligible, and a collision would
// only merge two usernames' failure counters, which throttles
// sooner rather than later.
failureKeyHashBytes = 8
)
// failureWindow counts failed credential verifications for one
// bucket, and records when that count lapses.
type failureWindow struct {
count int
resetAt time.Time
}
// loginGuard is what replaced the pre-emptive rate limiter on the
// login POST.
//
// A limiter that spends budget on arrival cannot protect a
// single-admin product: behind the reverse proxy the deployment
// requires, with TRUSTED_PROXIES unset, every client keys on the
// proxy, so a stranger trickling five POSTs a minute keeps the one
// bucket full and the operator's own correct password is answered 429
// forever. There is no second administrative path.
//
// So budget is spent only by a FAILED verification. A correct
// password is never throttled, whatever the counters say, which is
// the only shape that guarantees the operator can get in. Two
// consequences follow and are handled here:
//
// - Every login request now costs an Argon2id hash, so the number
// running concurrently is bounded by slots. Without that bound
// this trades an admin lockout for memory exhaustion, which is
// strictly worse.
// - Counting per (client, username) makes the key set
// attacker-influenced, so both key sets are capped. Beyond the
// per-username cap, failures fall back to a counter keyed on the
// client alone; beyond that cap too, a failure is answered as
// throttled without being recorded, since refusing to answer a
// wrong password costs the operator nothing.
type loginGuard struct {
mu sync.Mutex
byUser map[string]*failureWindow
byAddr map[string]*failureWindow
slots chan struct{}
// queue holds one token per request waiting for a slot. A token
// is taken non-blockingly, so a request that finds it full is
// shed rather than queued, and is given up as soon as the wait
// ends however it ends.
queue chan struct{}
limit int
interval time.Duration
maxKeys int
wait time.Duration
// now is time.Now outside tests.
now func() time.Time
}
// newLoginGuard builds a guard with the given failure limit per
// interval, key-set cap, verification concurrency, queue depth and
// slot wait.
func newLoginGuard(
limit int,
interval time.Duration,
maxKeys, concurrency, maxWaiters int,
wait time.Duration,
) *loginGuard {
return &loginGuard{
byUser: make(map[string]*failureWindow),
byAddr: make(map[string]*failureWindow),
slots: make(chan struct{}, concurrency),
queue: make(chan struct{}, maxWaiters),
limit: limit,
interval: interval,
maxKeys: maxKeys,
wait: wait,
now: time.Now,
}
}
// acquire reserves a verification slot, waiting up to the guard's
// wait for one. It reports false when the queue of waiters is
// already full, when no slot became available in time, or when the
// request was cancelled while waiting; the caller must then answer
// 503 without verifying anything. The returned function releases the
// slot and must be called exactly once.
//
// ctx is consulted only once the request has to wait: a slot that is
// free on arrival is handed out without looking at it, so an
// already-cancelled request can be granted one. That is deliberate
// and matches lifecycle.waitDone — the caller abandons the work on
// its own ctx and releases the slot immediately, so nothing is spent
// on it, and refusing instead would mean shedding a request with
// capacity standing free.
func (g *loginGuard) acquire(ctx context.Context) (func(), bool) {
// A free slot is taken before any timer is armed, and before a
// queue place is claimed: a request that never waits is not a
// waiter. Without this preamble the bounded select below can find
// its slot send and an already-expired timer ready at the same
// time, and Go picks among ready cases uniformly at random — so a
// process descheduled for longer than the wait sheds a request
// with slots standing free, which is precisely when shedding is
// least defensible.
//
// This cannot let a late arrival barge past a queued waiter. A
// waiter can only be parked on a FULL buffer, and a release
// refills that buffer from the head of the send queue under the
// channel lock, so the buffer never appears non-full while anyone
// is parked and this send fails whenever there is a waiter.
select {
case g.slots <- struct{}{}:
return func() { <-g.slots }, true
default:
}
// Shedding past the queue depth is what keeps waiting memory
// bounded; the wait alone only bounds how long one waiter holds
// its parsed form, not how many hold one at once.
select {
case g.queue <- struct{}{}:
default:
return nil, false
}
// Held only for the wait. A request that gets a slot gives its
// queue token back before it starts hashing, so the depth is a
// bound on waiters rather than on requests in the handler.
defer func() { <-g.queue }()
timer := time.NewTimer(g.wait)
defer timer.Stop()
// The blocking send is deliberate: a receive on a full buffered
// channel hands the slot straight to the head of the send queue,
// so slots go out in arrival order and a later arrival cannot
// barge past a request already waiting.
select {
case g.slots <- struct{}{}:
return func() { <-g.slots }, true
case <-timer.C:
return nil, false
case <-ctx.Done():
return nil, false
}
}
// fail records one failed credential verification by clientKey
// against username, and reports whether this client has now spent
// its failure budget and should be answered 429.
func (g *loginGuard) fail(clientKey, username string) bool {
g.mu.Lock()
defer g.mu.Unlock()
now := g.now()
window := g.window(
g.byUser, userFailureKey(clientKey, username), now,
)
if window == nil {
window = g.window(g.byAddr, clientKey, now)
}
if window == nil {
// Both key sets are full and neither already tracks this
// client, so nothing can be counted without unbounded
// growth. Answering the failure as throttled is the safe
// direction: it never touches a correct password.
return true
}
window.count++
return window.count >= g.limit
}
// succeed forgives clientKey's failures against username. A correct
// password clears the counters, so an operator who mistypes several
// times and then gets it right is not throttled afterwards.
func (g *loginGuard) succeed(clientKey, username string) {
g.mu.Lock()
defer g.mu.Unlock()
delete(g.byUser, userFailureKey(clientKey, username))
delete(g.byAddr, clientKey)
}
// window returns the live counter for key in set, resetting a lapsed
// one and creating a missing one when the cap allows. It returns nil
// only when key is absent and set is full even after lapsed entries
// are swept.
func (g *loginGuard) window(
set map[string]*failureWindow,
key string,
now time.Time,
) *failureWindow {
window, ok := set[key]
if ok {
if !now.Before(window.resetAt) {
window.count = 0
window.resetAt = now.Add(g.interval)
}
return window
}
if len(set) >= g.maxKeys {
sweepLapsed(set, now)
}
if len(set) >= g.maxKeys {
return nil
}
window = &failureWindow{resetAt: now.Add(g.interval)}
set[key] = window
return window
}
// sweepLapsed drops counters whose interval has elapsed.
func sweepLapsed(set map[string]*failureWindow, now time.Time) {
for key, window := range set {
if !now.Before(window.resetAt) {
delete(set, key)
}
}
}
// userFailureKey identifies one (client, submitted username) pair.
// The username is hashed rather than embedded: a submitted username
// is attacker-controlled text of attacker-chosen length, and hashing
// makes every key the same size whatever was sent.
func userFailureKey(clientKey, username string) string {
sum := sha256.Sum256([]byte(username))
return clientKey + "|" +
hex.EncodeToString(sum[:failureKeyHashBytes])
}
// guard returns the middleware's login guard, building it on first
// use so that every construction path — fx and the test constructor
// alike — gets one.
func (m *Middleware) guard() *loginGuard {
m.loginGuardOnce.Do(func() {
m.loginGuard = newLoginGuard(
loginRateLimit,
loginRateInterval,
loginFailureMaxKeys,
passwordVerifyConcurrency,
passwordVerifyMaxWaiters,
passwordVerifyWait,
)
})
return m.loginGuard
}
// BeginPasswordVerification reserves one of the bounded Argon2id
// verification slots. It reports false when the queue of waiting
// requests is already at passwordVerifyMaxWaiters, or when no slot
// became free within passwordVerifyWait; in either case the caller
// must answer 503 and must not verify a password. The returned
// function releases the slot and must be called exactly once.
//
// Every endpoint that hashes a password on request must go through
// this, or the bound has a hole: the memory is committed per hash,
// not per endpoint.
func (m *Middleware) BeginPasswordVerification(
ctx context.Context,
) (func(), bool) {
return m.guard().acquire(ctx)
}
// RecordLoginFailure counts a failed credential verification for the
// request's client against the submitted username, and reports
// whether the response should be 429 rather than 401.
func (m *Middleware) RecordLoginFailure(
r *http.Request,
username string,
) bool {
throttled := m.guard().fail(m.clientKey(r), username)
if throttled {
// Truncated even though chi pins this route's path to
// the 12-byte constant "/pages/login": RecordLoginFailure
// is exported and takes any *http.Request, so a caller on
// a route with a URL parameter would otherwise widen this
// line. logbound_test.go pins the cap by making exactly
// that call, since no request through the mux can.
m.log.Warn(
"login failure limit exceeded",
"path", logfield.Truncate(
r.URL.Path, logfield.MaxBytes,
),
)
}
return throttled
}
// ForgiveLoginFailures clears the failure counters for the request's
// client and the submitted username after a successful
// authentication.
func (m *Middleware) ForgiveLoginFailures(
r *http.Request,
username string,
) {
m.guard().succeed(m.clientKey(r), username)
}
// LoginFailureInterval is how long a spent login failure budget
// takes to refill, which is what a throttled login answers as
// Retry-After.
func (m *Middleware) LoginFailureInterval() time.Duration {
return m.guard().interval
}

View File

@@ -1,629 +0,0 @@
package middleware_test
import (
"context"
"fmt"
"sync"
"sync/atomic"
"testing"
"time"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
"sneak.berlin/go/webhooker/internal/database"
"sneak.berlin/go/webhooker/internal/middleware"
)
// mib converts the Argon2id memory parameter, which is in KiB, to MB.
const mib = 1024
const (
// guardInterval is the failure window these tests use. It is
// long enough that nothing lapses mid-test on its own; tests
// that need a lapse drive the clock instead.
guardInterval = time.Minute
// guardWait is the slot wait for tests that expect to get a
// slot. Tests that expect to be refused set their own.
guardWait = 2 * time.Second
guardClient = "198.51.100.7"
guardUser = "admin"
// racePasses is how many times a both-cases-ready select race is
// run. A pass can only go the wrong way once the zero-duration
// timer has fired, so the per-pass detection probability is
// somewhere below 1/2 rather than exactly it; the bound that
// matters is that passes are independent, so a regression that
// survives is exponentially unlikely in N. The test still waits
// on nothing.
racePasses = 1000
)
// newGuard builds a guard with production-shaped defaults and the
// given key-set cap and verification concurrency.
func newGuard(maxKeys, concurrency int) *middleware.LoginGuard {
return middleware.NewLoginGuardForTest(
middleware.LoginRateLimitConst,
guardInterval,
maxKeys,
concurrency,
middleware.PasswordVerifyMaxWaitersConst,
guardWait,
)
}
// TestLoginGuard_ThrottlesRepeatedFailures is the brute-force half:
// wrong passwords for one username from one client key still run out
// of budget and are answered 429.
func TestLoginGuard_ThrottlesRepeatedFailures(t *testing.T) {
t.Parallel()
g := newGuard(middleware.LoginFailureMaxKeysConst, 1)
for i := range middleware.LoginRateLimitConst - 1 {
assert.False(
t, g.FailForTest(guardClient, guardUser),
"failure %d is still inside the budget", i,
)
}
assert.True(
t, g.FailForTest(guardClient, guardUser),
"the last failure of the budget must throttle",
)
assert.True(
t, g.FailForTest(guardClient, guardUser),
"failures past the budget must stay throttled",
)
}
// TestLoginGuard_SuccessForgivesFailures pins the forgiveness rule:
// an operator who mistypes several times and then gets it right must
// not be left throttled.
func TestLoginGuard_SuccessForgivesFailures(t *testing.T) {
t.Parallel()
g := newGuard(middleware.LoginFailureMaxKeysConst, 1)
for range middleware.LoginRateLimitConst {
g.FailForTest(guardClient, guardUser)
}
g.SucceedForTest(guardClient, guardUser)
assert.False(
t, g.FailForTest(guardClient, guardUser),
"a success must reset the counter, so the next mistake "+
"starts a fresh budget",
)
}
// TestLoginGuard_FailuresAreKeyedPerUsername proves the second half
// of the keying: one username's spent budget does not throttle
// another's from the same client.
func TestLoginGuard_FailuresAreKeyedPerUsername(t *testing.T) {
t.Parallel()
g := newGuard(middleware.LoginFailureMaxKeysConst, 1)
for range middleware.LoginRateLimitConst {
g.FailForTest(guardClient, guardUser)
}
assert.True(t, g.FailForTest(guardClient, guardUser))
assert.False(
t, g.FailForTest(guardClient, "someone-else"),
"a different submitted username must have its own budget",
)
}
// TestLoginGuard_WindowLapses covers the interval: a counter that has
// gone quiet for the whole window starts again from zero.
func TestLoginGuard_WindowLapses(t *testing.T) {
t.Parallel()
g := newGuard(middleware.LoginFailureMaxKeysConst, 1)
var now atomic.Int64
now.Store(time.Now().UnixNano())
g.SetNowForTest(func() time.Time {
return time.Unix(0, now.Load())
})
for range middleware.LoginRateLimitConst {
g.FailForTest(guardClient, guardUser)
}
assert.True(t, g.FailForTest(guardClient, guardUser))
now.Add(int64(guardInterval) + 1)
assert.False(
t, g.FailForTest(guardClient, guardUser),
"a lapsed window must start a fresh budget",
)
}
// TestLoginGuard_UsernameKeySetIsBounded is the memory bound. The
// submitted username is attacker-controlled, so an attacker rotating
// usernames must not be able to grow the guard without limit: past
// the cap, tracking falls back to a counter keyed on the client
// address alone.
func TestLoginGuard_UsernameKeySetIsBounded(t *testing.T) {
t.Parallel()
const (
maxKeys = 8
attempts = 500
)
g := newGuard(maxKeys, 1)
for i := range attempts {
g.FailForTest(guardClient, fmt.Sprintf("user-%d", i))
}
byUser, byAddr := g.TrackedKeysForTest()
assert.LessOrEqual(
t, byUser, maxKeys,
"the per-username key set must not grow past its cap",
)
assert.LessOrEqual(
t, byAddr, maxKeys,
"the fallback key set must not grow past its cap either",
)
assert.Positive(
t, byAddr,
"past the cap, failures must fall back to the address "+
"bucket rather than being dropped",
)
assert.Less(
t, byUser+byAddr, attempts,
"memory must not grow with the number of distinct "+
"usernames submitted",
)
}
// TestLoginGuard_BeyondBothCapsStaysThrottled covers the hard stop.
// When both key sets are full of live counters and the client is in
// neither, there is nothing to count without unbounded growth, so the
// failure is answered as throttled. That costs the operator nothing:
// a correct password never reaches this path.
func TestLoginGuard_BeyondBothCapsStaysThrottled(t *testing.T) {
t.Parallel()
const maxKeys = 4
g := newGuard(maxKeys, 1)
// Fill the per-username set from one client, then fill the
// address set from distinct clients.
for i := range maxKeys {
g.FailForTest(guardClient, fmt.Sprintf("user-%d", i))
}
for i := range maxKeys {
g.FailForTest(fmt.Sprintf("203.0.113.%d", i), "whoever")
}
assert.True(
t, g.FailForTest("203.0.113.200", "brand-new"),
"a client that fits in neither full key set must be "+
"answered as throttled rather than tracked",
)
byUser, byAddr := g.TrackedKeysForTest()
assert.LessOrEqual(t, byUser, maxKeys)
assert.LessOrEqual(t, byAddr, maxKeys)
}
// TestLoginGuard_SemaphoreBoundsConcurrentVerifications is the memory
// bound on the hashing itself. Verifying credentials before spending
// limiter budget means an attacker can force one Argon2id hash per
// request, and each allocates 64 MB; without this bound the fix for
// an admin lockout would be a memory-exhaustion DoS instead.
func TestLoginGuard_SemaphoreBoundsConcurrentVerifications(
t *testing.T,
) {
t.Parallel()
const (
concurrency = 2
workers = 12
// rendezvousDeadlock is the deadlock guard described below.
// It is orders of magnitude longer than any scheduling delay,
// so it never decides the result, and well inside script/test's
// 30s timeout, so a wedge fails on the assertion instead of
// blowing the package timeout.
rendezvousDeadlock = 5 * time.Second
)
g := newGuard(middleware.LoginFailureMaxKeysConst, concurrency)
var (
mu sync.Mutex
inside int
highest int
wg sync.WaitGroup
recorded sync.WaitGroup
once sync.Once
)
// Slot holders rendezvous instead of sleeping, and they hold until
// every worker has been answered. A sleep only makes overlap
// likely — on a host loaded enough to deschedule a goroutine for
// longer than the sleep the workers serialise and the maximum
// observed comes back as 1 — so the rendezvous is what makes the
// overlap a fact rather than a race won.
//
// The barrier must not open at the concurrency-th holder, which
// would fix the lower bound at the cost of the upper one this test
// exists to enforce: holders would leave as soon as the count
// reached concurrency, so a guard admitting extra requests would
// let them arrive after the first holders had already left and
// highest would report concurrency however many were really let
// in. It opens instead once every worker's acquire has returned
// and any slot it won has been counted, so under a broken guard
// every admitted worker is inside simultaneously and highest is
// the true maximum. Under a correct guard the refused workers
// return within the guard's own wait, which decides nothing beyond
// how long that takes.
overlapped := make(chan struct{})
closeOverlapped := func() {
once.Do(func() { close(overlapped) })
}
// Deadlock guard, not a timing margin: no assertion depends on its
// length, and the only way to reach it is a worker that never
// returns from acquire at all. It is here so that such a wedge
// fails legibly on the assertion below instead of hanging until
// the package test timeout.
abandon := time.AfterFunc(rendezvousDeadlock, closeOverlapped)
defer abandon.Stop()
recorded.Add(workers)
go func() {
recorded.Wait()
closeOverlapped()
}()
for range workers {
wg.Go(func() {
release, ok := g.AcquireForTest(context.Background())
if !ok {
recorded.Done()
return
}
defer release()
mu.Lock()
inside++
if inside > highest {
highest = inside
}
mu.Unlock()
// Counted before signalling, so the barrier can never open
// while an admitted worker is still on its way to being
// counted.
recorded.Done()
<-overlapped
mu.Lock()
inside--
mu.Unlock()
})
}
wg.Wait()
mu.Lock()
defer mu.Unlock()
assert.Equal(
t, concurrency, highest,
"no more than %d verifications may run at once", concurrency,
)
}
// TestLoginGuard_SaturatedSemaphoreRefusesRatherThanQueueing pins
// what happens when every slot is taken for longer than the wait: the
// request is refused, so the caller answers 503 without allocating
// another 64 MB hash.
//
// Neither half of this rides on the wait being long enough. The
// refusal holds the only slot across the whole of the second call, so
// there is no wait it could get lucky with — the wait fixes only how
// long the refusal takes, not whether it happens. The reuse after
// release is settled by acquire's non-blocking preamble, which is
// pinned separately by TestLoginGuard_FreeSlotBeatsAnExpiredWait. So
// the wait below is sized to keep the test quick, not to win a race.
func TestLoginGuard_SaturatedSemaphoreRefusesRatherThanQueueing(
t *testing.T,
) {
t.Parallel()
g := middleware.NewLoginGuardForTest(
middleware.LoginRateLimitConst,
guardInterval,
middleware.LoginFailureMaxKeysConst,
1,
middleware.PasswordVerifyMaxWaitersConst,
10*time.Millisecond,
)
release, ok := g.AcquireForTest(context.Background())
require.True(t, ok, "the first acquire must get the only slot")
_, ok = g.AcquireForTest(context.Background())
assert.False(
t, ok,
"with the only slot held, a second request must be refused "+
"rather than wait indefinitely",
)
release()
release, ok = g.AcquireForTest(context.Background())
// require, not assert: acquire returns a nil release alongside a
// false ok, so calling it after a non-fatal assertion turns one
// failed test into a segfault that takes the whole package test
// binary down. Every assertion whose value is dereferenced later
// has to stop the test.
require.True(
t, ok, "the slot must be reusable once released",
)
release()
}
// TestLoginGuard_FreeSlotBeatsAnExpiredWait is the determinism this
// file used to lack. acquire selects over a slot send and a wait
// timer, and Go chooses among ready cases uniformly at random, so a
// call made after the timer had already fired was a coin flip: on a
// loaded host the previous test's third acquire could be refused
// with its slot standing free, and then dereference the nil release
// it got back.
//
// The wait here is already elapsed on arrival, which is the worst
// case that scheduling can produce, so a free slot must still be
// granted every time. Without acquire's non-blocking preamble each
// pass is an independent coin flip and the loop fails within a few
// passes; with it the property holds by construction and no wall
// clock is involved.
func TestLoginGuard_FreeSlotBeatsAnExpiredWait(t *testing.T) {
t.Parallel()
g := middleware.NewLoginGuardForTest(
middleware.LoginRateLimitConst,
guardInterval,
middleware.LoginFailureMaxKeysConst,
1,
middleware.PasswordVerifyMaxWaitersConst,
0,
)
for pass := range racePasses {
release, ok := g.AcquireForTest(context.Background())
require.Truef(
t, ok,
"pass %d was refused a slot that was free; an expired "+
"wait must never beat an available slot",
pass,
)
release()
}
}
// TestLoginGuard_AcquireHonoursCancellation proves a client that
// disconnects while queued frees its place immediately instead of
// holding it for the full wait.
func TestLoginGuard_AcquireHonoursCancellation(t *testing.T) {
t.Parallel()
g := newGuard(middleware.LoginFailureMaxKeysConst, 1)
release, ok := g.AcquireForTest(context.Background())
require.True(t, ok)
defer release()
ctx, cancel := context.WithCancel(context.Background())
cancel()
_, ok = g.AcquireForTest(ctx)
assert.False(
t, ok, "a cancelled request must not wait for a slot",
)
}
// TestPasswordVerifyConcurrency_MatchesMemoryBudget pins the
// concurrency constant to the arithmetic behind it: the number of
// slots is the hashing budget divided by what one Argon2id hash
// actually costs.
//
// The per-hash figure is read out of the shipped password
// parameters rather than copied here. A guard that asserts a literal
// against a literal cannot see the thing it guards: raising
// argon2Memory would leave it green while the real ceiling doubled.
func TestPasswordVerifyConcurrency_MatchesMemoryBudget(t *testing.T) {
t.Parallel()
// Memory is the real argon2Memory, in KiB.
perHashMB := int(database.DefaultPasswordConfig().Memory) / mib
require.Positive(
t, perHashMB,
"the Argon2id memory parameter must be readable in MB",
)
// The memory this service commits to password hashing.
const budgetMB = 128
assert.Equal(
t,
middleware.PasswordVerifyConcurrencyConst,
budgetMB/perHashMB,
"the verification concurrency must be the %d MB hashing "+
"budget divided by the %d MB one Argon2id hash costs; "+
"if the Argon2id parameters changed, the slot count "+
"must change with them",
budgetMB, perHashMB,
)
}
// TestLoginGuard_ShedsPastTheQueueCap pins the memory bound on
// waiting, as distinct from the bound on hashing. A waiter arrives
// with its form already parsed, and the retained parse plus its
// header block cost several MB — far more than maxFormBodySize
// suggests, since that caps only the raw body read — so an unbounded
// queue would hold that much per waiting request for the whole wait;
// past the cap the guard must refuse instantly rather than grow.
func TestLoginGuard_ShedsPastTheQueueCap(t *testing.T) {
t.Parallel()
const (
maxWaiters = 2
// Long enough that a queued waiter never times out on its
// own, so anything the test observes leaving the queue left
// because it was shed.
neverElapses = time.Minute
// The probe carries its own deadline, so a guard that queues
// the probe instead of shedding it fails here rather than
// hanging until the package test timeout.
//
// This is a patience budget, not a margin to be won. A shed
// returns in microseconds and a probe that queued instead
// would not return for neverElapses, so the two are a whole
// minute apart and any budget between them separates them. It
// is set far above any scheduling stall a loaded host can
// produce, because the previous 200 ms — and the 100 ms
// elapsed-time assertion it fed — bounded the latency of a
// goroutine hand-off, which is a false red waiting to happen
// on the machine this suite runs on. What actually proves the
// probe was not queued is the queue depth asserted below.
probePatience = 5 * time.Second
)
g := middleware.NewLoginGuardForTest(
middleware.LoginRateLimitConst,
guardInterval,
middleware.LoginFailureMaxKeysConst,
1,
maxWaiters,
neverElapses,
)
// Occupy the only slot, so everything after this queues.
release, ok := g.AcquireForTest(context.Background())
require.True(t, ok)
defer release()
defer fillQueue(t, g, maxWaiters)()
granted, answered := probeQueueCap(g, probePatience)
require.True(
t, answered,
"a request arriving past the queue cap is still waiting to "+
"be queued; it must have been shed",
)
assert.False(
t, granted,
"a request arriving past the queue cap must be shed",
)
assert.Equal(
t, maxWaiters, g.QueuedWaitersForTest(),
"a shed request must not have grown the queue",
)
}
// fillQueue starts n waiters on g and returns once all of them are
// queued for a slot. The returned function releases them and waits
// for them to exit.
func fillQueue(
t *testing.T,
g *middleware.LoginGuard,
n int,
) func() {
t.Helper()
ctx, cancel := context.WithCancel(context.Background())
var wg sync.WaitGroup
for range n {
wg.Go(func() {
done, got := g.AcquireForTest(ctx)
if got {
done()
}
})
}
// Patience budget, not a margin: the waiters park in microseconds
// and nothing releases them, so the only way to exhaust this is a
// guard that never queues. One second is the same order as the
// scheduling stalls this suite has to survive, so it is not one.
require.Eventually(
t,
func() bool { return g.QueuedWaitersForTest() == n },
5*time.Second, time.Millisecond,
"the waiters must reach the queue before the cap is tested",
)
return func() {
cancel()
wg.Wait()
}
}
// probeQueueCap acquires from another goroutine. It reports, in
// order, whether the call was granted a slot and whether it was
// answered at all within wait; a call that never returned reports
// false for both.
//
// It runs off the test goroutine deliberately. Joining a full queue
// is not cancellable by context — refusing to join is the property
// under test — so a guard that fails this would otherwise hang the
// package until the test timeout instead of failing here.
//
// It reports no elapsed time. Timing a goroutine hand-off measures
// the host, not the guard, and the caller distinguishes shedding from
// queueing by the queue depth instead.
func probeQueueCap(
g *middleware.LoginGuard,
wait time.Duration,
) (bool, bool) {
probed := make(chan bool, 1)
go func() {
release, ok := g.AcquireForTest(context.Background())
if ok {
release()
}
probed <- ok
}()
select {
case result := <-probed:
return result, true
case <-time.After(wait):
return false, false
}
}

View File

@@ -6,11 +6,9 @@ import (
"log/slog"
"net"
"net/http"
"sync"
"time"
basicauth "github.com/99designs/basicauth-go"
"github.com/go-chi/chi"
"github.com/go-chi/chi/middleware"
"github.com/go-chi/cors"
metrics "github.com/slok/go-http-metrics/metrics/prometheus"
@@ -19,7 +17,6 @@ import (
"go.uber.org/fx"
"sneak.berlin/go/webhooker/internal/config"
"sneak.berlin/go/webhooker/internal/globals"
"sneak.berlin/go/webhooker/internal/logfield"
"sneak.berlin/go/webhooker/internal/logger"
"sneak.berlin/go/webhooker/internal/session"
)
@@ -28,118 +25,6 @@ const (
// corsMaxAge is the maximum time (in seconds) that a
// preflight response can be cached.
corsMaxAge = 300
// unmatchedRoute is logged in the access log's url field when a
// redirected or rejected request matched no route pattern at
// all. Every byte of such a path is client-chosen, so none of it
// is logged.
unmatchedRoute = "(unmatched)"
// redactedQuery stands in for the query string on the access log
// branches that keep the concrete URL. The query is client-chosen
// on every route, including the ones that answer an
// unauthenticated 200, so logging it verbatim would let a client
// pick the size of the line it writes.
redactedQuery = "?(redacted)"
// maxLogRequestIDBytes bounds the request id, which is also
// client-supplied: chi's RequestID middleware passes an inbound
// X-Request-Id header through verbatim. Its generated form is an
// order of magnitude shorter than this.
maxLogRequestIDBytes = 128
// maxLogMethodBytes bounds the method. Go accepts any RFC 7230
// token there, bounded only by the header size limit, so it is
// client-chosen text like the rest. The longest registered method
// is half this.
maxLogMethodBytes = 32
// MaxAccessLogLineBytes is the ceiling on one JSON access log line,
// and the number an operator multiplies by the request rate to size
// log storage. It is not an observation of a sample: it is the sum
// of the budgets above, each of which logfield.Truncate enforces in
// ENCODED bytes, plus the part of the line no client can influence.
//
// url, useragent, referer 3*(512+11) = 1569
// request_id 128+11 = 139
// method 32+11 = 43
// fixed portion = 336
// ----
// 2087
//
// The 512 is logfield.MaxBytes; the 11 is the truncation marker,
// charged on top of each budget rather than inside it.
//
// The fixed portion is the JSON punctuation, the field names, the
// level and the message, both timestamps at their longest, an IPv6
// remoteIP with a zone, a three-digit status and a full-width int64
// latency. Stated at 2560 so the figure carries headroom rather
// than sitting on the arithmetic.
//
// The tty text handler in internal/logger is covered by the same
// figure. logfield.EncodedBytes charges every rune at least what
// the wider of the two handlers emits for it — including the ten
// bytes strconv.Quote spends on a non-printable rune at or above
// U+10000, which is four more than the JSON handler ever spends —
// so each budget bounds the encoded field under either handler.
// The text handler's fixed portion is 286, the smaller of the two,
// which puts its worst case at 2037.
//
// It is also the ceiling on every OTHER line this service writes
// THROUGH SLOG that carries text an UNAUTHENTICATED client
// supplies. Those lines — the MaxBodySize rejection, the CSRF
// rejection, the rate-limit rejection, the unauthenticated-request
// and unknown-entrypoint DEBUG lines, the failed-login DEBUG
// lines, and the two login-throttle WARN lines ("login failure
// limit exceeded" in loginguard.go and "password verification
// capacity exhausted" in internal/handlers/auth.go) — spend the
// same per-field budgets, and each carries strictly fewer
// client-supplied fields than the access log does,
// so none of them can reach a width the access log cannot. That is
// asserted directly, per line and under both handlers, rather than
// left to the reasoning: see logbound_test.go in this package and
// in internal/handlers.
//
// The two login-throttle lines are capped defensively: chi pins
// their route to the constant path "/pages/login", so no request
// through the mux can widen either one. Their assertions call
// RecordLoginFailure and the login handler directly with the path
// a caller on a parameterised route would supply, which is the
// only way those caps can be pinned at all.
//
// One writer in that set is not a handler's own slog call. The
// GORM adapter in internal/gormlog logs the statement with the
// client-chosen parameter already interpolated into it, which on
// the receiver and login lookups is exactly the value the budgets
// above exist for. Its widest line spends two logfield.MaxBytes
// budgets — the statement and the driver error — against a fixed
// portion smaller than this one's, and
// internal/gormlog/gormlog_test.go asserts every line it emits
// against this constant directly rather than leaving it as
// arithmetic.
//
// What it does NOT cover, so that the figure above is not read as
// more than it is:
//
// - Lines carrying an AUTHENTICATED operator's own input, which
// are not truncated at all: the webhook name on "webhook
// created" and the target host on "target URL blocked by SSRF
// protection" (both internal/handlers/source_management.go),
// and target_name in internal/delivery/engine.go and
// target_http.go. Each is bounded only by the 1 MB form body
// cap, so a 100 KB name writes one line of roughly 600 KB.
// Deliberate: truncating the operator's own configuration
// echoed back costs debuggability against no adversary.
// - The "log" delivery target, which exists to write the whole
// inbound event to the log. Deliberate; see
// internal/delivery/target_log.go.
// - The record a recovered panic writes, which is not an access
// log line: its client-supplied fields are charged the same
// budgets, but it carries a whole goroutine stack as well and
// is wider than this figure. It has its own stated ceiling,
// MaxPanicLogLineBytes in recoverer.go, and is written once
// per recovered panic rather than once per request.
MaxAccessLogLineBytes = 2560
)
//nolint:revive // MiddlewareParams is a standard fx naming convention.
@@ -158,12 +43,6 @@ type Middleware struct {
log *slog.Logger
params *MiddlewareParams
session *session.Session
// loginGuard counts failed credential verifications and bounds
// concurrent password hashing. It is built on first use so that
// every construction path gets one; see guard().
loginGuardOnce sync.Once
loginGuard *loginGuard
}
// New creates a Middleware from the provided fx parameters.
@@ -215,70 +94,6 @@ func (lrw *loggingResponseWriter) WriteHeader(code int) {
lrw.ResponseWriter.WriteHeader(code)
}
// concreteLogURL renders the request's own URL for the access log
// branches that keep it, with the query string replaced by a fixed
// marker.
//
// The path on those branches is bounded by the service's routes or by
// the operator's data — a 2xx on the receiver means the UUID named a
// stored entrypoint, a 2xx under /s means the file is in the embedded
// tree. The query is not bounded by anything: /.well-known/healthcheck
// and /s/* take no authentication and sit behind no rate limiter, and
// /pages/login behind only the login limiter, so any of them will
// answer 200 to a URL carrying an arbitrary number of arbitrary bytes
// after the '?'. Keeping the path and dropping the query is what makes
// this branch as bounded as the pattern branches below.
//
// Nothing debuggable is lost. One route in the service reads a query
// parameter at all — `page`, on the authenticated pagination links in
// internal/handlers/source_management.go — and the alternatives that
// would preserve more (a key count, a key allowlist) all require
// parsing an attacker-sized query on every request, which is work an
// unauthenticated client would then be choosing for us.
func concreteLogURL(r *http.Request) string {
path := r.URL.EscapedPath()
if r.URL.RawQuery == "" && !r.URL.ForceQuery {
return path
}
return path + redactedQuery
}
// accessLogURL returns the value for the access log's url field.
//
// 2xx and 5xx responses get the concrete path (see concreteLogURL). A
// success resolved against a static route or against the operator's
// own data — on the receiver, a 2xx means the UUID named a stored
// entrypoint — and a server error is our own bug, where the exact URL
// is the primary evidence and which no client can provoke at will.
//
// 3xx and 4xx responses get the chi route pattern instead. Those are
// the outcomes an unauthenticated client drives for free: 404 or 429
// on any invented /webhook/ path, 303 to the login page on any
// invented /user/ path. Logging the concrete URL there lets a flood
// write attacker-chosen text, of attacker-chosen length, into the
// operator's log at one line per request. The pattern comes from the
// router's own table, so it is bounded by the service's routes while
// still naming which class of request was rejected.
//
// The pattern is only populated once routing has run, so this must be
// called after the handler returns, not before.
func accessLogURL(r *http.Request, status int) string {
if status < http.StatusMultipleChoices ||
status >= http.StatusInternalServerError {
return concreteLogURL(r)
}
if rc := chi.RouteContext(r.Context()); rc != nil {
if pattern := rc.RoutePattern(); pattern != "" {
return pattern
}
}
return unmatchedRoute
}
// Logging returns middleware that logs each HTTP request with
// timing and metadata.
func (s *Middleware) Logging() func(http.Handler) http.Handler {
@@ -303,27 +118,13 @@ func (s *Middleware) Logging() func(http.Handler) http.Handler {
}
}
// Every field below that a client can influence is
// truncated to a fixed budget, so the size of this
// line does not track the size of the request.
s.log.Info("http request",
"request_start", start,
"method", logfield.Truncate(
r.Method, maxLogMethodBytes,
),
"url", logfield.Truncate(
accessLogURL(r, lrw.statusCode),
logfield.MaxBytes,
),
"useragent", logfield.Truncate(
r.UserAgent(), logfield.MaxBytes,
),
"request_id", logfield.Truncate(
requestID, maxLogRequestIDBytes,
),
"referer", logfield.Truncate(
r.Referer(), logfield.MaxBytes,
),
"method", r.Method,
"url", r.URL.String(),
"useragent", r.UserAgent(),
"request_id", requestID,
"referer", r.Referer(),
"proto", r.Proto,
"remoteIP", ipFromHostPort(r.RemoteAddr),
"status", lrw.statusCode,
@@ -390,21 +191,10 @@ func (s *Middleware) RequireAuth() func(http.Handler) http.Handler {
// session lands here and is sent back to the login
// page.
if !s.session.IsAuthenticated(sess) {
// This is the unauthenticated branch, so both
// fields are entirely client-chosen and neither
// is bounded by anything the router did. DEBUG
// is off by default, but turning it on to
// diagnose a problem must not hand a client an
// unbounded write into the log, so the same
// budgets apply here as in the access log.
s.log.Debug(
"auth middleware: unauthenticated request",
"path", logfield.Truncate(
r.URL.Path, logfield.MaxBytes,
),
"method", logfield.Truncate(
r.Method, maxLogMethodBytes,
),
"path", r.URL.Path,
"method", r.Method,
)
http.Redirect(
w, r, "/pages/login", http.StatusSeeOther,
@@ -564,26 +354,10 @@ func (s *Middleware) MaxBodySize(
}
if r.ContentLength > maxBytes {
// This runs ahead of RequireAuth (see
// setupUserRoutes and friends in
// internal/server/routes.go), so an
// unauthenticated client reaches it with a path
// of its own choosing and its own length —
// POST /source/<8 KB>/edit with an oversize
// declared Content-Length costs nothing to
// send. At WARN, on by default, that is a
// write into the operator's log sized by the
// attacker unless the path is capped. Same
// budgets as the access log, so this line
// cannot be wider than that one.
s.log.Warn(
"request body exceeds limit",
"method", logfield.Truncate(
r.Method, maxLogMethodBytes,
),
"path", logfield.Truncate(
r.URL.Path, logfield.MaxBytes,
),
"method", r.Method,
"path", r.URL.Path,
"content_length", r.ContentLength,
"limit", maxBytes,
)

View File

@@ -57,21 +57,7 @@ func testMiddlewareWithSessionClock(
SessionIdleTimeout: idleTimeout,
}
sessManager := newTestSessionManager(cfg, log, clock)
m := middleware.NewForTest(log, cfg, sessManager)
return m, sessManager, clock
}
// newTestSessionManager builds the real session.Session the
// middleware tests run against: an in-memory cookie store with a
// known key, and optionally a manually advanced clock.
func newTestSessionManager(
cfg *config.Config,
log *slog.Logger,
clock *fakeClock,
) *session.Session {
// Create a real session manager with a known key
key := make([]byte, testKeySize)
for i := range key {
@@ -93,7 +79,11 @@ func newTestSessionManager(
now = clock.Now
}
return session.NewForTest(store, cfg, log, key, now)
sessManager := session.NewForTest(store, cfg, log, key, now)
m := middleware.NewForTest(log, cfg, sessManager)
return m, sessManager, clock
}
// fakeClock is a manually advanced clock, so session expiry can be

View File

@@ -9,19 +9,14 @@ import (
"time"
"github.com/go-chi/httprate"
"sneak.berlin/go/webhooker/internal/logfield"
)
const (
// loginRateLimit is the maximum number of FAILED login attempts
// one client may make against one submitted username per
// interval before further failures are answered 429. Successful
// attempts are never counted and never throttled — see
// loginGuard.
// loginRateLimit is the maximum number of login attempts
// per interval.
loginRateLimit = 5
// loginRateInterval is the time window for the login failure
// limit.
// loginRateInterval is the time window for the rate limit.
loginRateInterval = 1 * time.Minute
// passwordChangeRateLimit is the maximum number of password
@@ -53,12 +48,6 @@ const (
// bound every request pays a walk proportional to whatever the
// client sent.
maxForwardedHops = 64
// ipv6BucketBits is the prefix length IPv6 clients are bucketed
// on. A routed /64 is the normal residential and mobile
// allocation, so it is the unit an attacker gets addresses in
// and therefore the unit worth limiting.
ipv6BucketBits = 64
)
// normalizeAddr strips the IPv4-in-IPv6 wrapper and any zone from
@@ -67,40 +56,6 @@ func normalizeAddr(addr netip.Addr) netip.Addr {
return addr.Unmap().WithZone("")
}
// bucketKey is the rate-limit bucket identity of a client address.
// IPv4 keys on the full address; IPv6 keys on its /64 prefix,
// because keying IPv6 per /128 lets one ordinary subscriber rotate
// source addresses inside its own routed /64 and mint a fresh bucket
// per request — evading every limiter here at the network layer,
// with no spoofing and nothing to detect.
//
// An IPv4-mapped address (::ffff:1.2.3.4) is keyed as the IPv4
// address it carries, never masked to a /64: mapped form all shares
// the ::ffff:0:0/96 prefix, so masking would collapse every IPv4
// client reaching a proxy that emits it into one bucket. Callers
// pass addresses through normalizeAddr, which already unmaps; the
// unmap here keeps the property true of the key function itself.
//
// The two families cannot collide: an IPv4 key is a bare dotted
// quad, and an IPv6 key always carries a "/64" suffix.
func bucketKey(addr netip.Addr) string {
addr = addr.Unmap()
if addr.Is4() {
return addr.String()
}
// Prefix errors only on a negative bit count, on over 32 bits
// for an IPv4 address, or on over 128 for IPv6. The count here
// is the constant 64 and the IPv4 case returned above, so the
// error is unreachable. (The zero Addr does not error either: it
// yields the zero Prefix. Neither call site can produce one,
// since both parse the address first.)
prefix, _ := addr.Prefix(ipv6BucketBits)
return prefix.String()
}
// isTrustedProxy reports whether addr belongs to a network the
// operator listed in TRUSTED_PROXIES. The list is empty by default,
// so by default nothing is trusted.
@@ -188,9 +143,6 @@ func (m *Middleware) forwardedClientAddr(
// another client's bucket, by picking an X-Forwarded-For value —
// which makes every limit here decorative against a deliberate
// attacker.
//
// The address that identifies the client is then reduced to a bucket
// by bucketKey: full address for IPv4, /64 prefix for IPv6.
func (m *Middleware) rateLimitKey(r *http.Request) (string, error) {
return m.clientKey(r), nil
}
@@ -200,48 +152,35 @@ func (m *Middleware) clientKey(r *http.Request) string {
peer, err := netip.ParseAddr(ipFromHostPort(r.RemoteAddr))
if err != nil {
// Not an address we can reason about; key on the raw
// value, the most specific identity left. Distinct
// RemoteAddr values stay in distinct buckets, so this
// path cannot silently collapse unrelated clients
// together. On a Unix-socket listener every peer
// carries the same RemoteAddr and so shares one bucket,
// which is the fail-closed direction.
// value, the most specific identity left. On a
// Unix-socket listener every peer carries the same
// RemoteAddr and so shares one bucket, which is the
// fail-closed direction.
return r.RemoteAddr
}
peer = normalizeAddr(peer)
if !m.isTrustedProxy(peer) {
return bucketKey(peer)
return peer.String()
}
if addr, ok := m.forwardedClientAddr(r); ok {
return bucketKey(addr)
return addr.String()
}
return bucketKey(peer)
return peer.String()
}
// tooManyRequests returns the 429 handler used by the
// tooManyRequests returns the 429 handler used by the login,
// password-change and per-entrypoint receiver limiters: it logs the
// rejection with logMessage and answers with responseMessage.
// httprate adds the Retry-After header (RFC 6585). The aggregate
// receiver limiter uses floodTooManyRequests instead.
//
// The path is capped against the same budget as the access log's url
// field. The per-entrypoint receiver limiter is unauthenticated and
// its path is a client-chosen segment of client-chosen length, so at
// WARN an uncapped path would let a sender pick the size of the line
// it writes — the same defect the access log capping closed.
func (m *Middleware) tooManyRequests(
logMessage, responseMessage string,
) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
m.log.Warn(
logMessage,
"path", logfield.Truncate(
r.URL.Path, logfield.MaxBytes,
),
)
m.log.Warn(logMessage, "path", r.URL.Path)
http.Error(w, responseMessage, http.StatusTooManyRequests)
}
}
@@ -271,15 +210,26 @@ func (m *Middleware) floodTooManyRequests(
}
}
// LoginRateLimit returns middleware that enforces per-IP rate
// limiting on login attempts using go-chi/httprate. Only POST
// requests are rate-limited; GET requests (rendering the login
// form) pass through unaffected. When the rate limit is exceeded,
// a 429 Too Many Requests response is returned. Clients are
// identified by rateLimitKey.
func (m *Middleware) LoginRateLimit() func(http.Handler) http.Handler {
return m.postRateLimit(
loginRateLimit,
loginRateInterval,
"login rate limit exceeded",
"Too many login attempts. Please try again later.",
)
}
// PasswordChangeRateLimit returns middleware that enforces
// per-IP rate limiting on password change attempts. The change
// endpoint verifies the current password, so without a limit a
// stolen session could be used to brute-force it.
//
// Unlike the login POST this limit is still spent on arrival, which
// is safe here: RequireAuth runs ahead of it, so only a request
// already carrying a valid session can reach the bucket, and an
// operator locked out of changing a password can still log in.
// stolen session could be used to brute-force it; the limit
// matches the login endpoint's.
func (m *Middleware) PasswordChangeRateLimit() func(http.Handler) http.Handler {
return m.postRateLimit(
passwordChangeRateLimit,

View File

@@ -15,19 +15,18 @@ import (
"time"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
"sneak.berlin/go/webhooker/internal/config"
"sneak.berlin/go/webhooker/internal/middleware"
)
func TestPostRateLimit_AllowsGET(t *testing.T) {
func TestLoginRateLimit_AllowsGET(t *testing.T) {
t.Parallel()
m, _ := testMiddleware(t, config.EnvironmentDev)
var callCount int
handler := m.PasswordChangeRateLimit()(http.HandlerFunc(
handler := m.LoginRateLimit()(http.HandlerFunc(
func(w http.ResponseWriter, _ *http.Request) {
callCount++
@@ -39,7 +38,7 @@ func TestPostRateLimit_AllowsGET(t *testing.T) {
for i := range 20 {
req := httptest.NewRequestWithContext(
context.Background(),
http.MethodGet, "/user/admin/password", nil,
http.MethodGet, "/pages/login", nil,
)
req.RemoteAddr = "192.168.1.1:12345"
@@ -110,6 +109,20 @@ func runPostLimitTest(
assert.Equal(t, limit, callCount)
}
func TestLoginRateLimit_LimitsPOST(t *testing.T) {
t.Parallel()
m, _ := testMiddleware(t, config.EnvironmentDev)
runPostLimitTest(
t,
m.LoginRateLimit(),
middleware.LoginRateLimitConst,
"/pages/login",
"10.0.0.1:12345",
)
}
func TestPasswordChangeRateLimit_LimitsPOST(t *testing.T) {
t.Parallel()
@@ -124,19 +137,19 @@ func TestPasswordChangeRateLimit_LimitsPOST(t *testing.T) {
)
}
func TestPostRateLimit_IndependentPerIP(t *testing.T) {
func TestLoginRateLimit_IndependentPerIP(t *testing.T) {
t.Parallel()
m, _ := testMiddleware(t, config.EnvironmentDev)
handler := m.PasswordChangeRateLimit()(http.HandlerFunc(
handler := m.LoginRateLimit()(http.HandlerFunc(
func(w http.ResponseWriter, _ *http.Request) {
w.WriteHeader(http.StatusOK)
},
))
// Exhaust limit for IP1
for range middleware.PasswordChangeRateLimitConst {
for range middleware.LoginRateLimitConst {
req := httptest.NewRequestWithContext(
context.Background(),
http.MethodPost, "/pages/login", nil,
@@ -353,41 +366,10 @@ func TestReceiverRateLimit_CountsEveryMethod(t *testing.T) {
}
const (
// limitedPath is the endpoint these tests drive the shared POST
// rate limiter through. It is the password-change path: since
// the login POST verifies credentials before spending any
// budget, the password-change limiter is the only pre-emptive
// POST limiter left, and it is what pins the shared key
// function's behaviour here.
limitedPath = "/user/admin/password"
loginPath = "/pages/login"
headerXFF = "X-Forwarded-For"
headerReal = "X-Real-IP"
headerTrue = "True-Client-IP"
// clientIPv4 is the sample IPv4 client address these tests key
// on, both directly and in IPv4-mapped form. clientIPv4Alt is
// its neighbour, used to show the two do not share a bucket.
clientIPv4 = "198.51.100.7"
clientIPv4Alt = "198.51.100.8"
// clientIPv6 and clientIPv6Same are two addresses inside one
// routed /64, so both must key on clientBucketV6.
// clientIPv6Other is a different allocation and must key on
// clientOtherBucketV6.
clientIPv6 = "2001:db8:1:2:3:4:5:6"
clientIPv6Same = "2001:db8:1:2:aaaa:bbbb:cccc:dddd"
clientIPv6Other = "2001:db8:1:3::1"
clientBucketV6 = "2001:db8:1:2::/64"
clientOtherBucketV6 = "2001:db8:1:3::/64"
// trustedProxyCIDR is the proxy network the forwarded-path
// tests configure, and trustedPeer an address inside it. A
// production deployment is required to run behind a reverse
// proxy with TRUSTED_PROXIES set, so this is the shape the
// bucketing has to hold in.
trustedProxyCIDR = "10.0.0.0/8"
trustedPeer = "10.0.0.1:44444"
)
// assertSharedBucket drives the login limiter from peer with the
@@ -408,10 +390,10 @@ func assertSharedBucket(
m := rateLimitMiddleware(
t, &config.Config{TrustedProxies: proxies},
)
handler := m.PasswordChangeRateLimit()(okHandler())
handler := m.LoginRateLimit()(okHandler())
for i := range middleware.PasswordChangeRateLimitConst {
w := postWithHeaders(handler, peer, limitedPath, headers(i))
for i := range middleware.LoginRateLimitConst {
w := postWithHeaders(handler, peer, loginPath, headers(i))
assert.Equal(
t, http.StatusOK, w.Code,
"request %d should pass", i,
@@ -419,8 +401,8 @@ func assertSharedBucket(
}
w := postWithHeaders(
handler, peer, limitedPath,
headers(middleware.PasswordChangeRateLimitConst),
handler, peer, loginPath,
headers(middleware.LoginRateLimitConst),
)
assert.Equal(t, http.StatusTooManyRequests, w.Code, msg)
}
@@ -476,8 +458,8 @@ func TestRateLimitKey_SingleValuedHeadersIgnoredFromTrustedPeer(
t.Parallel()
assertSharedBucket(
t, trustedProxies(trustedProxyCIDR),
trustedPeer,
t, trustedProxies("10.0.0.0/8"),
"10.0.0.1:44444",
func(i int) map[string]string {
return map[string]string{
header: fmt.Sprintf(
@@ -513,8 +495,8 @@ func TestRateLimitKey_MalformedRightmostHopFallsBackToPeer(
t.Parallel()
assertSharedBucket(
t, trustedProxies(trustedProxyCIDR),
trustedPeer,
t, trustedProxies("10.0.0.0/8"),
"10.0.0.1:44444",
func(i int) map[string]string {
return map[string]string{
headerXFF: fmt.Sprintf(
@@ -540,27 +522,27 @@ func TestRateLimitKey_ForwardedHonouredFromTrustedPeer(
t.Parallel()
m := rateLimitMiddleware(t, &config.Config{
TrustedProxies: trustedProxies(trustedProxyCIDR),
TrustedProxies: trustedProxies("10.0.0.0/8"),
})
handler := m.PasswordChangeRateLimit()(okHandler())
handler := m.LoginRateLimit()(okHandler())
const peer = trustedPeer
const peer = "10.0.0.1:44444"
first := map[string]string{headerXFF: clientIPv4}
first := map[string]string{headerXFF: "198.51.100.7"}
for range middleware.PasswordChangeRateLimitConst {
postWithHeaders(handler, peer, limitedPath, first)
for range middleware.LoginRateLimitConst {
postWithHeaders(handler, peer, loginPath, first)
}
w := postWithHeaders(handler, peer, limitedPath, first)
w := postWithHeaders(handler, peer, loginPath, first)
assert.Equal(
t, http.StatusTooManyRequests, w.Code,
"the forwarded client's own bucket must fill up",
)
w = postWithHeaders(
handler, peer, limitedPath,
map[string]string{headerXFF: clientIPv4Alt},
handler, peer, loginPath,
map[string]string{headerXFF: "198.51.100.8"},
)
assert.Equal(
t, http.StatusOK, w.Code,
@@ -577,7 +559,7 @@ func TestRateLimitKey_ChainWalkSkipsClientPrepended(t *testing.T) {
t.Parallel()
assertSharedBucket(
t, trustedProxies(trustedProxyCIDR), trustedPeer,
t, trustedProxies("10.0.0.0/8"), "10.0.0.1:44444",
func(i int) map[string]string {
return map[string]string{
headerXFF: fmt.Sprintf(
@@ -612,7 +594,7 @@ func TestRateLimitKey_LongChainCapsWalkAndFallsBackToPeer(
start := time.Now()
assertSharedBucket(
t, trustedProxies(trustedProxyCIDR), trustedPeer,
t, trustedProxies("10.0.0.0/8"), "10.0.0.1:44444",
func(i int) map[string]string {
return map[string]string{
headerXFF: fmt.Sprintf("9.9.9.%d%s", i+1, padding),
@@ -651,13 +633,13 @@ func TestRateLimitKey_LongChainAllocationIsBounded(t *testing.T) {
)
m := rateLimitMiddleware(t, &config.Config{
TrustedProxies: trustedProxies(trustedProxyCIDR),
TrustedProxies: trustedProxies("10.0.0.0/8"),
})
req := httptest.NewRequestWithContext(
context.Background(), http.MethodPost, limitedPath, nil,
context.Background(), http.MethodPost, loginPath, nil,
)
req.RemoteAddr = trustedPeer
req.RemoteAddr = "10.0.0.1:44444"
req.Header.Set(
headerXFF, "9.9.9.9"+strings.Repeat(", 10.0.0.2", hops),
)
@@ -853,435 +835,3 @@ func TestReceiverRateLimit_IgnoresForwardedFromUntrustedPeer(
"not mint a fresh receiver bucket",
)
}
// clientKeyFor returns the bucket key m computes for a request whose
// direct peer is remoteAddr and which carries no forwarded headers.
func clientKeyFor(
t *testing.T, m *middleware.Middleware, remoteAddr string,
) string {
t.Helper()
req := httptest.NewRequestWithContext(
context.Background(), http.MethodPost, limitedPath, nil,
)
req.RemoteAddr = remoteAddr
return middleware.ClientKeyForTest(m, req)
}
// TestRateLimitKey_IPv6BucketsByPrefix pins the key function's
// address-family behaviour. IPv6 clients must bucket by /64 — a
// routed /64 is the normal residential and mobile allocation, so
// per-/128 keying lets one subscriber rotate source addresses and
// mint a fresh bucket per request — while IPv4 keeps keying on the
// full address and IPv4-mapped form is keyed as the IPv4 address it
// carries.
func TestRateLimitKey_IPv6BucketsByPrefix(t *testing.T) {
t.Parallel()
m := rateLimitMiddleware(t, &config.Config{})
for _, tc := range []struct {
name string
peer string
want string
about string
}{{
name: "ipv6",
peer: "[" + clientIPv6 + "]:44444",
want: clientBucketV6,
about: "an IPv6 peer must key on its /64",
}, {
name: "ipv6-other-in-same-64",
peer: "[" + clientIPv6Same + "]:1",
want: clientBucketV6,
about: "another address in the same /64 must key the same",
}, {
name: "ipv6-different-64",
peer: "[" + clientIPv6Other + "]:44444",
want: clientOtherBucketV6,
about: "a different /64 must key differently",
}, {
name: "ipv4",
peer: clientIPv4 + ":44444",
want: clientIPv4,
about: "IPv4 must keep keying on the full address",
}, {
name: "ipv4-neighbour",
peer: clientIPv4Alt + ":44444",
want: clientIPv4Alt,
about: "adjacent IPv4 addresses must not share a bucket",
}, {
name: "ipv4-mapped",
peer: "[::ffff:" + clientIPv4 + "]:44444",
want: clientIPv4,
about: "IPv4-mapped form must key as the IPv4 address, " +
"not be masked to a /64: mapped addresses all share " +
"::ffff:0:0/96, so masking would collapse every IPv4 " +
"client behind a mapping proxy into one bucket",
}} {
t.Run(tc.name, func(t *testing.T) {
t.Parallel()
assert.Equal(
t, tc.want, clientKeyFor(t, m, tc.peer), tc.about,
)
})
}
}
// TestRateLimitKey_FamiliesDoNotCollide pins the structure the
// no-collision property rests on, rather than one sample pair: every
// IPv4 key is a bare address and every IPv6 key is a /64 in CIDR
// form, so the two name spaces are disjoint by shape. Dropping the
// masking strips the suffix that guarantees it, which is why this
// asserts the form of each key and not just that two of them differ.
func TestRateLimitKey_FamiliesDoNotCollide(t *testing.T) {
t.Parallel()
// Restated here rather than imported from the package under
// test, so that changing the production bucket width fails this
// test instead of silently moving with it.
const wantBits = 64
m := rateLimitMiddleware(t, &config.Config{})
v4Keys := map[string]bool{}
for _, peer := range []string{
clientIPv4 + ":44444",
clientIPv4Alt + ":44444",
"[::ffff:" + clientIPv4 + "]:44444",
} {
key := clientKeyFor(t, m, peer)
addr, err := netip.ParseAddr(key)
require.NoError(
t, err, "%s: an IPv4 key must be a bare address", peer,
)
assert.True(
t, addr.Is4(),
"%s: an IPv4 key must be a dotted quad, got %q", peer, key,
)
v4Keys[key] = true
}
for _, peer := range []string{
"[" + clientIPv6 + "]:44444",
"[" + clientIPv6Same + "]:44444",
"[" + clientIPv6Other + "]:44444",
"[2001:db8::" + clientIPv4 + "]:44444",
} {
key := clientKeyFor(t, m, peer)
prefix, err := netip.ParsePrefix(key)
require.NoError(
t, err, "%s: an IPv6 key must be a CIDR prefix", peer,
)
assert.Equal(
t, wantBits, prefix.Bits(),
"%s: an IPv6 key must name a /64", peer,
)
assert.False(
t, v4Keys[key],
"%s: an IPv6 key must never equal an IPv4 key", peer,
)
}
}
// TestRateLimitKey_UnparseablePeerKeepsDistinctBuckets covers the
// fallback path. A RemoteAddr that is not an address must not panic,
// and must not drop unrelated clients into one shared bucket by
// accident: the raw value is the most specific identity left, so
// distinct values stay in distinct buckets.
func TestRateLimitKey_UnparseablePeerKeepsDistinctBuckets(
t *testing.T,
) {
t.Parallel()
m := rateLimitMiddleware(t, &config.Config{})
first := clientKeyFor(t, m, "not-an-address")
second := clientKeyFor(t, m, "also-not-an-address:1234")
assert.NotEmpty(t, first)
assert.NotEqual(
t, first, second,
"unparseable peers must not collapse into one bucket",
)
}
// TestPostRateLimit_IPv6SharesBucketWithinSlash64 is the behavioural
// half, and the regression test for the bypass itself: a client that
// rotates source addresses inside its own routed /64 must stay in one
// bucket. Reverting the masking makes this test fail, because each
// rotated address would mint a fresh bucket and nothing would be
// rejected.
func TestPostRateLimit_IPv6SharesBucketWithinSlash64(t *testing.T) {
t.Parallel()
m := rateLimitMiddleware(t, &config.Config{})
handler := m.PasswordChangeRateLimit()(okHandler())
for i := range middleware.PasswordChangeRateLimitConst {
w := postWithHeaders(
handler,
fmt.Sprintf("[2001:db8:1:2::%d]:44444", i+1),
limitedPath, nil,
)
assert.Equal(
t, http.StatusOK, w.Code, "request %d should pass", i,
)
}
w := postWithHeaders(
handler, "[2001:db8:1:2::ffff]:44444", limitedPath, nil,
)
assert.Equal(
t, http.StatusTooManyRequests, w.Code,
"rotating source addresses inside one routed /64 must not "+
"mint fresh buckets",
)
}
// TestPostRateLimit_IPv6IndependentAcrossSlash64 is the other side
// of the trade: bucketing by /64 must not merge separate allocations,
// so a client in a different /64 keeps its own limit.
func TestPostRateLimit_IPv6IndependentAcrossSlash64(t *testing.T) {
t.Parallel()
m := rateLimitMiddleware(t, &config.Config{})
handler := m.PasswordChangeRateLimit()(okHandler())
for range middleware.PasswordChangeRateLimitConst + 1 {
postWithHeaders(
handler, "[2001:db8:1:2::1]:44444", limitedPath, nil,
)
}
w := postWithHeaders(
handler, "[2001:db8:1:3::1]:44444", limitedPath, nil,
)
assert.Equal(
t, http.StatusOK, w.Code,
"a different /64 must have its own bucket",
)
}
// TestPostRateLimit_IPv4IndependentPerAddress guards against the
// masking leaking into IPv4: two addresses one apart must still hold
// separate buckets.
func TestPostRateLimit_IPv4IndependentPerAddress(t *testing.T) {
t.Parallel()
m := rateLimitMiddleware(t, &config.Config{})
handler := m.PasswordChangeRateLimit()(okHandler())
for range middleware.PasswordChangeRateLimitConst + 1 {
postWithHeaders(
handler, clientIPv4+":44444", limitedPath, nil,
)
}
w := postWithHeaders(
handler, clientIPv4Alt+":44444", limitedPath, nil,
)
assert.Equal(
t, http.StatusOK, w.Code,
"a second IPv4 address must have its own bucket",
)
}
// forwardedKeyFor returns the bucket key m computes for a request
// that arrives from trustedPeer — a configured trusted proxy — and
// names forwarded as its client in X-Forwarded-For. That is the
// production path: a deployment is required to run behind a reverse
// proxy with TRUSTED_PROXIES set, so the forwarded address, not the
// peer, is what the limiters bucket on there.
func forwardedKeyFor(
t *testing.T, m *middleware.Middleware, forwarded string,
) string {
t.Helper()
req := httptest.NewRequestWithContext(
context.Background(), http.MethodPost, limitedPath, nil,
)
req.RemoteAddr = trustedPeer
req.Header.Set(headerXFF, forwarded)
return middleware.ClientKeyForTest(m, req)
}
// TestRateLimitKey_ForwardedIPv6BucketsByPrefix pins the /64
// bucketing on the trusted-proxy branch. The direct-peer tests above
// cannot reach it, so without this the masking could be reverted for
// forwarded clients alone — the only shape a production deployment
// runs in — and the rest of the suite would stay green.
func TestRateLimitKey_ForwardedIPv6BucketsByPrefix(t *testing.T) {
t.Parallel()
m := rateLimitMiddleware(t, &config.Config{
TrustedProxies: trustedProxies(trustedProxyCIDR),
})
for _, tc := range []struct {
name string
forwarded string
want string
about string
}{{
name: "ipv6",
forwarded: clientIPv6,
want: clientBucketV6,
about: "a forwarded IPv6 client must key on its /64",
}, {
name: "ipv6-other-in-same-64",
forwarded: clientIPv6Same,
want: clientBucketV6,
about: "another forwarded address in the same /64 must " +
"key the same",
}, {
name: "ipv6-different-64",
forwarded: clientIPv6Other,
want: clientOtherBucketV6,
about: "a forwarded address in another /64 must differ",
}, {
name: "ipv4",
forwarded: clientIPv4,
want: clientIPv4,
about: "a forwarded IPv4 client must key on the address",
}, {
name: "ipv4-mapped",
forwarded: "::ffff:" + clientIPv4,
want: clientIPv4,
about: "a proxy that forwards IPv4-mapped form must key as " +
"the IPv4 address it carries, not be masked to a /64: " +
"mapped addresses all share ::ffff:0:0/96",
}} {
t.Run(tc.name, func(t *testing.T) {
t.Parallel()
assert.Equal(
t, tc.want,
forwardedKeyFor(t, m, tc.forwarded), tc.about,
)
})
}
}
// TestRateLimitKey_TrustedPeerUnusableForwardedMasksPeer covers the
// third bucketKey call site: the peer IS a trusted proxy, but the
// forwarded chain cannot name a client, so the key falls back to the
// peer address — and that fallback owes the same /64 masking every
// other key gets.
//
// Every existing test of this fallback uses an IPv4 proxy, where
// bucketKey is the identity function, so replacing the call with
// peer.String() leaves the whole suite green. Only operator-listed
// addresses reach this line and the fallback is fail-closed, so this
// pins behaviour rather than fixing a defect.
func TestRateLimitKey_TrustedPeerUnusableForwardedMasksPeer(
t *testing.T,
) {
t.Parallel()
const (
proxyCIDR = "2001:db8:ffff::/48"
proxyPeer = "[2001:db8:ffff:1::5]:44444"
wantKey = "2001:db8:ffff:1::/64"
)
m := rateLimitMiddleware(t, &config.Config{
TrustedProxies: trustedProxies(proxyCIDR),
})
for _, tc := range []struct {
name string
forwarded string
about string
}{{
name: "absent",
about: "no X-Forwarded-For at all falls back to the peer",
}, {
name: "unreadable-hop",
forwarded: "unknown",
about: "a hop that is not a bare address ends the walk " +
"and falls back to the peer",
}, {
name: "all-hops-trusted",
forwarded: "2001:db8:ffff:2::9",
about: "a chain naming only trusted proxies names no " +
"client, so the peer is used",
}} {
t.Run(tc.name, func(t *testing.T) {
t.Parallel()
req := httptest.NewRequestWithContext(
context.Background(),
http.MethodPost, limitedPath, nil,
)
req.RemoteAddr = proxyPeer
if tc.forwarded != "" {
req.Header.Set(headerXFF, tc.forwarded)
}
assert.Equal(
t, wantKey,
middleware.ClientKeyForTest(m, req),
"%s, masked to its /64", tc.about,
)
})
}
}
// TestPostRateLimit_ForwardedIPv6SharesBucketWithinSlash64 is the
// behavioural half on the production path: behind a trusted proxy, a
// client rotating source addresses inside its own routed /64 must
// stay in one bucket.
func TestPostRateLimit_ForwardedIPv6SharesBucketWithinSlash64(
t *testing.T,
) {
t.Parallel()
assertSharedBucket(
t, trustedProxies(trustedProxyCIDR), trustedPeer,
func(i int) map[string]string {
return map[string]string{
headerXFF: fmt.Sprintf("2001:db8:1:2::%d", i+1),
}
},
"rotating forwarded source addresses inside one routed /64 "+
"must not mint fresh buckets",
)
}
// TestPostRateLimit_ForwardedIPv6IndependentAcrossSlash64 is the
// other side of that trade on the same path: bucketing by /64 must
// not merge two allocations reaching the proxy.
func TestPostRateLimit_ForwardedIPv6IndependentAcrossSlash64(
t *testing.T,
) {
t.Parallel()
m := rateLimitMiddleware(t, &config.Config{
TrustedProxies: trustedProxies(trustedProxyCIDR),
})
handler := m.PasswordChangeRateLimit()(okHandler())
spent := map[string]string{headerXFF: clientIPv6}
for range middleware.PasswordChangeRateLimitConst + 1 {
postWithHeaders(handler, trustedPeer, limitedPath, spent)
}
w := postWithHeaders(
handler, trustedPeer, limitedPath,
map[string]string{headerXFF: clientIPv6Other},
)
assert.Equal(
t, http.StatusOK, w.Code,
"a forwarded client in a different /64 must have its own "+
"bucket",
)
}

View File

@@ -1,209 +0,0 @@
package middleware
import (
"errors"
"fmt"
"net/http"
"runtime/debug"
"github.com/go-chi/chi/middleware"
"sneak.berlin/go/webhooker/internal/logfield"
)
const (
// maxPanicValueBytes bounds the recovered panic value. The value
// is our own text, but a handler is free to build one out of the
// request — panic(fmt.Sprintf("bad %q", r.URL.Path)) — so it is
// charged the same budget the access log gives a field the
// client supplies outright.
maxPanicValueBytes = logfield.MaxBytes
// maxPanicStackBytes bounds the stack, in the same ENCODED bytes
// logfield.Truncate charges everywhere else. Nothing a client
// sends chooses the depth of our own call stack, so this is not
// a safety limit; it is what makes MaxPanicLogLineBytes an
// arithmetic ceiling rather than an observation. A stack is cut
// at its far end, which is net/http's accept frames — the panic
// site and the handler that reached it are at the near end and
// are always kept.
//
// For scale: a handler panicking under the full shipped
// middleware chain produces a stack of roughly 3,690 bytes in a
// record of roughly 3,960, so this budget holds better than
// twice the depth that case reaches. Neither number is an
// invariant — debug.Stack() embeds absolute source paths, so
// both move with where the tree sits, and four checkouts have
// reported records of 3,959, 3,961, 3,984 and 4,026 bytes.
// internal/server's TestPanicThroughProductionRouter asserts the
// ceiling and that the stack arrived uncut, not the figures.
maxPanicStackBytes = 8192
// MaxPanicLogLineBytes is the ceiling on the single line a
// recovered panic writes. It is wider than
// MaxAccessLogLineBytes, which bounds a line written once per
// request, where this one is written once per recovered panic.
//
// panic 512+11 = 523
// stack 8192+11 = 8203
// request_id 128+11 = 139
// fixed portion = 256
// ----
// 9121
//
// The fixed portion is the JSON punctuation, the field names,
// the level, the message, the timestamp at its longest and the
// response_committed boolean.
//
// Stated at 10240 so the figure carries headroom rather than
// sitting on the arithmetic, exactly as MaxAccessLogLineBytes
// is. Both handlers internal/logger can install are covered, for
// the reason given there: logfield.EncodedBytes charges every
// rune the wider of the two.
//
// The 9121 and the 10240 are the invariants here. What follows
// is illustration: with all three growable fields — the stack,
// the panic value and a client-supplied X-Request-Id — driven
// past their budgets at once, TestRecovererBoundsTheStack
// measured 9,009 bytes on the JSON handler and 8,982 to 8,983 on
// the text one in this checkout. Neither is fixed: the stack's
// own content decides where its cut lands, so the figures move by
// a byte or so between runs and with the checkout. The test
// asserts the ceiling and that every growable field was cut,
// never the figures.
MaxPanicLogLineBytes = 10240
)
// recoverResponseWriter records whether the response has been
// committed, which is the one thing the recoverer cannot learn from
// the panic itself: a handler that panics after writing a status has
// already spent the response, and a second WriteHeader would only
// draw net/http's "superfluous response.WriteHeader" complaint
// without changing what the client received.
type recoverResponseWriter struct {
http.ResponseWriter
committed bool
}
func (w *recoverResponseWriter) WriteHeader(code int) {
w.committed = true
w.ResponseWriter.WriteHeader(code)
}
func (w *recoverResponseWriter) Write(b []byte) (int, error) {
// An unheralded Write commits the response just as surely as
// WriteHeader does: net/http sends 200 in front of it.
w.committed = true
//nolint:wrapcheck // Pass the writer's own error through unchanged.
return w.ResponseWriter.Write(b)
}
// Unwrap lets http.ResponseController reach the writer underneath, so
// a handler can still flush or set a write deadline through this
// wrapper.
func (w *recoverResponseWriter) Unwrap() http.ResponseWriter {
return w.ResponseWriter
}
// Recoverer returns middleware that turns a handler panic into one
// structured ERROR record and a 500, rather than a dropped
// connection.
//
// It replaces chi's middleware.Recoverer, which does neither on a
// current Go release. chi v1.5.5's pretty-printer scans the stack for
// a frame beginning "panic(0x", which the runtime has not emitted
// since it started printing "panic({0x...}"; the scan therefore never
// terminates early, every line reaches decorateFuncCallLine, and that
// function slices pkg[strings.Index(pkg, "."):] without checking for
// -1. The resulting second panic escapes chi's own deferred function,
// so its WriteHeader(500) never runs and net/http closes the
// connection reporting its own crash instead of the original one.
// See https://git.eeqj.de/sneak/webhooker/issues/187.
//
// chi v5.3.1 has since fixed both halves of that — it scans for
// "panic(" and guards the index — so upgrading would restore the 500.
// It would not give what this does: v5 still writes an ANSI-coloured
// pretty stack straight to os.Stderr, outside internal/logger, outside
// any budget, at no level the operator set.
//
// Where this sits in the chain is load-bearing, and routes.go states
// it: inside everything that observes the response, so the 500 is
// what the access log records and the metrics count, and outside the
// sentryhttp handler, whose Repanic option depends on something
// further out recovering what it re-raises.
func (s *Middleware) Recoverer() func(http.Handler) http.Handler {
return func(next http.Handler) http.Handler {
return http.HandlerFunc(func(
w http.ResponseWriter,
r *http.Request,
) {
rw := &recoverResponseWriter{ResponseWriter: w}
defer func() {
rvr := recover()
if rvr == nil {
return
}
// http.ErrAbortHandler is a handler stating that it
// is abandoning the connection on purpose, not a
// fault. net/http special-cases it, suppressing both
// the stack trace and any response, so it is passed
// straight back out rather than logged and answered.
err, isError := rvr.(error)
if isError &&
errors.Is(err, http.ErrAbortHandler) {
panic(rvr)
}
s.logPanic(r, rvr, rw.committed)
if rw.committed {
return
}
http.Error(
rw,
http.StatusText(
http.StatusInternalServerError,
),
http.StatusInternalServerError,
)
}()
next.ServeHTTP(rw, r)
})
}
}
// logPanic writes the record. Every field it can grow is truncated to
// a fixed budget, so MaxPanicLogLineBytes holds.
//
// The request is identified by request_id alone rather than by
// repeating the method, URL and address: the access log line for the
// same request carries all of those, already bounded, and — because
// the recoverer runs inside the logging middleware — now carries the
// 500 as its status too. Repeating them here would double those
// budgets against a line already wider than the access log's ceiling,
// to say a second time what one join already says.
func (s *Middleware) logPanic(
r *http.Request,
rvr any,
committed bool,
) {
s.log.Error("handler panic",
"panic", logfield.Truncate(
fmt.Sprint(rvr), maxPanicValueBytes,
),
"stack", logfield.Truncate(
string(debug.Stack()), maxPanicStackBytes,
),
"request_id", logfield.Truncate(
middleware.GetReqID(r.Context()),
maxLogRequestIDBytes,
),
"response_committed", committed,
)
}

View File

@@ -1,645 +0,0 @@
package middleware_test
import (
"bytes"
"encoding/json"
"io"
"log"
"net/http"
"net/http/httptest"
"strings"
"testing"
"github.com/go-chi/chi"
chimw "github.com/go-chi/chi/middleware"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
"sneak.berlin/go/webhooker/internal/middleware"
)
// panicMarker is the panic value the probe handlers raise. The
// recoverer's whole job is to put this string, and not some second
// panic's, in front of an operator.
const panicMarker = "QQORIGINALPANICVALUEQQ"
// probeFuncName appears in the stack of every panic raised below,
// since that is the function raising it. Its presence is how these
// tests tell a real stack from an empty field.
const probeFuncName = "panicProbe"
// committedStatus is the status a handler sends before panicking in
// the already-committed case. It is deliberately not 200, so a test
// cannot pass on net/http's implicit default.
const committedStatus = http.StatusMultiStatus
// recovererProbe is a test server carrying one panicking route,
// behind the production recoverer.
type recovererProbe struct {
server *httptest.Server
// logs holds every record the middleware wrote.
logs *bytes.Buffer
// serverErrors holds everything net/http wrote to its own error
// log. A working recoverer leaves it empty: net/http only reports
// a request when a panic escapes the handler chain, which is the
// failure this issue is about.
serverErrors *bytes.Buffer
}
// newRecovererProbe stands up a real HTTP server — a real listener, a
// real connection, a real client — behind the production recoverer.
//
// A real server rather than an httptest.ResponseRecorder, because a
// recorder cannot express the outcome that made this a defect: chi's
// Recoverer left net/http to close the connection, which a recorder
// records as an ordinary unwritten response while a client sees EOF.
// The status a client actually receives is only observable over a
// socket.
func newRecovererProbe(
t *testing.T,
textHandler bool,
handler http.HandlerFunc,
) *recovererProbe {
t.Helper()
newMiddleware := capturingMiddleware
if textHandler {
newMiddleware = capturingTextMiddleware
}
m, logs := newMiddleware(t)
router := chi.NewRouter()
// The registration order the production router uses: RequestID
// outside so the recoverer's record can name the request,
// Logging outside so the recovered 500 is the status it records.
router.Use(chimw.RequestID)
router.Use(m.Logging())
router.Use(m.Recoverer())
router.Get("/probe", handler)
serverErrors := new(bytes.Buffer)
server := httptest.NewUnstartedServer(router)
server.Config.ErrorLog = log.New(serverErrors, "", 0)
server.Start()
t.Cleanup(server.Close)
return &recovererProbe{
server: server,
logs: logs,
serverErrors: serverErrors,
}
}
// get drives one request at the probe route and returns the response,
// or the transport error if the connection was dropped instead.
func (p *recovererProbe) get(t *testing.T) (*http.Response, error) {
t.Helper()
return p.getWithRequestID(t, "")
}
// getWithRequestID drives the same request carrying a client-supplied
// X-Request-Id. chi's RequestID middleware adopts that header verbatim
// when it is present and only generates a value when it is absent, so
// this is the third growable field on the panic record and the only
// one a client fills outright.
func (p *recovererProbe) getWithRequestID(
t *testing.T,
requestID string,
) (*http.Response, error) {
t.Helper()
req, err := http.NewRequestWithContext(
t.Context(), http.MethodGet, p.server.URL+"/probe", nil,
)
require.NoError(t, err)
if requestID != "" {
req.Header.Set(chimw.RequestIDHeader, requestID)
}
return p.server.Client().Do(req)
}
// wait shuts the server down and blocks until every in-flight request
// has finished, which is what makes the log buffer safe to read.
//
// A client returns as soon as the response is complete — or, for a
// deliberately aborted connection, as soon as it is closed — while the
// access log line for the same request is still being written on the
// server goroutine. It is idempotent, so a test may call it directly
// before reading the buffer itself.
func (p *recovererProbe) wait() {
p.server.Close()
}
// records decodes every JSON log line the probe captured.
func (p *recovererProbe) records(t *testing.T) []map[string]any {
t.Helper()
p.wait()
var out []map[string]any
for line := range strings.SplitSeq(
strings.TrimSpace(p.logs.String()), "\n",
) {
if line == "" {
continue
}
record := map[string]any{}
require.NoError(t, json.Unmarshal([]byte(line), &record))
out = append(out, record)
}
return out
}
// panicRecord returns the single "handler panic" record, failing if
// there is not exactly one.
func (p *recovererProbe) panicRecord(t *testing.T) map[string]any {
t.Helper()
var found []map[string]any
for _, record := range p.records(t) {
if record["msg"] == "handler panic" {
found = append(found, record)
}
}
require.Len(
t, found, 1,
"exactly one panic record expected, log was:\n%s",
p.logs.String(),
)
return found[0]
}
// panicProbe panics with the marker. It is a named function so the
// stack assertions have something to look for.
func panicProbe(http.ResponseWriter, *http.Request) {
panic(panicMarker)
}
func TestRecovererAnswers500AndLogsTheOriginalPanic(t *testing.T) {
t.Parallel()
probe := newRecovererProbe(t, false, panicProbe)
resp, err := probe.get(t)
require.NoError(
t, err,
"a panicking handler must answer, not drop the connection",
)
defer func() { _ = resp.Body.Close() }()
body, err := io.ReadAll(resp.Body)
require.NoError(t, err)
assert.Equal(t, http.StatusInternalServerError, resp.StatusCode)
assert.Contains(t, string(body), "Internal Server Error")
record := probe.panicRecord(t)
assert.Equal(t, "ERROR", record["level"])
assert.Equal(t, panicMarker, record["panic"])
assert.Equal(t, false, record["response_committed"])
stack, ok := record["stack"].(string)
require.True(t, ok, "the record must carry a stack")
assert.Contains(
t, stack, probeFuncName,
"the stack must reach the function that panicked",
)
assert.NotContains(
t, stack, "slice bounds out of range",
"a secondary panic must not have occurred",
)
assert.Empty(
t, probe.serverErrors.String(),
"net/http must not have had to report anything",
)
}
// TestRecovererStatusReachesTheAccessLog pins the placement. The
// recoverer runs inside the logging middleware precisely so the status
// it writes is the one the access log records; registered outside it,
// as chi's Recoverer was, the same request is logged as a 200 that the
// client never received.
func TestRecovererStatusReachesTheAccessLog(t *testing.T) {
t.Parallel()
probe := newRecovererProbe(t, false, panicProbe)
resp, err := probe.get(t)
require.NoError(t, err)
require.NoError(t, resp.Body.Close())
require.Equal(t, http.StatusInternalServerError, resp.StatusCode)
var access map[string]any
for _, record := range probe.records(t) {
if record["msg"] == "http request" {
access = record
}
}
require.NotNil(t, access, "the request must still be logged")
assert.EqualValues(
t, http.StatusInternalServerError, access["status"],
"the access log must record the status the client got",
)
// The panic record identifies its request by request_id alone,
// so that join has to work.
assert.Equal(
t, access["request_id"],
probe.panicRecord(t)["request_id"],
)
assert.NotEmpty(t, access["request_id"])
}
// TestRecovererRepanicsErrAbortHandler covers the one panic value that
// must not be turned into a 500. net/http documents it as the way a
// handler abandons a connection deliberately and special-cases it,
// suppressing both the response and its own stack report.
func TestRecovererRepanicsErrAbortHandler(t *testing.T) {
t.Parallel()
probe := newRecovererProbe(
t, false,
func(http.ResponseWriter, *http.Request) {
panic(http.ErrAbortHandler)
},
)
resp, err := probe.get(t)
if err == nil {
_ = resp.Body.Close()
}
require.Error(
t, err,
"an aborted handler must not answer with a status",
)
for _, record := range probe.records(t) {
assert.NotEqual(
t, "handler panic", record["msg"],
"a deliberate abort is not a fault to report",
)
}
assert.Empty(
t, probe.serverErrors.String(),
"net/http suppresses ErrAbortHandler; it must still see it",
)
}
// TestRecovererKeepsAnAlreadyCommittedResponse covers a handler that
// panics after sending its status. The bytes are already on the wire,
// so a second WriteHeader would change nothing the client sees and
// would draw net/http's "superfluous response.WriteHeader" report.
func TestRecovererKeepsAnAlreadyCommittedResponse(t *testing.T) {
t.Parallel()
probe := newRecovererProbe(
t, false,
func(w http.ResponseWriter, _ *http.Request) {
w.WriteHeader(committedStatus)
_, _ = w.Write([]byte("partial"))
panic(panicMarker)
},
)
resp, err := probe.get(t)
require.NoError(t, err)
defer func() { _ = resp.Body.Close() }()
body, err := io.ReadAll(resp.Body)
require.NoError(t, err)
assert.Equal(t, committedStatus, resp.StatusCode)
assert.Equal(t, "partial", string(body))
record := probe.panicRecord(t)
assert.Equal(t, panicMarker, record["panic"])
assert.Equal(
t, true, record["response_committed"],
"the record must say why no 500 was sent",
)
assert.NotContains(
t, probe.serverErrors.String(),
"superfluous response.WriteHeader",
)
}
// TestRecovererKeepsAnImplicitlyCommittedResponse is the same case
// without an explicit WriteHeader: a bare Write commits the response
// to 200 just as surely.
func TestRecovererKeepsAnImplicitlyCommittedResponse(t *testing.T) {
t.Parallel()
probe := newRecovererProbe(
t, false,
func(w http.ResponseWriter, _ *http.Request) {
_, _ = w.Write([]byte("partial"))
panic(panicMarker)
},
)
resp, err := probe.get(t)
require.NoError(t, err)
require.NoError(t, resp.Body.Close())
assert.Equal(t, http.StatusOK, resp.StatusCode)
assert.Equal(
t, true, probe.panicRecord(t)["response_committed"],
)
assert.NotContains(
t, probe.serverErrors.String(),
"superfluous response.WriteHeader",
)
}
// panicLogHandler names one of the two handlers internal/logger can
// install. The recoverer's probe selects between them with a bool
// rather than by constructing one, which is why this does not reuse
// logHandlers() the way the fills reuse escapeFills().
type panicLogHandler struct {
name string
text bool
}
func panicLogHandlers() []panicLogHandler {
return []panicLogHandler{{"json", false}, {"text", true}}
}
// TestRecovererBoundsThePanicRecord holds the record to its stated
// ceiling with a panic value the size of a request. A handler is free
// to build a panic value out of what the client sent, so the value is
// charged a client-sized budget even though the stack is not.
func TestRecovererBoundsThePanicRecord(t *testing.T) {
t.Parallel()
for _, handler := range panicLogHandlers() {
// The fills are internal/middleware's own access log fills,
// shared rather than restated: plain text, the characters
// both handlers escape to two bytes, a bare C0 control, and
// an astral non-printable the text handler spells with a
// ten-byte \U escape.
for fillName, fillRune := range escapeFills() {
t.Run(handler.name+"/"+fillName, func(t *testing.T) {
t.Parallel()
value := strings.Repeat(
fillRune, oversizedSegmentBytes,
) + tailMarker
probe := newRecovererProbe(
t, handler.text,
func(http.ResponseWriter, *http.Request) {
panic(value)
},
)
resp, err := probe.get(t)
require.NoError(t, err)
require.NoError(t, resp.Body.Close())
require.Equal(
t, http.StatusInternalServerError,
resp.StatusCode,
)
probe.wait()
for line := range strings.SplitSeq(
strings.TrimSpace(probe.logs.String()), "\n",
) {
assert.LessOrEqual(
t, len(line),
middleware.MaxPanicLogLineBytes,
"log line exceeded its stated bound",
)
assert.NotContains(
t, line, tailMarker,
"the far end of the panic value reached "+
"the log, so nothing truncated it",
)
}
})
}
}
}
// deepPanic recurses to depth and then panics, so the stack itself
// overruns its budget. It is the only way to exercise the stack cut:
// the shipped middleware chain does not come close (see
// TestPanicThroughProductionRouter in internal/server).
func deepPanic(depth int, value string) int {
if depth == 0 {
panic(value)
}
return deepPanic(depth-1, value) + 1
}
// assertEveryFieldWasCut holds each of the record's three growable
// fields to its own budget, which is what the ceiling is the sum of.
// The stack is cut at its far end, so its near end — the panic site —
// has to survive; the request id is the client's own bytes, so its
// cut is the one that bounds an attacker rather than our own call
// depth.
func assertEveryFieldWasCut(t *testing.T, record map[string]any) {
t.Helper()
stack, ok := record["stack"].(string)
require.True(t, ok)
assert.True(
t, strings.HasSuffix(stack, truncationSuffix),
"an oversized stack must be marked as cut",
)
assert.Contains(
t, stack, "deepPanic",
"the near end of the stack must survive the cut",
)
assert.NotContains(
t, stack, "net/http.(*conn).serve",
"the far end is what a cut discards",
)
id, ok := record["request_id"].(string)
require.True(t, ok)
assert.True(
t, strings.HasSuffix(id, truncationSuffix),
"an oversized request id must be marked as cut",
)
assert.LessOrEqual(
t, len(id), maxRequestIDBytes+len(truncationSuffix),
"the request id must be held to its own budget",
)
value, ok := record["panic"].(string)
require.True(t, ok)
assert.True(
t, strings.HasSuffix(value, truncationSuffix),
"an oversized panic value must be marked as cut",
)
}
// TestRecovererBoundsTheStack drives every growable field on the
// record past its budget at once — an oversized stack, an oversized
// panic value and an oversized client-supplied X-Request-Id — over
// both log handlers. It holds that line to the stated ceiling and
// reports what it measured, and it pins that a cut stack keeps its
// near end — the panic site — rather than its far one.
//
// The two fields the test picks the content of — the panic value and
// the request id — are filled with the quotation mark. Both handlers
// escape it to two bytes, which is exactly what logfield charges for
// it, so each of those fields emits every byte of its budget; no fill
// emits more, since logfield charges each rune the wider of the two
// handlers and a field can therefore never emit more than it spent.
// The stack is not a fill: recursion drives it past its budget and
// the cut lands wherever its own content puts it, which is why the
// measured widths move by a byte between runs.
func TestRecovererBoundsTheStack(t *testing.T) {
t.Parallel()
for _, handler := range panicLogHandlers() {
t.Run(handler.name, func(t *testing.T) {
t.Parallel()
value := strings.Repeat(`"`, oversizedSegmentBytes) +
tailMarker
requestID := strings.Repeat(`"`, oversizedSegmentBytes) +
tailMarker
probe := newRecovererProbe(
t, handler.text,
func(http.ResponseWriter, *http.Request) {
_ = deepPanic(512, value)
},
)
resp, err := probe.getWithRequestID(t, requestID)
require.NoError(t, err)
require.NoError(t, resp.Body.Close())
require.Equal(
t, http.StatusInternalServerError, resp.StatusCode,
)
probe.wait()
// The text handler does not emit JSON, so the field-level
// assertions run on the JSON one; the line bound below
// is asserted on both, which is the point of the sweep.
if !handler.text {
assertEveryFieldWasCut(t, probe.panicRecord(t))
}
widest := 0
for line := range strings.SplitSeq(
strings.TrimSpace(probe.logs.String()), "\n",
) {
assert.LessOrEqual(
t, len(line), middleware.MaxPanicLogLineBytes,
)
assert.NotContains(t, line, tailMarker)
widest = max(widest, len(line))
}
t.Logf(
"widest line measured: %d bytes (ceiling %d)",
widest, middleware.MaxPanicLogLineBytes,
)
})
}
}
// TestRecovererIgnoresANonPanickingHandler is the negative control:
// the middleware must be inert on the ordinary path.
func TestRecovererIgnoresANonPanickingHandler(t *testing.T) {
t.Parallel()
probe := newRecovererProbe(
t, false,
func(w http.ResponseWriter, _ *http.Request) {
w.WriteHeader(http.StatusTeapot)
},
)
resp, err := probe.get(t)
require.NoError(t, err)
require.NoError(t, resp.Body.Close())
assert.Equal(t, http.StatusTeapot, resp.StatusCode)
for _, record := range probe.records(t) {
assert.NotEqual(t, "handler panic", record["msg"])
}
}
// TestRecovererKeepsResponseControllerWorking pins the Unwrap method.
// The middleware wraps the ResponseWriter to learn whether the
// response was committed, and a wrapper without Unwrap hides
// net/http's own writer from http.ResponseController, so a handler
// that flushes or sets a deadline starts failing.
//
// The recoverer is the only middleware in the chain here. The access
// logger's own wrapper does not implement Unwrap, so a chain
// containing it fails this regardless of what the recoverer does;
// what is being pinned is that the recoverer adds no such opacity of
// its own.
func TestRecovererKeepsResponseControllerWorking(t *testing.T) {
t.Parallel()
m, _ := capturingMiddleware(t)
handler := m.Recoverer()(http.HandlerFunc(
func(w http.ResponseWriter, _ *http.Request) {
_, _ = w.Write([]byte("chunk"))
flushErr := http.NewResponseController(w).Flush()
if flushErr != nil {
http.Error(
w, "flush failed",
http.StatusInternalServerError,
)
return
}
},
))
server := httptest.NewServer(handler)
t.Cleanup(server.Close)
req, err := http.NewRequestWithContext(
t.Context(), http.MethodGet, server.URL, nil,
)
require.NoError(t, err)
resp, err := server.Client().Do(req)
require.NoError(t, err)
defer func() { _ = resp.Body.Close() }()
body, err := io.ReadAll(resp.Body)
require.NoError(t, err)
assert.Equal(t, http.StatusOK, resp.StatusCode)
assert.Equal(t, "chunk", string(body))
}

View File

@@ -4,7 +4,6 @@ import (
"log/slog"
"net/http"
"github.com/getsentry/sentry-go"
"sneak.berlin/go/webhooker/internal/config"
"sneak.berlin/go/webhooker/internal/handlers"
"sneak.berlin/go/webhooker/internal/middleware"
@@ -14,25 +13,6 @@ import (
// build requests that sit exactly at, below, and above it.
const MaxFormBodySizeForTest = maxFormBodySize
// ScrubSentryRequestForTest exposes the BeforeSend hook that
// enableSentry installs, so a test can assert on what it leaves in an
// event without standing up a Sentry client.
func ScrubSentryRequestForTest(
event *sentry.Event,
hint *sentry.EventHint,
) *sentry.Event {
return scrubSentryRequest(event, hint)
}
// SentryClientOptionsForTest exposes the exact options enableSentry
// initialises the SDK with, so a test can capture events through the
// production hook wiring rather than a hand-built equivalent.
func SentryClientOptionsForTest(
dsn, release string,
) sentry.ClientOptions {
return sentryClientOptions(dsn, release)
}
// NewRouterForTest builds the real route tree via SetupRoutes with
// the supplied middleware and handlers, bypassing the fx lifecycle
// and the HTTP listener. Tests use it so that route-group middleware
@@ -54,40 +34,3 @@ func NewRouterForTest(
return s.router
}
// ProbePattern is the route NewRouterWithProbeForTest adds to the
// production route tree.
const ProbePattern = "/probe"
// NewRouterWithProbeForTest builds the production route tree exactly
// as NewRouterForTest does and then registers probe at ProbePattern,
// so a test can drive a handler that panics through the shipped
// global middleware chain rather than a hand-assembled one. Nothing
// about the chain is rebuilt here: the probe is an extra leaf under
// the same Use() registrations every other route gets.
//
// sentryEnabled selects whether the sentryhttp handler is registered,
// which in production a configured SENTRY_DSN decides. It is a
// parameter because the relationship between that handler's Repanic
// option and the recoverer registered outside it is the thing a test
// has to be able to pin.
func NewRouterWithProbeForTest(
log *slog.Logger,
cfg *config.Config,
mw *middleware.Middleware,
h *handlers.Handlers,
sentryEnabled bool,
probe http.HandlerFunc,
) http.Handler {
s := &Server{
log: log,
mw: mw,
h: h,
params: ServerParams{Config: cfg},
sentryEnabled: sentryEnabled,
}
s.SetupRoutes()
s.router.Handle(ProbePattern, probe)
return s.router
}

View File

@@ -1,285 +0,0 @@
package server_test
import (
"bytes"
"context"
"encoding/json"
"fmt"
"net/http"
"net/http/httptest"
"os"
"os/exec"
"strings"
"testing"
"github.com/getsentry/sentry-go"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
"sneak.berlin/go/webhooker/internal/middleware"
"sneak.berlin/go/webhooker/internal/server"
)
// panicProbeMarker is the value the probe handler panics with. The
// defect this pins lost it entirely: what reached the operator was the
// recoverer's own secondary panic, naming chi's decorateFuncCallLine
// and nothing about the fault that caused it.
const panicProbeMarker = "QQPRODUCTIONPANICVALUEQQ"
// panicChildEnv, when set, tells the re-executed test binary to run
// the child half of the fd-level probe below.
const panicChildEnv = "WEBHOOKER_PANIC_PROBE_CHILD"
// panicChildResultPrefix labels the child's own one-line report of
// what the HTTP client saw, so the parent can find it among whatever
// else lands on the child's standard output.
const panicChildResultPrefix = "PANIC-PROBE-RESULT "
// stackTruncationMarker mirrors what internal/middleware appends to a
// field it cut. It is duplicated rather than exported, as the access
// log's budgets are, so that changing it has to be restated here
// deliberately.
const stackTruncationMarker = "[truncated]"
// TestPanicThroughProductionRouter drives a handler panic through the
// shipped router, over a real server, in a subprocess whose actual
// file descriptors are captured.
//
// Every part of that is load-bearing.
//
// A subprocess, because the question is what reaches fd 1 and fd 2 of
// the process an operator runs. The defect's signature was 0 bytes on
// standard error and a 2,772-byte record on standard output describing
// chi's own crash, and neither is visible to a test that swaps the
// logger for a buffer.
//
// A real server, because a panicking handler under chi's Recoverer
// dropped the connection: the client got EOF, not a status. An
// httptest.ResponseRecorder has no connection to drop and would have
// recorded the same unwritten response either way, which is why this
// defect survived the existing suite.
//
// The production router, because the placement of the recoverer among
// the other global middleware is part of the fix.
func TestPanicThroughProductionRouter(t *testing.T) {
t.Parallel()
if os.Getenv(panicChildEnv) != "" {
t.Skip("child half; run by the parent below")
}
//nolint:gosec // Re-executing this test binary, with a fixed arg.
cmd := exec.CommandContext(
t.Context(), os.Args[0],
"-test.run", "^TestPanicProbeChild$",
)
cmd.Env = append(os.Environ(), panicChildEnv+"=1")
var stdout, stderr bytes.Buffer
cmd.Stdout = &stdout
cmd.Stderr = &stderr
require.NoError(
t, cmd.Run(),
"child failed\nstdout:\n%s\nstderr:\n%s",
stdout.String(), stderr.String(),
)
assertPanicProbeOutput(t, stdout.String(), stderr.String())
}
// assertPanicProbeOutput holds the child's descriptors to what a
// working recoverer produces.
func assertPanicProbeOutput(t *testing.T, stdout, stderr string) {
t.Helper()
result := ""
var record map[string]any
for line := range strings.SplitSeq(stdout, "\n") {
if after, found := strings.CutPrefix(
line, panicChildResultPrefix,
); found {
result = after
continue
}
if !strings.HasPrefix(line, `{"time"`) {
continue
}
decoded := map[string]any{}
if json.Unmarshal([]byte(line), &decoded) != nil {
continue
}
if decoded["msg"] == "handler panic" {
require.Nil(
t, record, "one panic record expected, got two",
)
record = decoded
assert.LessOrEqual(
t, len(line), middleware.MaxPanicLogLineBytes,
"the panic record must hold its stated ceiling",
)
t.Logf(
"panic record through the shipped chain: %d bytes "+
"(ceiling %d)",
len(line), middleware.MaxPanicLogLineBytes,
)
}
}
// What the client got. Under the defect this read
// `status=0 err=... EOF`.
require.Equal(
t, "status=500 err=<nil>", result,
"the client must receive a 500, not a dropped connection",
)
// What the operator got. Under the defect there was no such
// record: standard output carried net/http reporting chi's own
// crash, at INFO, with the original panic value nowhere in it.
require.NotNil(
t, record,
"no structured panic record reached standard output",
)
assert.Equal(t, "ERROR", record["level"])
assert.Equal(t, panicProbeMarker, record["panic"])
assert.Equal(t, false, record["response_committed"])
stack, ok := record["stack"].(string)
require.True(t, ok)
assert.Contains(t, stack, "panicProbeHandler")
assert.NotContains(
t, stack, stackTruncationMarker,
"the shipped middleware chain's own stack must fit the "+
"stack budget without being cut",
)
t.Logf("stack through the shipped chain: %d bytes", len(stack))
// The secondary panic, in every form it took. net/http's report
// is the tell: it only logs a request when something escaped the
// handler chain.
assert.NotContains(t, stdout, "http: panic serving")
assert.NotContains(t, stdout, "slice bounds out of range")
assert.NotContains(t, stdout, "decorateFuncCallLine")
assert.Empty(
t, strings.TrimSpace(stderr),
"nothing may reach standard error",
)
}
// panicProbeHandler is the panicking route the child installs. It is a
// named function so the stack assertion has something to look for.
func panicProbeHandler(http.ResponseWriter, *http.Request) {
panic(panicProbeMarker)
}
// TestPanicProbeChild is the child half of the probe above. It runs
// only when re-executed with panicChildEnv set; in an ordinary run it
// returns immediately.
//
// It writes its result to standard output with a prefix rather than
// asserting, because the assertions belong to the parent, which is the
// only side that can see both descriptors.
func TestPanicProbeChild(t *testing.T) {
t.Parallel()
if os.Getenv(panicChildEnv) == "" {
return
}
env := newTestEnv(t)
router := server.NewRouterWithProbeForTest(
env.log.Get(), env.cfg, env.mw, env.hnd,
false, panicProbeHandler,
)
srv := httptest.NewServer(router)
defer srv.Close()
req, err := http.NewRequestWithContext(
context.Background(), http.MethodGet,
srv.URL+server.ProbePattern, nil,
)
require.NoError(t, err)
status := 0
resp, err := srv.Client().Do(req)
if err == nil {
status = resp.StatusCode
_ = resp.Body.Close()
}
// Written to the descriptor rather than through the testing
// package's own output, because fd 1 is exactly what the parent
// is measuring.
_, writeErr := fmt.Fprintf(
os.Stdout, "%sstatus=%d err=%v\n",
panicChildResultPrefix, status, err,
)
require.NoError(t, writeErr)
}
// TestSentryStillSeesAPanic pins the relationship the recoverer's
// placement has to preserve. sentryhttp is registered with
// Repanic: true, inside the recoverer, so an operator with SENTRY_DSN
// set keeps the report and the client still gets a 500. Registered the
// other way round, the SDK would swallow the panic and the recoverer
// would never see it.
func TestSentryStillSeesAPanic(t *testing.T) {
t.Parallel()
env := newTestEnv(t)
transport := &captureTransport{}
opts := server.SentryClientOptionsForTest(
"https://public@sentry.invalid/1", "webhooker-test",
)
opts.Transport = transport
client, err := sentry.NewClient(opts)
require.NoError(t, err)
router := server.NewRouterWithProbeForTest(
env.log.Get(), env.cfg, env.mw, env.hnd,
true, panicProbeHandler,
)
req := httptest.NewRequestWithContext(
sentry.SetHubOnContext(
context.Background(),
sentry.NewHub(client, sentry.NewScope()),
),
http.MethodGet, server.ProbePattern, nil,
)
w := httptest.NewRecorder()
router.ServeHTTP(w, req)
assert.Equal(
t, http.StatusInternalServerError, w.Code,
"the recoverer must still answer what sentryhttp re-raised",
)
events := transport.events
require.Len(t, events, 1, "Sentry must still see the panic")
assert.Equal(t, sentry.LevelFatal, events[0].Level)
// The SDK renders a string panic value as the event message
// rather than an exception, so the whole payload is checked for
// the value rather than one field of it.
assert.Contains(
t, marshalEvent(t, events[0]), panicProbeMarker,
)
}

View File

@@ -14,28 +14,6 @@ import (
// maxFormBodySize is the maximum allowed request body size (in
// bytes) for form POST endpoints. 1 MB is generous for any form
// submission while preventing abuse from oversized payloads.
//
// Every route group below installs MaxBodySize(maxFormBodySize) as
// its FIRST middleware, ahead of both CSRF and RequireAuth. Both
// orderings are deliberate.
//
// Ahead of CSRF because gorilla/csrf parses the form. The cap has to
// be installed before anything reads the body, or the parse runs
// under net/http's 10 MB default instead of this one.
//
// Ahead of RequireAuth because an oversize body should be refused
// before the request buys a cookie decrypt, a session load and the
// database read behind it. Rejecting first is the cheaper failure,
// and it is the ordering that keeps an unauthenticated flood from
// choosing how much session work the process does.
//
// What that ordering costs: the 413 branch is reachable
// unauthenticated, at a URL of the client's choosing and of the
// client's chosen length. So is the CSRF rejection, which sits in
// front of RequireAuth for the same reason. Both log that path, so
// both cap it — see the log calls in Middleware.MaxBodySize and
// Middleware.CSRF, which spend the same per-field budget as the
// access log.
const maxFormBodySize int64 = 1 * 1024 * 1024 // 1 MB
// requestTimeout is the maximum time allowed for a single HTTP
@@ -51,6 +29,7 @@ func (s *Server) SetupRoutes() {
}
func (s *Server) setupGlobalMiddleware() {
s.router.Use(middleware.Recoverer)
s.router.Use(middleware.RequestID)
s.router.Use(s.mw.SecurityHeaders())
s.router.Use(s.mw.Logging())
@@ -63,21 +42,8 @@ func (s *Server) setupGlobalMiddleware() {
s.router.Use(s.mw.CORS())
s.router.Use(middleware.Timeout(requestTimeout))
// Panic recovery, deliberately here rather than first. It has to
// run inside every middleware that observes the response, so the
// 500 it writes is the status the access log records and the
// metrics count, and outside the sentryhttp handler below, whose
// Repanic option needs something further out to catch what it
// re-raises. chi's own middleware.Recoverer held the first slot
// until it was measured: on a current Go release it crashes
// inside its stack pretty-printer instead of recovering, so the
// connection dropped and the original panic was never reported.
// See https://git.eeqj.de/sneak/webhooker/issues/187.
s.router.Use(s.mw.Recoverer())
// Sentry error reporting (if SENTRY_DSN is set). Repanic is
// true so panics still bubble up to the Recoverer middleware
// registered immediately above.
// true so panics still bubble up to the Recoverer middleware.
if s.sentryEnabled {
sentryHandler := sentryhttp.New(sentryhttp.Options{
Repanic: true,
@@ -124,20 +90,17 @@ func (s *Server) setupRoutes() {
func (s *Server) setupPageRoutes() {
s.router.Route("/pages", func(r chi.Router) {
// MaxBodySize precedes CSRF and RequireAuth deliberately;
// see maxFormBodySize for why, and for what it costs.
// MaxBodySize must precede CSRF: gorilla/csrf parses the
// form, so the cap has to be installed before it runs.
r.Use(s.mw.MaxBodySize(maxFormBodySize))
r.Use(s.mw.CSRF())
r.Use(s.mw.NoCache())
// The login POST carries no pre-emptive rate limiter. Behind
// the reverse proxy production requires, with TRUSTED_PROXIES
// unset, every client shares one bucket, so a limiter spent
// on arrival lets any stranger deny the operator the only
// administrative path. The handler verifies credentials first
// and charges only failures; see Handlers.authenticateUser.
r.Get("/login", s.h.HandleLoginPage())
r.Post("/login", s.h.HandleLoginSubmit())
r.Group(func(r chi.Router) {
r.Use(s.mw.LoginRateLimit())
r.Get("/login", s.h.HandleLoginPage())
r.Post("/login", s.h.HandleLoginSubmit())
})
r.Post("/logout", s.h.HandleLogout())
})
@@ -145,8 +108,8 @@ func (s *Server) setupPageRoutes() {
func (s *Server) setupUserRoutes() {
s.router.Route("/user/{username}", func(r chi.Router) {
// MaxBodySize precedes CSRF and RequireAuth deliberately;
// see maxFormBodySize for why, and for what it costs.
// MaxBodySize must precede CSRF: gorilla/csrf parses the
// form, so the cap has to be installed before it runs.
r.Use(s.mw.MaxBodySize(maxFormBodySize))
r.Use(s.mw.CSRF())
r.Use(s.mw.NoCache())
@@ -160,8 +123,8 @@ func (s *Server) setupUserRoutes() {
func (s *Server) setupSourceRoutes() {
s.router.Route("/sources", func(r chi.Router) {
// MaxBodySize precedes CSRF and RequireAuth deliberately;
// see maxFormBodySize for why, and for what it costs.
// MaxBodySize must precede CSRF: gorilla/csrf parses the
// form, so the cap has to be installed before it runs.
r.Use(s.mw.MaxBodySize(maxFormBodySize))
r.Use(s.mw.CSRF())
r.Use(s.mw.NoCache())
@@ -172,8 +135,8 @@ func (s *Server) setupSourceRoutes() {
})
s.router.Route("/source/{sourceID}", func(r chi.Router) {
// MaxBodySize precedes CSRF and RequireAuth deliberately;
// see maxFormBodySize for why, and for what it costs.
// MaxBodySize must precede CSRF: gorilla/csrf parses the
// form, so the cap has to be installed before it runs.
r.Use(s.mw.MaxBodySize(maxFormBodySize))
r.Use(s.mw.CSRF())
r.Use(s.mw.NoCache())
@@ -183,15 +146,6 @@ func (s *Server) setupSourceRoutes() {
r.Post("/edit", s.h.HandleSourceEditSubmit())
r.Post("/delete", s.h.HandleSourceDelete())
r.Get("/logs", s.h.HandleSourceLogs())
// The log page renders each body only up to its cap, so
// this is the only route that serves a whole one. It
// belongs to this group for its RequireAuth and
// NoCache; see HandleEventBodyDownload for the headers
// that keep the bytes it returns inert.
r.Get(
"/logs/{eventID}/body",
s.h.HandleEventBodyDownload(),
)
r.Post(
"/entrypoints",
s.h.HandleEntrypointCreate(),

View File

@@ -7,7 +7,6 @@ import (
"net/http/httptest"
"net/url"
"regexp"
"strconv"
"strings"
"testing"
@@ -15,7 +14,6 @@ import (
"github.com/stretchr/testify/require"
"go.uber.org/fx"
"go.uber.org/fx/fxtest"
"gorm.io/gorm/clause"
"sneak.berlin/go/webhooker/internal/config"
"sneak.berlin/go/webhooker/internal/database"
"sneak.berlin/go/webhooker/internal/delivery"
@@ -51,16 +49,6 @@ type testEnv struct {
router http.Handler
sess *session.Session
db *database.Database
dbMgr *database.WebhookDBManager
// The collaborators the router was built from, kept so a test
// that needs a second router over the same graph — one carrying
// a panicking probe route, or one with Sentry registered — can
// build it without wiring the graph again.
log *logger.Logger
cfg *config.Config
mw *middleware.Middleware
hnd *handlers.Handlers
}
// newTestEnv wires the dependency graph with fx and builds the
@@ -70,13 +58,12 @@ func newTestEnv(t *testing.T) *testEnv {
t.Helper()
var (
log *logger.Logger
cfg *config.Config
mw *middleware.Middleware
hnd *handlers.Handlers
sess *session.Session
db *database.Database
dbMgr *database.WebhookDBManager
log *logger.Logger
cfg *config.Config
mw *middleware.Middleware
hnd *handlers.Handlers
sess *session.Session
db *database.Database
)
app := fxtest.New(
@@ -99,7 +86,7 @@ func newTestEnv(t *testing.T) *testEnv {
middleware.New,
handlers.New,
),
fx.Populate(&log, &cfg, &mw, &hnd, &sess, &db, &dbMgr),
fx.Populate(&log, &cfg, &mw, &hnd, &sess, &db),
)
app.RequireStart()
t.Cleanup(app.RequireStop)
@@ -108,11 +95,6 @@ func newTestEnv(t *testing.T) *testEnv {
router: server.NewRouterForTest(log.Get(), cfg, mw, hnd),
sess: sess,
db: db,
dbMgr: dbMgr,
log: log,
cfg: cfg,
mw: mw,
hnd: hnd,
}
}
@@ -251,49 +233,6 @@ func (e *testEnv) seedUser(
return user.ID, hash
}
// seedWebhook creates a webhook owned by the given user.
func (e *testEnv) seedWebhook(
t *testing.T,
userID string,
) *database.Webhook {
t.Helper()
wh := &database.Webhook{UserID: userID, Name: "routed"}
require.NoError(
t,
e.db.DB().Omit(clause.Associations).Create(wh).Error,
)
return wh
}
// seedEvent records one event with the given body in a webhook's
// own database.
func (e *testEnv) seedEvent(
t *testing.T,
webhookID, body string,
) *database.Event {
t.Helper()
webhookDB, err := e.dbMgr.GetDB(webhookID)
require.NoError(t, err)
event := &database.Event{
WebhookID: webhookID,
Method: http.MethodPost,
Body: body,
ContentType: "application/octet-stream",
}
require.NoError(
t,
webhookDB.Omit(clause.Associations).Create(event).Error,
)
return event
}
// storedHash reads the current password hash for a username.
func (e *testEnv) storedHash(t *testing.T, username string) string {
t.Helper()
@@ -433,78 +372,6 @@ func TestPagesLogin_UnderLimit_ValidToken_ReachesHandler(
)
}
// TestPagesLogin_CorrectPasswordSurvivesASpentBudget pins the
// routing half of the fix, which every other login test misses by
// driving the handler directly: no pre-emptive limiter sits in front
// of POST /pages/login on the real route tree.
//
// A limiter registered there would answer the last request 429
// however correct its password is, because the wrong passwords
// before it have already spent the bucket — which is the lockout
// this endpoint exists to not have. CSRF and the body cap still run,
// since every request here carries a harvested token.
func TestPagesLogin_CorrectPasswordSurvivesASpentBudget(
t *testing.T,
) {
t.Parallel()
const (
username = "operator"
password = "correct-horse-battery-staple"
)
env := newTestEnv(t)
env.seedUser(t, username, password)
submit := func(t *testing.T, pw string) *httptest.ResponseRecorder {
t.Helper()
token, cookies := env.csrfFrom(t, "/pages/login", nil)
form := url.Values{}
form.Set("csrf_token", token)
form.Set("username", username)
form.Set("password", pw)
return env.post("/pages/login", form, cookies)
}
// Spend the failure budget against this username. The exact
// limit belongs to the middleware; this waits for the throttle
// to appear rather than restating it, under a ceiling well
// above it so a broken limiter fails the test instead of
// looping.
const maxAttempts = 20
spent := false
for range maxAttempts {
code := submit(t, "wrong").Code
if code == http.StatusTooManyRequests {
spent = true
break
}
require.Equal(
t, http.StatusUnauthorized, code,
"a wrong password must be rejected, not accepted",
)
}
require.True(
t, spent,
"repeated wrong passwords must eventually be throttled",
)
assert.Equal(
t, http.StatusSeeOther, submit(t, password).Code,
"a correct password must be accepted on the routed "+
"endpoint even with the failure budget spent: the "+
"operator has no second administrative path",
)
}
// --- /user/{username} group ---
// TestPasswordChange_OversizeBody_RejectedAndPasswordUnchanged
@@ -565,95 +432,3 @@ func TestPasswordChange_UnderLimit_Succeeds(t *testing.T) {
"an under-limit password change should still apply",
)
}
// --- /source/{sourceID} group ---
// TestSourceLogs_TruncationLinkDownloadsTheBody walks the whole
// feature the way a user does: render the event log page through
// the production router, take the download URL out of the markup
// the template emitted, and fetch that URL through the router
// again. Nothing here is hand-written, so a typo in either the
// route pattern or the template href fails this test — the
// handler-level tests cannot catch that, because they forge
// their own route context and assert a URL string they wrote
// themselves.
func TestSourceLogs_TruncationLinkDownloadsTheBody(t *testing.T) {
t.Parallel()
env := newTestEnv(t)
userID, _ := env.seedUser(t, "loguser", "somepassword")
cookies := env.authCookies(t, userID, "loguser")
// Comfortably over the event log page's render cap, so the
// page truncates the body and renders the download link at
// all. The exact cap is the handlers package's business and
// is pinned by its own tests; this only needs to exceed it.
stored := strings.Repeat("Z", 64*1024)
wh := env.seedWebhook(t, userID)
env.seedEvent(t, wh.ID, stored)
page := env.get("/source/"+wh.ID+"/logs", cookies)
require.Equal(t, http.StatusOK, page.Code)
link := regexp.MustCompile(
`href="(/source/[^"]+/body)"`,
).FindStringSubmatch(page.Body.String())
require.Len(
t, link, 2,
"truncated body should render a download link",
)
w := env.get(html.UnescapeString(link[1]), cookies)
require.Equal(
t, http.StatusOK, w.Code,
"the link the page emits must be a live route",
)
assert.Equal(t, stored, w.Body.String())
assert.Equal(
t, strconv.Itoa(len(stored)),
w.Header().Get("Content-Length"),
)
assert.Equal(
t, "application/octet-stream",
w.Header().Get("Content-Type"),
)
assert.Contains(
t, w.Header().Get("Content-Disposition"), "attachment",
)
assert.Equal(
t, "nosniff", w.Header().Get("X-Content-Type-Options"),
)
}
// TestSourceLogsBody_OtherUser404s pins that the download route
// as registered is behind the auth the group provides and the
// ownership check the handler applies: another logged-in user
// asking the real router for the same URL gets a 404, and an
// unauthenticated request never reaches the handler at all.
func TestSourceLogsBody_OtherUser404s(t *testing.T) {
t.Parallel()
env := newTestEnv(t)
ownerID, _ := env.seedUser(t, "owner", "somepassword")
wh := env.seedWebhook(t, ownerID)
const payload = "OWNERS-PAYLOAD-77c1"
evt := env.seedEvent(t, wh.ID, payload)
path := "/source/" + wh.ID + "/logs/" + evt.ID + "/body"
intruderID, _ := env.seedUser(t, "intruder", "somepassword")
intruder := env.authCookies(t, intruderID, "intruder")
w := env.get(path, intruder)
assert.Equal(t, http.StatusNotFound, w.Code)
assert.NotContains(t, w.Body.String(), payload)
anon := env.get(path, nil)
assert.Equal(t, http.StatusSeeOther, anon.Code)
assert.Equal(t, "/pages/login", anon.Header().Get("Location"))
}

View File

@@ -1,236 +0,0 @@
package server
import (
"net/http"
"net/url"
"strings"
"github.com/getsentry/sentry-go"
"github.com/go-chi/chi"
)
// sentryRedacted stands in for a withheld field on every event shipped
// to Sentry. It is a marker rather than an empty string so a reader
// can tell a suppressed value from an absent one.
const sentryRedacted = "(redacted)"
// sentryRedactedPath is what stands in for the request path when the
// route pattern is not reachable. It is deliberately not the concrete
// path: on the receiver route that path carries the entrypoint UUID,
// which is a write capability rather than an identifier.
const sentryRedactedPath = "/" + sentryRedacted
// sentryClientOptions builds the options the SDK is initialised with.
// It is its own function so a test can stand up a client wired exactly
// as production is, with only the transport swapped.
func sentryClientOptions(dsn, release string) sentry.ClientOptions {
return sentry.ClientOptions{
Dsn: dsn,
Release: release,
// Both hooks, because the SDK runs one for error events
// and the other for transactions.
BeforeSend: scrubSentryRequest,
BeforeSendTransaction: scrubSentryRequest,
}
}
// scrubSentryRequest strips client-supplied content from an event's
// request context before it leaves the process.
//
// sentryhttp attaches the whole *http.Request to the scope
// (sentryhttp.go:113), and Scope.ApplyToEvent fills the event's
// Request from it inside prepareEvent, which runs before this hook.
// Two of the fields it fills are copied with no SendDefaultPII guard:
//
// - QueryString, verbatim from r.URL.RawQuery.
// - Data, the first 10 KiB of the request body, teed off r.Body by
// SetRequest and filled precisely because the handlers call
// ParseForm.
//
// Since every form field in this service is read with PostFormValue,
// the body is the only place a credential is submitted: a target's
// destination URL, whose path segments are the bearer token, plus the
// login password and both password-change fields. None of that may
// reach a third-party service.
//
// URL is the third such field. NewRequest builds it as
// scheme://host/path (interfaces.go:183), and on the receiver route
// that path is /webhook/<uuid> in full — a write capability, not an
// identifier. It is rebuilt here from the chi route pattern, on every
// route, keeping the scheme and the host.
//
// This hook is a floor, not a default: the fields it clears stay
// cleared even if SendDefaultPII is ever turned on.
func scrubSentryRequest(
event *sentry.Event,
hint *sentry.EventHint,
) *sentry.Event {
if event == nil {
return event
}
pattern := sentryRoutePattern(hint)
// Only transaction events carry a Transaction name, and the SDK
// builds it from the concrete path too (sentryhttp.go:105 via
// tracing.go:553). Rewritten on the same terms.
if event.Transaction != "" {
event.Transaction = sentryTransactionName(
event.Transaction, pattern,
)
}
if event.Request == nil {
return event
}
req := event.Request
if req.URL != "" {
req.URL = sentryRouteURL(req.URL, pattern)
}
if req.QueryString != "" {
req.QueryString = sentryRedacted
}
if req.Data != "" {
req.Data = sentryRedacted
}
req.Cookies = ""
req.Env = nil
req.Headers = keptSentryHeaders(req.Headers)
return event
}
// sentryRoutePattern returns the chi route pattern for the request the
// hint carries, or "" when it is not reachable.
//
// The request is reachable on the error dispatch only. sentryhttp's
// recover path calls RecoverWithContext with the request on the
// context under sentry.RequestContextKey (sentryhttp.go:124-125), and
// the client copies that context onto the hint (client.go:484-485)
// before handing it to BeforeSend (client.go:631). chi's routing
// context is a pointer placed on the request context before the
// middleware chain runs (chi mux.go:84) and filled in as the mux
// routes, so by the time a handler panics it names the matched route.
//
// The transaction dispatch has no such request: Span.doFinish calls
// hub.CaptureEvent (tracing.go:356), which passes a nil hint that the
// client replaces with an empty one (client.go:620-622). The pattern
// is therefore always "" there, and the callers fall back.
func sentryRoutePattern(hint *sentry.EventHint) string {
if hint == nil || hint.Context == nil {
return ""
}
req, ok := hint.Context.Value(
sentry.RequestContextKey,
).(*http.Request)
if !ok || req == nil {
return ""
}
rctx := chi.RouteContext(req.Context())
if rctx == nil {
return ""
}
// Empty when no route matched, which is the fallback case too.
return rctx.RoutePattern()
}
// sentryRouteURL rebuilds an event's request URL with the route
// pattern in place of the concrete path.
//
// The scheme is load-bearing and is kept: the SDK derives it from
// r.TLS != nil || r.Header.Get("X-Forwarded-Proto") == "https"
// (interfaces.go:180), byte for byte the predicate
// internal/middleware/csrf.go uses, so it is the CSRF TLS decision and
// the reason dropping X-Forwarded-Proto from the header allowlist
// costs nothing. The host is parsed.Host of the SDK's
// scheme://r.Host/path, so it is whatever the client's Host header
// carried: this service validates no hostname. It is kept because that
// same header is on the allowlist, so scrubbing it here would withhold
// nothing that is not sent anyway.
//
// Everything else in the URL is discarded rather than edited, so a
// future SDK that starts appending a query string cannot widen this.
func sentryRouteURL(rawURL, pattern string) string {
parsed, err := url.Parse(rawURL)
if err != nil || parsed.Scheme == "" {
// Not a shape this can safely take apart.
return sentryRedacted
}
if pattern == "" {
pattern = sentryRedactedPath
}
return parsed.Scheme + "://" + parsed.Host + pattern
}
// sentryTransactionName rebuilds the SDK's "METHOD /path" transaction
// name with the route pattern in place of the concrete path. Method is
// kept for the same reason Request.Method is: net/http admits only a
// bounded token there. A name in any other shape is withheld whole,
// since nothing can be said about which part of it is a path.
func sentryTransactionName(name, pattern string) string {
method, _, found := strings.Cut(name, " ")
if !found {
return sentryRedacted
}
if pattern == "" {
pattern = sentryRedactedPath
}
return method + " " + pattern
}
// keptSentryHeaders returns the subset of headers an event may carry
// off-host. Dropping by allowlist rather than by blocklist is what
// makes an unrecognised header safe: the SDK's own filter removes four
// names and passes everything else, so X-Csrf-Token — which
// gorilla/csrf accepts in place of the form field — and the shared
// secrets senders put on the receiver route (X-Gitlab-Token and the
// per-provider signature headers) would otherwise ship verbatim.
func keptSentryHeaders(headers map[string]string) map[string]string {
if len(headers) == 0 {
return headers
}
kept := make(map[string]string, len(headers))
for name, value := range headers {
if sentryKeepsHeader(name) {
kept[name] = value
}
}
return kept
}
// sentryKeepsHeader reports whether a request header is routing or
// content metadata rather than client-chosen payload. Referer is kept
// on the reasoning that it is browser-set, that this service emits
// only ?page= in its own links, and that Referrer-Policy is set to
// strict-origin-when-cross-origin. X-Request-Id ties the event to the
// local access log line, which holds the rest of the detail.
func sentryKeepsHeader(name string) bool {
switch http.CanonicalHeaderKey(name) {
case "Accept",
"Content-Length",
"Content-Type",
"Host",
"Origin",
"Referer",
"User-Agent",
"X-Request-Id":
return true
default:
return false
}
}

View File

@@ -1,486 +0,0 @@
package server_test
import (
"context"
"encoding/json"
"net/http"
"net/http/httptest"
"net/url"
"strings"
"sync"
"testing"
"time"
"github.com/getsentry/sentry-go"
sentryhttp "github.com/getsentry/sentry-go/http"
"github.com/go-chi/chi"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
"sneak.berlin/go/webhooker/internal/server"
)
// The four markers below are the credentials a captured event could
// carry off-host, one per field of sentry.Request that the SDK fills
// from the request without a SendDefaultPII guard.
const (
// sentryBodyMarker is submitted as a form value. Since every
// handler reads its fields with PostFormValue, the body is the
// only place a password or a target URL is ever supplied.
sentryBodyMarker = "QQSENTRYBODYMARKERQQ"
// sentryQueryMarker rides the request line.
sentryQueryMarker = "T00000000/B00000000/QQSENTRYQUERYMARKERQQ"
// sentryHeaderMarker rides X-Csrf-Token, which gorilla/csrf
// accepts in place of the form field.
sentryHeaderMarker = "QQSENTRYHEADERMARKERQQ"
// sentryReceiverUUID is the entrypoint identifier in the path of
// a receiver request. It is a write capability: anyone holding
// it can POST events this service accepts and its targets then
// deliver, so it may not reach a third-party tracker.
sentryReceiverUUID = "6d1f9c2a-3b7e-4f58-9a0d-c0ffeebadc0d"
)
// sentryKeptUserAgent is a non-secret header value planted so the
// assertions below cannot pass by the event carrying no headers at
// all.
const sentryKeptUserAgent = "webhooker-test-agent"
// captureTransport records events instead of shipping them, so a test
// sees exactly the payload the SDK would have put on the wire.
type captureTransport struct {
mu sync.Mutex
events []*sentry.Event
}
func (c *captureTransport) Configure(sentry.ClientOptions) {}
func (c *captureTransport) Flush(time.Duration) bool { return true }
func (c *captureTransport) SendEvent(event *sentry.Event) {
c.mu.Lock()
defer c.mu.Unlock()
c.events = append(c.events, event)
}
// sentryCase drives one request through the real sentryhttp middleware
// inside a real chi router and returns the events the SDK produced.
//
// Routing through a chi mux is load-bearing, not decoration. chi puts
// its routing context on the request context before the middleware
// chain runs and fills it in as it matches, so a hand-built request
// carries no route pattern at all and could not distinguish the hook
// working from the hook falling back.
//
// This is also the only construction path on which Request.Data
// appears: sentryhttp calls Scope.SetRequest, which tees r.Body into a
// 10 KiB buffer, ParseForm drains the tee, and Scope.ApplyToEvent
// copies the buffer into the event inside prepareEvent — before
// BeforeSend runs. A hand-built sentry.NewRequest never reads the body
// and so cannot regress-test any of it.
type sentryCase struct {
// scrub selects whether the production BeforeSend hooks are
// installed, so the same path shows both what the SDK collects
// and what survives.
scrub bool
// tracing enables the transaction dispatch, which the service
// leaves off. With it on, a served request produces a
// transaction event through BeforeSendTransaction.
tracing bool
// panics selects the error dispatch, via BeforeSend.
panics bool
// request builds the request to serve, given the client whose
// hub it must carry.
request func(*sentry.Client) *http.Request
}
func (c sentryCase) capture(t *testing.T) []*sentry.Event {
t.Helper()
transport := &captureTransport{}
opts := server.SentryClientOptionsForTest(
"https://public@sentry.invalid/1", "webhooker-test",
)
opts.Transport = transport
if !c.scrub {
opts.BeforeSend = nil
opts.BeforeSendTransaction = nil
}
if c.tracing {
opts.EnableTracing = true
opts.TracesSampleRate = 1.0
}
client, err := sentry.NewClient(opts)
require.NoError(t, err)
c.router().ServeHTTP(httptest.NewRecorder(), c.request(client))
return transport.events
}
// router mirrors the one ordering these tests depend on, over the two
// route patterns they need: a recovering middleware outside, then the
// sentryhttp handler registered with Use and Repanic set, exactly as
// routes.go orders the two. The bare recover stands in for
// Middleware.Recoverer, which holds that outer slot in production; it
// is here only to keep panic stacks out of the test output. That the
// production one really does catch what sentryhttp re-raises is
// pinned separately, by TestSentryStillSeesAPanic.
func (c sentryCase) router() http.Handler {
handler := func(_ http.ResponseWriter, r *http.Request) {
// This call is what drains the body tee and fills the
// buffer. Its success is asserted by the unscrubbed case
// below, which sees the body in the event.
_ = r.ParseForm()
if c.panics {
panic("boom")
}
}
router := chi.NewRouter()
router.Use(recoveringMiddleware)
router.Use(
sentryhttp.New(sentryhttp.Options{Repanic: true}).Handle,
)
router.HandleFunc("/pages/login", handler)
router.HandleFunc("/webhook/{uuid}", handler)
return router
}
func recoveringMiddleware(next http.Handler) http.Handler {
return http.HandlerFunc(
func(w http.ResponseWriter, r *http.Request) {
defer func() { _ = recover() }()
next.ServeHTTP(w, r)
},
)
}
// sentryLoginRequest builds the password POST most cases drive, with a
// credential planted in the body, the query and a header.
func sentryLoginRequest(client *sentry.Client) *http.Request {
form := url.Values{}
form.Set("username", "admin")
form.Set("password", sentryBodyMarker)
req := sentryRequest(
client,
"/pages/login?url=https://hooks.slack.com/services/"+
sentryQueryMarker,
form.Encode(),
)
req.Header.Set("X-Csrf-Token", sentryHeaderMarker)
return req
}
// sentryReceiverRequest builds a POST to the receiver route, whose
// concrete path carries the entrypoint capability.
func sentryReceiverRequest(client *sentry.Client) *http.Request {
return sentryRequest(
client, "/webhook/"+sentryReceiverUUID, "payload=hello",
)
}
func sentryRequest(
client *sentry.Client,
target, body string,
) *http.Request {
req := httptest.NewRequestWithContext(
sentry.SetHubOnContext(
context.Background(),
sentry.NewHub(client, sentry.NewScope()),
),
http.MethodPost,
target,
strings.NewReader(body),
)
req.Header.Set(
"Content-Type", "application/x-www-form-urlencoded",
)
req.Header.Set("User-Agent", sentryKeptUserAgent)
return req
}
// marshalEvent encodes an event the way the transport does.
func marshalEvent(t *testing.T, event *sentry.Event) string {
t.Helper()
encoded, err := json.Marshal(event)
require.NoError(t, err)
return string(encoded)
}
// onlyEvent asserts a single event was captured and returns it.
func onlyEvent(t *testing.T, events []*sentry.Event) *sentry.Event {
t.Helper()
require.Len(t, events, 1)
require.NotNil(t, events[0].Request)
return events[0]
}
// TestSentryScrub_SDKCollectsTheRequestUnscrubbed pins the premise the
// hook exists for. Without it the SDK ships the whole POST body, the
// raw query, the CSRF header and the concrete request path, none of
// which SendDefaultPII=false suppresses.
func TestSentryScrub_SDKCollectsTheRequestUnscrubbed(t *testing.T) {
t.Parallel()
event := onlyEvent(t, sentryCase{
panics: true,
request: sentryLoginRequest,
}.capture(t))
assert.Contains(
t, event.Request.Data, sentryBodyMarker,
"the SDK is expected to collect the POST body; if it no "+
"longer does, the scrub hook's premise changed",
)
assert.Contains(t, event.Request.QueryString, sentryQueryMarker)
assert.Contains(
t, marshalEvent(t, event), sentryHeaderMarker,
)
receiver := onlyEvent(t, sentryCase{
panics: true,
request: sentryReceiverRequest,
}.capture(t))
assert.Contains(
t, receiver.Request.URL, sentryReceiverUUID,
"the SDK is expected to build Request.URL from the "+
"concrete path; if it no longer does, the route "+
"pattern rewrite's premise changed",
)
}
// TestSentryScrub_RedactsTheCapturedRequest is the regression test: no
// byte of any planted credential may survive into the marshalled event
// that leaves the process.
func TestSentryScrub_RedactsTheCapturedRequest(t *testing.T) {
t.Parallel()
event := onlyEvent(t, sentryCase{
scrub: true,
panics: true,
request: sentryLoginRequest,
}.capture(t))
encoded := marshalEvent(t, event)
assert.NotContains(t, encoded, sentryBodyMarker)
assert.NotContains(t, encoded, sentryQueryMarker)
assert.NotContains(t, encoded, sentryHeaderMarker)
assert.NotContains(t, encoded, "hooks.slack.com")
assert.Equal(t, "(redacted)", event.Request.Data)
assert.Equal(t, "(redacted)", event.Request.QueryString)
assert.Empty(t, event.Request.Cookies)
assert.Empty(t, event.Request.Env)
}
// TestSentryScrub_ReplacesTheCapabilityPathWithTheRoutePattern is the
// regression test for the receiver URL: the entrypoint UUID is a write
// capability and may not reach the tracker, while the route it names
// must still be readable there.
func TestSentryScrub_ReplacesTheCapabilityPathWithTheRoutePattern(
t *testing.T,
) {
t.Parallel()
event := onlyEvent(t, sentryCase{
scrub: true,
panics: true,
request: sentryReceiverRequest,
}.capture(t))
assert.NotContains(
t, marshalEvent(t, event), sentryReceiverUUID,
)
assert.Equal(
t, "http://example.com/webhook/{uuid}", event.Request.URL,
)
}
// TestSentryScrub_KeepsTheRoutingContext checks the hook does not cost
// the debugging signal: the route, its scheme and host, the method and
// the metadata headers still identify what failed. On a static route
// the pattern is the path, so the URL is unchanged there.
func TestSentryScrub_KeepsTheRoutingContext(t *testing.T) {
t.Parallel()
event := onlyEvent(t, sentryCase{
scrub: true,
panics: true,
request: sentryLoginRequest,
}.capture(t))
assert.Equal(
t, "http://example.com/pages/login", event.Request.URL,
)
assert.Equal(t, http.MethodPost, event.Request.Method)
assert.Equal(
t,
sentryKeptUserAgent,
event.Request.Headers["User-Agent"],
)
assert.Equal(
t,
"application/x-www-form-urlencoded",
event.Request.Headers["Content-Type"],
)
}
// TestSentryScrub_RedactsTheTransactionDispatch covers the other hook.
// Span.doFinish captures with a nil hint, so BeforeSendTransaction
// gets one with no context and no request: the route pattern is out of
// reach and both the URL and the SDK-built transaction name have to
// fall back. Tracing is off in this service, so no transaction event
// is produced today; the hook is a floor against that changing.
func TestSentryScrub_RedactsTheTransactionDispatch(t *testing.T) {
t.Parallel()
events := sentryCase{
scrub: true,
tracing: true,
request: sentryReceiverRequest,
}.capture(t)
event := onlyEvent(t, events)
require.Equal(t, "transaction", event.Type)
assert.NotContains(
t, marshalEvent(t, event), sentryReceiverUUID,
)
assert.Equal(
t, "http://example.com/(redacted)", event.Request.URL,
)
assert.Equal(t, "POST /(redacted)", event.Transaction)
}
// TestSentryScrub_TransactionDispatchIsUnscrubbedWithoutTheHook pins
// that dispatch's premise the same way, since it is the one the
// service does not exercise today.
func TestSentryScrub_TransactionDispatchIsUnscrubbedWithoutTheHook(
t *testing.T,
) {
t.Parallel()
event := onlyEvent(t, sentryCase{
tracing: true,
request: sentryReceiverRequest,
}.capture(t))
require.Equal(t, "transaction", event.Type)
assert.Contains(t, event.Request.URL, sentryReceiverUUID)
assert.Contains(t, event.Transaction, sentryReceiverUUID)
}
// TestSentryScrub_FallsBackWithoutARoutePattern covers every way the
// pattern can be missing. None of them may fall back to the concrete
// path, and all of them keep the scheme, which is the CSRF TLS
// decision.
func TestSentryScrub_FallsBackWithoutARoutePattern(t *testing.T) {
t.Parallel()
concrete := "https://example.com/webhook/" + sentryReceiverUUID
// A request with no chi routing context on it at all, which is
// what an event captured outside the router would carry.
unrouted := httptest.NewRequestWithContext(
context.Background(), http.MethodPost, concrete, nil,
)
for name, hint := range map[string]*sentry.EventHint{
"no hint": nil,
"no context": {},
"no request": {Context: context.Background()},
"unrouted request": {
Context: context.WithValue(
context.Background(),
sentry.RequestContextKey,
unrouted,
),
},
} {
t.Run(name, func(t *testing.T) {
t.Parallel()
event := sentry.NewEvent()
event.Request = &sentry.Request{URL: concrete}
event.Transaction = "POST /webhook/" +
sentryReceiverUUID
scrubbed := server.ScrubSentryRequestForTest(
event, hint,
)
require.NotNil(t, scrubbed)
assert.Equal(
t,
"https://example.com/(redacted)",
scrubbed.Request.URL,
)
assert.Equal(
t, "POST /(redacted)", scrubbed.Transaction,
)
assert.NotContains(
t,
marshalEvent(t, scrubbed),
sentryReceiverUUID,
)
})
}
}
// TestSentryScrub_WithholdsUnparseableValues covers the shapes the
// rewrite cannot take apart. Withholding them whole is the safe
// answer, since nothing can be said about which part is a path.
func TestSentryScrub_WithholdsUnparseableValues(t *testing.T) {
t.Parallel()
event := sentry.NewEvent()
event.Request = &sentry.Request{
URL: "/webhook/" + sentryReceiverUUID,
}
event.Transaction = "/webhook/" + sentryReceiverUUID
scrubbed := server.ScrubSentryRequestForTest(event, nil)
require.NotNil(t, scrubbed)
assert.Equal(t, "(redacted)", scrubbed.Request.URL)
assert.Equal(t, "(redacted)", scrubbed.Transaction)
}
// TestSentryScrub_ToleratesEventsWithoutARequest covers the events the
// hook sees outside an HTTP handler, where no request is attached.
func TestSentryScrub_ToleratesEventsWithoutARequest(t *testing.T) {
t.Parallel()
scrubbed := server.ScrubSentryRequestForTest(
sentry.NewEvent(), nil,
)
require.NotNil(t, scrubbed)
assert.Nil(t, scrubbed.Request)
assert.Empty(t, scrubbed.Transaction)
assert.Nil(t, server.ScrubSentryRequestForTest(nil, nil))
}

View File

@@ -141,14 +141,14 @@ func (s *Server) enableSentry() {
return
}
err := sentry.Init(sentryClientOptions(
s.params.Config.SentryDSN,
fmt.Sprintf(
err := sentry.Init(sentry.ClientOptions{
Dsn: s.params.Config.SentryDSN,
Release: fmt.Sprintf(
"%s-%s",
s.params.Globals.Appname,
s.params.Globals.Version,
),
))
})
if err != nil {
s.log.Error("sentry init failure", "error", err)
// Don't use fatal since we still want the service to run

View File

@@ -3,14 +3,20 @@
# this repo. Idempotent: every install is guarded by a check so already
# installed tools are skipped. Base tooling comes from nix, apt, brew,
# or apk (detected in that order); assumes NOTHING is present (not git,
# make, or go). golangci-lint is deliberately not installed: linting runs
# only in docker, via script/lint and Dockerfile.lint. Finishes by running
# script/fetch-assets, which installs the hash-pinned third-party browser
# assets the repo does not commit.
# make, or go). golangci-lint is packaged in nix, brew, and apk; on apt
# it is installed from a hash-verified GitHub release archive (never
# curl | sh). Finishes by running script/fetch-assets, which installs the
# hash-pinned third-party browser assets the repo does not commit.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
# Pinned versions, 2026-08-07. Never "latest"; exact versions only.
GOLANGCI_LINT_VERSION="2.12.2"
# sha256 of golangci-lint-2.12.2-linux-<arch>.tar.gz release archives
GOLANGCI_LINT_SHA256_AMD64="8df580d2670fed8fa984aac0507099af8df275e665215f5c7a2ae3943893a553"
GOLANGCI_LINT_SHA256_ARM64="44cd40a8c76c86755375adfeea52cfd3533cb43d7bd647771e0ae065e166df3a"
PKGMGR=""
SUDO=""
@@ -51,6 +57,52 @@ missing() {
! command -v "$1" >/dev/null 2>&1
}
# verify_sha256 <file> <expected-hash>
verify_sha256() {
if command -v sha256sum >/dev/null 2>&1; then
actual="$(sha256sum "$1" | cut -d' ' -f1)"
else
actual="$(shasum -a 256 "$1" | cut -d' ' -f1)"
fi
if [ "$actual" != "$2" ]; then
echo "bootstrap: sha256 mismatch for $1" >&2
echo " expected: $2" >&2
echo " actual: $actual" >&2
exit 1
fi
}
# apt has no golangci-lint package: install a pinned release archive
# from GitHub, verified by hardcoded sha256 (never curl | sh).
install_golangci_lint_release() {
case "$(uname -m)" in
x86_64) goarch="amd64"; sha="$GOLANGCI_LINT_SHA256_AMD64" ;;
aarch64|arm64) goarch="arm64"; sha="$GOLANGCI_LINT_SHA256_ARM64" ;;
*)
echo "bootstrap: unsupported architecture $(uname -m)" >&2
exit 1
;;
esac
if missing curl; then pkg_install curl curl curl curl; fi
name="golangci-lint-${GOLANGCI_LINT_VERSION}-linux-${goarch}"
tmp="$(mktemp -d)"
curl -fsSL -o "$tmp/$name.tar.gz" \
"https://github.com/golangci/golangci-lint/releases/download/v${GOLANGCI_LINT_VERSION}/${name}.tar.gz"
verify_sha256 "$tmp/$name.tar.gz" "$sha"
tar -xzf "$tmp/$name.tar.gz" -C "$tmp"
$SUDO install -m 0755 "$tmp/$name/golangci-lint" /usr/local/bin/golangci-lint
rm -rf "$tmp"
}
ensure_golangci_lint() {
if ! missing golangci-lint; then return 0; fi
detect_pkgmgr
case "$PKGMGR" in
apt) install_golangci_lint_release ;;
*) pkg_install golangci-lint golangci-lint golangci-lint golangci-lint ;;
esac
}
main() {
cd "$ROOT"
@@ -58,14 +110,9 @@ main() {
if missing git; then pkg_install git git git git; fi
if missing make; then pkg_install gnumake make make make; fi
# Go toolchain
# Go toolchain and linter
if missing go; then pkg_install go golang go go; fi
# Not installed here: docker is platform-specific and out of scope for a
# package-manager bootstrap, but script/lint needs it.
if missing docker; then
echo "bootstrap: docker not found; script/lint requires it" >&2
fi
ensure_golangci_lint
go mod download

View File

@@ -1,152 +0,0 @@
#!/bin/sh
# script/ci-mark-superseded: record an honest status on commits whose CI
# run Gitea cancelled because a newer commit landed on the same branch.
# Gitea writes `failure` / "Has been cancelled" for such a run, which
# reads as a test result on a commit nothing ever tested. Cancellation is
# unconditional server-side for push events, so the superseding run
# rewrites those statuses to `failure` with a description that says the
# commit was never tested. `skipped` cannot be used: Gitea's combined
# status folds `skipped` into `success`, so a never-tested commit would
# report green. Genuine failures and successes are never touched.
#
# Called by the Gitea Actions workflow, which supplies GITHUB_API_URL,
# GITHUB_REPOSITORY, GITHUB_SHA, GITHUB_WORKFLOW, GITHUB_JOB,
# GITHUB_EVENT_NAME and GITEA_TOKEN. ANCESTOR_LIMIT (default 20) caps how
# far back the walk looks; a value that is set but not a positive integer
# aborts rather than silently disabling the walk.
set -eu
SUPERSEDED_DESC='Superseded by a newer commit; never tested'
# Gitea builds the commit-status context as
# "<workflow name> / <job name> (<event>)", so derive it rather than
# hardcoding the result.
#
# The derivation is deliberately not byte-exact with Gitea's own rule and
# must not be "fixed" into a silent fallback. Gitea uses the job's `name:`
# (falling back to the job id) and the workflow's `name:` (falling back to
# the workflow filename), while the runner exports GITHUB_JOB as the job
# *id* and GITHUB_WORKFLOW as the parsed workflow `name:`. So giving the
# job a display `name:`, or dropping the workflow's `name:`, makes the
# derived context stop matching --- and require_own_context below then
# turns every push red with a message. That loud failure is the point
# (https://git.eeqj.de/sneak/webhooker/issues/147 item 2); guessing at a
# fallback would restore the silent no-op it replaced.
context() {
printf '%s / %s (%s)' \
"$GITHUB_WORKFLOW" "$GITHUB_JOB" "$GITHUB_EVENT_NAME"
}
# ANCESTOR_LIMIT is a documented knob, so a value that is set but
# unusable must fail loudly instead of defaulting
# (https://git.eeqj.de/sneak/webhooker/issues/80). Passing it straight to
# git would print `fatal: not an integer` into a discarded exit status
# and mark nothing.
ancestor_limit() {
# `-` and not `:-`: an explicitly empty value is set-but-unusable
# config, so it aborts like any other bad value rather than silently
# running at the default.
_limit="${ANCESTOR_LIMIT-20}"
case "$_limit" in
'' | *[!0-9]* | 0*)
echo "ANCESTOR_LIMIT must be a positive integer," \
"got '${_limit}'" >&2
return 1
;;
esac
printf '%s' "$_limit"
}
# The status Gitea created for this very job proves which context string
# it uses. If the derived one is missing, the workflow or the job was
# renamed and the match below would silently stop firing, restoring the
# false-red bug with no signal. Fail loudly instead.
require_own_context() {
if ! _body="$(curl -sf --retry 3 --retry-delay 2 --max-time 30 \
"${1}/commits/${GITHUB_SHA}/status")"; then
echo "cannot read commit statuses for ${GITHUB_SHA}" >&2
return 1
fi
_found="$(printf '%s' "$_body" | jq -r '(.statuses // [])[].context')"
if printf '%s\n' "$_found" | grep -qxF "$2"; then
return 0
fi
echo "no commit status with context '${2}' on ${GITHUB_SHA}:" >&2
echo "workflow or job renamed? contexts present:" >&2
printf '%s\n' "$_found" >&2
return 1
}
# Latest status for our context on a commit, as "state|description".
# The read is retried and bounded, and a read that still fails aborts the
# step: a laundered commit that cannot be read is not the same as one
# with nothing to do, and piping curl into jq would discard the
# difference.
status_of() {
if ! _sbody="$(curl -sf --retry 3 --retry-delay 2 --max-time 30 \
"${1}/commits/${2}/status")"; then
echo "cannot read commit statuses for ${2}" >&2
return 1
fi
printf '%s' "$_sbody" | jq -r --arg c "$3" \
'[(.statuses // [])[] | select(.context == $c)][0] // empty
| "\(.status)|\(.description)"'
}
mark_superseded() {
curl -sf -X POST "${1}/statuses/${2}" \
-H "Authorization: token ${GITEA_TOKEN}" \
-H 'Content-Type: application/json' \
-d "$(jq -nc --arg c "$3" --arg d "$SUPERSEDED_DESC" \
'{context: $c, state: "failure", description: $d}')" \
>/dev/null
}
main() {
_api="${GITHUB_API_URL}/repos/${GITHUB_REPOSITORY}"
_ctx="$(context)"
_limit="$(ancestor_limit)"
require_own_context "$_api" "$_ctx"
# A shallow clone cannot resolve the parent, so it looks exactly like
# a root commit to rev-parse below and would exit 0 having walked
# nothing (or, at depth > 1, only the ancestors that happen to be
# present). The workflow checks out with `fetch-depth: 0`; verify
# that here rather than depend on it silently.
if [ "$(git rev-parse --is-shallow-repository)" = 'true' ]; then
echo "shallow repository: the ancestor walk needs full history" >&2
return 1
fi
# A root commit legitimately has no ancestors and is not an error.
# A SHA this repository does not have lands here too, since its
# parent is equally unresolvable, but require_own_context above has
# already aborted on the 404 for it. The walk itself carries no
# `|| true`, so a rev-list failure aborts.
if ! git rev-parse -q --verify "${GITHUB_SHA}^" >/dev/null; then
echo "no ancestor of ${GITHUB_SHA} to check"
return 0
fi
_walk="$(git rev-list --max-count="$_limit" "${GITHUB_SHA}^")"
for _sha in $_walk; do
_latest="$(status_of "$_api" "$_sha" "$_ctx")"
# A run that was cancelled, or one an earlier revision of this
# script laundered into `skipped`. Anything else stands.
case "$_latest" in
'failure|Has been cancelled' | "skipped|${SUPERSEDED_DESC}") ;;
*) continue ;;
esac
mark_superseded "$_api" "$_sha" "$_ctx"
echo "marked superseded: ${_sha}"
done
}
main "$@"

View File

@@ -1,55 +1,12 @@
#!/bin/sh
# script/lint: run the linter. golangci-lint is never installed locally: it
# runs via docker only, one way, everywhere — script/lint builds
# Dockerfile.lint, which COPYs the repo into the pinned golangci-lint image
# and lints as a build step. This works even when the docker daemon is remote
# and bind mounts are impossible, and it removes the host linter's shared
# cache, which has attributed other checkouts' findings to this one.
#
# --no-cache-filter=lint forces the lint stage to re-execute on every run; a
# cached lint stage exits 0 in under a second having linted nothing. The deps
# stage keeps its cache, so module downloads are not repeated.
# --progress=plain keeps the linter's own output visible on success, so a
# passing run shows the issue count rather than nothing.
# --output=type=cacheonly leaves no image behind to clean up.
#
# docker silently ignores --no-cache-filter for a stage name that does not
# match, so a rename or a typo would restore the cached false green with no
# warning and a fast exit 0. The flag is therefore not trusted: the build
# output is teed to a log and a run is only a pass if golangci-lint's own
# summary line ("N issues." / "N issues:") is in it. No summary, no lint,
# whatever the exit code says.
# script/lint: run the linter.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
main() {
cd "$ROOT"
log="$(mktemp -t webhooker-lint.XXXXXXXX)"
rcfile="$(mktemp -t webhooker-lint-rc.XXXXXXXX)"
trap 'rm -f "$log" "$rcfile"' EXIT INT TERM
# The pipeline's status is tee's, and POSIX sh has no pipefail, so the
# build's status travels via a file. Output still streams live.
{
docker build \
-f Dockerfile.lint \
--no-cache-filter=lint \
--progress=plain \
--output=type=cacheonly \
. 2>&1 && echo 0 >"$rcfile" || echo $? >"$rcfile"
} | tee "$log" >&2
rc="$(cat "$rcfile")"
[ "$rc" -eq 0 ] || exit "$rc"
if ! grep -qE '[0-9]+ issues[.:]' "$log"; then
echo "script/lint: golangci-lint printed no summary line; the linter" >&2
echo " did not run. Check that the stage named in --no-cache-filter" >&2
echo " still matches a stage in Dockerfile.lint." >&2
exit 1
fi
golangci-lint run --config .golangci.yml ./...
}
main "$@"

View File

@@ -1,34 +1,12 @@
#!/bin/sh
# script/test: run the test suite.
#
# -timeout is applied by `go test` per package, not to the run as a whole, so
# it only has to clear the slowest single package. That is internal/handlers,
# measured in a cache-defeated builder stage on the 48-core shared build host
# (2026-08-18); load- and host-dependent, not invariants:
#
# 16.9s host load 5-20, GOMAXPROCS 48
# 45.9s / 47.3s / 49.0s three runs at deliberate host load 31-73
# 30.6s / 39.7s host load 5-20, GOMAXPROCS 6 / 4
# 67.3s / 97.5s host load 5-20, GOMAXPROCS 2 / 1
# 67.3s GOMAXPROCS 4 at deliberate host load 52-68
#
# The old 30s budget was breached by every loaded run and by every GOMAXPROCS
# at or below 6; at GOMAXPROCS 4 it failed outright ("panic: test timed out
# after 30s"), reproduced on 33e4fa4 with no other change.
#
# 90s matches the org-wide backstop in REPO_POLICIES.md and is sized here
# against the figures above: the worst case under native parallelism is 49.0s,
# and the compound GOMAXPROCS-4-under-load case at 67.3s sits at 75% of it.
# The one figure above 90s is GOMAXPROCS 1, a synthetic core floor rather than
# a condition CI runs under. If a CPU-limited runner ever puts a real run near
# 67s, that is the datum to revisit the org figure with.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
main() {
cd "$ROOT"
go test -v -race -timeout 90s ./...
go test -v -race -timeout 30s ./...
}
main "$@"

View File

@@ -38,7 +38,7 @@
<div x-show="open" x-cloak class="mt-3 p-3 bg-gray-50 rounded-md">
<pre class="text-xs text-gray-700 overflow-x-auto whitespace-pre-wrap break-all">{{.Body}}</pre>
{{if .BodyTruncated}}
<p class="mt-2 text-xs text-gray-500">Body truncated for display: showing {{.BodyShownBytes}} of {{.BodyBytes}} bytes. The stored body is unchanged &mdash; <a href="/source/{{$.Webhook.ID}}/logs/{{.ID}}/body" class="text-primary-600 hover:text-primary-700 underline">download the full body</a>.</p>
<p class="mt-2 text-xs text-gray-500">Body truncated for display: showing {{.BodyShownBytes}} of {{.BodyBytes}} bytes. The stored body is unchanged.</p>
{{end}}
</div>
</div>