The database target type is called an archive on its badge, in the add target form's type list and on its edit page, and its settings read "Archive expiry" and "Archive rotation" everywhere, the new webhook page included. The retry field is labelled "Delivery attempts", with its help text and error messages to match, on both target forms and in the target list, where a stored 0 shows as one attempt. The target list's other labels take the same capitalisation. The navbar says "Sign out" and the sign-in page "Sign in", and the README follows. The resubmit notice says "webhook". The stored values (`database`, `max_retries`) and their meaning are unchanged.
Model: opus-5-5
A database target's rotation setting (none, monthly, daily or hourly) puts the UTC period of each event's receive time in its archive file name, so each file holds exactly its period's events. It is on the new webhook page and both target forms, and shown in the target list. Renames move every one of a target's files and move them all back if one fails. The sweep prunes one file at a time under the target's lock and deletes a rotated file it leaves empty. Download opens one file at a time, oldest period first, finding each again under the target's current name. The target list names the current file and totals all of them.
Model: opus-5-5
A target whose circuit breaker had tripped still showed as Active, and its deliveries sat at a bare "retrying" with no attempts. Its row on the webhook page now says deliveries are paused after repeated failures and until when the cooldown ends, adding that one waiting delivery is then sent to test the target; while half-open it says deliveries are held while one tests it, with no time. Waiting deliveries show "next try no earlier than" the later of the cooldown and their own backoff, with the date when not today. The engine gains one read of a breaker's state and remaining cooldown under one lock, and shares the backoff formula.
Model: opus-5-5
An http target's max_queue_size was stored and shown in the target list as "Max Queue Size", but nothing in the delivery engine read it, so an operator who set it expecting deliveries to be bounded got nothing. It is removed from the target, the target list and the README's target table; neither form had a field for it. Nothing checks for a leftover value. An existing database keeps its old column, which is no longer read.
Model: opus-5-5
A database target's archive expiry was typed by hand as never or a raw duration such as 720h, and the target list showed it back raw. Adding or editing a database target now offers the new-webhook page's list of choices (never, 1h, 12h, 24h, 30d, 90d, 365d), defined once and shared by all three forms. The edit form starts on the stored expiry, or on the submitted one after a refused save; a stored value outside the choices is listed under its own value, so saving unchanged keeps it. The target list shows the expiry in plain units: "30 days", "12 hours", "never".
Model: opus-5-5
A slack target's edit page offers Max Retries and the delivery engine honours it, but the target list showed only its masked webhook URL, so setting retries changed nothing visible. The list now shows a slack target's Max Retries line exactly as an http target's, from the one function both use, so the label and the "0 (fire-and-forget)" wording cannot drift apart. The Max Queue Size line stays on http targets only. Tests cover a slack target with retries set, and one with a queue size stored that shows no queue-size line.
Model: opus-5-5
An audit of every file webhooker reads configuration or required state from found cases that carried on silently. A zero-length webhooker.db, and a missing or zero-length per-webhook database, now log the "created a new, empty database" warning naming the file; restart recovery opens every live webhook's database, checking under the manager's lock that it still exists, so a missing one is reported at start. The main database's open errors name webhooker.db, for the server and webhooker resetpw; resetpw refuses a zero-length webhooker.db. A directory in place of any database file or its -wal or -shm is refused naming it. The README says how each case is treated. Also closes#459.
Model: opus-5-5
For a database target, the target list showed only its expiry, so the archive file the README's backup and move-away advice depend on could only be found from a shell on the host. Each database target now shows its archive file's name, its size on disk (the file and its -wal together) and when it was last written, relative with the full UTC time on hover, all from the files' metadata without opening the archive. Before the first write, and after the file has been moved away, it shows the name and "not created yet". Tests cover all three states.
Model: opus-5-5
Each database target on the webhook page has a Download button that streams its archive as gzipped JSON, archive-WEBHOOKNAME-TARGETNAME-TIME.json.gz, with names made safe by delivery.ArchiveFileName's function. The export reads one consistent snapshot through one cursor in a read-only transaction, so archive writes carry on, and holds the rename lock only while it reads the stored names and opens the file. It extends its write deadline as it writes, so a large archive downloads for as long as the client reads; a failure after the response has started aborts the connection so the browser marks the download failed. The request limit is now the service's own middleware, which no longer writes a 504 over a response already started.
Model: opus-5-5
Six test-only gorm.Open calls in internal/delivery passed a bare gorm.Config, leaving the unfiltered idiom in the tree to be copied into production code, where every gorm.Open goes through gormlog.New. They now pass gormlog.New over a logger that discards, so no gorm.Open in the tree uses a bare gorm.Config. The stale sentence saying the tree has one test-only (*gorm.DB).Scan caller is corrected in the README and in the ParamsFilter comment: only tests call Scan, and what a test binds is fixture data. Test and documentation change only.
Model: opus-5-5
Seven error checks in internal/delivery/target_http.go could be removed without any test noticing, among them withRetry's check on writing the delivery result, the branch that leaves a sent delivery retrying and recoverable when its bookkeeping write fails. Each now fails a test when removed. The "send succeeded" case starts from a tripped circuit breaker, so a probe whose send succeeds but whose result write fails must still close the breaker. The checks in remainingBackoff and backoffElapsed stay unpinned: removing them gives the same answer, and they state a rule a reader needs. Test change only.
Model: opus-5-5
The archive sweep closes a target's archive connection before it reopens the file, and no test noticed if that close was removed, which would leak one SQLite connection per target per sweep. A test now keeps the connection from before a sweep that reopens the archive and requires it closed afterwards. The sweeper's query for database targets takes the sweep's context; a sweep whose context is done returns without an error line, so stopping is not logged as a failure, and a test pins that. The comment on the sweeper's cancel function gives the true reason it needs no lock: fx calls the stop hook only after the start hook has returned.
Model: opus-5-5
TestProcessRetryTask_LargeBody_FetchFromDB and TestProcessRetryTask_SuccessfulRetry checked only that the delivery ended delivered, so deleting the event-body fetch on the retry path, the behaviour the first is named for, left both green while a retry could deliver an empty or truncated body. Both now compare the body the target received with the stored event body byte for byte, and both fail when that fetch is deleted. The other retry-path tests are not about the body and are unchanged. Test change only.
Model: opus-5-5
The archive writer's reopen debounce test made two writes that had to land inside the real 2-second window, then slept 2.1 seconds to cross it, so a slow host could turn correct code red. The archive writer now reads the time for its reopen debounce from a clock field, time.Now in production, and the test moves that clock instead of sleeping: two writes at one instant open the file once, and a write one debounce later closes and reopens it once. Removing the debounce check fails the test. This was the last test whose result depended on real elapsed time.
Model: opus-5-5
Each database target now writes its own archive file, archive-WEBHOOKNAME-TARGETNAME-TARGETID.db, named by delivery.ArchiveFileName, in place of one archive per webhook keyed on its UUID. Renaming a webhook or a target renames its archive files (with any -wal and -shm) before the new name is saved, never over an existing file, and moves every one back if a rename or the save fails. The webhook edit, the target edit and target creation share one lock so no two interleave. Deleting a webhook or target leaves its files on disk. Nothing looks for the old archive-WEBHOOKID.db files. The README gives the naming and the recovery steps.
Model: opus-5-5
On the build host a connection to [::] reaches a listener on ::1, and one to 0.0.0.0 reaches 127.0.0.1, so a delivery target at either reached this host's loopback past the guard. 0.0.0.0/32 and ::/128 are now in alwaysBlockedNetworks, which no allowlist opens; an allowlist reaches loopback only through an entry covering a loopback address. IPv6 multicast (ff00::/8) and documentation space (2001:db8::/32) are refused by default and reopen when listed.
Every default blocklist entry has a one-line comment, each list is pinned on its own, and tests refuse each address at target creation and at delivery. The README and the rules above each list match.
Model: opus-5-5
The http and slack targets sent the constant User-Agent webhooker/1.0. They now send webhooker/ followed by the version the build stamped, the same value the footer shows, built in one place on the delivery engine. User-Agent stays a reserved header and is still set after the target's configured headers, so a configured one cannot override it. Tests check the header each target sends against a known version, and that a configured User-Agent is replaced.
Model: opus-5-5
A second metrics-enabled router in one process panicked on a duplicate collector registration, because every collector registered on Prometheus's global default registry. metrics.NewRegistry now builds one registry with the Go runtime and process collectors; fx provides it and the delivery metric set built on it. The middleware builds its HTTP recorder once on that registry (NewForTest on a fresh one), the engine and handlers take the metric set from fx, and nothing registers on the global default any more.
/metrics is served from the new registry with the same series names, labels and auth. A test builds two metrics-enabled routers in one process.
Model: opus-5-5
Refusing an http or slack target whose address is private or reserved, on add or edit, now adds one sentence: such addresses are refused by default, and the server's ALLOWED_EGRESS_CIDRS setting allows named networks, with a pointer to the README section. It suggests no value, so it never points at allowing everything.
Metadata refusals get no such sentence. To tell them apart, the default blocklist's public addresses now have their own list, blockedPublicNetworks, still checked after the allowlist; a test pins which addresses each list refuses and how listing opens them, unchanged from before.
Model: opus-5-5
The webhook page opens with a statistics pane: entrypoints and targets, deliveries in progress, the last arrival, the retention period; events, deliveries and failures, lifetime and within retention; and events, failures and failure percentage over 10 minutes and 24 hours.
Running totals, one row per target plus one for events, sit in each webhook's event database and are written in the same transaction as the rows they count; retention prunes in paused, stoppable batches and subtracts what it removes. Window figures are index-range counts on a new finished_at column. An existing event database must be recreated.
Model: opus-5-5
The README's egress section and the comment above blockedNetworks now state what the default blocklist covers: the IPv4 private and reserved ranges, IPv6 loopback, unique local and link-local addresses, and certain public addresses, each added only because it hands credentials, user data or bootstrap material to whatever can reach it without the caller presenting anything. That is the same material as the rule above alwaysBlockedNetworks, so a future candidate can be accepted or refused against one rule.
A provider's other public addresses, such as 161.26.0.0/16 and 166.8.0.0/14, are not refused. Docs and a comment only; the lists are unchanged.
Model: opus-5-5
Add 168.63.129.16 (Azure WireServer) to blockedNetworks, the default
blocklist, not alwaysBlockedNetworks: it is public unicast, so an
operator who lists it in ALLOWED_EGRESS_CIDRS can reach it again. The
refusal message, the allowlist startup warning, the README and the
comments no longer call every blocked address private/reserved, and
no longer claim the allowlist cannot open any metadata endpoint.
Sources:
- https://learn.microsoft.com/en-us/azure/virtual-network/what-is-ip-address-168-63-129-16
- https://learn.microsoft.com/en-us/azure/virtual-machines/metadata-security-protocol/overview
Deviation: 147.75.207.243 (Equinix Metal) is not added; Equinix
documents only a hostname, and the service was sunset on 2026-06-30.
Model: opus-5-5
The engine cached archive writers and never closed them at shutdown,
so after a clean stop an archive's rows could sit in its -wal while
the .db held no table. The engine's stop hook now evicts every cached
writer once its workers have returned, the same way deleting a webhook
does, so a clean stop leaves each archive as one file and a late write
is refused. If the workers do not return within the stop budget, the
writers are left open as a kill would leave them: closing would wait
on a write in progress, and a still-running worker would open new
ones.
The README no longer says archives keep their sidecars across a clean
stop.
Model: opus-5-5
A delivery set Content-Type from the event's ContentType and then
added the inbound Content-Type from the event's stored headers, so a
target could receive two values. The inbound Content-Type is no
longer forwarded from the stored headers; the receiver already saves
it as the event's ContentType.
Which value is sent is now stated at applyRequestHeaders: a
Content-Type configured on the target, otherwise the event's
ContentType, otherwise none. A configured one still survives a
cross-origin 307/308 with its body.
Model: opus-5-5
Restart recovery and the pending sweep skipped a pending delivery
whose target was missing from the batch's target map, every minute,
for the life of the database. A miss now asks loadTarget: no row
fails the delivery terminally with a recorded reason; any other error
leaves it pending, since the map is also empty when its query failed;
a target found there is used.
The failure goes through the ownership-gated function the retrying
paths already used, now failMissingTarget. Once it owns the delivery
it re-reads the row and fails it only if the status is unchanged, so
a delivery sent and settled in between is left alone.
Model: opus-5-5
While the breaker was half-open, Allow refused every delivery but the
probe and CooldownRemaining returned zero, so each queued task for the
target went straight back onto the retry channel and rewrote its status
on every pass until the probe finished.
CooldownRemaining now returns the whole cooldown while half-open, so a
refused delivery waits that long. A refused delivery already at
retrying is not written again, so the retry counter now moves only
when a refusal moves a delivery into retrying.
Model: opus-5-5
Restart recovery could find a just-written delivery pending, send it
and release it before the receiver's Notify queued the same delivery.
Notify's claim then succeeded on the released id, and the worker sent
it again because the new-task path never read the delivery's row.
Before sending a new task the worker now reads the delivery's status
by primary key and skips the task unless the row still says pending,
as the retry path already does for retrying. Nothing else can change
the row while the worker owns the delivery. A row left pending by a
failed bookkeeping write is still sent again.
loadRetryDelivery is renamed loadDelivery now that both paths use it.
Model: opus-5-5
Closes #301. Docs-only apart from one test comment; no behaviour change.
The receiver has authenticated on the entrypoint UUID alone since inbound signature verification was removed in #279. The README described that as the current state. It did not say it is the decision, which leaves a future contributor free to propose HMAC as an improvement rather than as a reversal.
What changed:
- `## The entrypoint URL is the authentication secret` now states the rule: the v4 UUID at `/webhook/{uuid}` is the credential and the only one; no shared secret, HMAC signature, bearer token or second factor will be added, including as defence in depth. It names the removal that settled it, and it says explicitly that signature headers a sender sends anyway are stored and forwarded but never checked — the previous text left that ambiguous.
- The same section carries the two consequences an operator has to act on: the URL is a capability, so keep it out of logs, tickets and screenshots; and rotation means minting a new entrypoint, not changing a key.
- It also handles the case the rule will next be argued from: a sender that only supports signed payloads to a well-known URL is a constraint on that integration, to be raised on its own terms, not grounds to reintroduce shared secrets.
- The rule is reachable without scrolling 1,100 lines: a pointer in the intro, a new first bullet under Authentication (which previously listed the web UI, the API and `/metrics` and said nothing about the receiver at all), and a sharpened bullet under Security.
Stale language found and corrected: one, in `internal/delivery/redirect_test.go`. Its comment justified same-origin header retention partly by "the inbound signature the receiver verifies" — in this repo's vocabulary "the receiver" is `/webhook/{uuid}`, which verifies nothing. The endpoint that verifies it is the delivery target's, and the comment now says so.
Two places that read like stale signing language were checked and left alone as accurate: `internal/delivery/redirect.go` and `internal/server/sentry.go` describe signature headers senders put on the receiver route, which do arrive and are forwarded — neither claims webhooker checks them.
`REPO_POLICIES.md` was deliberately not touched. It is the cross-project policy document synced from `sneak/prompts` and carries `last_modified` front matter for that purpose, so a webhooker-specific carve-out does not belong in it. Worth knowing: its hardening section ends "if a standard security hardening measure exists for HTTP services and is not listed here, it is still expected. When in doubt, harden" — that is the sentence a future HMAC proposal will cite, and only the README now answers it.
`TODO.md` is untouched per its own Workflow section (issue branches do not touch it).
Co-authored-by: sneak <sneak@sneak.berlin>
Reviewed-on: #302
Co-authored-by: clawbot <clawbot@noreply.example.org>
Co-committed-by: clawbot <clawbot@noreply.example.org>
delivery_results stored status_code, response_body, error, duration and
attempt_num, and no template rendered any of it, so a failure read as
"target: failed" and diagnosing it meant opening the per-webhook SQLite
file by hand.
An expanded delivery now lists its attempts with attempt number, status
code, duration, error and response body. The body is bounded in the query
rather than read whole and truncated in Go (#135), and a body the engine
itself cut is no longer presented as complete.
The response body and error are untrusted remote content, so target
credentials are removed before rendering. Two cases needed care: a secret
severed by the 4096-byte cut matches nothing as a whole string, and the
engine's io.LimitReader cuts at the same constant the renderer uses, so
the guard keys on the body reaching the cap rather than on the stored size
exceeding it. Empty secrets are filtered where the secret list is built,
because an empty string passed to strings.ReplaceAll inserts the marker at
every byte boundary.
loadTargetMap builds the redactor half unscoped, so a soft-deleted
target's historical deliveries still render redacted.
Also regenerates static/css/tailwind.css, which had drifted from the
templates: hover:text-red-700, text-red-500, underline and w-28 were in
use but absent from the served stylesheet (#236).
check / check (push) Superseded by a newer commit; never tested
The receiver had no inbound authentication of any kind: /webhook/{uuid}
was mounted behind a rate limiter alone, so the only thing protecting an
entrypoint was the secrecy of a v4 UUID in a URL path. Inbound headers are
forwarded almost verbatim to the target, so anyone who learned the URL
also chose the headers the downstream service received.
Adds an optional per-entrypoint secret with two schemes: github
(X-Hub-Signature-256, HMAC-SHA256 hex over the raw body) and gitlab
(X-Gitlab-Token, a plain shared token). Comparison is constant-time, the
HMAC is computed over the raw body before any parsing, and rejection
happens before persistence -- an unauthenticated request creates no event
row. An entrypoint with no secret behaves exactly as before, including
every row that predates this change.
The scheme's credential header is stripped from the header map before it
is marshalled into Event.Headers, so the GitLab token reaches neither the
event store nor any delivery target. SchemeInfo.HeaderIsDigest defaults to
false meaning strip, so a scheme added later is protected unless its
header is positively declared a digest.
GORM's association upsert copied whole targets rows -- plaintext
credential-bearing config -- into the per-webhook event databases with an
empty webhook_id. The leak was in updateDeliveryStatus, not the create
path: Update leaves Statement.Model pointing at a Delivery whose Target
the engine populated, so save_before_associations upserts it.
A connection-level callback now appends clause.Associations to
Statement.Omits on the create and update chains of every per-webhook
connection, so every write path is covered rather than one call site.
Existing files are swept on first open: the leaked rows are deleted and
the file is VACUUMed, because DELETE alone only unlinks the pages and
leaves the credential recoverable in the file's free space. The sweep is
recorded in PRAGMA user_version only after the VACUUM returns, so a sweep
that fails or is interrupted fails the open and is retried on the next
one, rather than being marked done.
Encryption of target config at rest is deliberately out of scope and
deferred to #212.
check / check (push) Superseded by a newer commit; never tested
Both queue-depth reads used (*gorm.DB).Scan, which swaps GORM's own trace
recorder in for the logging adapter. That recorder does not implement
gorm.ParamsFilter, so those statements logged their bound values
interpolated, bypassing the suppression added for #207.
The scan guard from #222 and the queue-depth sampler from #224 each gated
green against a next that lacked the other; both landed and next went red.
Converted to Find. The emitted SQL is identical apart from placeholders,
and both paths parse the anonymous dest schema the same way, so the
queue-depth gauges are unchanged.
check / check (push) Superseded by a newer commit; never tested
The http target's destination URL can itself be a bearer credential, and
the source detail page rendered it in full. Render it through the
existing MaskURL instead, matching the rule already applied to slack
targets.
Independently reviewed: mutation-verified (reverting to the raw value
fails the absence assertions, not merely the masked-form ones), MaskURL
probed against userinfo, query, fragment, port, IPv6 literal and
non-http schemes, and every sibling path that surfaces target data
re-walked and found clean.
check / check (push) Superseded by a newer commit; never tested
The README env table was missing RETENTION_SWEEP_INTERVAL, TODO.md omitted five landed units, and three passages sold manual redelivery in the present tense when nothing implements it. The same false claim was corrected in the doc comment on failUnretryableRetry, which was its source text. Also removes a console.log from the shipped static asset.
Go embeds the request URL in *url.Error, so any transport failure — DNS,
TLS, refused, timeout, SSRF dial block — persisted the full Slack webhook
URL into the per-webhook SQLite database via DeliveryResult.Error. That
field is tagged json:"error,omitempty", so a future REST API would have
served it.
maskURLError rebuilds the error preserving Op and the wrapped cause, so DNS
vs TLS vs timeout still read differently and errors.Is/As and Timeout()
keep working; only path, query and userinfo are dropped. Applied where the
errors are born, which covers both the Slack and HTTP targets. url.Parse
embeds the URL too, so ValidateTargetURL's parse branch gets the same
treatment.
The SSRF rejection log now logs the masked URL, and source_logs.html
receives view types rather than raw rows, so no config blob is reachable
from that template.
MaskURL is now the single masker for the whole tree.
check / check (push) Superseded by a newer commit; never tested
The page rendered the stored target config verbatim, exposing the Slack
incoming-webhook URL, which is a bearer credential: anyone holding it can
post to the channel indefinitely, and it cannot be scoped or revoked
per-holder.
Target config now reaches the template only as a TargetView carrying
labelled fields, so no code path can render the raw blob. maskURL keeps
scheme and host and elides the path, and drops query, fragment and
userinfo; every parse failure yields a neutral placeholder rather than
falling back to the stored string. HTTP header values are never rendered,
only a count.
Rendering change only: the stored config format and the delivery path are
unchanged.
check / check (push) Superseded by a newer commit; never tested
A delivery left in `retrying` whose target type was edited to a fire-and-forget
or unknown type was skipped forever by both restart recovery and the retry
sweep. Both paths now record a result row and mark it `failed`.