Expose delivery metrics on /metrics (closes #209)
All checks were successful
check / check (push) Successful in 4m13s
All checks were successful
check / check (push) Successful in 4m13s
/metrics carried only the inbound HTTP surface, so a destination failing for an hour, a growing retry backlog and a stuck-open circuit breaker were all invisible: the receive side stays healthy in each case because it is. New internal/metrics registers, on the existing default registry that the go-http-metrics recorder and the promhttp handler already share: - webhooker_events_received_total - webhooker_delivery_attempts_total - webhooker_deliveries_succeeded_total - webhooker_deliveries_failed_total - webhooker_delivery_retries_total - webhooker_delivery_duration_seconds - webhooker_deliveries_pending / _retrying - webhooker_circuit_breakers_open The route mounting is untouched. Every delivery metric carries one label, target_type, whose domain is the four target-type constants; anything outside it collapses to "unknown" so no series can be minted from a UUID. Target ids, event ids and entrypoint ids are deliberately not labels. Instrumentation sits at the points every target type already passes through: processDelivery for the attempt counter and the duration histogram, updateDeliveryStatus for the outcome counters. The queue-depth gauges are counted out of the per-webhook databases by a 30s sampler rather than tracked as deltas, which would need seeding at startup and would drift on any transition that failed to persist. The open-breaker gauge is recounted from the target's breaker registry on every state change.
This commit is contained in:
41
README.md
41
README.md
@@ -1122,6 +1122,44 @@ delivery as `retrying` and schedules a retry timer for after the
|
||||
remaining cooldown period. This ensures no deliveries are lost — they're
|
||||
just delayed until the target is healthy again.
|
||||
|
||||
### Metrics
|
||||
|
||||
`/metrics` serves one Prometheus registry behind basic auth (see
|
||||
[Infrastructure Endpoints](#infrastructure-endpoints)). Alongside the
|
||||
inbound HTTP metrics recorded by the middleware, it exposes the
|
||||
delivery pipeline — the part of the service that can be failing while
|
||||
the receive side looks perfectly healthy, because it is: events are
|
||||
arriving and being stored, they are just not getting anywhere.
|
||||
|
||||
| Metric | Type | Meaning |
|
||||
| ------ | ---- | ------- |
|
||||
| `webhooker_events_received_total` | counter | Events received and durably stored. Compare against the delivery counters on one dashboard |
|
||||
| `webhooker_delivery_attempts_total` | counter | Delivery attempts dispatched to a target |
|
||||
| `webhooker_deliveries_succeeded_total` | counter | Deliveries that reached `delivered` |
|
||||
| `webhooker_deliveries_failed_total` | counter | Deliveries that failed terminally and will not be retried |
|
||||
| `webhooker_delivery_retries_total` | counter | Deliveries put back into `retrying` |
|
||||
| `webhooker_delivery_duration_seconds` | histogram | Wall time of a single delivery attempt |
|
||||
| `webhooker_deliveries_pending` | gauge | Deliveries currently in `pending` |
|
||||
| `webhooker_deliveries_retrying` | gauge | Deliveries currently in `retrying` |
|
||||
| `webhooker_circuit_breakers_open` | gauge | Circuit breakers currently open |
|
||||
|
||||
Every delivery metric carries exactly one label, `target_type`, and
|
||||
cardinality is the whole reason for that restriction. A target type is
|
||||
one of four compile-time constants, so the label domain is bounded by
|
||||
construction; a value outside that set collapses to `unknown` rather
|
||||
than minting a series of its own. Target ids, event ids and entrypoint
|
||||
ids are deliberately not labels: they are UUIDs minted per operator
|
||||
action or per inbound request, a series is never reclaimed once it
|
||||
exists, and labelling by any of them would make `/metrics` a memory
|
||||
leak that grows with traffic.
|
||||
|
||||
The two queue-depth gauges are counted out of the databases by a
|
||||
sampler that runs every 30 seconds for as long as the delivery engine
|
||||
does, rather than tracked as deltas alongside the status transitions: a
|
||||
delta would have to be seeded at startup from rows a previous process
|
||||
wrote, and would drift permanently on any transition that failed to
|
||||
persist.
|
||||
|
||||
### Rate Limiting
|
||||
|
||||
Global blanket rate limiting middleware (e.g., a per-IP throttle shared
|
||||
@@ -1747,6 +1785,7 @@ webhooker/
|
||||
│ │ ├── target_log.go # Log target (stdout)
|
||||
│ │ ├── target_config_view.go # Masked target config for templates
|
||||
│ │ ├── archive_sweeper.go # Periodic pruning of idle archives
|
||||
│ │ ├── queue_depth.go # Periodic sampler behind the queue-depth gauges
|
||||
│ │ ├── url_mask.go # Strips credentials from *url.Error
|
||||
│ │ └── ssrf.go # SSRF prevention (IP validation, safe HTTP transport)
|
||||
│ ├── handlers/
|
||||
@@ -1764,6 +1803,8 @@ webhooker/
|
||||
│ │ └── lifecycle.go # Shared stop-hook waiter, bounded by the stop context
|
||||
│ ├── logger/
|
||||
│ │ └── logger.go # slog setup with TTY detection
|
||||
│ ├── metrics/
|
||||
│ │ └── metrics.go # Delivery Prometheus collectors, labelled by target type
|
||||
│ ├── middleware/
|
||||
│ │ ├── middleware.go # Logging, CORS, Auth, Metrics, MetricsAuth, SecurityHeaders, MaxBodySize
|
||||
│ │ ├── csrf.go # CSRF protection middleware (gorilla/csrf)
|
||||
|
||||
Reference in New Issue
Block a user