Compare commits

..
1 Commits
Author SHA1 Message Date
clawbot 860940246d Anomaly thresholds: alerts for unusual traffic, nothing refused (closes #101)
check / check (push) Canceled after 0s
SWWAF_ANOMALY_CLIENT_*, _NET_*, _ASN_*, _TOTAL_* and SWWAF_WATCH_* with
SWWAF_WATCH_NETS: requests and bytes per minute and per hour, each off by
default. Every request but the health check is counted, allow-listed and
exempt ones included, in the buckets the rate limits use; a count over its
threshold raises an anomaly alert, with a cooldown per scope. At most
20,000 counters, kept in alerts.json. A per-AS-number threshold with
lookups off, or a malformed SWWAF_WATCH_NETS, stops the start.

Judgement call: refused requests are counted too.
Judgement call: per-client counters are kept in alerts.json, which SPEC.md does not list.
Judgement call: a request counts for an AS number only if the lookup answered before it ended.

Model: opus-5-5
2026-10-07 13:04:15 +00:00
5 changed files with 54 additions and 185 deletions
+13 -17
View File
@@ -820,7 +820,7 @@ is sent on one line:
is the file it was renamed to, and the `error`, which for a file that does not
parse names where in it the error is.
- `suppressed_repeats` is how many repeats the cooldown held back before this
alert, and for a `summary`, those no other alert gives (see below).
alert.
Slack and ntfy are each sent the alert as a message: a title, the instance and
the event, such as `fsn1app1/gitea: ban`, and a text, the `reason`, then a line
@@ -868,18 +868,15 @@ An alert for the same event as the last one sent, on the same netblock, or for a
source, or for an `anomaly` in the same scope, with the same netblock, AS number
or name, whatever its window and kind, less than `SWWAF_ALERT_COOLDOWN` after
it, is a repeat: it is held back and counted, and the next alert sent for them
gives that count as `suppressed_repeats`. As each hour of the clock, in UTC,
ends, the cooldowns that have run out are dropped, and the repeats they held
back, which no alert sent since has given, go in that hour's summary.
gives that count as `suppressed_repeats`.
Past `SWWAF_ALERT_MAX_PER_HOUR` alerts in an hour, the hour's other alerts are
held back and counted by event. An alert held back this way starts no cooldown.
Once the hour has ended, one alert sums up the alerts held back and the repeats
of the cooldowns dropped: its `event` is `summary`, its `reason` says how many
of each were held back, its `detail` gives the `hour` as when it started, the
`count` of alerts held back, and the count for each event, as `events`, and its
`suppressed_repeats` gives the repeats. An hour with neither ends without a
summary.
Past `SWWAF_ALERT_MAX_PER_HOUR` alerts in an hour of the clock, in UTC, the
hour's other alerts are held back and counted by event. Once the hour has ended,
one alert sums them up: its `event` is `summary`, its `reason` says how many
were held back, and its `detail` gives the `hour` as when it started, the
`count`, and the count for each event, as `events`. An alert held back this way
starts no cooldown, and the repeats held back before it are given by the next
alert sent for the same event and netblock, file, source or scope.
Each destination has a queue of its own, of at most 1000 alerts, from which they
are sent to it one at a time, the oldest first, so a destination that is slow or
@@ -941,11 +938,10 @@ which are listed by scope first, with times in UTC.
number and the `name` of a named netblock, and its two buckets of requests in
the minute and the hour, `minute` and `hour`, and of bytes, `minute_bytes` and
`hour_bytes`, each left out while it is empty. As an hour ends, the cooldowns
that have run out are dropped, and the hour's summary gives the repeats they
held back. As the file is read, the alerts waiting for a destination you no
longer name are dropped. A file whose `waiting` is a list, as it was before
alerts went to Slack and ntfy too, stops the start: put the list under
`"webhook"`, or remove the file.
that have run out with no repeat held back are dropped. As the file is read,
the alerts waiting for a destination you no longer name are dropped. A file
whose `waiting` is a list, as it was before alerts went to Slack and ntfy too,
stops the start: put the list under `"webhook"`, or remove the file.
`bans.json` is written `SWWAF_STATE_WRITE_DELAY` after a ban is made, lifted
through `DELETE /_smallwebwaf/bans/<client>`, or made permanent, with every such
+15 -36
View File
@@ -161,9 +161,7 @@ type Alert struct {
Reason string `json:"reason"`
Detail map[string]any `json:"detail"`
// SuppressedRepeats is how many repeats of the alert the cooldown
// held back since the last one let through. For a summary, it is how
// many the cooldowns dropped as the hour ended had held back that no
// alert let through gave.
// held back since the last one let through.
SuppressedRepeats int `json:"suppressed_repeats"`
}
@@ -315,8 +313,7 @@ func New(params Params) *Queue {
// alerts let through in the hour under way, by the clock, an alert is
// held back for that hour's summary instead, which is sent once the hour
// has ended; it starts no cooldown, and the repeats held back before it
// are given by that summary, as their cooldown, which has run out, is
// dropped when the hour ends. Raise never waits: an alert
// are given by the next alert let through. Raise never waits: an alert
// let through joins the queue of each destination, from which Run sends
// it, and with queueSize alerts waiting for a destination, the oldest is
// dropped.
@@ -565,60 +562,42 @@ func (q *Queue) startCooldown(alert *Alert, now time.Time) {
}
}
// endHour ends the hour under way, if now is past it. It drops the
// cooldowns that have run out, whatever repeats they held back, so that
// they do not pile up, and queues that hour's summary when alerts were
// held back in it past MaxPerHour, or when a cooldown dropped had held
// back repeats, which no alert let through has given: the summary gives
// them.
// endHour ends the hour under way, if now is past it: it queues that
// hour's summary when alerts were held back in it past MaxPerHour, and
// forgets the cooldowns that have run out with no repeat held back, which
// no alert needs any more.
func (q *Queue) endHour(now time.Time) {
start := now.Truncate(time.Hour)
if !start.After(q.hour.Start) {
return
}
repeats := 0
for key, cooldown := range q.cooldowns {
if now.Sub(cooldown.Sent) >= q.params.Cooldown {
repeats += cooldown.SuppressedRepeats
delete(q.cooldowns, key)
}
}
heldBack := 0
for _, count := range q.hour.HeldBack {
heldBack += count
}
var reasons []string
if heldBack > 0 {
reasons = append(reasons, fmt.Sprintf("%d alerts held back in the hour from %s, "+
"past the %d an hour SWWAF_ALERT_MAX_PER_HOUR allows", heldBack,
q.hour.Start.Format(time.RFC3339), q.params.MaxPerHour))
}
if repeats > 0 {
reasons = append(reasons, fmt.Sprintf("%d repeats held back by "+
"SWWAF_ALERT_COOLDOWN that no later alert gives", repeats))
}
if len(reasons) > 0 {
q.queue(&Alert{
Instance: q.params.Instance,
Time: now,
Event: EventSummary,
Reason: strings.Join(reasons, "; "),
Reason: fmt.Sprintf("%d alerts held back in the hour from %s, past the %d "+
"an hour SWWAF_ALERT_MAX_PER_HOUR allows", heldBack,
q.hour.Start.Format(time.RFC3339), q.params.MaxPerHour),
Detail: map[string]any{
"hour": q.hour.Start, "count": heldBack, "events": q.hour.HeldBack,
},
SuppressedRepeats: repeats,
})
}
q.hour = Hour{Start: start, HeldBack: map[string]int{}}
for key, cooldown := range q.cooldowns {
if now.Sub(cooldown.Sent) >= q.params.Cooldown && cooldown.SuppressedRepeats == 0 {
delete(q.cooldowns, key)
}
}
}
// queue adds alert to the alerts waiting for each destination.
+25 -87
View File
@@ -381,7 +381,7 @@ func TestAlertsPastTheHourlyLimitAreRolledIntoOneSummary(t *testing.T) {
})
}
func TestRepeatsBeforeAnAlertPastTheHourlyLimitAreGivenByTheSummary(t *testing.T) {
func TestRepeatsBeforeAnAlertPastTheHourlyLimitAreGivenByTheNextSent(t *testing.T) {
t.Parallel()
synctest.Test(t, func(t *testing.T) {
@@ -401,8 +401,8 @@ func TestRepeatsBeforeAnAlertPastTheHourlyLimitAreGivenByTheSummary(t *testing.T
time.Sleep(cooldown)
raise()
// The summary gives the alert past the limit and the two repeats,
// and the next hour's first alert none.
// The next hour's first alert gives the two repeats, and the summary
// the alert past the limit.
time.Sleep(time.Hour - cooldown)
synctest.Wait()
raise()
@@ -413,14 +413,11 @@ func TestRepeatsBeforeAnAlertPastTheHourlyLimitAreGivenByTheSummary(t *testing.T
got := webhook.received()
if len(got) == 3 {
detail, _ := got[1].alert["detail"].(map[string]any)
summaryRepeats := got[1].alert["suppressed_repeats"]
lastRepeats := got[2].alert["suppressed_repeats"]
repeats := got[2].alert["suppressed_repeats"]
if detail["count"] != float64(1) || summaryRepeats != float64(2) ||
lastRepeats != float64(0) {
t.Errorf("the summary counts %v alerts and %v repeats, and the last "+
"alert gives %v repeats, want 1, 2 and 0", detail["count"],
summaryRepeats, lastRepeats)
if detail["count"] != float64(1) || repeats != float64(2) {
t.Errorf("the summary counts %v alerts, and the last alert gives %v "+
"repeats, want 1 and 2", detail["count"], repeats)
}
}
@@ -428,70 +425,6 @@ func TestRepeatsBeforeAnAlertPastTheHourlyLimitAreGivenByTheSummary(t *testing.T
})
}
func TestCooldownsThatHaveRunOutAreDroppedAndTheirRepeatsSummedUp(t *testing.T) {
t.Parallel()
for name, maxPerHour := range map[string]int{"limit off": 0, "limit set": 60} {
t.Run(name, func(t *testing.T) {
t.Parallel()
synctest.Test(t, func(t *testing.T) {
params := newParams()
params.MaxPerHour = maxPerHour
webhook, q := start(t, params)
// For four hours, a netblock of its own each minute is over an
// anomaly threshold twice: an alert, and a repeat the cooldown
// holds back.
netblocks := 0
for range 4 {
for range 60 {
anomaly := alerts.Alert{
Event: alerts.EventAnomaly, Netblock: netblock(netblocks),
Detail: map[string]any{"scope": "net"},
}
q.Raise(anomaly)
q.Raise(anomaly)
netblocks++
time.Sleep(time.Minute)
}
// As the hour ends, only the cooldowns started less than the
// cooldown before are kept, in memory and for alerts.json.
synctest.Wait()
kept := len(q.Snapshot().Cooldowns)
if kept > int(cooldown/time.Minute) {
t.Errorf("after %d netblocks, %d cooldowns are kept, want at most %d",
netblocks, kept, int(cooldown/time.Minute))
}
}
// An hour on, every cooldown has been dropped, and the summaries
// have given every repeat.
time.Sleep(time.Hour)
synctest.Wait()
repeats := 0.0
for _, request := range webhook.received() {
count, _ := request.alert["suppressed_repeats"].(float64)
repeats += count
}
kept := len(q.Snapshot().Cooldowns)
if kept != 0 || repeats != float64(netblocks) {
t.Errorf("%d cooldowns are kept and the webhook was given %v repeats, "+
"want 0 and %d", kept, repeats, netblocks)
}
})
})
}
}
func TestFailedRequestIsSentAgainWithBackoff(t *testing.T) {
t.Parallel()
@@ -702,8 +635,7 @@ func TestStateLoadedIntoANewQueueCarriesOn(t *testing.T) {
after.Load(roundTrip(t, before.Snapshot()))
// The new queue sends the alert waiting, holds back the repeat as
// the cooldown still runs, and sends the summary of the hour, which
// gives both repeats, as the cooldown has run out.
// the cooldown still runs, and sends the summary of the hour.
after.Raise(alerts.Alert{Event: alerts.EventBan, Netblock: netblock(1)})
synctest.Wait()
wantEvents(t, webhook, alerts.EventBan)
@@ -712,12 +644,18 @@ func TestStateLoadedIntoANewQueueCarriesOn(t *testing.T) {
synctest.Wait()
wantEvents(t, webhook, alerts.EventBan, alerts.EventSummary)
summary := webhook.received()[1].alert
detail, _ := summary["detail"].(map[string]any)
detail, _ := webhook.received()[1].alert["detail"].(map[string]any)
if detail["count"] != float64(1) {
t.Errorf("the summary counts %v alerts, want 1", detail["count"])
}
if detail["count"] != float64(1) || summary["suppressed_repeats"] != float64(2) {
t.Errorf("the summary counts %v alerts and %v repeats, want 1 and 2",
detail["count"], summary["suppressed_repeats"])
// The cooldown has run out, and the next one gives both repeats.
after.Raise(alerts.Alert{Event: alerts.EventBan, Netblock: netblock(1)})
synctest.Wait()
got := webhook.received()
if repeats := got[len(got)-1].alert["suppressed_repeats"]; repeats != float64(2) {
t.Errorf("the last alert gives %v repeats, want 2", repeats)
}
})
}
@@ -763,8 +701,8 @@ func TestSlackAndNtfyAreSentTheSummaryAndTheRepeatsHeldBack(t *testing.T) {
ban := alerts.Alert{Event: alerts.EventBan, Netblock: netblock(1), Reason: "a ban"}
// The hour's one alert, a repeat of it the cooldown holds back, and
// an alert past the limit; once the hour has ended, its summary,
// which gives the repeat, and the next alert, which gives none.
// an alert past the limit; once the hour has ended, its summary, and
// the next alert, which gives the repeat.
q.Raise(ban)
q.Raise(ban)
q.Raise(alerts.Alert{Event: alerts.EventFileError, Reason: "a file error"})
@@ -782,14 +720,14 @@ func TestSlackAndNtfyAreSentTheSummaryAndTheRepeatsHeldBack(t *testing.T) {
}
const summary = "1 alerts held back in the hour from 2000-01-01T00:00:00Z, " +
"past the 1 an hour SWWAF_ALERT_MAX_PER_HOUR allows; 1 repeats held back " +
"by SWWAF_ALERT_COOLDOWN that no later alert gives\nsuppressed repeats: 1"
"past the 1 an hour SWWAF_ALERT_MAX_PER_HOUR allows"
wantSlackMessage(t, slack[1], "*"+instance+": summary*\n"+summary)
wantNtfyMessage(t, ntfy[1], instance+": summary", "default bar_chart", summary)
wantSlackMessage(t, slack[2], "*"+instance+": ban*\na ban\nnetblock: 203.0.113.1/32")
wantSlackMessage(t, slack[2],
"*"+instance+": ban*\na ban\nnetblock: 203.0.113.1/32\nsuppressed repeats: 1")
wantNtfyMessage(t, ntfy[2], instance+": ban", "default no_entry",
"a ban\nnetblock: 203.0.113.1/32")
"a ban\nnetblock: 203.0.113.1/32\nsuppressed repeats: 1")
})
}
-30
View File
@@ -1,30 +0,0 @@
package proxy
import (
"net/netip"
"testing"
"sneak.berlin/go/smallwebwaf/internal/config"
)
func TestWithEveryAnomalyThresholdOffARequestIsNotCounted(t *testing.T) {
t.Parallel()
// A request from a client looked up through GeoJS, with every anomaly
// threshold off. Its handler has neither GeoJS's answers nor the
// anomaly counters, nor a clock, and the request no response: reading
// any of them to count the request panics.
rq := &request{
h: &handler{config: &config.Config{LookupSource: "geojs"}},
client: netip.MustParseAddr("203.0.113.9"),
lookedUp: true,
}
defer func() {
if r := recover(); r != nil {
t.Errorf("counting the request did work, with every threshold off: %v", r)
}
}()
rq.countAnomalies()
}
+1 -15
View File
@@ -17,7 +17,6 @@ import (
"time"
"sneak.berlin/go/smallwebwaf/internal/anomaly"
"sneak.berlin/go/smallwebwaf/internal/config"
"sneak.berlin/go/smallwebwaf/internal/lookup"
"sneak.berlin/go/smallwebwaf/internal/ratelimit"
"sneak.berlin/go/smallwebwaf/internal/requestlog"
@@ -542,13 +541,8 @@ func (rq *request) addToHistory() {
// with it: a request refused, one from a client in SWWAF_ALLOW_NETS or
// SWWAF_RATE_LIMIT_EXEMPT_NETS, and one for a path in
// SWWAF_RATE_LIMIT_EXEMPT_PATHS are counted too. It is counted for its
// client's AS number when answerAtTheEnd gives one. With every anomaly
// threshold off, the default, it does nothing.
// client's AS number when answerAtTheEnd gives one.
func (rq *request) countAnomalies() {
if !anomalyThresholdsSet(rq.h.config) {
return
}
answer, _ := rq.answerAtTheEnd()
rq.h.anomalies.Count(rq.h.now(), anomaly.Request{
@@ -561,14 +555,6 @@ func (rq *request) countAnomalies() {
})
}
// anomalyThresholdsSet reports whether any anomaly threshold is set.
func anomalyThresholdsSet(cfg *config.Config) bool {
off := anomaly.Thresholds{}
return cfg.AnomalyClient != off || cfg.AnomalyNet != off || cfg.AnomalyASN != off ||
cfg.AnomalyTotal != off || cfg.AnomalyWatch != off
}
// answerAtTheEnd returns, for a client that was looked up, the lookup's
// answer about it as the request ends, and whether there is one: the
// lookup database's, which was there at once, or the one GeoJS has given