Compare commits

..
1 Commits
Author SHA1 Message Date
clawbot 0b2ff56417 Blocklists and an AS percentage file fetched by URL (closes #29)
check / check (push) Canceled after 0s
SWWAF_BLOCKLIST_URLS names lists of addresses and netblocks, fetched every
SWWAF_BLOCKLIST_REFRESH (24h, never under 1h). The last good copy is kept
whole in reputation.json and used while a fetch fails and across restarts.
SWWAF_BLOCKLIST_ACTION denies, limits or only logs a listed client; the log
line names the lists, each raises reputation_hit, and a failed fetch raises
source_failure. SWWAF_ASN_LIMIT_PERCENT_URL is fetched the same way and
counts as SWWAF_ASN_LIMIT_PERCENT does, the lower winning.

Judgement call: a restart fetches only lists whose copy is due.
Judgement call: a failed fetch is retried after the refresh, not sooner.
Not done: ban notes do not name the lists yet.

Model: opus-5-5
2026-10-07 15:10:41 +00:00
6 changed files with 124 additions and 274 deletions
+32 -40
View File
@@ -132,16 +132,15 @@ in `bin/state` unless `SWWAF_STATE_DIR` is set, and the default rule file of
`SWWAF_COUNTRY_LIMIT_PERCENT`, the percentage they give of every rate limit `SWWAF_COUNTRY_LIMIT_PERCENT`, the percentage they give of every rate limit
and byte limit, so that the same rules ban them after fewer requests, and, and byte limit, so that the same rules ban them after fewer requests, and,
while `SWWAF_UNKNOWN_LIMIT_PERCENT` is below 100, every client without a while `SWWAF_UNKNOWN_LIMIT_PERCENT` is below 100, every client without a
country that percentage. A client to which several apply gets the lowest. For country that percentage. A client to which several apply gets the lowest.
the byte limits, `SWWAF_ASN_BYTES_PERCENT` gives an AS number it lists a `SWWAF_ASN_BYTES_PERCENT` and `SWWAF_COUNTRY_BYTES_PERCENT` give the AS
percentage in place of those `SWWAF_ASN_LIMIT_PERCENT` and the file give it, numbers and countries they list a percentage of the byte limits in place of
and `SWWAF_COUNTRY_BYTES_PERCENT` gives a country it lists one in place of the the other two. Each client is counted on its own, against its own lowered
one `SWWAF_COUNTRY_LIMIT_PERCENT` gives it. Each client is counted on its own, limits: no budget is shared by a whole AS number or country, which one abuser
against its own lowered limits: no budget is shared by a whole AS number or could use up and so lock out everyone else there. The log line of each request
country, which one abuser could use up and so lock out everyone else there. the rate limits count gives its client's percentages below 100 and the
The log line of each request the rate limits count gives its client's settings that gave them, and so do the notes of a ban for a lowered limit, and
percentages below 100 and the settings that gave them, and so do the notes of its alert.
a ban for a lowered limit, and its alert.
- Bans a client that breaks a rate limit or a byte limit, as "Bans" in - Bans a client that breaks a rate limit or a byte limit, as "Bans" in
[`SPEC.md`](SPEC.md) describes: the first ban lasts an hour, and a limit [`SPEC.md`](SPEC.md) describes: the first ban lasts an hour, and a limit
broken again within a day of a ban ending bans for three times as long as that broken again within a day of a ban ending bans for three times as long as that
@@ -986,12 +985,11 @@ with times in UTC.
`grep` shows everything about one. `grep` shows everything about one.
- `lookups.json`: GeoJS's answers, one to a line, each with the client's AS - `lookups.json`: GeoJS's answers, one to a line, each with the client's AS
number, AS name and country, when GeoJS gave it and when it was last used. number, AS name and country, when GeoJS gave it and when it was last used.
- `reputation.json`: each list fetched from a URL (see "Blocklists" below), - `reputation.json`: the last good copy of each list fetched from a URL (see
indented to be read, under `lists`: its `url`, when it was last `tried`, the "Blocklists" below), indented to be read, under `lists`: its `url`, when it
fetch failed or not, and its last good copy: when that was `fetched`, and its was `fetched`, and its `lines`, as fetched, comment lines included, each on a
`lines`, as fetched, comment lines included, each on a line of its own, both line of its own. As the file is read, the copies of lists the settings no
left out while no fetch of it has succeeded. As the file is read, the lists longer name are dropped.
the settings no longer name are dropped.
- `alerts.json`: the state of the alerts (see "Alerts" above), indented to be - `alerts.json`: the state of the alerts (see "Alerts" above), indented to be
read: under `cooldowns`, for each event and netblock, with the `source` too read: under `cooldowns`, for each event and netblock, with the `source` too
for a `reputation_hit`, or event and `file` or `source`, or for an `anomaly`, for a `reputation_hit`, or event and `file` or `source`, or for an `anomaly`,
@@ -1614,28 +1612,24 @@ GeoJS.
one to a line, written as the Spamhaus DROP list, one to a line, written as the Spamhaus DROP list,
`https://www.spamhaus.org/drop/drop.txt`, is. Anything after a `;` or a `#` on a `https://www.spamhaus.org/drop/drop.txt`, is. Anything after a `;` or a `#` on a
line is left out, and so is a line left blank. A bare address stands for itself line is left out, and so is a line left blank. A bare address stands for itself
alone, as in the settings. An IPv4-mapped address or netblock, such as alone, as in the settings. None is named by default: a list judges a client by
`::ffff:192.0.2.0/120`, is read as the IPv4 one it stands for, here what others saw it do, while the defaults judge it by what it does to your
`192.0.2.0/24`, since a client's IPv4 address is checked as IPv4; a mapped service.
netblock shorter than `/96` stands for none, and is not a netblock. None is
named by default: a list judges a client by what others saw it do, while the
defaults judge it by what it does to your service.
`smallwebwaf` fetches each list, and the file `SWWAF_ASN_LIMIT_PERCENT_URL` `smallwebwaf` fetches each list, and the file `SWWAF_ASN_LIMIT_PERCENT_URL`
names, `SWWAF_BLOCKLIST_REFRESH` after it last fetched it or tried to, 24 hours names, `SWWAF_BLOCKLIST_REFRESH` after it last fetched it or tried to, 24 hours
by default and never less than one, one list after another. Since by default and never less than one, one list after another. At start it fetches
`reputation.json` keeps when each list was last tried, the fetch failed or not, at once a list it keeps no copy of, and one whose copy is that old; a younger
at start `smallwebwaf` fetches at once only a list it has never tried, and one copy waits its turn, so that restarts do not fetch a list more often. A fetch
it last tried that long ago; any other waits its turn, so that restarts do not fails when the server answers other than `200`, when it does not finish within a
fetch a list more often. A fetch fails when the server answers other than `200`, minute, when the list is longer than 16 MiB, or when a line of it is not an
when it does not finish within a minute, when the list is longer than 16 MiB, or address or a netblock, or for `SWWAF_ASN_LIMIT_PERCENT_URL`, not an AS number,
when a line of it is not an address or a netblock, or for `:` and a percentage. The copy fetched before then stays in use, and the failure
`SWWAF_ASN_LIMIT_PERCENT_URL`, not an AS number, `:` and a percentage. The copy is counted, logged and raised as a `source_failure` alert. The last good copy of
fetched before then stays in use, and the failure is counted, logged and raised each list is kept whole, comment lines included, in `reputation.json` (see
as a `source_failure` alert. The last good copy of each list is kept whole, "State files" above), so that a restart keeps it in use too. Each list is named
comment lines included, in `reputation.json` (see "State files" above), so that by its URL, in the request log, the alerts and the metrics, so keep a secret out
a restart keeps it in use too. Each list is named by its URL, in the request of it.
log, the alerts and the metrics, so keep a secret out of it.
A client in `SWWAF_ALLOW_NETS` is not checked. Any other is checked by its own A client in `SWWAF_ALLOW_NETS` is not checked. Any other is checked by its own
address after the country lists, and `SWWAF_BLOCKLIST_ACTION` says what is done address after the country lists, and `SWWAF_BLOCKLIST_ACTION` says what is done
@@ -1651,11 +1645,9 @@ and keep the list's date and copyright lines with the data, which the copy in
from The Spamhaus Project, and should say so. They also ask that it be fetched from The Spamhaus Project, and should say so. They also ask that it be fetched
automatically no more than once an hour, once a day being more than enough in automatically no more than once an hour, once a day being more than enough in
most cases, and Spamhaus may block an address that fetches it more often. Each most cases, and Spamhaus may block an address that fetches it more often. Each
`smallwebwaf` fetches its own copy, on its own schedule, so on a host where `smallwebwaf` fetches its own copy, so on a host where several apps run it,
several apps run it, sharing one address, their fetches can come less than an sharing one address, set `SWWAF_BLOCKLIST_REFRESH` long enough that their
hour apart whatever `SWWAF_BLOCKLIST_REFRESH` is: each fetches a list again that fetches come at least an hour apart.
long after its own last try, so those first started within the same hour, with
the same refresh, keep fetching within the same hour.
## How the code is laid out ## How the code is laid out
+36 -66
View File
@@ -2,8 +2,8 @@
// blocklists of SWWAF_BLOCKLIST_URLS, and the file of AS:percent lines // blocklists of SWWAF_BLOCKLIST_URLS, and the file of AS:percent lines
// SWWAF_ASN_LIMIT_PERCENT_URL names. It keeps the last good copy of each, // SWWAF_ASN_LIMIT_PERCENT_URL names. It keeps the last good copy of each,
// whole, comment lines included, which is used while a fetch fails, and // whole, comment lines included, which is used while a fetch fails, and
// when each was last tried, which the state package writes to // which the state package writes to reputation.json and reads from it, so
// reputation.json and reads from it, so that a restart keeps them too. // that a restart keeps it too.
package reputation package reputation
import ( import (
@@ -29,9 +29,6 @@ const (
maxListBytes = 16 << 20 maxListBytes = 16 << 20
// fetchTimeout bounds one fetch of a list. // fetchTimeout bounds one fetch of a list.
fetchTimeout = time.Minute fetchTimeout = time.Minute
// mappedBits is the length of ::ffff:0.0.0.0/96, the netblock of every
// IPv4-mapped address.
mappedBits = 96
) )
var ( var (
@@ -42,15 +39,13 @@ var (
"is not an AS number, : and a percentage, such as AS64496:50") "is not an AS number, : and a percentage, such as AS64496:50")
) )
// List is a list as reputation.json holds it: the URL it is fetched from, // List is the last good copy of a list, as reputation.json holds it: the
// when it was last tried, the fetch failed or not, and its last good copy: // URL it was fetched from, when it was fetched, and its lines, as fetched,
// when that was fetched, and its lines, as fetched, comment lines // comment lines included.
// included, both left out while no fetch of it has succeeded.
type List struct { type List struct {
URL string `json:"url"` URL string `json:"url"`
Tried time.Time `json:"tried"` Fetched time.Time `json:"fetched"`
Fetched time.Time `json:"fetched,omitzero"` Lines []string `json:"lines"`
Lines []string `json:"lines,omitzero"`
} }
// Params are what New needs. // Params are what New needs.
@@ -82,12 +77,13 @@ type Lists struct {
lists map[string]*list lists map[string]*list
} }
// list is one list: what reputation.json keeps of it, its last try, zero // list is one list: its last good copy, what that says, when the list was
// before the first, and its last good copy, what that copy says, and how // last fetched or tried, zero before the first try, and how many fetches
// many fetches of it failed. // of it failed.
type list struct { type list struct {
kept List kept List
entries entries entries entries
tried time.Time
failures int failures int
} }
@@ -105,7 +101,7 @@ func New(params Params) *Lists {
l := &Lists{params: params, httpClient: &http.Client{}, lists: map[string]*list{}} l := &Lists{params: params, httpClient: &http.Client{}, lists: map[string]*list{}}
for _, listURL := range l.URLs() { for _, listURL := range l.URLs() {
l.lists[listURL] = &list{kept: List{URL: listURL}} l.lists[listURL] = &list{}
} }
return l return l
@@ -172,9 +168,9 @@ func (l *Lists) Failures(listURL string) int {
} }
// Run fetches each list once Refresh has passed since it was last fetched // Run fetches each list once Refresh has passed since it was last fetched
// or tried, the later of the two, until ctx is done. A list never tried is // or tried, the later of the two, until ctx is done. A list without a copy
// fetched at once, and so is one whose last try or copy, read from // has never been fetched, so it is fetched at once, and so is one whose
// reputation.json, is that old. // copy, read from reputation.json, is that old.
func (l *Lists) Run(ctx context.Context) { func (l *Lists) Run(ctx context.Context) {
if len(l.lists) == 0 { if len(l.lists) == 0 {
return return
@@ -193,35 +189,35 @@ func (l *Lists) Run(ctx context.Context) {
} }
} }
// Snapshot returns each list that has been tried, with its copy, if it // Snapshot returns the copy of each list that has one, sorted by URL, as
// has one, sorted by URL, as reputation.json lists them. // reputation.json lists them.
func (l *Lists) Snapshot() []List { func (l *Lists) Snapshot() []List {
l.mu.Lock() l.mu.Lock()
tried := make([]List, 0, len(l.lists)) copies := make([]List, 0, len(l.lists))
for _, held := range l.lists { for _, held := range l.lists {
if !held.kept.Tried.IsZero() { if !held.kept.Fetched.IsZero() {
tried = append(tried, held.kept) copies = append(copies, held.kept)
} }
} }
l.mu.Unlock() l.mu.Unlock()
slices.SortFunc(tried, func(a, b List) int { slices.SortFunc(copies, func(a, b List) int {
return strings.Compare(a.URL, b.URL) return strings.Compare(a.URL, b.URL)
}) })
return tried return copies
} }
// Load puts lists, read from reputation.json, in place of the last tries // Load puts copies, read from reputation.json, in place of the copies
// and copies held. A list Params does not name is dropped. A copy with a // held. A copy of a list Params does not name is dropped. A copy with a
// line that parse refuses is an error, and then nothing changes. // line that parse refuses is an error, and then nothing changes.
func (l *Lists) Load(lists []List) error { func (l *Lists) Load(copies []List) error {
found := make(map[string]entries, len(lists)) found := make(map[string]entries, len(copies))
for _, kept := range lists { for _, kept := range copies {
if _, named := l.lists[kept.URL]; !named { if _, named := l.lists[kept.URL]; !named {
continue continue
} }
@@ -237,11 +233,11 @@ func (l *Lists) Load(lists []List) error {
l.mu.Lock() l.mu.Lock()
defer l.mu.Unlock() defer l.mu.Unlock()
for listURL, held := range l.lists { for _, held := range l.lists {
held.kept, held.entries = List{URL: listURL}, entries{} held.kept, held.entries = List{}, entries{}
} }
for _, kept := range lists { for _, kept := range copies {
read, named := found[kept.URL] read, named := found[kept.URL]
if named { if named {
l.lists[kept.URL].kept, l.lists[kept.URL].entries = kept, read l.lists[kept.URL].kept, l.lists[kept.URL].entries = kept, read
@@ -280,8 +276,8 @@ func (l *Lists) due(listURL string) time.Time {
held := l.lists[listURL] held := l.lists[listURL]
last := held.kept.Fetched last := held.kept.Fetched
if held.kept.Tried.After(last) { if held.tried.After(last) {
last = held.kept.Tried last = held.tried
} }
return last.Add(l.params.Refresh) return last.Add(l.params.Refresh)
@@ -308,10 +304,10 @@ func (l *Lists) fetch(ctx context.Context, listURL string) {
l.mu.Lock() l.mu.Lock()
held := l.lists[listURL] held := l.lists[listURL]
held.kept.Tried = now held.tried = now
if err == nil { if err == nil {
held.kept.Fetched, held.kept.Lines = now, lines held.kept = List{URL: listURL, Fetched: now, Lines: lines}
held.entries = found held.entries = found
} else { } else {
held.failures++ held.failures++
@@ -403,8 +399,8 @@ func parseNetblocks(lines []string) (entries, error) {
continue continue
} }
netblock, ok := parseNetblock(text) netblock, err := config.ParseNetblock(text)
if !ok { if err != nil {
return entries{}, fmt.Errorf("line %d %w", i+1, errNotNetblock) return entries{}, fmt.Errorf("line %d %w", i+1, errNotNetblock)
} }
@@ -418,32 +414,6 @@ func parseNetblocks(lines []string) (entries, error) {
return found, nil return found, nil
} }
// parseNetblock reads text, a line of a blocklist, and reports whether it
// is an address or a netblock as the settings take them. A client's IPv4
// address is checked as IPv4, never IPv4-mapped, so an IPv4-mapped line,
// such as ::ffff:192.0.2.0/120, is read as the IPv4 address or netblock it
// stands for, 192.0.2.0/24, and a mapped netblock shorter than /96, which
// stands for none, is refused.
func parseNetblock(text string) (netip.Prefix, bool) {
netblock, err := config.ParseNetblock(text)
if err != nil {
return netip.Prefix{}, false
}
// The address as written: ParseNetblock's has the bits past the
// netblock's length cleared, the ::ffff among them below /96.
written, _, _ := strings.Cut(text, "/")
if addr, _ := netip.ParseAddr(written); !addr.Is4In6() {
return netblock, true
}
if netblock.Bits() < mappedBits {
return netip.Prefix{}, false
}
return netip.PrefixFrom(netblock.Addr().Unmap(), netblock.Bits()-mappedBits), true
}
// parsePercents reads the lines of the file of AS:percent lines, each an // parsePercents reads the lines of the file of AS:percent lines, each an
// AS number, : and a percentage, as SWWAF_ASN_LIMIT_PERCENT takes them. An // AS number, : and a percentage, as SWWAF_ASN_LIMIT_PERCENT takes them. An
// AS number listed more than once gets the lowest of its percentages. // AS number listed more than once gets the lowest of its percentages.
+6 -83
View File
@@ -72,41 +72,6 @@ func TestListedAddressesAndNetblocksWithTheCommentsLeftOut(t *testing.T) {
}) })
} }
func TestIPv4MappedLineListsTheIPv4AddressOrNetblockItStandsFor(t *testing.T) {
t.Parallel()
now := time.Now()
lists := reputation.New(params(dropURL))
err := lists.Load([]reputation.List{{
URL: dropURL, Tried: now, Fetched: now,
Lines: []string{"::ffff:192.0.2.9", "::ffff:203.0.113.0/120"},
}})
if err != nil {
t.Fatalf("load: %v", err)
}
for addr, want := range map[string][]string{
"192.0.2.9": {dropURL},
"203.0.113.255": {dropURL},
"192.0.2.8": nil,
"203.0.114.0": nil,
} {
wantListedBy(t, lists, addr, want...)
}
// A mapped netblock shorter than /96 stands for no IPv4 one.
err = lists.Load([]reputation.List{{
URL: dropURL, Tried: now, Fetched: now, Lines: []string{"::ffff:198.51.100.0/88"},
}})
const want = "the copy of " + dropURL +
": line 1 is not an address or a netblock, such as 192.0.2.0/24"
if err == nil || err.Error() != want {
t.Errorf("error %v, want %s", err, want)
}
}
func TestClientIsListedByEachBlocklistThatListsIt(t *testing.T) { func TestClientIsListedByEachBlocklistThatListsIt(t *testing.T) {
t.Parallel() t.Parallel()
@@ -202,11 +167,8 @@ func TestFailedFetchKeepsTheLastGoodCopyAndAlertsOncePerCooldown(t *testing.T) {
wantFetches(t, servers, 3) wantFetches(t, servers, 3)
wantListedBy(t, lists, "203.0.113.9", dropURL) wantListedBy(t, lists, "203.0.113.9", dropURL)
want := kept[0] if got := lists.Snapshot(); !reflect.DeepEqual(got, kept) {
want.Tried = time.Now() t.Errorf("copies %+v, want the first %+v", got, kept)
if got := lists.Snapshot(); !reflect.DeepEqual(got, []reputation.List{want}) {
t.Errorf("lists %+v, want the first copy, last tried now, %+v", got, want)
} }
if lists.Failures(dropURL) != 2 { if lists.Failures(dropURL) != 2 {
@@ -289,43 +251,6 @@ func TestKeptCopyIsFetchedAgainOnceRefreshHasPassedSinceItWasFetched(t *testing.
}) })
} }
func TestRestartWaitsRefreshAfterTheLastTryEvenOneThatFailed(t *testing.T) {
t.Parallel()
synctest.Test(t, func(t *testing.T) {
// drop.txt is fetched, and a refresh later the fetch downloads it
// whole but fails on a line that does not read.
servers := &standIn{lists: map[string]string{dropURL: "198.51.100.1\n"}}
lists := start(t, servers, params(dropURL))
servers.set(dropURL, "198.51.100.2\n<html>\n")
time.Sleep(refresh)
wantFetches(t, servers, 2)
// Restarted with what reputation.json keeps, it waits a refresh
// after the failed try, as it does while it runs.
restarted := &standIn{lists: map[string]string{dropURL: "198.51.100.2\n"}}
again := reputation.New(params(dropURL))
again.SetTransport(restarted)
err := again.Load(lists.Snapshot())
if err != nil {
t.Fatalf("load: %v", err)
}
run(t, again)
wantFetches(t, restarted, 0)
wantListedBy(t, again, "198.51.100.1", dropURL)
time.Sleep(refresh - time.Nanosecond)
wantFetches(t, restarted, 0)
time.Sleep(time.Nanosecond)
wantFetches(t, restarted, 1)
wantListedBy(t, again, "198.51.100.2", dropURL)
})
}
func TestFetchCutOffAsItStopsIsNoFailure(t *testing.T) { func TestFetchCutOffAsItStopsIsNoFailure(t *testing.T) {
t.Parallel() t.Parallel()
@@ -395,14 +320,12 @@ func TestLoadDropsCopiesOfListsNotNamedAndRefusesOnesThatDoNotRead(t *testing.T)
t.Parallel() t.Parallel()
fetched := time.Date(2026, 10, 6, 0, 0, 0, 0, time.UTC) fetched := time.Date(2026, 10, 6, 0, 0, 0, 0, time.UTC)
kept := reputation.List{ kept := reputation.List{URL: dropURL, Fetched: fetched, Lines: []string{"203.0.113.9"}}
URL: dropURL, Tried: fetched, Fetched: fetched, Lines: []string{"203.0.113.9"},
}
lists := reputation.New(params(dropURL)) lists := reputation.New(params(dropURL))
err := lists.Load([]reputation.List{kept, { err := lists.Load([]reputation.List{
URL: torURL, Tried: fetched, Fetched: fetched, Lines: []string{"198.51.100.9"}, kept, {URL: torURL, Fetched: fetched, Lines: []string{"198.51.100.9"}},
}}) })
if err != nil { if err != nil {
t.Fatalf("load: %v", err) t.Fatalf("load: %v", err)
} }
+21 -34
View File
@@ -574,39 +574,31 @@ func TestLookupDatabaseReplacedWhileRunningTakesEffect(t *testing.T) {
} }
} }
func TestBlocklistTriesAndCopiesKeptInReputationJSONAcrossRestarts(t *testing.T) { func TestBlocklistCopyKeptInReputationJSONAcrossARestart(t *testing.T) {
t.Parallel() t.Parallel()
const ( const (
token = "0123456789abcdef0123456789abcdef" token = "0123456789abcdef0123456789abcdef"
torPath = "/tor.txt" // drop is the blocklist, with the date and copyright line the
dropPath = "/drop.txt" // Spamhaus DROP list starts with.
// copyright is the DROP list's date and copyright line.
copyright = "; Spamhaus DROP List 2026/10/07 - (c) 2026 The Spamhaus Project SLL" copyright = "; Spamhaus DROP List 2026/10/07 - (c) 2026 The Spamhaus Project SLL"
drop = copyright + "\n" + placed + " ; SBL1\n"
) )
failing, torFetches := new(atomic.Bool), new(atomic.Int32) var failing atomic.Bool
lists := map[string]string{
torPath: "198.51.100.0/24\n",
dropPath: copyright + "\n" + placed + " ; SBL1\n",
}
server := httptest.NewServer(http.HandlerFunc(
func(w http.ResponseWriter, r *http.Request) {
if r.URL.Path == torPath {
torFetches.Add(1)
}
list := httptest.NewServer(http.HandlerFunc(
func(w http.ResponseWriter, _ *http.Request) {
if failing.Load() { if failing.Load() {
w.WriteHeader(http.StatusServiceUnavailable) w.WriteHeader(http.StatusServiceUnavailable)
return return
} }
_, _ = io.WriteString(w, lists[r.URL.Path]) _, _ = io.WriteString(w, drop)
})) }))
t.Cleanup(server.Close) t.Cleanup(list.Close)
torURL, dropURL := server.URL+torPath, server.URL+dropPath
dir := t.TempDir() dir := t.TempDir()
env := map[string]string{ env := map[string]string{
listenAddr: localhost + ":0", listenAddr: localhost + ":0",
@@ -616,48 +608,43 @@ func TestBlocklistTriesAndCopiesKeptInReputationJSONAcrossRestarts(t *testing.T)
trustedProxies: localhost + "/32", trustedProxies: localhost + "/32",
metricsToken: token, metricsToken: token,
instanceName: instance, instanceName: instance,
"SWWAF_BLOCKLIST_URLS": torURL, "SWWAF_BLOCKLIST_URLS": list.URL,
// The requests sent until the list takes effect, and those for the // The requests sent until the list takes effect, and those for the
// metrics, must not break a rate limit, whose ban would refuse them // metrics, must not break a rate limit, whose ban would refuse them
// too. // too.
rateLimitExemptNets: placed + "," + localhost, rateLimitExemptNets: placed + "," + localhost,
} }
failures := `smallwebwaf_reputation_failures_total{instance="fsn1app1/gitea",` + failures := `smallwebwaf_reputation_failures_total{instance="fsn1app1/gitea",` +
`source="` + torURL + `"} ` `source="` + list.URL + `"} `
// tor.txt cannot be fetched at first, which is counted. // The list cannot be fetched at first, which leaves the client let
// through, and is counted.
failing.Store(true) failing.Store(true)
runUntilStopped(t, env, func(url string) { runUntilStopped(t, env, func(url string) {
metricsWith(t, url+"_smallwebwaf/metrics", token, failures+"1") metricsWith(t, url+"_smallwebwaf/metrics", token, failures+"1")
wantStatus(t, url, placed, http.StatusOK)
}) })
// Restarted with drop.txt named after it and the server answering, tor.txt // With no copy kept, it is fetched as smallwebwaf starts again, and from
// waits SWWAF_BLOCKLIST_REFRESH after its failed try, kept in reputation.json, // then on the list refuses the client.
// while drop.txt, never tried, is fetched at once. Lists are fetched in the
// order named, so once drop.txt refuses the client, tor.txt has had its turn.
failing.Store(false) failing.Store(false)
env["SWWAF_BLOCKLIST_URLS"] = torURL + "," + dropURL
out := runUntilStopped(t, env, func(url string) { out := runUntilStopped(t, env, func(url string) {
for statusFrom(t, url, placed) != http.StatusForbidden { for statusFrom(t, url, placed) != http.StatusForbidden {
time.Sleep(pollInterval) time.Sleep(pollInterval)
} }
}) })
wantDeniedByList(t, out.line(t, "action", "denied"), dropURL) wantDeniedByList(t, out.line(t, "action", "denied"), list.URL)
if fetches := torFetches.Load(); fetches != 1 { // After another restart, with the list's server failing, the copy kept
t.Errorf("tor.txt fetched %d times, want once, before the restart", fetches) // in reputation.json, its copyright line included, refuses the client
} // from the first request.
// After another restart, with the server failing, the copy of drop.txt
// kept in reputation.json, its copyright line included, refuses the
// client from the first request.
failing.Store(true) failing.Store(true)
out = runUntilStopped(t, env, func(url string) { out = runUntilStopped(t, env, func(url string) {
wantStatus(t, url, placed, http.StatusForbidden) wantStatus(t, url, placed, http.StatusForbidden)
}) })
wantDeniedByList(t, out.line(t, "type", "request"), dropURL) wantDeniedByList(t, out.line(t, "type", "request"), list.URL)
path := filepath.Join(dir, "reputation.json") path := filepath.Join(dir, "reputation.json")
+9 -12
View File
@@ -1,10 +1,10 @@
// Package state keeps smallwebwaf's state in JSON files in // Package state keeps smallwebwaf's state in JSON files in
// SWWAF_STATE_DIR, as the "Persistent state" section of SPEC.md describes: // SWWAF_STATE_DIR, as the "Persistent state" section of SPEC.md describes:
// bans.json holds the bans, clients.json each client's counters and // bans.json holds the bans, clients.json each client's counters and
// history, lookups.json GeoJS's answers, reputation.json the last try and // history, lookups.json GeoJS's answers, reputation.json the last good
// last good copy of each list fetched from a URL, and alerts.json the // copy of each list fetched from a URL, and alerts.json the cooldowns, the
// cooldowns, the hour under way, the alerts waiting for each destination // hour under way, the alerts waiting for each destination and the anomaly
// and the anomaly counters. Load // counters. Load
// reads them at start, Watch takes in an admin's edit of one while // reads them at start, Watch takes in an admin's edit of one while
// smallwebwaf runs, and Run and WriteAll write them. The disk is read and // smallwebwaf runs, and Run and WriteAll write them. The disk is read and
// written outside the parts' locks, which are held only to take a // written outside the parts' locks, which are held only to take a
@@ -726,20 +726,17 @@ func (f *lookupsFile) check(data []byte) error {
return nil return nil
} }
// check refuses a list without its URL, which would name no list, or the // check refuses a list's copy without its URL, which would name no list,
// time it was last tried, which would have it fetched at once, and a copy // the time it was fetched, which would have the list fetched at once, or
// of it without the time it was fetched, or without its lines, which hold // its lines, which hold the list.
// the list.
func (f *reputationFile) check([]byte) error { func (f *reputationFile) check([]byte) error {
for i, kept := range f.Lists { for i, kept := range f.Lists {
switch { switch {
case kept.URL == "": case kept.URL == "":
return missing(i, "url") return missing(i, "url")
case kept.Tried.IsZero(): case kept.Fetched.IsZero():
return missing(i, "tried")
case kept.Fetched.IsZero() && kept.Lines != nil:
return missing(i, "fetched") return missing(i, "fetched")
case kept.Lines == nil && !kept.Fetched.IsZero(): case kept.Lines == nil:
return missing(i, "lines") return missing(i, "lines")
} }
} }
+19 -38
View File
@@ -38,9 +38,8 @@ const (
lookupsJSON = "lookups.json" lookupsJSON = "lookups.json"
reputationJSON = "reputation.json" reputationJSON = "reputation.json"
alertsJSON = "alerts.json" alertsJSON = "alerts.json"
// blocklistURL and torURL are the blocklists the tests' lists name. // blocklistURL is the blocklist the tests' lists name.
blocklistURL = "https://lists.example/drop.txt" blocklistURL = "https://lists.example/drop.txt"
torURL = "https://lists.example/tor.txt"
// The AS number and AS name the tests' clients are looked up in. // The AS number and AS name the tests' clients are looked up in.
asn = "AS64496" asn = "AS64496"
asName = "Example Net" asName = "Example Net"
@@ -209,24 +208,19 @@ const filledAlertsJSON = `{
} }
` `
// filledReputationJSON is reputation.json holding the blocklists' last // filledReputationJSON is reputation.json holding the copy of the
// tries and the copy of one, with its comment line, as fill puts them in. // blocklist fill puts in, with its comment line.
const filledReputationJSON = `{ const filledReputationJSON = `{
"version": 1, "version": 1,
"lists": [ "lists": [
{ {
"url": "https://lists.example/drop.txt", "url": "https://lists.example/drop.txt",
"tried": "2026-10-06T00:00:00Z",
"fetched": "2026-10-05T23:00:00Z", "fetched": "2026-10-05T23:00:00Z",
"lines": [ "lines": [
"; Spamhaus DROP List 2026/10/05 - (c) 2026 The Spamhaus Project SLL", "; Spamhaus DROP List 2026/10/05 - (c) 2026 The Spamhaus Project SLL",
"203.0.113.0/24 ; SBL1", "203.0.113.0/24 ; SBL1",
"2001:db8::/32 ; SBL2" "2001:db8::/32 ; SBL2"
] ]
},
{
"url": "https://lists.example/tor.txt",
"tried": "2026-10-06T00:00:00Z"
} }
] ]
} }
@@ -485,8 +479,7 @@ func TestFileThatDoesNotParseStopsTheStart(t *testing.T) {
{ {
"a copy of a list with a line that does not read", reputationJSON, "a copy of a list with a line that does not read", reputationJSON,
`{"version": 1, "lists": [{"url": "` + blocklistURL + `", ` + `{"version": 1, "lists": [{"url": "` + blocklistURL + `", ` +
`"tried": "2026-10-06T00:00:00Z", "fetched": "2026-10-06T00:00:00Z", ` + `"fetched": "2026-10-06T00:00:00Z", "lines": ["; DROP", "203.0.113.300"]}]}`,
`"lines": ["; DROP", "203.0.113.300"]}]}`,
`: the copy of ` + blocklistURL + `: line 2 is not an address or a netblock, ` + `: the copy of ` + blocklistURL + `: line 2 is not an address or a netblock, ` +
`such as 192.0.2.0/24`, `such as 192.0.2.0/24`,
}, },
@@ -588,11 +581,7 @@ func TestEntryWithoutAFieldItNeedsStopsTheStart(t *testing.T) {
func TestReputationJSONEntryWithoutAFieldItNeedsStopsTheStart(t *testing.T) { func TestReputationJSONEntryWithoutAFieldItNeedsStopsTheStart(t *testing.T) {
t.Parallel() t.Parallel()
const ( const fetched = `"fetched": "2026-10-06T00:00:00Z"`
drop = `"url": "` + blocklistURL + `", `
tried = `"tried": "2026-10-06T00:00:00Z", `
fetched = `"fetched": "2026-10-06T00:00:00Z"`
)
for _, tc := range []struct { for _, tc := range []struct {
name, content string name, content string
@@ -600,25 +589,20 @@ func TestReputationJSONEntryWithoutAFieldItNeedsStopsTheStart(t *testing.T) {
want string want string
}{ }{
{ {
"a list without its URL", "a copy of a list without its URL",
`{"version": 1, "lists": [{` + tried + fetched + `, "lines": []}]}`, `{"version": 1, "lists": [{` + fetched + `, "lines": []}]}`,
`: entry 1 has no "url"`, `: entry 1 has no "url"`,
}, },
{
"a list without the time it was last tried",
`{"version": 1, "lists": [{` + drop + fetched + `, "lines": []}]}`,
`: entry 1 has no "tried"`,
},
{ {
"a copy of a list without the time it was fetched", "a copy of a list without the time it was fetched",
`{"version": 1, "lists": [{` + drop + tried + `"lines": []}]}`, `{"version": 1, "lists": [{"url": "` + blocklistURL + `", "lines": []}]}`,
`: entry 1 has no "fetched"`, `: entry 1 has no "fetched"`,
}, },
{ {
// An empty list has no lines, which is not having none. // An empty list has no lines, which is not having none.
"a copy of a list without its lines", "a copy of a list without its lines",
`{"version": 1, "lists": [{` + drop + tried + fetched + `, "lines": []}, ` + `{"version": 1, "lists": [{"url": "` + blocklistURL + `", ` + fetched +
`{"url": "` + torURL + `", ` + tried + fetched + `}]}`, `, "lines": []}, {"url": "https://lists.example/tor.txt", ` + fetched + `}]}`,
`: entry 2 has no "lines"`, `: entry 2 has no "lines"`,
}, },
} { } {
@@ -1132,8 +1116,7 @@ func TestEditOfEachFileTakenIn(t *testing.T) {
[]lookup.Answer{{Client: client, Country: "FR", Answered: midnight()}}) []lookup.Answer{{Client: client, Country: "FR", Answered: midnight()}})
edit(t, dir, reputationJSON, `{"version": 1, "lists": [{"url": "`+blocklistURL+`", `+ edit(t, dir, reputationJSON, `{"version": 1, "lists": [{"url": "`+blocklistURL+`", `+
`"tried": "2026-10-06T00:00:00Z", "fetched": "2026-10-06T00:00:00Z", `+ `"fetched": "2026-10-06T00:00:00Z", "lines": ["198.51.100.7"]}]}`)
`"lines": ["198.51.100.7"]}]}`)
wantTakenIn(t, lines, dir, reputationJSON) wantTakenIn(t, lines, dir, reputationJSON)
listedBy := params.Lists.ListedBy(client.Addr()) listedBy := params.Lists.ListedBy(client.Addr())
@@ -1523,8 +1506,8 @@ func midnight() time.Time {
} }
// newParams returns Params for the state files in dir, with parts that // newParams returns Params for the state files in dir, with parts that
// hold nothing yet. GeoJS is never asked, the lists, two blocklists, are // hold nothing yet. GeoJS is never asked, the lists, of which there is one
// never fetched, and the alerts, at most two an hour, are // blocklist, are never fetched, and the alerts, at most two an hour, are
// never sent. The anomaly counters count the scopes fill counts, with // never sent. The anomaly counters count the scopes fill counts, with
// thresholds fill does not reach. // thresholds fill does not reach.
func newParams(dir string) state.Params { func newParams(dir string) state.Params {
@@ -1555,7 +1538,7 @@ func newParams(dir string) state.Params {
Now: midnight, ProcessLog: discard, Metrics: m, Now: midnight, ProcessLog: discard, Metrics: m,
}), }),
Lists: reputation.New(reputation.Params{ Lists: reputation.New(reputation.Params{
BlocklistURLs: []string{blocklistURL, torURL}, Refresh: 24 * time.Hour, BlocklistURLs: []string{blocklistURL}, Refresh: 24 * time.Hour,
Now: midnight, ProcessLog: discard, Alerts: queue, Now: midnight, ProcessLog: discard, Alerts: queue,
}), }),
Alerts: queue, Alerts: queue,
@@ -1582,9 +1565,9 @@ func office() netip.Prefix {
// fill puts a permanent ban an admin made, a ban for a broken limit and // fill puts a permanent ban an admin made, a ban for a broken limit and
// one for a clear sign of attack, clients with counts and histories, // one for a clear sign of attack, clients with counts and histories,
// GeoJS answers, the blocklists' last tries and the copy of one, as // GeoJS answers, a copy of the blocklist, as filledReputationJSON holds
// filledReputationJSON holds them, and alerts and anomaly counters, as // it, and alerts and anomaly counters, as filledAlertsJSON holds them,
// filledAlertsJSON holds them, into the parts of params. // into the parts of params.
func fill(params state.Params) { func fill(params state.Params) {
now := midnight() now := midnight()
client := netip.MustParsePrefix("203.0.113.9/32") client := netip.MustParsePrefix("203.0.113.9/32")
@@ -1617,16 +1600,14 @@ func fill(params state.Params) {
}, },
}) })
// The copy of drop.txt was fetched an hour ago, and the fetches of it
// and of tor.txt tried since failed.
err := params.Lists.Load([]reputation.List{{ err := params.Lists.Load([]reputation.List{{
URL: blocklistURL, Tried: now, Fetched: now.Add(-time.Hour), URL: blocklistURL, Fetched: now.Add(-time.Hour),
Lines: []string{ Lines: []string{
"; Spamhaus DROP List 2026/10/05 - (c) 2026 The Spamhaus Project SLL", "; Spamhaus DROP List 2026/10/05 - (c) 2026 The Spamhaus Project SLL",
"203.0.113.0/24 ; SBL1", "203.0.113.0/24 ; SBL1",
"2001:db8::/32 ; SBL2", "2001:db8::/32 ; SBL2",
}, },
}, {URL: torURL, Tried: now}}) }})
if err != nil { if err != nil {
panic(err) // the copy reads panic(err) // the copy reads
} }