diff --git a/README.md b/README.md index 7740a42..2dd9256 100644 --- a/README.md +++ b/README.md @@ -26,29 +26,33 @@ Slack and ntfy, and remote log sending. So is the stage after that: the AS number and country of every client, looked up through GeoJS or in the IPinfo Lite database file, the byte limits, the biased thresholds, lower limits for the AS numbers and countries you list, and the anomaly thresholds, alerts for -unusual traffic that refuse nothing. `smallwebwaf` passes each request to the -app and the app's answer back, unchanged, within its timeouts and size limits, -works out each client's address, looks up its AS number and country unless you -switch that off, bans a client that sends too many requests or too many bytes, -not counting those for the paths you choose, with lower limits for the clients -of the AS numbers and countries you list, refuses a client that comes from a -country you refuse or from a network you refuse, lets the networks you choose +unusual traffic that refuse nothing. So is the first part of the stage after +that: the blocklists you name by URL, which it fetches and keeps, and a file of +AS numbers' percentages, fetched the same way. `smallwebwaf` passes each request +to the app and the app's answer back, unchanged, within its timeouts and size +limits, works out each client's address, looks up its AS number and country +unless you switch that off, bans a client that sends too many requests or too +many bytes, not counting those for the paths you choose, with lower limits for +the clients of the AS numbers and countries you list, refuses a client that +comes from a country you refuse or from a network you refuse, refuses, limits or +only notes a client a blocklist you name lists, lets the networks you choose through, checks each request against the rule files and bans a client whose request is a clear sign of attack, keeps its bans, each client's counters and -history, and GeoJS's answers in JSON files across restarts, takes in your edits -of those files, such as a ban you make, keep or lift, and of the rule files -while it runs, writes a JSON log line for every request, sends its log lines to -a syslog server too if you name one, sends an alert to a webhook, to Slack and -to ntfy, each if you name one, for each ban it makes or makes permanent, for -traffic over an anomaly threshold you set, for GeoJS failing, for a rule file or -state file with an error and for a replacement of the lookup database it cannot -read, serves Prometheus metrics to a scraper that holds the metrics token, lets -an admin who holds the admin token list, add and lift bans and ask what it knows -of a client, and in `observe` mode passes on the requests it would refuse, -logging what it would have done with them. It comes as the image the app's own -image is built on. The rest of the design comes after that, in the order of the -build order in [`SPEC.md`](SPEC.md). The survey of existing tools that led to -the design is in [`EVALUATION.md`](EVALUATION.md). +history, GeoJS's answers and the last good copy of each list it fetches in JSON +files across restarts, takes in your edits of those files, such as a ban you +make, keep or lift, and of the rule files while it runs, writes a JSON log line +for every request, sends its log lines to a syslog server too if you name one, +sends an alert to a webhook, to Slack and to ntfy, each if you name one, for +each ban it makes or makes permanent, for traffic over an anomaly threshold you +set, for a client a blocklist lists, for GeoJS failing or a list it cannot +fetch, for a rule file or state file with an error and for a replacement of the +lookup database it cannot read, serves Prometheus metrics to a scraper that +holds the metrics token, lets an admin who holds the admin token list, add and +lift bans and ask what it knows of a client, and in `observe` mode passes on the +requests it would refuse, logging what it would have done with them. It comes as +the image the app's own image is built on. The rest of the design comes after +that, in the order of the build order in [`SPEC.md`](SPEC.md). The survey of +existing tools that led to the design is in [`EVALUATION.md`](EVALUATION.md). ## Getting started @@ -124,18 +128,20 @@ in `bin/state` unless `SWWAF_STATE_DIR` is set, and the default rule file of or `SWWAF_RATE_LIMIT_EXEMPT_NETS`, and a request for a path `SWWAF_RATE_LIMIT_EXEMPT_PATHS` exempts. - Gives the clients of the AS numbers and countries the biased thresholds list, - `SWWAF_ASN_LIMIT_PERCENT` and `SWWAF_COUNTRY_LIMIT_PERCENT`, the percentage - they give of every rate limit and byte limit, so that the same rules ban them - after fewer requests, and, while `SWWAF_UNKNOWN_LIMIT_PERCENT` is below 100, - every client without a country that percentage. A client to which several - apply gets the lowest. `SWWAF_ASN_BYTES_PERCENT` and - `SWWAF_COUNTRY_BYTES_PERCENT` give the AS numbers and countries they list a - percentage of the byte limits in place of the other two. Each client is - counted on its own, against its own lowered limits: no budget is shared by a - whole AS number or country, which one abuser could use up and so lock out - everyone else there. The log line of each request the rate limits count gives - its client's percentages below 100 and the settings that gave them, and so do - the notes of a ban for a lowered limit, and its alert. + `SWWAF_ASN_LIMIT_PERCENT`, the file `SWWAF_ASN_LIMIT_PERCENT_URL` names and + `SWWAF_COUNTRY_LIMIT_PERCENT`, the percentage they give of every rate limit + and byte limit, so that the same rules ban them after fewer requests, and, + while `SWWAF_UNKNOWN_LIMIT_PERCENT` is below 100, every client without a + country that percentage. A client to which several apply gets the lowest. For + the byte limits, `SWWAF_ASN_BYTES_PERCENT` gives an AS number it lists a + percentage in place of those `SWWAF_ASN_LIMIT_PERCENT` and the file give it, + and `SWWAF_COUNTRY_BYTES_PERCENT` gives a country it lists one in place of the + one `SWWAF_COUNTRY_LIMIT_PERCENT` gives it. Each client is counted on its own, + against its own lowered limits: no budget is shared by a whole AS number or + country, which one abuser could use up and so lock out everyone else there. + The log line of each request the rate limits count gives its client's + percentages below 100 and the settings that gave them, and so do the notes of + a ban for a lowered limit, and its alert. - Bans a client that breaks a rate limit or a byte limit, as "Bans" in [`SPEC.md`](SPEC.md) describes: the first ban lasts an hour, and a limit broken again within a day of a ban ending bans for three times as long as that @@ -191,34 +197,43 @@ in `bin/state` unless `SWWAF_STATE_DIR` is set, and the default rule file of link-local address has no country and is never looked up: `SWWAF_EXCLUSIVELY_ALLOWED_COUNTRIES` refuses it unless it is in `SWWAF_ALLOW_NETS`, and `SWWAF_DENIED_COUNTRIES` does not refuse it. +- Checks the client's own address against the blocklists `SWWAF_BLOCKLIST_URLS` + names, after the country lists and before the rate limits (see "Blocklists" + below). With `SWWAF_BLOCKLIST_ACTION` at `deny`, its default, a request from a + client a blocklist lists is refused with `SWWAF_BAN_RESPONSE` before its body + is read; it is not counted for the rate limits, and makes no ban. With + `limit:` the client gets that percentage of every rate limit and byte + limit, the lowest of its percentages applying, as for a biased threshold, and + with `log` nothing more is done. Whatever the action, the request's log line + names the lists, and each raises an alert. - Checks the client's own address against the static lists, the three netblock settings below, before anything else, its lookup included. A client in - `SWWAF_ALLOW_NETS` skips bans, the country lists, the rate limits, the byte - limits and the rule files, and is not looked up; the timeouts and size limits - still apply. A client in `SWWAF_DENY_NETS` is refused with + `SWWAF_ALLOW_NETS` skips bans, the country lists, the blocklists, the rate + limits, the byte limits and the rule files, and is not looked up; the timeouts + and size limits still apply. A client in `SWWAF_DENY_NETS` is refused with `SWWAF_BAN_RESPONSE` before its body is read, and the request is not counted for the rate limits; an address in `SWWAF_ALLOW_NETS` too is let through. A client in `SWWAF_RATE_LIMIT_EXEMPT_NETS` is neither counted nor refused by the rate limits, and has no bytes counted by the byte limits; the country lists, the rule files and bans still apply to it. - In `observe` mode, with `SWWAF_MODE=observe`, refuses none of the requests - that `SWWAF_DENY_NETS`, a ban, the country lists, a rate limit or a rule would - refuse: it passes them to the app, and their log lines name what `enforce` - mode would have done (see `would_action` in "Request log" below). The checks - run, and requests and bytes are counted, as in `enforce` mode, with three - differences: neither a broken rate limit or byte limit nor a `ban` rule makes - a ban; a broken limit does not set the client's counters back to zero, so each - request over a rate limit is logged as one that would be refused, and each - whose bytes keep the client over a byte limit as breaking it; and a request - under a ban does not make it permanent. As in `enforce` mode, the bytes - counted are only those of the requests `enforce` mode would have passed to the - app. A ban it would have made, or made permanent, raises the alert `enforce` - mode would have raised, marked as what would have happened (see "Alerts" - below). The bans in `bans.json` are kept, and refuse requests again when - `smallwebwaf` next runs in `enforce` mode, as long as they last. The timeouts - and size limits still apply, since they protect `smallwebwaf` and the app - themselves, and a request for one of `smallwebwaf`'s own endpoints without its - token is still answered `401`. It is for trying a configuration before + that `SWWAF_DENY_NETS`, a ban, the country lists, a blocklist, a rate limit or + a rule would refuse: it passes them to the app, and their log lines name what + `enforce` mode would have done (see `would_action` in "Request log" below). + The checks run, and requests and bytes are counted, as in `enforce` mode, with + three differences: neither a broken rate limit or byte limit nor a `ban` rule + makes a ban; a broken limit does not set the client's counters back to zero, + so each request over a rate limit is logged as one that would be refused, and + each whose bytes keep the client over a byte limit as breaking it; and a + request under a ban does not make it permanent. As in `enforce` mode, the + bytes counted are only those of the requests `enforce` mode would have passed + to the app. A ban it would have made, or made permanent, raises the alert + `enforce` mode would have raised, marked as what would have happened (see + "Alerts" below). The bans in `bans.json` are kept, and refuse requests again + when `smallwebwaf` next runs in `enforce` mode, as long as they last. The + timeouts and size limits still apply, since they protect `smallwebwaf` and the + app themselves, and a request for one of `smallwebwaf`'s own endpoints without + its token is still answered `401`. It is for trying a configuration before enforcing it. - Answers `GET /_smallwebwaf/healthz` itself with `200` and `ok`, before any check and without asking the app, for the image's health check. @@ -239,13 +254,14 @@ in `bin/state` unless `SWWAF_STATE_DIR` is set, and the default rule file of `SWWAF_LOG_REMOTE_URL` names one (see "Sending the log to a syslog server" below). - Sends an alert for each ban it makes or makes permanent, for a count over an - anomaly threshold, for GeoJS failing, for a rule file or state file with an - error, and for a replacement of the lookup database it cannot read, holding - back repeats and, past an hourly limit, rolling the rest into one summary, to - each destination you name: as a JSON object to the webhook - `SWWAF_ALERT_WEBHOOK_URL` names, as a message to the Slack incoming webhook - `SWWAF_ALERT_SLACK_WEBHOOK_URL` names, and as a message to the ntfy topic - `SWWAF_ALERT_NTFY_URL` names (see "Alerts" below). + anomaly threshold, for a client a blocklist lists, for GeoJS failing or a list + it cannot fetch, for a rule file or state file with an error, and for a + replacement of the lookup database it cannot read, holding back repeats and, + past an hourly limit, rolling the rest into one summary, to each destination + you name: as a JSON object to the webhook `SWWAF_ALERT_WEBHOOK_URL` names, as + a message to the Slack incoming webhook `SWWAF_ALERT_SLACK_WEBHOOK_URL` names, + and as a message to the ntfy topic `SWWAF_ALERT_NTFY_URL` names (see "Alerts" + below). - Counts requests and their bytes over a minute and an hour, per client, per netblock around a client, per AS number, for the whole service and per named netblock, and sends an `anomaly` alert for a count over the anomaly threshold @@ -308,8 +324,8 @@ effective settings are logged at start. - `SWWAF_REQUEST_MAX_BYTES` (default `100M`): the largest request body. - `SWWAF_RESPONSE_MAX_BYTES` (default `5G`): the largest response body. - `SWWAF_ALLOW_NETS` (default empty): netblocks whose clients skip bans, the - country lists, the rate limits, the byte limits and the rule files, such as - your monitoring or your own networks. + country lists, the blocklists, the rate limits, the byte limits and the rule + files, such as your monitoring or your own networks. - `SWWAF_RATE_LIMIT_EXEMPT_NETS` (default empty): netblocks whose clients the rate limits and the byte limits do not apply to, such as a machine that talks to the app all day. @@ -349,9 +365,10 @@ effective settings are logged at start. `SWWAF_LOOKUP_DB_PATH` names, or `off`, which looks up no client and sends no address to GeoJS. With `off`, a country list that is not empty, `SWWAF_ADD_LOOKUP_HEADERS` set to `true`, a biased threshold that lowers a - limit, a list of them that is not empty or `SWWAF_UNKNOWN_LIMIT_PERCENT` below - 100, or an anomaly threshold per AS number that is not `off`, stops the start, - with a message naming it and `SWWAF_LOOKUP_SOURCE`. + limit, a list of them that is not empty, `SWWAF_ASN_LIMIT_PERCENT_URL` set or + `SWWAF_UNKNOWN_LIMIT_PERCENT` below 100, or an anomaly threshold per AS number + that is not `off`, stops the start, with a message naming it and + `SWWAF_LOOKUP_SOURCE`. - `SWWAF_LOOKUP_DB_PATH` (default empty): the IPinfo Lite database file, in its `.mmdb` form, for `SWWAF_LOOKUP_SOURCE=file`. `file` without it, or it with any other `SWWAF_LOOKUP_SOURCE`, the default included, stops the start, with a @@ -389,11 +406,33 @@ effective settings are logged at start. client without a country gets: one the lookup cannot place, one on a private, loopback or link-local address, which is never looked up, and one whose answer from GeoJS has not come in time. +- `SWWAF_ASN_LIMIT_PERCENT_URL` (default unset): the `http` or `https` URL of a + file of AS numbers, each with its percentage of every rate limit and byte + limit, one such as `AS14061:50` to a line, whose percentages count as those of + `SWWAF_ASN_LIMIT_PERCENT` do, so that one list can serve several instances. + For an AS number both give a percentage, the lower applies, and an AS number + the file lists twice gets the lower of its two. It is fetched and kept as a + blocklist is (see "Blocklists" below). A URL `SWWAF_BLOCKLIST_URLS` names too + stops the start. +- `SWWAF_BLOCKLIST_URLS` (default empty): the blocklists, as `http` or `https` + URLs without a user or a fragment, such as + `https://www.spamhaus.org/drop/drop.txt` (see "Blocklists" below). A URL + listed twice stops the start. +- `SWWAF_BLOCKLIST_REFRESH` (default `24h`): how long after a list was last + fetched, or a fetch of it failed, it is fetched again, for the blocklists and + `SWWAF_ASN_LIMIT_PERCENT_URL`. Less than `1h` stops the start: the Spamhaus + lists may be fetched no more than once an hour. +- `SWWAF_BLOCKLIST_ACTION` (default `deny`): what is done with a client a + blocklist lists: `deny` refuses its requests with `SWWAF_BAN_RESPONSE`, + `limit:`, such as `limit:25`, gives it that percentage of every rate + limit and byte limit, and `log` does nothing more than note the lists in the + log line and raise the alert. - `SWWAF_BAN_RESPONSE` (default `403`): how a refused client is answered, one that is banned, breaks a rate limit, matches a `ban` rule, is in - `SWWAF_DENY_NETS` or comes from a refused country: `403`, `429`, or `close` to - close the connection without an answer. Behind traefik, `close` does not leave - the client unanswered: traefik answers `502`, as it does whenever its backend + `SWWAF_DENY_NETS`, comes from a refused country or is in a blocklist while + `SWWAF_BLOCKLIST_ACTION` is `deny`: `403`, `429`, or `close` to close the + connection without an answer. Behind traefik, `close` does not leave the + client unanswered: traefik answers `502`, as it does whenever its backend drops a connection. A `block` rule always answers `403`. - `SWWAF_LIMIT_BAN_DURATION` (default `1h`): the ban for a first broken rate limit or byte limit. @@ -486,8 +525,8 @@ effective settings are logged at start. `SWWAF_INSTANCE_NAME`, which ntfy is sent in the title. - `SWWAF_ALERT_EVENTS` (default `ban,permanent_ban,waf_block,anomaly,reputation_hit,source_failure,file_error`): - the events alerts are sent for. `waf_block` and `reputation_hit` come with the - features that raise them; nothing raises them yet. + the events alerts are sent for. `waf_block` comes with the Core Rule Set; + nothing raises it yet. - `SWWAF_ALERT_COOLDOWN` (default `15m`): how long a repeat of an alert is held back (see "Alerts" below). - `SWWAF_ALERT_MAX_PER_HOUR` (default `60`): the most alerts sent in an hour; @@ -536,9 +575,10 @@ an AS number or a country, `:` and a percentage; an AS number or a country listed twice in one of them stops the start. `off` switches a timeout, a size limit, a rate limit, a byte limit, an anomaly threshold, `SWWAF_ALERT_COOLDOWN` or `SWWAF_ALERT_MAX_PER_HOUR` off; `SWWAF_CLIENT_REQUEST_HEADER_MAX_BYTES`, -`SWWAF_LOOKUP_TIMEOUT`, `SWWAF_UNKNOWN_LIMIT_PERCENT`, the ban settings, the -state settings, `SWWAF_METRICS_TOP_N`, `SWWAF_LOG_REMOTE_BUFFER`, -`SWWAF_ANOMALY_NET_V4_PREFIX` and `SWWAF_ANOMALY_NET_V6_PREFIX` cannot be off. +`SWWAF_LOOKUP_TIMEOUT`, `SWWAF_UNKNOWN_LIMIT_PERCENT`, +`SWWAF_BLOCKLIST_REFRESH`, the ban settings, the state settings, +`SWWAF_METRICS_TOP_N`, `SWWAF_LOG_REMOTE_BUFFER`, `SWWAF_ANOMALY_NET_V4_PREFIX` +and `SWWAF_ANOMALY_NET_V6_PREFIX` cannot be off. Several limits are fixed rather than settings. At most 20,000 clients are kept, with their counters and history, and an IPv6 client is counted by its /64. At @@ -626,26 +666,29 @@ which every line has. app's, as passed on, or those of `smallwebwaf`'s own answer. - `request_bytes` and `response_bytes` count body bytes. - `action` is `forward` for a request passed to the app, `denied` for one - refused because its client is in `SWWAF_DENY_NETS`, `banned` for one refused - because a ban covers its client or because it matched a `ban` rule, which bans - its client, `country_denied` for one refused for its client's country, - `rate_limited` for one that broke a rate limit and banned its client, - `rule_blocked` for one a `block` rule refused, `too_large` for a request or - response over its size limit, `timed_out` for one that ran out of time, - `upstream_error` when the app could not be reached or its answer broke off, - and `admin` for one `smallwebwaf` answered at its own endpoint. + refused because its client is in `SWWAF_DENY_NETS`, or in a blocklist while + `SWWAF_BLOCKLIST_ACTION` is `deny`, `banned` for one refused because a ban + covers its client or because it matched a `ban` rule, which bans its client, + `country_denied` for one refused for its client's country, `rate_limited` for + one that broke a rate limit and banned its client, `rule_blocked` for one a + `block` rule refused, `too_large` for a request or response over its size + limit, `timed_out` for one that ran out of time, `upstream_error` when the app + could not be reached or its answer broke off, and `admin` for one + `smallwebwaf` answered at its own endpoint. - `would_action` is there in `observe` mode for a request that - `SWWAF_DENY_NETS`, a ban, the country lists, a rate limit or a rule would have - refused in `enforce` mode, and names the action that refusal would have had: - `denied`, `banned`, `country_denied`, `rate_limited` or `rule_blocked`. - `action` then names what was done: `forward` for a request passed to the app, - and another action, such as `too_large`, for one a size or time limit refused. + `SWWAF_DENY_NETS`, a ban, the country lists, a blocklist, a rate limit or a + rule would have refused in `enforce` mode, and names the action that refusal + would have had: `denied`, `banned`, `country_denied`, `rate_limited` or + `rule_blocked`. `action` then names what was done: `forward` for a request + passed to the app, and another action, such as `too_large`, for one a size or + time limit refused. - `limit_percent` is there for a request the rate limits count whose client a - biased threshold gives less than the whole of the rate limits, and gives the - percentage it gets, with `limit_percent_setting` naming the setting that gave - it, such as `SWWAF_ASN_LIMIT_PERCENT`. `bytes_percent` and - `bytes_percent_setting` are the same for the byte limits. Each is left out - when the client gets the whole of those limits. + biased threshold, or `SWWAF_BLOCKLIST_ACTION` for a blocklist that lists it, + gives less than the whole of the rate limits, and gives the percentage it + gets, with `limit_percent_setting` naming the setting that gave it, such as + `SWWAF_ASN_LIMIT_PERCENT`, or `SWWAF_ASN_LIMIT_PERCENT_URL` for the file it + names. `bytes_percent` and `bytes_percent_setting` are the same for the byte + limits. Each is left out when the client gets the whole of those limits. - `counts` gives the client's requests in the minute, the hour and the day as the rate limits count them, this request included: in each window, those in the bucket under way and a share of those in the bucket before, so a count can @@ -653,11 +696,11 @@ which every line has. broke it. It is left out for a request the rate limits do not count: the health check, one from a client in `SWWAF_ALLOW_NETS` or `SWWAF_RATE_LIMIT_EXEMPT_NETS`, one for a path that - `SWWAF_RATE_LIMIT_EXEMPT_PATHS` exempts, and one that `SWWAF_DENY_NETS`, a ban - or the country lists refuse, or would refuse in `observe` mode. Its - `minute_bytes`, `hour_bytes` and `day_bytes` give the client's bytes in each - window as the byte limits count them, in the same way: for a request whose - bytes they count, with its own, once it has ended; for any other, those + `SWWAF_RATE_LIMIT_EXEMPT_PATHS` exempts, and one that `SWWAF_DENY_NETS`, a + ban, the country lists or a blocklist refuse, or would refuse in `observe` + mode. Its `minute_bytes`, `hour_bytes` and `day_bytes` give the client's bytes + in each window as the byte limits count them, in the same way: for a request + whose bytes they count, with its own, once it has ended; for any other, those counted before it. - `rule_ids` is there for a request that matched rules of the rule files, and lists their ids in the order they matched, up to the one that refused it. @@ -668,6 +711,12 @@ which every line has. several. `offence` is then `limit`. A request whose bytes broke a byte limit is not refused: its `action` is what it would have been otherwise, such as `forward`. +- `reputation` is there for a request whose client a blocklist lists, and gives + the URLs of the blocklists that list it, in the order `SWWAF_BLOCKLIST_URLS` + names them, whatever `SWWAF_BLOCKLIST_ACTION` says. It is left out for a + client the blocklists are not checked for: one in `SWWAF_ALLOW_NETS`, and one + `SWWAF_DENY_NETS`, a ban or the country lists refuse first, or would in + `observe` mode. - `ban_expires` is there for a request that made a ban or was refused under one, or in `observe` mode would have been refused under one, and gives when the ban ends, in the same form as `time`, or `permanent`. @@ -737,7 +786,11 @@ it, as below. An alert is for one of these events, and is sent when - `anomaly`: a count of requests or bytes over an anomaly threshold, raised by each request that ends with the count over it, in `observe` mode as in `enforce` mode. It refuses and bans nothing. -- `source_failure`: GeoJS failing or refusing `smallwebwaf`. +- `reputation_hit`: a request whose client a blocklist lists, one alert for each + blocklist that lists it, whatever `SWWAF_BLOCKLIST_ACTION` says, in `observe` + mode as in `enforce` mode. +- `source_failure`: GeoJS failing or refusing `smallwebwaf`, or a fetch of a + list failing (see "Blocklists" below). - `file_error`: a rule file edited while it runs that has an error, an edit of a state file set aside as `.bad`, a state file it could not write while running, or a replacement of the lookup database it could not read, which it @@ -800,25 +853,29 @@ is sent on one line: - `client` is the address of the client whose request raised the alert, and `netblock` the netblock of the ban, or for an `anomaly`, the netblock counted: the client's own, the netblock around it or a named netblock, and none for an - AS number or the whole service; both are empty for `source_failure` and + AS number or the whole service, or for a `reputation_hit`, the client's own, + as `client_group` gives it; both are empty for `source_failure` and `file_error`. `asn`, `as_name` and `country` are, for a ban, the client's as the ban's notes give them when the alert is raised: empty, as in this alert, when GeoJS had not answered about the client by then; for an `anomaly`, the - client's as the lookup gave them by the time its request ended. + client's as the lookup gave them by the time its request ended; for a + `reputation_hit`, the client's as its request's log line gives them. - `reason` is a short sentence; for a ban, the ban's `reason` in `bans.json`; for an `anomaly`, what was counted over which threshold, such as - `requests per minute of the netblock 203.0.113.0/24 over the threshold of 1000`. + `requests per minute of the netblock 203.0.113.0/24 over the threshold of 1000`; + for a `reputation_hit`, `listed by a blocklist`. - `detail` is what is particular to the event: for a ban, its `cause`, when it ends as `ban_expires`, in the form the request log gives it, and its `notes`, as `bans.json` gives them; for an `anomaly`, the `scope`, `client`, `net`, `asn`, `total` or `watch`, as the settings name them, the `asn` counted for `asn` and the `name` of the named netblock for `watch`, the `window`, `minute` or `hour`, the `kind`, `requests` or `bytes`, the `count`, which is weighted - as the rate limits weigh theirs, and the `threshold`; for `source_failure`, - the `source`, `geojs`, the `error`, and when GeoJS is asked again, - `asking_again_in`; for `file_error`, the `file`, which for an edit set aside - is the file it was renamed to, and the `error`, which for a file that does not - parse names where in it the error is. + as the rate limits weigh theirs, and the `threshold`; for a `reputation_hit`, + the `source`, the URL of the blocklist; for `source_failure`, the `source`, + `geojs` or the URL of the list, the `error`, and for GeoJS, when it is asked + again, `asking_again_in`; for `file_error`, the `file`, which for an edit set + aside is the file it was renamed to, and the `error`, which for a file that + does not parse names where in it the error is. - `suppressed_repeats` is how many repeats the cooldown held back before this alert, and for a `summary`, those no other alert gives (see below). @@ -863,14 +920,15 @@ and Slack this JSON object, shown indented; it is sent on one line: } ``` -An alert for the same event as the last one sent, on the same netblock, or for a -`file_error` about the same file, or for a `source_failure` about the same -source, or for an `anomaly` in the same scope, with the same netblock, AS number -or name, whatever its window and kind, less than `SWWAF_ALERT_COOLDOWN` after -it, is a repeat: it is held back and counted, and the next alert sent for them -gives that count as `suppressed_repeats`. As each hour of the clock, in UTC, -ends, the cooldowns that have run out are dropped, and the repeats they held -back, which no alert sent since has given, go in that hour's summary. +An alert for the same event as the last one sent, on the same netblock, and for +a `reputation_hit` about the same blocklist, or for a `file_error` about the +same file, or for a `source_failure` about the same source, or for an `anomaly` +in the same scope, with the same netblock, AS number or name, whatever its +window and kind, less than `SWWAF_ALERT_COOLDOWN` after it, is a repeat: it is +held back and counted, and the next alert sent for them gives that count as +`suppressed_repeats`. As each hour of the clock, in UTC, ends, the cooldowns +that have run out are dropped, and the repeats they held back, which no alert +sent since has given, go in that hour's summary. Past `SWWAF_ALERT_MAX_PER_HOUR` alerts in an hour, the hour's other alerts are held back and counted by event. An alert held back this way starts no cooldown. @@ -897,11 +955,12 @@ restart the alerts waiting are sent, and the cooldowns go on. ## State files -`smallwebwaf` keeps its state in memory and a copy of it in four JSON files in +`smallwebwaf` keeps its state in memory and a copy of it in five JSON files in `SWWAF_STATE_DIR`, `/var/lib/smallwebwaf` by default, as "Persistent state" in [`SPEC.md`](SPEC.md) describes. Each has a top-level `version`, 1, and lists its -entries by client address, but for the alerts waiting, and the anomaly counters, -which are listed by scope first, with times in UTC. +entries by client address, but for the alerts waiting, the anomaly counters, +which are listed by scope first, and the copies of the lists, listed by URL, +with times in UTC. - `bans.json`: every ban with its notes, indented to be read. A permanent ban's `expires` is `null`. A ban's `cause` is `limit` for a broken rate limit or @@ -927,25 +986,32 @@ which are listed by scope first, with times in UTC. `grep` shows everything about one. - `lookups.json`: GeoJS's answers, one to a line, each with the client's AS number, AS name and country, when GeoJS gave it and when it was last used. +- `reputation.json`: each list fetched from a URL (see "Blocklists" below), + indented to be read, under `lists`: its `url`, when it was last `tried`, the + fetch failed or not, and its last good copy: when that was `fetched`, and its + `lines`, as fetched, comment lines included, each on a line of its own, both + left out while no fetch of it has succeeded. As the file is read, the lists + the settings no longer name are dropped. - `alerts.json`: the state of the alerts (see "Alerts" above), indented to be - read: under `cooldowns`, for each event and netblock, or event and `file` or - `source`, or for an `anomaly`, its `scope` with its `netblock`, `asn` or - `name`, or event alone, when the last alert was sent, `sent`, and the repeats - held back since, `suppressed_repeats`; under `hour`, the hour under way, from - its `start`, the alerts `sent` in it and those `held_back` for its summary, by - event; under `waiting`, for each destination you name, `webhook`, `slack` or - `ntfy`, the alerts still waiting to be sent to it, the oldest first, each as - the webhook is sent it; and under `anomaly_counters`, each anomaly counter: - its `scope`, as an `anomaly` alert names it, with the `netblock` of a client, - of a netblock around a client or of a named netblock, the `asn` of an AS - number and the `name` of a named netblock, and its two buckets of requests in - the minute and the hour, `minute` and `hour`, and of bytes, `minute_bytes` and - `hour_bytes`, each left out while it is empty. As an hour ends, the cooldowns - that have run out are dropped, and the hour's summary gives the repeats they - held back. As the file is read, the alerts waiting for a destination you no - longer name are dropped. A file whose `waiting` is a list, as it was before - alerts went to Slack and ntfy too, stops the start: put the list under - `"webhook"`, or remove the file. + read: under `cooldowns`, for each event and netblock, with the `source` too + for a `reputation_hit`, or event and `file` or `source`, or for an `anomaly`, + its `scope` with its `netblock`, `asn` or `name`, or event alone, when the + last alert was sent, `sent`, and the repeats held back since, + `suppressed_repeats`; under `hour`, the hour under way, from its `start`, the + alerts `sent` in it and those `held_back` for its summary, by event; under + `waiting`, for each destination you name, `webhook`, `slack` or `ntfy`, the + alerts still waiting to be sent to it, the oldest first, each as the webhook + is sent it; and under `anomaly_counters`, each anomaly counter: its `scope`, + as an `anomaly` alert names it, with the `netblock` of a client, of a netblock + around a client or of a named netblock, the `asn` of an AS number and the + `name` of a named netblock, and its two buckets of requests in the minute and + the hour, `minute` and `hour`, and of bytes, `minute_bytes` and `hour_bytes`, + each left out while it is empty. As an hour ends, the cooldowns that have run + out are dropped, and the hour's summary gives the repeats they held back. As + the file is read, the alerts waiting for a destination you no longer name are + dropped. A file whose `waiting` is a list, as it was before alerts went to + Slack and ntfy too, stops the start: put the list under `"webhook"`, or remove + the file. `bans.json` is written `SWWAF_STATE_WRITE_DELAY` after a ban is made, lifted through `DELETE /_smallwebwaf/bans/`, or made permanent, with every such @@ -957,27 +1023,30 @@ whole. A write that fails is logged, raised as a `file_error` alert while changed since the last write. At start the files are read back: each client keeps its counts, so a restart -gives it no fresh allowance, each anomaly counter keeps its counts, and each ban +gives it no fresh allowance, each anomaly counter keeps its counts, each ban keeps refusing every client in its netblock until it ends, even after -`SWWAF_BAN_SCOPE_V4_PREFIX` has changed. A netblock whose address has bits past -its length, such as `203.0.113.9/24`, is read as the netblock it is in, -`203.0.113.0/24`. Buckets and answers whose time has passed are dropped, and so -is an anomaly counter left with no bucket. A missing file is empty state, as on -a first start. A file that does not parse, or has another `version`, stops the -start with a message naming the file, and the line and column where Go's JSON -decoder gives them; so does a state directory `smallwebwaf` cannot write. So -does an entry without a field it needs, named with the entry's place in the -file: a ban's `netblock`, `start` or `expires`, which is `null` for a permanent -ban; a client's `client`, or the `start` of a window in which it has requests or -bytes; an answer's `client`, `country`, which is `""` for a client GeoJS cannot -place, or `answered`; a cooldown's `event` or `sent`; an alert waiting's `event` -or `time`; an anomaly counter's `netblock`, unless it counts an AS number or the +`SWWAF_BAN_SCOPE_V4_PREFIX` has changed, and the copy of each list stays in use +until a fetch of it succeeds. A netblock whose address has bits past its length, +such as `203.0.113.9/24`, is read as the netblock it is in, `203.0.113.0/24`. +Buckets and answers whose time has passed are dropped, and so is an anomaly +counter left with no bucket. A missing file is empty state, as on a first start. +A file that does not parse, or has another `version`, stops the start with a +message naming the file, and the line and column where Go's JSON decoder gives +them; so does a state directory `smallwebwaf` cannot write. So does an entry +without a field it needs, named with the entry's place in the file: a ban's +`netblock`, `start` or `expires`, which is `null` for a permanent ban; a +client's `client`, or the `start` of a window in which it has requests or bytes; +an answer's `client`, `country`, which is `""` for a client GeoJS cannot place, +or `answered`; a cooldown's `event` or `sent`; an alert waiting's `event` or +`time`; an anomaly counter's `netblock`, unless it counts an AS number or the whole service, its `asn`, for an AS number, its `name`, for a named netblock, or -the `start` of a window in which it has requests or bytes. So does a ban whose +the `start` of a window in which it has requests or bytes; a list's `url`, +`fetched` or `lines`, which is `[]` for an empty list. So does a ban whose `cause` is not `limit`, `attack` or `admin`, alerts waiting for a destination -that is not `webhook`, `slack` or `ntfy`, and an anomaly counter whose `scope` -is not `client`, `net`, `asn`, `total` or `watch`. An answer's `asn` or -`as_name` left out reads as empty. +that is not `webhook`, `slack` or `ntfy`, an anomaly counter whose `scope` is +not `client`, `net`, `asn`, `total` or `watch`, and a copy of a list with a line +that would make its fetch fail. An answer's `asn` or `as_name` left out reads as +empty. While it runs, `smallwebwaf` watches `SWWAF_STATE_DIR` and takes in your edit of a state file as soon as you save it: what the file then holds replaces what @@ -987,14 +1056,14 @@ writes a file it takes in any edit made since, so your edit is not overwritten; a change `smallwebwaf` made after you opened the file, such as a new ban, is lost when you save over it. An edit that would stop the start, because it does not parse, has another `version`, leaves out a field an entry needs, gives a ban -another `cause`, names another destination or gives an anomaly counter another -`scope`, does not stop the running `smallwebwaf`: it keeps what it holds, and at -the file's next write renames your file to `.bad`, such as -`bans.json.bad`, writes the file again from memory, logs the file and where the -error is, and raises a `file_error` alert for it. It waits for that write -because an editor's file can be read before the editor has finished writing it. -Mend the `.bad` file and move it back. A file you remove is written again at its -next write. +another `cause`, names another destination, gives an anomaly counter another +`scope` or gives a list's copy a line that would make its fetch fail, does not +stop the running `smallwebwaf`: it keeps what it holds, and at the file's next +write renames your file to `.bad`, such as `bans.json.bad`, writes the +file again from memory, logs the file and where the error is, and raises a +`file_error` alert for it. It waits for that write because an editor's file can +be read before the editor has finished writing it. Mend the `.bad` file and move +it back. A file you remove is written again at its next write. To ban a netblock, add an entry to `bans.json` with its `netblock`, its `start` and its `expires`, `null` for a ban that never ends; its `reason` and its @@ -1163,6 +1232,12 @@ scraped, and keeps this one as `exported_instance` unless the scrape sets database in use was read; and `smallwebwaf_lookup_database_read_failures_total`: the replacements of it that could not be read. +- By `source`, the URL of each list `SWWAF_BLOCKLIST_URLS` or + `SWWAF_ASN_LIMIT_PERCENT_URL` names: `smallwebwaf_reputation_hits_total`: the + requests whose client the blocklist lists, a series that comes with the first; + `smallwebwaf_reputation_failures_total`: the fetches of the list that failed; + and `smallwebwaf_reputation_last_fetch_timestamp_seconds`: when the copy of it + in use was fetched, `0` while there is none. - `smallwebwaf_tracked_clients`: the clients in the table of clients. - `smallwebwaf_state_file_writes_total`, `smallwebwaf_state_file_write_failures_total`, @@ -1354,9 +1429,9 @@ goes through the candidates one by one. readable JSON files, written regularly and at every stop, so a restart loses nothing. Edit a file, or add a rule file, and the running `smallwebwaf` picks up the change. Nothing is read from disk while serving a request. The files - for the bans, the clients, the GeoJS answers and the alerts are built, with an - edit taken in while running (see "State files" above); the others come with - their features. + for the bans, the clients, the GeoJS answers, the copies of the lists and the + alerts are built, with an edit taken in while running (see "State files" + above); the rest comes with its features. - Health checks, the metrics, and listing, adding and lifting bans or asking why a given address was refused, all on the one port every request uses: under `/_smallwebwaf/` on the app's own address, through traefik like any other @@ -1376,7 +1451,7 @@ For each request `smallwebwaf`: the deny list or currently banned; - looks up its AS number and country, and refuses it if that country is denied, or is not among the only ones allowed; -- checks for a cached reputation verdict; +- checks it against the blocklists, and for a cached reputation verdict; - picks the client's limit percentage from those; - checks the minute, hour and day request counters against the limits, and bans the client if it breaks one; @@ -1533,6 +1608,57 @@ country: `SWWAF_EXCLUSIVELY_ALLOWED_COUNTRIES` refuses it unless you list it in `SWWAF_UNKNOWN_LIMIT_PERCENT` sets its limits. Such addresses are never sent to GeoJS. +## Blocklists + +`SWWAF_BLOCKLIST_URLS` names blocklists: text files of addresses and netblocks, +one to a line, written as the Spamhaus DROP list, +`https://www.spamhaus.org/drop/drop.txt`, is. Anything after a `;` or a `#` on a +line is left out, and so is a line left blank. A bare address stands for itself +alone, as in the settings. An IPv4-mapped address or netblock, such as +`::ffff:192.0.2.0/120`, is read as the IPv4 one it stands for, here +`192.0.2.0/24`, since a client's IPv4 address is checked as IPv4; a mapped +netblock shorter than `/96` stands for none, and is not a netblock. None is +named by default: a list judges a client by what others saw it do, while the +defaults judge it by what it does to your service. + +`smallwebwaf` fetches each list, and the file `SWWAF_ASN_LIMIT_PERCENT_URL` +names, `SWWAF_BLOCKLIST_REFRESH` after it last fetched it or tried to, 24 hours +by default and never less than one, one list after another. Since +`reputation.json` keeps when each list was last tried, the fetch failed or not, +even one cut off as `smallwebwaf` stopped, whose request the server may have +had, at start `smallwebwaf` fetches at once only a list it has never tried, and +one it last tried that long ago; any other waits its turn, so that restarts do +not fetch a list more often. A fetch fails when the server answers other than +`200`, when it does not finish within a minute, when the list is longer than 16 +MiB, or when a line of it is not an address or a netblock, or for +`SWWAF_ASN_LIMIT_PERCENT_URL`, not an AS number, `:` and a percentage. The copy +fetched before then stays in use, and the failure is counted, logged and raised +as a `source_failure` alert; a fetch cut off as `smallwebwaf` stops is not a +failure. The last good copy of each list is kept whole, comment lines included, +in `reputation.json` (see "State files" above), so that a restart keeps it in +use too. Each list is named by its URL, in the request log, the alerts and the +metrics, so keep a secret out of it. + +A client in `SWWAF_ALLOW_NETS` is not checked. Any other is checked by its own +address after the country lists, and `SWWAF_BLOCKLIST_ACTION` says what is done +with one a list lists, as "What it does so far" above describes. `deny` suits +lists of networks that send nothing legitimate, such as DROP; `limit:` +suits lists of addresses shared with ordinary visitors, such as those of Tor's +exits. + +The Spamhaus DROP list is The Spamhaus Project's, https://www.spamhaus.org. Its +terms, on its DROP page, ask that a product using it credit The Spamhaus Project +and keep the list's date and copyright lines with the data, which the copy in +`reputation.json` does; a service that uses DROP through `smallwebwaf` uses data +from The Spamhaus Project, and should say so. They also ask that it be fetched +automatically no more than once an hour, once a day being more than enough in +most cases, and Spamhaus may block an address that fetches it more often. Each +`smallwebwaf` fetches its own copy, on its own schedule, so on a host where +several apps run it, sharing one address, their fetches can come less than an +hour apart whatever `SWWAF_BLOCKLIST_REFRESH` is: each fetches a list again that +long after its own last try, so those first started within the same hour, with +the same refresh, keep fetching within the same hour. + ## How the code is laid out - `cmd/smallwebwaf`: the binary, which only calls `internal/smallwebwaf`. @@ -1546,15 +1672,15 @@ GeoJS. standard library's `httputil.ReverseProxy` within the timeouts and size limits, and writes the request's log line. Its `check` method is where a request is refused before anything reaches the app: for `SWWAF_DENY_NETS`, for - a ban, for the country lists, for a rate limit, which bans the client, for a - `block` or `ban` rule, the latter banning the client, and for an announced - body over the size limit; in `observe` mode, only for the size limit, with - what it would have refused for noted in the log line. A request under - `/_smallwebwaf/` that `check` lets through is answered by `answerAdmin` - instead of reaching the app. Once the answer to a request passed to the app - has ended, `countBytes` counts its bytes for the byte limits, and once any - request but the health check has ended, `countAnomalies` counts it for the - anomaly thresholds. + a ban, for the country lists, for a blocklist, for a rate limit, which bans + the client, for a `block` or `ban` rule, the latter banning the client, and + for an announced body over the size limit; in `observe` mode, only for the + size limit, with what it would have refused for noted in the log line. A + request under `/_smallwebwaf/` that `check` lets through is answered by + `answerAdmin` instead of reaching the app. Once the answer to a request passed + to the app has ended, `countBytes` counts its bytes for the byte limits, and + once any request but the health check has ended, `countAnomalies` counts it + for the anomaly thresholds. - `internal/metrics`: the metrics, counted as the other parts tell it what happened, and served in the Prometheus text format. - `internal/bans`: the ban ledger: each netblock's bans with their notes, how @@ -1567,6 +1693,10 @@ GeoJS. client's history and to the notes of its bans; or in the lookup database, which it reads again when the file is replaced. `internal/lookup/lookuptest` writes lookup databases for the tests. +- `internal/reputation`: fetches the blocklists and the file + `SWWAF_ASN_LIMIT_PERCENT_URL` names as they are due, keeps the last good copy + of each, and tells which blocklists list an address and what percentage the + file gives an AS number. - `internal/ratelimit`: the table of clients: counts each client's requests and bytes, tells when they take it over a rate limit or a byte limit, and keeps each client's history. diff --git a/internal/alerts/alerts.go b/internal/alerts/alerts.go index 253a59d..8c72282 100644 --- a/internal/alerts/alerts.go +++ b/internal/alerts/alerts.go @@ -44,11 +44,12 @@ const ( // EventAnomaly is a count of requests or bytes over an anomaly // threshold. EventAnomaly = "anomaly" - // EventWAFBlock and EventReputationHit come with the Core Rule Set and - // the reputation sources; nothing raises them yet. - EventWAFBlock = "waf_block" + // EventWAFBlock comes with the Core Rule Set; nothing raises it yet. + EventWAFBlock = "waf_block" + // EventReputationHit is a request whose client a blocklist lists. EventReputationHit = "reputation_hit" - // EventSourceFailure is GeoJS failing or refusing smallwebwaf. + // EventSourceFailure is GeoJS failing or refusing smallwebwaf, or a + // fetch of a list failing. EventSourceFailure = "source_failure" // EventFileError is a rule file or state file edited while smallwebwaf // runs that does not parse, a replacement of the lookup database that diff --git a/internal/config/config.go b/internal/config/config.go index b8f4ec6..0343286 100644 --- a/internal/config/config.go +++ b/internal/config/config.go @@ -131,12 +131,27 @@ type Config struct { // percentages from 0 to 100, by AS number, written as AS64496, or by // country, a two-letter code in capitals, as the lookup gives them. // UnknownLimitPercent is the percentage of every limit a client without - // a country gets (SWWAF_UNKNOWN_LIMIT_PERCENT). + // a country gets (SWWAF_UNKNOWN_LIMIT_PERCENT). ASNLimitPercentURL is + // where a file of AS:percent lines is fetched from, whose percentages + // count as those of ASNLimitPercent do (SWWAF_ASN_LIMIT_PERCENT_URL), "" + // while it is unset. ASNLimitPercent map[string]int64 CountryLimitPercent map[string]int64 ASNBytesPercent map[string]int64 CountryBytesPercent map[string]int64 UnknownLimitPercent int64 + ASNLimitPercentURL string + // BlocklistURLs are where the blocklists are fetched from + // (SWWAF_BLOCKLIST_URLS). Each list, and ASNLimitPercentURL's, is + // fetched again BlocklistRefresh after it was last fetched or tried + // (SWWAF_BLOCKLIST_REFRESH), which is never less than an hour. + // BlocklistAction is what is done with a client a blocklist lists + // (SWWAF_BLOCKLIST_ACTION): deny, limit or log; for limit, + // BlocklistLimitPercent is the percentage of every limit it gets. + BlocklistURLs []string + BlocklistRefresh time.Duration + BlocklistAction string + BlocklistLimitPercent int64 // BanResponse is the status a refused client is answered with, 403 // or 429, or 0 to close the connection without an answer // (SWWAF_BAN_RESPONSE). It answers a banned client, a request that @@ -344,6 +359,13 @@ var ( "is not a code, : and a percentage, such as AS64496:50 or cn:25") errNotPercent = errors.New("is not a percentage, a whole number from 0 to 100") errListedTwice = errors.New("is listed twice") + errNotListURL = errors.New( + "is not an http or https URL without a user or a fragment, " + + "such as https://www.spamhaus.org/drop/drop.txt") + errInBlocklistURLs = errors.New("is in SWWAF_BLOCKLIST_URLS too") + errNotAnHourOrMore = errors.New("is not a duration of 1h or more, such as 24h") + errNotAction = errors.New( + "is not deny, limit: such as limit:25, or log") ) // FromEnvironment reads the settings with lookupEnv, normally @@ -388,11 +410,14 @@ func FromEnvironment(lookupEnv func(string) (string, bool)) (*Config, error) { DeniedCountries: env.countries("SWWAF_DENIED_COUNTRIES", ""), ExclusivelyAllowedCountries: env.countries( "SWWAF_EXCLUSIVELY_ALLOWED_COUNTRIES", ""), - ASNLimitPercent: env.percents("SWWAF_ASN_LIMIT_PERCENT", parseASN), + ASNLimitPercent: env.percents("SWWAF_ASN_LIMIT_PERCENT", ParseASN), CountryLimitPercent: env.percents("SWWAF_COUNTRY_LIMIT_PERCENT", parseCountry), - ASNBytesPercent: env.percents("SWWAF_ASN_BYTES_PERCENT", parseASN), + ASNBytesPercent: env.percents("SWWAF_ASN_BYTES_PERCENT", ParseASN), CountryBytesPercent: env.percents("SWWAF_COUNTRY_BYTES_PERCENT", parseCountry), UnknownLimitPercent: env.percent("SWWAF_UNKNOWN_LIMIT_PERCENT", "100"), + ASNLimitPercentURL: env.listURL("SWWAF_ASN_LIMIT_PERCENT_URL"), + BlocklistURLs: env.listURLs("SWWAF_BLOCKLIST_URLS"), + BlocklistRefresh: env.refresh("SWWAF_BLOCKLIST_REFRESH", "24h"), BanResponse: env.banResponse("SWWAF_BAN_RESPONSE", "403"), LimitBanDuration: env.durationNotOff("SWWAF_LIMIT_BAN_DURATION", "1h"), LimitBanRepeatWindow: env.durationNotOff("SWWAF_LIMIT_BAN_REPEAT_WINDOW", "24h"), @@ -435,10 +460,13 @@ func FromEnvironment(lookupEnv func(string) (string, bool)) (*Config, error) { cfg.LogRemoteAppName = env.appName("SWWAF_LOG_REMOTE_APP_NAME", cfg.InstanceName, cfg.LogRemoteURL != nil) + cfg.BlocklistAction, cfg.BlocklistLimitPercent = env.action( + "SWWAF_BLOCKLIST_ACTION", "deny") env.checkInstanceNameForNtfy(cfg.InstanceName, cfg.AlertNtfyURL != nil) env.checkLookupDBPath(cfg) env.checkCountriesAndLookups(cfg) + env.checkASNLimitPercentURL(cfg) if env.err != nil { return nil, env.err @@ -664,12 +692,66 @@ func (e *environment) percents( // percent reads a setting that is a percentage, from 0 to 100. func (e *environment) percent(name, defaultValue string) int64 { - percent, err := parsePercent(e.value(name, defaultValue)) + percent, err := ParsePercent(e.value(name, defaultValue)) e.check(name, err) return percent } +// listURL reads a setting that is the URL a list is fetched from, "" while +// it is unset or empty. +func (e *environment) listURL(name string) string { + value := e.value(name, "") + if value != "" && !isListURL(value) { + e.check(name, fmt.Errorf("%q %w", value, errNotListURL)) + } + + return value +} + +// listURLs reads a setting that is a list of the URLs lists are fetched +// from. It is empty by default. +func (e *environment) listURLs(name string) []string { + urls, err := parseListURLs(e.value(name, "")) + e.check(name, err) + + return urls +} + +// refresh reads the setting that is how long after a list was last +// fetched or tried it is fetched again: a duration of an hour or more, +// since the Spamhaus lists may be fetched no more often, which cannot be +// off. +func (e *environment) refresh(name, defaultValue string) time.Duration { + value := e.value(name, defaultValue) + + duration, err := parseDuration(value) + if err != nil || duration < time.Hour { + e.check(name, fmt.Errorf("%q %w", value, errNotAnHourOrMore)) + } + + return duration +} + +// action reads a setting that is what is done with a client a list names: +// deny, log, or limit:, which it returns as limit and the +// percentage. +func (e *environment) action(name, defaultValue string) (string, int64) { + value := e.value(name, defaultValue) + if value == "deny" || value == "log" { + return value, 0 + } + + percentText, isLimit := strings.CutPrefix(value, "limit:") + + percent, err := ParsePercent(percentText) + if !isLimit || err != nil { + e.check(name, fmt.Errorf("%q %w", value, errNotAction)) + } + + return "limit", percent +} + // lookupSource reads the setting that is where clients are looked up: // geojs, file, or off. func (e *environment) lookupSource(name, defaultValue string) string { @@ -698,8 +780,9 @@ func (e *environment) checkLookupDBPath(cfg *Config) { // checkCountriesAndLookups refuses a country on both country lists, and, // while SWWAF_LOOKUP_SOURCE is off, each setting that needs clients looked // up: the country lists, SWWAF_ADD_LOOKUP_HEADERS, the biased thresholds, -// of which SWWAF_UNKNOWN_LIMIT_PERCENT needs them only below 100, where it -// lowers a limit, and the anomaly thresholds per AS number. +// SWWAF_ASN_LIMIT_PERCENT_URL among them, of which +// SWWAF_UNKNOWN_LIMIT_PERCENT needs them only below 100, where it lowers a +// limit, and the anomaly thresholds per AS number. func (e *environment) checkCountriesAndLookups(cfg *Config) { for _, country := range cfg.ExclusivelyAllowedCountries { if slices.Contains(cfg.DeniedCountries, country) { @@ -724,6 +807,7 @@ func (e *environment) checkCountriesAndLookups(cfg *Config) { {"SWWAF_ASN_BYTES_PERCENT", len(cfg.ASNBytesPercent) > 0}, {"SWWAF_COUNTRY_BYTES_PERCENT", len(cfg.CountryBytesPercent) > 0}, {"SWWAF_UNKNOWN_LIMIT_PERCENT", cfg.UnknownLimitPercent < 100}, + {"SWWAF_ASN_LIMIT_PERCENT_URL", cfg.ASNLimitPercentURL != ""}, {"SWWAF_ANOMALY_ASN_REQUESTS_PER_MINUTE", cfg.AnomalyASN.RequestsPerMinute > 0}, {"SWWAF_ANOMALY_ASN_REQUESTS_PER_HOUR", cfg.AnomalyASN.RequestsPerHour > 0}, {"SWWAF_ANOMALY_ASN_BYTES_PER_MINUTE", cfg.AnomalyASN.BytesPerMinute > 0}, @@ -736,6 +820,15 @@ func (e *environment) checkCountriesAndLookups(cfg *Config) { } } +// checkASNLimitPercentURL refuses SWWAF_ASN_LIMIT_PERCENT_URL naming a +// blocklist too: the file at a URL is fetched as one list or the other. +func (e *environment) checkASNLimitPercentURL(cfg *Config) { + if slices.Contains(cfg.BlocklistURLs, cfg.ASNLimitPercentURL) { + e.check("SWWAF_ASN_LIMIT_PERCENT_URL", + fmt.Errorf("%q %w", cfg.ASNLimitPercentURL, errInBlocklistURLs)) + } +} + // headerNames reads a setting that is a list of header names, and // returns them in lower case. func (e *environment) headerNames(name, defaultValue string) []string { @@ -1187,7 +1280,7 @@ func parseNetblocks(value string) ([]netip.Prefix, error) { netblocks := make([]netip.Prefix, 0, len(items)) for _, item := range items { - netblock, err := parseNetblock(item) + netblock, err := ParseNetblock(item) if err != nil { return nil, err } @@ -1198,9 +1291,10 @@ func parseNetblocks(value string) ([]netip.Prefix, error) { return netblocks, nil } -// parseNetblock reads a netblock in CIDR form, such as 10.0.0.0/8. A bare -// address is a netblock of that address alone, a /32 or a /128. -func parseNetblock(value string) (netip.Prefix, error) { +// ParseNetblock reads a netblock in CIDR form, such as 10.0.0.0/8. A bare +// address is a netblock of that address alone, a /32 or a /128. A +// blocklist's lines are read with it too. +func ParseNetblock(value string) (netip.Prefix, error) { if strings.Contains(value, "/") { netblock, err := netip.ParsePrefix(value) if err != nil { @@ -1237,7 +1331,7 @@ func parseNamedNetblocks(value string) ([]anomaly.NamedNetblock, error) { return nil, fmt.Errorf("%q %w", item, errNotNamedNetblock) } - netblock, err := parseNetblock(strings.TrimSpace(netblockText)) + netblock, err := ParseNetblock(strings.TrimSpace(netblockText)) if err != nil { return nil, err } @@ -1337,10 +1431,11 @@ func parseCountry(value string) (string, error) { return country, nil } -// parseASN reads an AS number such as AS64496, in either case, and +// ParseASN reads an AS number such as AS64496, in either case, and // returns it as the lookup gives it: AS and the number, in capitals and -// without leading zeros. -func parseASN(value string) (string, error) { +// without leading zeros. The file SWWAF_ASN_LIMIT_PERCENT_URL names is +// read with it too. +func ParseASN(value string) (string, error) { digits, hasAS := strings.CutPrefix(strings.ToUpper(value), "AS") number, err := strconv.ParseUint(digits, 10, 32) @@ -1376,7 +1471,7 @@ func parsePercents( return nil, err } - percent, err := parsePercent(percentText) + percent, err := ParsePercent(percentText) if err != nil { return nil, err } @@ -1391,8 +1486,9 @@ func parsePercents( return percents, nil } -// parsePercent reads a percentage, a whole number from 0 to 100. -func parsePercent(value string) (int64, error) { +// ParsePercent reads a percentage, a whole number from 0 to 100. The file +// SWWAF_ASN_LIMIT_PERCENT_URL names is read with it too. +func ParsePercent(value string) (int64, error) { percent, err := strconv.ParseInt(value, 10, 64) if err != nil || percent < 0 || percent > 100 { return 0, fmt.Errorf("%q %w", value, errNotPercent) @@ -1548,16 +1644,7 @@ func parseWebhookURL(value string) (*url.URL, string, error) { } webhook, err := url.Parse(value) - if err != nil { - return nil, "", errNotWebhookURL - } - - port, err := strconv.ParseUint(webhook.Port(), 10, 16) - - valid := (webhook.Scheme == "http" || webhook.Scheme == "https") && - webhook.Hostname() != "" && (webhook.Port() == "" || (err == nil && port != 0)) && - webhook.User == nil && webhook.Opaque == "" && webhook.Fragment == "" - if !valid { + if err != nil || !isHTTPURL(webhook) { return nil, "", errNotWebhookURL } @@ -1569,6 +1656,47 @@ func parseWebhookURL(value string) (*url.URL, string, error) { return webhook, logged, nil } +// isHTTPURL reports whether u is http or https, with a host, and an +// optional port from 1 to 65535, path and query, without a user or a +// fragment. +func isHTTPURL(u *url.URL) bool { + port, err := strconv.ParseUint(u.Port(), 10, 16) + + return (u.Scheme == "http" || u.Scheme == "https") && u.Hostname() != "" && + (u.Port() == "" || (err == nil && port != 0)) && + u.User == nil && u.Opaque == "" && u.Fragment == "" +} + +// isListURL reports whether value is a URL a list can be fetched from, as +// isHTTPURL says. +func isListURL(value string) bool { + u, err := url.Parse(value) + + return err == nil && isHTTPURL(u) +} + +// parseListURLs reads a comma-separated list of the URLs lists are fetched +// from. A URL listed twice is an error: it would be fetched twice as +// often. +func parseListURLs(value string) ([]string, error) { + urls, err := parseList(value) + if err != nil { + return nil, err + } + + for i, listURL := range urls { + if !isListURL(listURL) { + return nil, fmt.Errorf("%q %w", listURL, errNotListURL) + } + + if slices.Contains(urls[:i], listURL) { + return nil, fmt.Errorf("%q %w", listURL, errListedTwice) + } + } + + return urls, nil +} + // parseWebhookHeaders reads a comma-separated list of headers, each its // name, :, and its value, and returns them, and how the log shows them, // with each value as ********. An error names the item by its place in diff --git a/internal/config/config_test.go b/internal/config/config_test.go index 1467277..78a8456 100644 --- a/internal/config/config_test.go +++ b/internal/config/config_test.go @@ -57,6 +57,10 @@ const ( asnBytesPercent = "SWWAF_ASN_BYTES_PERCENT" countryBytesPercent = "SWWAF_COUNTRY_BYTES_PERCENT" unknownLimitPercent = "SWWAF_UNKNOWN_LIMIT_PERCENT" + asnLimitPercentURL = "SWWAF_ASN_LIMIT_PERCENT_URL" + blocklistURLs = "SWWAF_BLOCKLIST_URLS" + blocklistRefresh = "SWWAF_BLOCKLIST_REFRESH" + blocklistAction = "SWWAF_BLOCKLIST_ACTION" banResponse = "SWWAF_BAN_RESPONSE" limitBanDuration = "SWWAF_LIMIT_BAN_DURATION" limitBanRepeatWindow = "SWWAF_LIMIT_BAN_REPEAT_WINDOW" @@ -953,6 +957,7 @@ func TestSettingNeedingLookupsStopsTheStartWhileTheyAreOff(t *testing.T) { deniedCountries: "kp", allowedCountries: "de", addLookupHeaders: enabled, + asnLimitPercentURL: asnURL, asnLimitPercent: "AS64496:50", countryLimitPercent: "cn:25", asnBytesPercent: "AS64496:50", @@ -978,11 +983,13 @@ func TestSettingNeedingLookupsStopsTheStartWhileTheyAreOff(t *testing.T) { // Set empty, the lists need nothing looked up, and nor does // SWWAF_UNKNOWN_LIMIT_PERCENT at 100, which lowers no limit, an anomaly - // threshold per AS number that is off, or any other anomaly threshold. + // threshold per AS number that is off, any other anomaly threshold, or + // a blocklist. env := environment{ lookupSource: off, deniedCountries: "", allowedCountries: "", asnLimitPercent: "", countryLimitPercent: "", asnBytesPercent: "", - countryBytesPercent: "", unknownLimitPercent: "100", + countryBytesPercent: "", unknownLimitPercent: "100", asnLimitPercentURL: "", + blocklistURLs: dropURL, } for _, name := range anomalyThresholds() { env[name] = "1000" @@ -1205,6 +1212,116 @@ func TestInvalidBiasedThresholdStopsTheStartSayingWhatIsWrong(t *testing.T) { } } +// dropURL and torURL are blocklists, and asnURL a file of AS:percent +// lines. +const ( + dropURL = "https://www.spamhaus.org/drop/drop.txt" + torURL = "https://lists.example/tor-exits.txt" + asnURL = "https://lists.example/asn.txt" +) + +// The actions of SWWAF_BLOCKLIST_ACTION, as Config gives them. +const ( + actionDeny = "deny" + actionLimit = "limit" + actionLog = "log" +) + +func TestReputationSettingsAsSet(t *testing.T) { + t.Parallel() + + cfg := fromEnvironment(t, environment{}) + if len(cfg.BlocklistURLs) != 0 || cfg.BlocklistRefresh != 24*time.Hour || + cfg.BlocklistAction != actionDeny || cfg.ASNLimitPercentURL != "" { + t.Errorf("%s, %s, %s and %s gave %v, %s, %s and %q by default, "+ + "want none, 24h, deny and none", blocklistURLs, blocklistRefresh, + blocklistAction, asnLimitPercentURL, cfg.BlocklistURLs, cfg.BlocklistRefresh, + cfg.BlocklistAction, cfg.ASNLimitPercentURL) + } + + for _, tc := range []struct { + value, action string + percent int64 + }{ + {actionDeny, actionDeny, 0}, + {actionLog, actionLog, 0}, + {"limit:25", actionLimit, 25}, + {"limit:0", actionLimit, 0}, + } { + // An hour, the shortest refresh allowed. + cfg := fromEnvironment(t, environment{ + blocklistURLs: dropURL + ", " + torURL, blocklistRefresh: "1h", + blocklistAction: tc.value, asnLimitPercentURL: asnURL, + }) + + if !slices.Equal(cfg.BlocklistURLs, []string{dropURL, torURL}) || + cfg.BlocklistRefresh != time.Hour || cfg.BlocklistAction != tc.action || + cfg.BlocklistLimitPercent != tc.percent || cfg.ASNLimitPercentURL != asnURL { + t.Errorf("%s=%s gave %v, %s, %s, %d and %s", blocklistAction, tc.value, + cfg.BlocklistURLs, cfg.BlocklistRefresh, cfg.BlocklistAction, + cfg.BlocklistLimitPercent, cfg.ASNLimitPercentURL) + } + } +} + +func TestInvalidReputationSettingStopsTheStartSayingWhatIsWrong(t *testing.T) { + t.Parallel() + + const ( + notURL = " is not an http or https URL without a user or a fragment, " + + "such as https://www.spamhaus.org/drop/drop.txt" + notAnHour = " is not a duration of 1h or more, such as 24h" + notAction = " is not deny, limit: such as limit:25, or log" + ) + + for _, tc := range []struct{ name, value, want string }{ + { + blocklistURLs, "ftp://lists.example/drop.txt", + `"ftp://lists.example/drop.txt"` + notURL, + }, + {blocklistURLs, "lists.example/drop.txt", `"lists.example/drop.txt"` + notURL}, + { + blocklistURLs, "https://me:secret@lists.example/drop.txt", + `"https://me:secret@lists.example/drop.txt"` + notURL, + }, + { + blocklistURLs, dropURL + "," + torURL + "," + dropURL, + `"` + dropURL + `" is listed twice`, + }, + {asnLimitPercentURL, asnURL + "#top", `"` + asnURL + `#top"` + notURL}, + {blocklistRefresh, "59m", `"59m"` + notAnHour}, + {blocklistRefresh, off, `"off"` + notAnHour}, + {blocklistRefresh, "a day", `"a day"` + notAnHour}, + {blocklistAction, "block", `"block"` + notAction}, + {blocklistAction, actionLimit, `"limit"` + notAction}, + {blocklistAction, "limit:101", `"limit:101"` + notAction}, + } { + t.Run(tc.name+"="+tc.value, func(t *testing.T) { + t.Parallel() + + _, err := config.FromEnvironment(environment{tc.name: tc.value}.lookupEnv) + + want := tc.name + ": " + tc.want + if err == nil || err.Error() != want { + t.Errorf("error %v, want %s", err, want) + } + }) + } +} + +func TestASNLimitPercentURLThatIsABlocklistStopsTheStart(t *testing.T) { + t.Parallel() + + _, err := config.FromEnvironment(environment{ + blocklistURLs: dropURL + "," + asnURL, asnLimitPercentURL: asnURL, + }.lookupEnv) + + want := asnLimitPercentURL + `: "` + asnURL + `" is in SWWAF_BLOCKLIST_URLS too` + if err == nil || err.Error() != want { + t.Errorf("error %v, want %s", err, want) + } +} + func TestSizesAndOff(t *testing.T) { t.Parallel() @@ -1628,6 +1745,10 @@ func TestLogsEachSettingWithItsValue(t *testing.T) { asnBytesPercent: "", countryBytesPercent: "", unknownLimitPercent: "100", + asnLimitPercentURL: "", + blocklistURLs: "", + blocklistRefresh: "24h", + blocklistAction: actionDeny, banResponse: "403", limitBanDuration: "1h", limitBanRepeatWindow: "24h", diff --git a/internal/metrics/metrics.go b/internal/metrics/metrics.go index 2dda76c..af5df88 100644 --- a/internal/metrics/metrics.go +++ b/internal/metrics/metrics.go @@ -16,6 +16,7 @@ import ( "sneak.berlin/go/smallwebwaf/internal/bans" "sneak.berlin/go/smallwebwaf/internal/ratelimit" "sneak.berlin/go/smallwebwaf/internal/remotelog" + "sneak.berlin/go/smallwebwaf/internal/reputation" "sneak.berlin/go/smallwebwaf/internal/requestlog" "sneak.berlin/go/smallwebwaf/internal/rules" ) @@ -35,10 +36,12 @@ type Metrics struct { rateLimitHits *prometheus.CounterVec sizeAndTimeLimitHits *prometheus.CounterVec offences *prometheus.CounterVec - // ruleMatches are made by AddRules. - ruleMatches *prometheus.CounterVec - countries *busiest - asns *busiest + // ruleMatches are made by AddRules, and reputationHits by + // AddReputation. + ruleMatches *prometheus.CounterVec + reputationHits *prometheus.CounterVec + countries *busiest + asns *busiest // GeoJSRequests are the requests to GeoJS, and GeoJSFailures those // that failed. GeoJSUnanswered are the requests that needed their @@ -253,6 +256,51 @@ func (m *Metrics) AddLookupFile(lastRead func() time.Time, readFailures func() i ) } +// AddReputation adds the metrics of the lists fetched from URLs, by +// source, each list's URL: the requests whose client a blocklist lists, +// which ReputationHit counts, and, read from lists as the metrics are +// asked for, the fetches that failed and when the copy in use was fetched. +// It is called once, before ReputationHit. +func (m *Metrics) AddReputation(lists *reputation.Lists) { + m.reputationHits = counterVec("smallwebwaf_reputation_hits_total", + "Requests whose client a blocklist lists, by the blocklist's URL.", + []string{"source"}) + m.registry.MustRegister(m.reputationHits) + + for _, listURL := range lists.URLs() { + source := prometheus.Labels{"source": listURL} + + m.registry.MustRegister( + prometheus.NewCounterFunc(prometheus.CounterOpts{ + Name: "smallwebwaf_reputation_failures_total", + Help: "Fetches of the list that failed.", + ConstLabels: source, + }, func() float64 { + return float64(lists.Failures(listURL)) + }), + prometheus.NewGaugeFunc(prometheus.GaugeOpts{ + Name: "smallwebwaf_reputation_last_fetch_timestamp_seconds", + Help: "When the copy of the list in use was fetched, in seconds since " + + "1970, or 0 while there is none.", + ConstLabels: source, + }, func() float64 { + fetched := lists.Fetched(listURL) + if fetched.IsZero() { + return 0 + } + + return float64(fetched.Unix()) + }), + ) + } +} + +// ReputationHit counts a request whose client the blocklist at source, its +// URL, lists. +func (m *Metrics) ReputationHit(source string) { + m.reputationHits.WithLabelValues(source).Inc() +} + // AddAlerts adds the metrics of the alerts sent to each destination set, // read from queue as the metrics are asked for, by destination: the // alerts sent, the requests to the destination that failed, the alerts diff --git a/internal/proxy/biased.go b/internal/proxy/biased.go index 2b879e9..aab89ca 100644 --- a/internal/proxy/biased.go +++ b/internal/proxy/biased.go @@ -17,33 +17,48 @@ type percentage struct { } // biasedThresholdsSet reports whether a biased threshold can lower a -// client's limits: one of its lists is not empty, or -// SWWAF_UNKNOWN_LIMIT_PERCENT is below 100. The client's lookup is then -// needed before its request goes on. +// client's limits: one of its lists is not empty, +// SWWAF_UNKNOWN_LIMIT_PERCENT is below 100, or SWWAF_ASN_LIMIT_PERCENT_URL +// is set. The client's lookup is then needed before its request goes on. func biasedThresholdsSet(cfg *config.Config) bool { return len(cfg.ASNLimitPercent) > 0 || len(cfg.CountryLimitPercent) > 0 || len(cfg.ASNBytesPercent) > 0 || len(cfg.CountryBytesPercent) > 0 || - cfg.UnknownLimitPercent < whole + cfg.UnknownLimitPercent < whole || cfg.ASNLimitPercentURL != "" } -// limitPercentages returns a client's limit percentages, for the rate +// limitPercentages returns the client's limit percentages, for the rate // limits and for the byte limits, by its AS number and country as looked -// up, each "" when unknown. Each is the lowest of those the settings give -// it, the first of them in the order below when several are lowest: the -// percentage SWWAF_ASN_LIMIT_PERCENT gives its AS number, the one -// SWWAF_COUNTRY_LIMIT_PERCENT gives its country, and, for a client -// without a country, SWWAF_UNKNOWN_LIMIT_PERCENT. For the byte limits, -// SWWAF_ASN_BYTES_PERCENT and SWWAF_COUNTRY_BYTES_PERCENT take the place -// of the first two for an AS number or a country they list. -func limitPercentages( - cfg *config.Config, asn, country string, -) (percentage, percentage) { +// up, each "" when unknown, and the blocklists that list it. Each is the +// lowest of those the settings give it, the first of them in the order +// below when several are lowest: the percentage SWWAF_ASN_LIMIT_PERCENT +// gives its AS number, the one the file SWWAF_ASN_LIMIT_PERCENT_URL names +// gives it, the one SWWAF_COUNTRY_LIMIT_PERCENT gives its country, for a +// client without a country, SWWAF_UNKNOWN_LIMIT_PERCENT, and for a client +// a blocklist lists, the percentage of SWWAF_BLOCKLIST_ACTION while it is +// limit. For the byte limits, SWWAF_ASN_BYTES_PERCENT and +// SWWAF_COUNTRY_BYTES_PERCENT take the place of the first three for an AS +// number or a country they list. +func (rq *request) limitPercentages() (percentage, percentage) { + cfg := rq.h.config + asn, country := rq.line.ASN, rq.line.Country + unknown := percentage{percent: whole} if country == "" { unknown = percentage{cfg.UnknownLimitPercent, "SWWAF_UNKNOWN_LIMIT_PERCENT"} } - asnRequests := given(cfg.ASNLimitPercent, asn, "SWWAF_ASN_LIMIT_PERCENT") + fetched := percentage{percent: whole} + if percent, listed := rq.h.lists.ASNLimitPercent(asn); listed { + fetched = percentage{percent, "SWWAF_ASN_LIMIT_PERCENT_URL"} + } + + listed := percentage{percent: whole} + if len(rq.line.Reputation) > 0 && cfg.BlocklistAction == "limit" { + listed = percentage{cfg.BlocklistLimitPercent, "SWWAF_BLOCKLIST_ACTION"} + } + + asnRequests := lowest(given(cfg.ASNLimitPercent, asn, "SWWAF_ASN_LIMIT_PERCENT"), + fetched) countryRequests := given(cfg.CountryLimitPercent, country, "SWWAF_COUNTRY_LIMIT_PERCENT") @@ -56,8 +71,8 @@ func limitPercentages( countryBytes = given(cfg.CountryBytesPercent, country, "SWWAF_COUNTRY_BYTES_PERCENT") } - return lowest(asnRequests, countryRequests, unknown), - lowest(asnBytes, countryBytes, unknown) + return lowest(asnRequests, countryRequests, unknown, listed), + lowest(asnBytes, countryBytes, unknown, listed) } // given returns the percentage percents, the setting named setting, gives diff --git a/internal/proxy/biased_test.go b/internal/proxy/biased_test.go index 9f4505f..03d4c33 100644 --- a/internal/proxy/biased_test.go +++ b/internal/proxy/biased_test.go @@ -310,6 +310,7 @@ func TestRequestWaitsForItsLookupWhileABiasedThresholdIsSet(t *testing.T) { {asnBytesPercent, asnDEHalf, true}, {countryBytesPercent, countryDEHalf, true}, {unknownLimitPercent, "99", true}, + {asnLimitPercentURL, asnURL, true}, // At 100, its default, it lowers no limit. {unknownLimitPercent, "100", false}, } { diff --git a/internal/proxy/proxy.go b/internal/proxy/proxy.go index 4df438f..170eb92 100644 --- a/internal/proxy/proxy.go +++ b/internal/proxy/proxy.go @@ -18,6 +18,7 @@ import ( "sneak.berlin/go/smallwebwaf/internal/lookup" "sneak.berlin/go/smallwebwaf/internal/metrics" "sneak.berlin/go/smallwebwaf/internal/ratelimit" + "sneak.berlin/go/smallwebwaf/internal/reputation" "sneak.berlin/go/smallwebwaf/internal/requestlog" "sneak.berlin/go/smallwebwaf/internal/rules" ) @@ -69,14 +70,16 @@ type Params struct { // against. Rules *rules.Files // Alerts receive the alert for each ban the proxy makes or makes - // permanent, for each count over an anomaly threshold, and for GeoJS - // failing. + // permanent, for each count over an anomaly threshold, for each request + // whose client a blocklist lists, and for GeoJS failing or a fetch of a + // list failing. Alerts *alerts.Queue } // Server is the server smallwebwaf runs, with the parts of the proxy // whose state the state files keep, the lookup database, nil unless -// SWWAF_LOOKUP_SOURCE is file, and the metrics. +// SWWAF_LOOKUP_SOURCE is file, the lists fetched from URLs, which its Run +// fetches, and the metrics. type Server struct { *http.Server @@ -85,6 +88,7 @@ type Server struct { GeoJS *lookup.GeoJS Anomalies *anomaly.Counters LookupFile *lookup.File + Lists *reputation.Lists Metrics *metrics.Metrics } @@ -132,8 +136,13 @@ func New(params Params) *Server { Alerts: params.Alerts, }), lookupFile: params.LookupFile, - rules: params.Rules, - alerts: params.Alerts, + lists: reputation.New(reputation.Params{ + BlocklistURLs: params.Config.BlocklistURLs, Refresh: params.Config.BlocklistRefresh, + ASNLimitPercentURL: params.Config.ASNLimitPercentURL, Now: params.Now, + ProcessLog: params.ProcessLog, Alerts: params.Alerts, + }), + rules: params.Rules, + alerts: params.Alerts, } h.geojs = lookup.New(lookup.Params{ URL: params.GeoJSURL, @@ -151,6 +160,7 @@ func New(params Params) *Server { }) m.AddBansAndClients(h.ledger, h.limiter, params.Now) m.AddRules(params.Rules) + m.AddReputation(h.lists) return &Server{ Server: &http.Server{ @@ -170,6 +180,7 @@ func New(params Params) *Server { GeoJS: h.geojs, Anomalies: h.anomalies, LookupFile: h.lookupFile, + Lists: h.lists, Metrics: m, } } @@ -189,6 +200,7 @@ type handler struct { geojs *lookup.GeoJS anomalies *anomaly.Counters lookupFile *lookup.File + lists *reputation.Lists rules *rules.Files alerts *alerts.Queue } diff --git a/internal/proxy/reputation.go b/internal/proxy/reputation.go new file mode 100644 index 0000000..2867b2f --- /dev/null +++ b/internal/proxy/reputation.go @@ -0,0 +1,32 @@ +package proxy + +import ( + "sneak.berlin/go/smallwebwaf/internal/alerts" +) + +// blocklistDenied notes in the log line the URLs of the blocklists that +// list the client, counts each of them in the metrics and raises a +// reputation_hit alert for it, and reports whether SWWAF_BLOCKLIST_ACTION, +// being deny, refuses the request. Being limit, it lowers the client's +// limits instead (see limitPercentages), and being log, it does nothing +// more. +func (rq *request) blocklistDenied() bool { + listedBy := rq.h.lists.ListedBy(rq.client) + rq.line.Reputation = listedBy + + for _, listURL := range listedBy { + rq.h.metrics.ReputationHit(listURL) + rq.h.alerts.Raise(alerts.Alert{ + Event: alerts.EventReputationHit, + Client: rq.client, + Netblock: clientGroup(rq.client), + ASN: rq.line.ASN, + ASName: rq.line.ASName, + Country: rq.line.Country, + Reason: "listed by a blocklist", + Detail: map[string]any{"source": listURL}, + }) + } + + return len(listedBy) > 0 && rq.h.config.BlocklistAction == "deny" +} diff --git a/internal/proxy/reputation_test.go b/internal/proxy/reputation_test.go new file mode 100644 index 0000000..9381f29 --- /dev/null +++ b/internal/proxy/reputation_test.go @@ -0,0 +1,333 @@ +package proxy_test + +import ( + "maps" + "net/http" + "net/netip" + "slices" + "testing" + "time" + + "sneak.berlin/go/smallwebwaf/internal/alerts" + "sneak.berlin/go/smallwebwaf/internal/proxy" + "sneak.berlin/go/smallwebwaf/internal/reputation" + "sneak.berlin/go/smallwebwaf/internal/requestlog" +) + +// The reputation settings. +const ( + blocklistURLs = "SWWAF_BLOCKLIST_URLS" + blocklistAction = "SWWAF_BLOCKLIST_ACTION" + asnLimitPercentURL = "SWWAF_ASN_LIMIT_PERCENT_URL" +) + +// The actions of SWWAF_BLOCKLIST_ACTION but limit, which has a +// percentage. +const ( + actionDeny = "deny" + actionLog = "log" +) + +// The lists these tests name, which are never fetched: each test puts in +// the copies it needs, as reputation.json would at start. +const ( + dropURL = "https://lists.example/drop.txt" + torURL = "https://lists.example/tor.txt" + asnURL = "https://lists.example/asn.txt" +) + +func TestEachBlocklistActionForAListedAddressAndAListedNetblock(t *testing.T) { + t.Parallel() + + forward, denied := requestlog.ActionForward, requestlog.ActionDenied + + for _, tc := range []struct { + action string + // statuses and actions are those of a listed client's three + // requests, and percent their limit_percent, as percentText gives it. + statuses []int + actions []string + percent string + }{ + { + actionDeny, []int{http.StatusForbidden, http.StatusForbidden, http.StatusForbidden}, + []string{denied, denied, denied}, none, + }, + { + // Half of 4 requests a minute: the third breaks the limit. + "limit:50", []int{http.StatusOK, http.StatusOK, http.StatusForbidden}, + []string{forward, forward, requestlog.ActionRateLimited}, + "50 from " + blocklistAction, + }, + { + actionLog, []int{http.StatusOK, http.StatusOK, http.StatusOK}, + []string{forward, forward, forward}, none, + }, + } { + t.Run(tc.action, func(t *testing.T) { + t.Parallel() + + s, server, _ := startWithLookups(t, map[string]string{ + rateLimitPerMinute: fourAMinute, blocklistURLs: dropURL, + blocklistAction: tc.action, + }) + // fromDE is listed as an address, and fromKP in a netblock. + loadLists(t, server, map[string][]string{ + dropURL: {"; DROP", fromDE, "198.51.100.0/24 ; SBL1"}, + }) + + for _, from := range []string{fromDE, fromKP} { + for i := range 3 { + line := s.get(from, tc.statuses[i], tc.actions[i]) + wantReputation(t, line, dropURL) + wantPercent(t, "limit_percent", line.LimitPercent, + line.LimitPercentSetting, tc.percent) + + // A request refused for the list is not counted. + counted := line.fields["counts"] != nil + if counted != (tc.actions[i] != denied) { + t.Errorf("request from %s counted %t, logged %s", from, counted, + tc.actions[i]) + } + } + } + + // A client no list lists has the whole limit. + for range 3 { + line := s.get(unplaced, http.StatusOK, forward) + wantReputation(t, line) + wantPercent(t, "limit_percent", line.LimitPercent, line.LimitPercentSetting, + none) + } + + // A refusal for the list makes no ban. + if held := server.Ledger.Snapshot(); tc.action == actionDeny && len(held) != 0 { + t.Errorf("bans %+v, want none", held) + } + }) + } +} + +func TestBlocklistsComeAfterTheCountryListsAndSkipAllowNets(t *testing.T) { + t.Parallel() + + s, server, queue := startWithLookups(t, map[string]string{ + blocklistURLs: dropURL, deniedCountries: "kp", allowNets: fromDE, + }) + loadLists(t, server, map[string][]string{dropURL: {fromDE, fromKP}}) + + // fromKP's country refuses it before the list is looked at, and fromDE, + // in SWWAF_ALLOW_NETS, is not checked at all: neither is noted, nor + // alerted. + wantReputation(t, s.get(fromKP, http.StatusForbidden, requestlog.ActionCountryDenied)) + wantReputation(t, s.get(fromDE, http.StatusOK, requestlog.ActionForward)) + wantAlerts(t, queue) +} + +func TestObserveModeForwardsAClientABlocklistDeniesAndAlertsIt(t *testing.T) { + t.Parallel() + + s, server, queue := startWithLookups(t, map[string]string{ + blocklistURLs: dropURL, mode: observe, + }) + loadLists(t, server, map[string][]string{dropURL: {fromDE}}) + + line := s.get(fromDE, http.StatusOK, requestlog.ActionForward) + wantWouldAction(t, line, requestlog.ActionDenied) + wantReputation(t, line, dropURL) + + waiting := queue.Snapshot().Waiting[alerts.DestinationWebhook] + if len(waiting) != 1 || waiting[0].Event != alerts.EventReputationHit { + t.Errorf("alerts waiting %+v, want a reputation_hit alert", waiting) + } +} + +func TestBlocklistLimitTakesPartInTheLowestPercentageOfEveryLimit(t *testing.T) { + t.Parallel() + + for _, tc := range []struct { + action, asnPercent string + // want is the upload's limit_percent and bytes_percent, as + // percentText gives them, and limitHit its limit_hit. + want, limitHit string + }{ + {"limit:50", asnDEQuarter, "25 from " + asnLimitPercent, minuteBytes}, + {"limit:25", asnDEHalf, "25 from " + blocklistAction, minuteBytes}, + // The AS number's, the first of two alike. + {"limit:25", asnDEQuarter, "25 from " + asnLimitPercent, minuteBytes}, + {actionLog, asnDE + ":100", none, ""}, + } { + t.Run(tc.action+" "+tc.asnPercent, func(t *testing.T) { + t.Parallel() + + s, server, _ := startWithLookups(t, map[string]string{ + bytesLimitPerMinute: twoUploads, blocklistURLs: dropURL, + blocklistAction: tc.action, asnLimitPercent: tc.asnPercent, + }) + loadLists(t, server, map[string][]string{dropURL: {fromDE}}) + + // The upload's 100 bytes are over 49, a quarter of 199, and 99, + // half of it, and within 199. + line := s.uploadFrom(fromDE) + wantPercent(t, "limit_percent", line.LimitPercent, line.LimitPercentSetting, + tc.want) + wantPercent(t, "bytes_percent", line.BytesPercent, line.BytesPercentSetting, + tc.want) + + if line.LimitHit != tc.limitHit { + t.Errorf("log line has limit_hit %q, want %q", line.LimitHit, tc.limitHit) + } + }) + } +} + +func TestASNLimitPercentFileCountsAsTheSettingDoesTheLowerWinning(t *testing.T) { + t.Parallel() + + const ( + fromURL = "25 from " + asnLimitPercentURL + fromSetting = "25 from " + asnLimitPercent + ) + + for _, tc := range []struct { + name string + env map[string]string + file string + // limitPercent and bytesPercent are the upload's, as percentText + // gives them. + limitPercent, bytesPercent string + }{ + {"the file's alone", nil, asnDEQuarter, fromURL, fromURL}, + { + "the file's, lower than the setting's", + map[string]string{asnLimitPercent: asnDEHalf}, asnDEQuarter, fromURL, fromURL, + }, + { + "the setting's, lower than the file's", + map[string]string{asnLimitPercent: asnDEQuarter}, asnDEHalf, + fromSetting, fromSetting, + }, + { + "the setting's, the first of two alike", + map[string]string{asnLimitPercent: asnDEQuarter}, asnDEQuarter, + fromSetting, fromSetting, + }, + {"none, for an AS number the file does not list", nil, asnKP + ":25", none, none}, + { + "SWWAF_ASN_BYTES_PERCENT's in place of the file's for the byte limits", + map[string]string{asnBytesPercent: asnDE + ":100"}, asnDEQuarter, fromURL, none, + }, + } { + t.Run(tc.name, func(t *testing.T) { + t.Parallel() + + env := map[string]string{asnLimitPercentURL: asnURL} + maps.Copy(env, tc.env) + + s, server, _ := startWithLookups(t, env) + loadLists(t, server, map[string][]string{asnURL: {"# by AS number", tc.file}}) + + line := s.uploadFrom(fromDE) + wantPercent(t, "limit_percent", line.LimitPercent, line.LimitPercentSetting, + tc.limitPercent) + wantPercent(t, "bytes_percent", line.BytesPercent, line.BytesPercentSetting, + tc.bytesPercent) + }) + } +} + +func TestEachBlocklistThatListsAClientRaisesAnAlertOncePerCooldownAndIsCounted( + t *testing.T, +) { + t.Parallel() + + const emptyURL = "https://lists.example/empty.txt" + + s, clk, server, queue := startWithLookupsAndClock(t, map[string]string{ + blocklistURLs: dropURL + "," + torURL + "," + emptyURL, + blocklistAction: actionLog, + metricsToken: token, + }) + loadLists(t, server, map[string][]string{dropURL: {fromDE}, torURL: {fromDE}}) + + // The second request's alerts are repeats, which the cooldown holds + // back. + for range 2 { + wantReputation(t, s.get(fromDE, http.StatusOK, requestlog.ActionForward), + dropURL, torURL) + } + + hit := func(source string) alerts.Alert { + return alerts.Alert{ + Instance: alertInstance, + Time: clk.Now(), + Event: alerts.EventReputationHit, + Client: netip.MustParseAddr(fromDE), + Netblock: netip.MustParsePrefix(fromDE + "/32"), + ASN: asnDE, + ASName: asNameDE, + Country: "DE", + Reason: "listed by a blocklist", + Detail: map[string]any{"source": source}, + } + } + wantAlerts(t, queue, hit(dropURL), hit(torURL)) + + if queue.Suppressed() != 2 { + t.Errorf("%d alerts held back, want the second request's 2", queue.Suppressed()) + } + + // Each list's hits, none of its fetches failed, and when its copy was + // fetched, 0 for the one without. + metrics := s.scrape(unplaced) + fetched := float64(listsFetched().Unix()) + + for listURL, want := range map[string]struct{ hits, fetched float64 }{ + dropURL: {2, fetched}, torURL: {2, fetched}, emptyURL: {0, 0}, + } { + labels := `{instance="` + alertInstance + `",source="` + listURL + `"}` + + if want.hits == 0 { + wantNoSeries(t, metrics, "smallwebwaf_reputation_hits_total"+labels) + } else { + wantMetric(t, metrics, "smallwebwaf_reputation_hits_total"+labels, want.hits) + } + + wantMetric(t, metrics, "smallwebwaf_reputation_failures_total"+labels, 0) + wantMetric(t, metrics, "smallwebwaf_reputation_last_fetch_timestamp_seconds"+labels, + want.fetched) + } +} + +// listsFetched is when loadLists has the copies fetched. +func listsFetched() time.Time { + return time.Date(2026, 10, 5, 0, 0, 0, 0, time.UTC) +} + +// loadLists puts copies of lists into server's lists, by URL, each with +// its lines, fetched at listsFetched, as reputation.json would at start. +func loadLists(t *testing.T, server *proxy.Server, copies map[string][]string) { + t.Helper() + + lists := make([]reputation.List, 0, len(copies)) + for listURL, lines := range copies { + lists = append(lists, reputation.List{ + URL: listURL, Fetched: listsFetched(), Lines: lines, + }) + } + + err := server.Lists.Load(lists) + if err != nil { + t.Fatalf("load the lists: %v", err) + } +} + +// wantReputation checks the URLs of the blocklists the log line names in +// its reputation. +func wantReputation(t *testing.T, line logLine, want ...string) { + t.Helper() + + if !slices.Equal(line.Reputation, want) { + t.Errorf("log line has reputation %v, want %v", line.Reputation, want) + } +} diff --git a/internal/proxy/request.go b/internal/proxy/request.go index 6d3c220..ad2c463 100644 --- a/internal/proxy/request.go +++ b/internal/proxy/request.go @@ -213,14 +213,14 @@ func (rq *request) check(ctx context.Context) *refusal { // client in SWWAF_ALLOW_NETS skips them, and is not looked up. For any // other client, SWWAF_DENY_NETS comes first, then a ban on its netblock, // so that a client either refuses is not looked up, then the lookup of -// its AS number and country, and then the country lists; a request any of -// them refuses is not counted for the rate limits. Then come the rate -// limits, unless the client is in SWWAF_RATE_LIMIT_EXEMPT_NETS or the -// request's path is exempt under SWWAF_RATE_LIMIT_EXEMPT_PATHS, so that -// every other request is counted, each of them by the client's limit -// percentages, and last the rule files. A request exempt from the rate -// limits is exempt from the byte limits too. ctx is the request's own -// context. +// its AS number and country, then the country lists, and then the +// blocklists; a request any of them refuses is not counted for the rate +// limits. Then come the rate limits, unless the client is in +// SWWAF_RATE_LIMIT_EXEMPT_NETS or the request's path is exempt under +// SWWAF_RATE_LIMIT_EXEMPT_PATHS, so that every other request is counted, +// each of them by the client's limit percentages, and last the rule +// files. A request exempt from the rate limits is exempt from the byte +// limits too. ctx is the request's own context. func (rq *request) checkClient(ctx context.Context) string { cfg := rq.h.config if isInside(rq.client, cfg.AllowNets) { @@ -243,10 +243,14 @@ func (rq *request) checkClient(ctx context.Context) string { return requestlog.ActionCountryDenied } + if rq.blocklistDenied() { + return requestlog.ActionDenied + } + rq.counted = !isInside(rq.client, cfg.RateLimitExemptNets) && !pathExempt(rq.in.URL, cfg.RateLimitExemptPaths) if rq.counted { - rq.limitPercent, rq.bytesPercent = limitPercentages(cfg, rq.line.ASN, rq.line.Country) + rq.limitPercent, rq.bytesPercent = rq.limitPercentages() rq.line.LimitPercent, rq.line.LimitPercentSetting = rq.limitPercent.logged() rq.line.BytesPercent, rq.line.BytesPercentSetting = rq.bytesPercent.logged() } diff --git a/internal/reputation/export_test.go b/internal/reputation/export_test.go new file mode 100644 index 0000000..4c4035d --- /dev/null +++ b/internal/reputation/export_test.go @@ -0,0 +1,9 @@ +package reputation + +import "net/http" + +// SetTransport has l's fetches go through transport instead of the +// network. +func (l *Lists) SetTransport(transport http.RoundTripper) { + l.httpClient.Transport = transport +} diff --git a/internal/reputation/reputation.go b/internal/reputation/reputation.go new file mode 100644 index 0000000..45872cd --- /dev/null +++ b/internal/reputation/reputation.go @@ -0,0 +1,499 @@ +// Package reputation fetches the lists the settings name by URL: the +// blocklists of SWWAF_BLOCKLIST_URLS, and the file of AS:percent lines +// SWWAF_ASN_LIMIT_PERCENT_URL names. It keeps the last good copy of each, +// whole, comment lines included, which is used while a fetch fails, and +// when each was last tried, which the state package writes to +// reputation.json and reads from it, so that a restart keeps them too. +package reputation + +import ( + "context" + "errors" + "fmt" + "io" + "log/slog" + "net/http" + "net/netip" + "slices" + "strings" + "sync" + "time" + + "sneak.berlin/go/smallwebwaf/internal/alerts" + "sneak.berlin/go/smallwebwaf/internal/config" +) + +const ( + // maxListBytes is the most of a list that is read. A longer one is a + // failure, so that a wrong URL cannot fill the memory. + maxListBytes = 16 << 20 + // fetchTimeout bounds one fetch of a list. + fetchTimeout = time.Minute + // mappedBits is the length of ::ffff:0.0.0.0/96, the netblock of every + // IPv4-mapped address. + mappedBits = 96 +) + +var ( + errStatus = errors.New("the server answered") + errTooLong = errors.New("the list is longer than 16 MiB") + errNotNetblock = errors.New("is not an address or a netblock, such as 192.0.2.0/24") + errNotASNPercent = errors.New( + "is not an AS number, : and a percentage, such as AS64496:50") +) + +// List is a list as reputation.json holds it: the URL it is fetched from, +// when it was last tried, the fetch failed or not, and its last good copy: +// when that was fetched, and its lines, as fetched, comment lines +// included, both left out while no fetch of it has succeeded. +type List struct { + URL string `json:"url"` + Tried time.Time `json:"tried"` + Fetched time.Time `json:"fetched,omitzero"` + Lines []string `json:"lines,omitzero"` +} + +// Params are what New needs. +type Params struct { + // BlocklistURLs are the blocklists (SWWAF_BLOCKLIST_URLS), and + // ASNLimitPercentURL the file of AS:percent lines + // (SWWAF_ASN_LIMIT_PERCENT_URL), "" while it is unset. + BlocklistURLs []string + ASNLimitPercentURL string + // Refresh is how long after a list was last fetched or tried it is + // fetched again (SWWAF_BLOCKLIST_REFRESH). + Refresh time.Duration + // Now tells the time, normally time.Now in UTC. + Now func() time.Time + // ProcessLog receives each fetch of a list, and why one failed. + ProcessLog *slog.Logger + // Alerts receive a source_failure alert for each fetch that fails. + Alerts *alerts.Queue +} + +// Lists are the lists Params names, each with its last good copy. They +// are safe for concurrent use. +type Lists struct { + params Params + httpClient *http.Client + + mu sync.Mutex + // lists are by URL, one for each URL Params names. + lists map[string]*list +} + +// list is one list: what reputation.json keeps of it, its last try, zero +// before the first, and its last good copy, what that copy says, and how +// many fetches of it failed. +type list struct { + kept List + entries entries + failures int +} + +// entries are what the lines of a copy say: for a blocklist, the netblocks +// it names, with the lengths among them, and for the file of AS:percent +// lines, the percentage it gives each AS number. +type entries struct { + netblocks map[netip.Prefix]bool + lengths []int + percents map[string]int64 +} + +// New returns the lists, without a copy of any yet. +func New(params Params) *Lists { + l := &Lists{params: params, httpClient: &http.Client{}, lists: map[string]*list{}} + + for _, listURL := range l.URLs() { + l.lists[listURL] = &list{kept: List{URL: listURL}} + } + + return l +} + +// URLs returns the URL of every list: the blocklists' in the order +// SWWAF_BLOCKLIST_URLS names them, then SWWAF_ASN_LIMIT_PERCENT_URL. +func (l *Lists) URLs() []string { + urls := slices.Clone(l.params.BlocklistURLs) + if l.params.ASNLimitPercentURL != "" { + urls = append(urls, l.params.ASNLimitPercentURL) + } + + return urls +} + +// ListedBy returns the URLs of the blocklists whose copy lists addr, in +// the order SWWAF_BLOCKLIST_URLS names them. +func (l *Lists) ListedBy(addr netip.Addr) []string { + l.mu.Lock() + defer l.mu.Unlock() + + var listedBy []string + + for _, listURL := range l.params.BlocklistURLs { + if l.lists[listURL].entries.contain(addr) { + listedBy = append(listedBy, listURL) + } + } + + return listedBy +} + +// ASNLimitPercent returns the percentage the copy of the file of +// AS:percent lines gives asn, and whether it lists asn. +func (l *Lists) ASNLimitPercent(asn string) (int64, bool) { + if l.params.ASNLimitPercentURL == "" { + return 0, false + } + + l.mu.Lock() + defer l.mu.Unlock() + + percent, listed := l.lists[l.params.ASNLimitPercentURL].entries.percents[asn] + + return percent, listed +} + +// Fetched returns when the copy in use of the list at listURL was +// fetched, or zero while there is none. +func (l *Lists) Fetched(listURL string) time.Time { + l.mu.Lock() + defer l.mu.Unlock() + + return l.lists[listURL].kept.Fetched +} + +// Failures returns how many fetches of the list at listURL failed. +func (l *Lists) Failures(listURL string) int { + l.mu.Lock() + defer l.mu.Unlock() + + return l.lists[listURL].failures +} + +// Run fetches each list once Refresh has passed since it was last fetched +// or tried, the later of the two, until ctx is done. A list never tried is +// fetched at once, and so is one whose last try or copy, read from +// reputation.json, is that old. +func (l *Lists) Run(ctx context.Context) { + if len(l.lists) == 0 { + return + } + + for ctx.Err() == nil { + next := l.fetchDue(ctx) + timer := time.NewTimer(next.Sub(l.params.Now())) + + select { + case <-ctx.Done(): + case <-timer.C: + } + + timer.Stop() + } +} + +// Snapshot returns each list that has been tried, with its copy, if it +// has one, sorted by URL, as reputation.json lists them. +func (l *Lists) Snapshot() []List { + l.mu.Lock() + + tried := make([]List, 0, len(l.lists)) + + for _, held := range l.lists { + if !held.kept.Tried.IsZero() { + tried = append(tried, held.kept) + } + } + + l.mu.Unlock() + + slices.SortFunc(tried, func(a, b List) int { + return strings.Compare(a.URL, b.URL) + }) + + return tried +} + +// Load puts lists, read from reputation.json, in place of the last tries +// and copies held. A list Params does not name is dropped. A copy with a +// line that parse refuses is an error, and then nothing changes. +func (l *Lists) Load(lists []List) error { + found := make(map[string]entries, len(lists)) + + for _, kept := range lists { + if _, named := l.lists[kept.URL]; !named { + continue + } + + read, err := l.parse(kept.URL, kept.Lines) + if err != nil { + return fmt.Errorf("the copy of %s: %w", kept.URL, err) + } + + found[kept.URL] = read + } + + l.mu.Lock() + defer l.mu.Unlock() + + for listURL, held := range l.lists { + held.kept, held.entries = List{URL: listURL}, entries{} + } + + for _, kept := range lists { + read, named := found[kept.URL] + if named { + l.lists[kept.URL].kept, l.lists[kept.URL].entries = kept, read + } + } + + return nil +} + +// fetchDue fetches each list that is due, one after another, and returns +// when the next is due. Once ctx has ended, it starts none, since a fetch +// cut off is noted as a try. +func (l *Lists) fetchDue(ctx context.Context) time.Time { + var next time.Time + + for _, listURL := range l.URLs() { + due := l.due(listURL) + if ctx.Err() == nil && !l.params.Now().Before(due) { + l.fetch(ctx, listURL) + due = l.due(listURL) + } + + if next.IsZero() || due.Before(next) { + next = due + } + } + + return next +} + +// due returns when the list at listURL is to be fetched: Refresh after it +// was last fetched or tried, the later of the two. +func (l *Lists) due(listURL string) time.Time { + l.mu.Lock() + defer l.mu.Unlock() + + held := l.lists[listURL] + + last := held.kept.Fetched + if held.kept.Tried.After(last) { + last = held.kept.Tried + } + + return last.Add(l.params.Refresh) +} + +// fetch fetches the list at listURL, and notes the try. A good copy takes +// the place of the one held. A failure leaves that in use, and is counted, +// logged and raised as a source_failure alert. A fetch cut off as ctx +// ends, as smallwebwaf stops, is no failure, but is still noted as a try, +// so that a restart waits for it: the server may have had its request. +func (l *Lists) fetch(ctx context.Context, listURL string) { + lines, err := l.get(ctx, listURL) + + var found entries + if err == nil { + found, err = l.parse(listURL, lines) + } + + cutOff := err != nil && ctx.Err() != nil + now := l.params.Now() + + l.mu.Lock() + + held := l.lists[listURL] + held.kept.Tried = now + + if err == nil { + held.kept.Fetched, held.kept.Lines = now, lines + held.entries = found + } else if !cutOff { + held.failures++ + } + + l.mu.Unlock() + + if cutOff { + return + } + + if err != nil { + const failed = "fetching a list failed" + + // Raised before it is logged, so that the alert is there once the + // log line is. + l.params.Alerts.Raise(alerts.Alert{ + Event: alerts.EventSourceFailure, + Reason: failed, + Detail: map[string]any{"source": listURL, "error": err.Error()}, + }) + l.params.ProcessLog.Warn(failed, "url", listURL, "error", err.Error()) + + return + } + + l.params.ProcessLog.Info("fetched a list", "url", listURL, "lines", len(lines)) +} + +// get fetches the list at listURL, and returns its lines. An answer other +// than 200, or a list longer than maxListBytes, is a failure. +func (l *Lists) get(ctx context.Context, listURL string) ([]string, error) { + ctx, cancel := context.WithTimeout(ctx, fetchTimeout) + defer cancel() + + req, err := http.NewRequestWithContext(ctx, http.MethodGet, listURL, http.NoBody) + if err != nil { + return nil, fmt.Errorf("make the request: %w", err) + } + + res, err := l.httpClient.Do(req) + if err != nil { + // Do's error names the URL, which the log line and the alert name + // already: only what went wrong is kept. + return nil, fmt.Errorf("fetch the list: %w", errors.Unwrap(err)) + } + + defer func() { + _ = res.Body.Close() + }() + + if res.StatusCode != http.StatusOK { + return nil, fmt.Errorf("%w %s", errStatus, res.Status) + } + + body, err := io.ReadAll(io.LimitReader(res.Body, maxListBytes+1)) + if err != nil { + return nil, fmt.Errorf("read the list: %w", err) + } + + if len(body) > maxListBytes { + return nil, errTooLong + } + + lines := []string{} + for line := range strings.Lines(string(body)) { + lines = append(lines, strings.TrimSuffix(line, "\n")) + } + + return lines, nil +} + +// parse reads the lines of the list at listURL: those of a blocklist, or +// of the file of AS:percent lines. Anything after a ; or a # on a line is +// left out, and so is a line left blank. Any other line that does not read +// is an error naming it by its number. +func (l *Lists) parse(listURL string, lines []string) (entries, error) { + if listURL == l.params.ASNLimitPercentURL { + return parsePercents(lines) + } + + return parseNetblocks(lines) +} + +// parseNetblocks reads a blocklist's lines, each an address or a netblock +// as the settings take them. +func parseNetblocks(lines []string) (entries, error) { + found := entries{netblocks: map[netip.Prefix]bool{}} + + for i, line := range lines { + text := withoutComment(line) + if text == "" { + continue + } + + netblock, ok := parseNetblock(text) + if !ok { + return entries{}, fmt.Errorf("line %d %w", i+1, errNotNetblock) + } + + found.netblocks[netblock] = true + + if !slices.Contains(found.lengths, netblock.Bits()) { + found.lengths = append(found.lengths, netblock.Bits()) + } + } + + return found, nil +} + +// parseNetblock reads text, a line of a blocklist, and reports whether it +// is an address or a netblock as the settings take them. A client's IPv4 +// address is checked as IPv4, never IPv4-mapped, so an IPv4-mapped line, +// such as ::ffff:192.0.2.0/120, is read as the IPv4 address or netblock it +// stands for, 192.0.2.0/24, and a mapped netblock shorter than /96, which +// stands for none, is refused. +func parseNetblock(text string) (netip.Prefix, bool) { + netblock, err := config.ParseNetblock(text) + if err != nil { + return netip.Prefix{}, false + } + + // The address as written: ParseNetblock's has the bits past the + // netblock's length cleared, the ::ffff among them below /96. + written, _, _ := strings.Cut(text, "/") + if addr, _ := netip.ParseAddr(written); !addr.Is4In6() { + return netblock, true + } + + if netblock.Bits() < mappedBits { + return netip.Prefix{}, false + } + + return netip.PrefixFrom(netblock.Addr().Unmap(), netblock.Bits()-mappedBits), true +} + +// parsePercents reads the lines of the file of AS:percent lines, each an +// AS number, : and a percentage, as SWWAF_ASN_LIMIT_PERCENT takes them. An +// AS number listed more than once gets the lowest of its percentages. +func parsePercents(lines []string) (entries, error) { + found := entries{percents: map[string]int64{}} + + for i, line := range lines { + text := withoutComment(line) + if text == "" { + continue + } + + asnText, percentText, _ := strings.Cut(text, ":") + asn, asnErr := config.ParseASN(asnText) + + percent, percentErr := config.ParsePercent(percentText) + if asnErr != nil || percentErr != nil { + return entries{}, fmt.Errorf("line %d %w", i+1, errNotASNPercent) + } + + earlier, listed := found.percents[asn] + if !listed || percent < earlier { + found.percents[asn] = percent + } + } + + return found, nil +} + +// withoutComment returns line without anything after a ; or a #, and +// without the spaces around what is left. +func withoutComment(line string) string { + text, _, _ := strings.Cut(line, ";") + text, _, _ = strings.Cut(text, "#") + + return strings.TrimSpace(text) +} + +// contain reports whether the netblocks of a blocklist's copy hold addr: +// whether addr, cut to one of their lengths, is one of them. +func (e entries) contain(addr netip.Addr) bool { + for _, length := range e.lengths { + netblock, err := addr.Prefix(length) + if err == nil && e.netblocks[netblock] { + return true + } + } + + return false +} diff --git a/internal/reputation/reputation_test.go b/internal/reputation/reputation_test.go new file mode 100644 index 0000000..b1bdb5c --- /dev/null +++ b/internal/reputation/reputation_test.go @@ -0,0 +1,610 @@ +package reputation_test + +import ( + "bytes" + "context" + "fmt" + "io" + "log/slog" + "net/http" + "net/netip" + "net/url" + "reflect" + "slices" + "strings" + "sync" + "testing" + "testing/synctest" + "time" + + "sneak.berlin/go/smallwebwaf/internal/alerts" + "sneak.berlin/go/smallwebwaf/internal/reputation" +) + +// The tests run in a synctest bubble, where the time package runs on a +// clock of the test's own: time.Sleep moves it on at once, and +// synctest.Wait returns once Run waits for the next list to be due, so +// that every fetch due by then has been made. The stand-in for the +// servers the lists are fetched from answers without the network, since a +// fetch waiting on the network would keep that clock from moving on. + +const ( + // dropURL and torURL are the blocklists, and asnURL the file of + // AS:percent lines. + dropURL = "https://lists.example/drop.txt" + torURL = "https://lists.example/tor.txt" + asnURL = "https://lists.example/asn.txt" + // refresh is the tests' SWWAF_BLOCKLIST_REFRESH, and cooldown their + // SWWAF_ALERT_COOLDOWN, longer than it. + refresh = 24 * time.Hour + cooldown = 48 * time.Hour + // drop is a blocklist as the Spamhaus DROP list is written, with an + // address and a netblock in each of its comments, which list nothing. + drop = "; Spamhaus DROP List 2026/10/07 - (c) 2026 The Spamhaus Project SLL\n" + + "; Last-Modified: Wed, 07 Oct 2026 00:00:00 GMT ; 192.0.2.1\n" + + "# 198.51.100.0/24\n" + + "\n" + + "203.0.113.0/24 ; SBL1\n" + + " 192.0.2.9 # one address\n" + + "2001:db8:1::/48 ; SBL2\n" +) + +func TestListedAddressesAndNetblocksWithTheCommentsLeftOut(t *testing.T) { + t.Parallel() + + synctest.Test(t, func(t *testing.T) { + servers := &standIn{lists: map[string]string{dropURL: drop}} + lists := start(t, servers, params(dropURL)) + + for addr, want := range map[string][]string{ + "203.0.113.0": {dropURL}, + "203.0.113.255": {dropURL}, + "192.0.2.9": {dropURL}, + "2001:db8:1::7": {dropURL}, + "203.0.114.0": nil, + "192.0.2.8": nil, + "192.0.2.1": nil, + "198.51.100.7": nil, + "2001:db8:2::7": nil, + } { + wantListedBy(t, lists, addr, want...) + } + }) +} + +func TestIPv4MappedLineListsTheIPv4AddressOrNetblockItStandsFor(t *testing.T) { + t.Parallel() + + now := time.Now() + lists := reputation.New(params(dropURL)) + + err := lists.Load([]reputation.List{{ + URL: dropURL, Tried: now, Fetched: now, + Lines: []string{"::ffff:192.0.2.9", "::ffff:203.0.113.0/120"}, + }}) + if err != nil { + t.Fatalf("load: %v", err) + } + + for addr, want := range map[string][]string{ + "192.0.2.9": {dropURL}, + "203.0.113.255": {dropURL}, + "192.0.2.8": nil, + "203.0.114.0": nil, + } { + wantListedBy(t, lists, addr, want...) + } + + // A mapped netblock shorter than /96 stands for no IPv4 one. + err = lists.Load([]reputation.List{{ + URL: dropURL, Tried: now, Fetched: now, Lines: []string{"::ffff:198.51.100.0/88"}, + }}) + + const want = "the copy of " + dropURL + + ": line 1 is not an address or a netblock, such as 192.0.2.0/24" + if err == nil || err.Error() != want { + t.Errorf("error %v, want %s", err, want) + } +} + +func TestClientIsListedByEachBlocklistThatListsIt(t *testing.T) { + t.Parallel() + + synctest.Test(t, func(t *testing.T) { + servers := &standIn{lists: map[string]string{ + dropURL: "203.0.113.0/24\n", torURL: "203.0.113.9\n", + }} + lists := start(t, servers, params(torURL, dropURL)) + + // In the order SWWAF_BLOCKLIST_URLS names them. + wantListedBy(t, lists, "203.0.113.9", torURL, dropURL) + wantListedBy(t, lists, "203.0.113.8", dropURL) + }) +} + +func TestListFetchedAgainOnceRefreshHasPassed(t *testing.T) { + t.Parallel() + + synctest.Test(t, func(t *testing.T) { + servers := &standIn{lists: map[string]string{dropURL: "203.0.113.9\n"}} + lists := start(t, servers, params(dropURL)) + began := time.Now() + + wantFetches(t, servers, 1) + + servers.set(dropURL, "203.0.113.10\n") + time.Sleep(refresh - time.Nanosecond) + wantFetches(t, servers, 1) + wantListedBy(t, lists, "203.0.113.9", dropURL) + + time.Sleep(time.Nanosecond) + wantFetches(t, servers, 2) + wantListedBy(t, lists, "203.0.113.9") + wantListedBy(t, lists, "203.0.113.10", dropURL) + + if fetched := lists.Fetched(dropURL); !fetched.Equal(began.Add(refresh)) { + t.Errorf("the copy in use was fetched at %s, want %s", fetched, + began.Add(refresh)) + } + }) +} + +func TestFailedFetchKeepsTheLastGoodCopyAndAlertsOncePerCooldown(t *testing.T) { + t.Parallel() + + for _, tc := range []struct { + name string + // fail has the stand-in answer the fetches after the first so that + // they fail with error. + fail func(servers *standIn) + error string + }{ + { + "an answer other than 200", + func(servers *standIn) { servers.set(dropURL, "") }, + "the server answered 503 Service Unavailable", + }, + { + "a line that does not read", + func(servers *standIn) { servers.set(dropURL, "203.0.113.10\n\n") }, + "line 2 is not an address or a netblock, such as 192.0.2.0/24", + }, + { + "a list longer than 16 MiB", + func(servers *standIn) { + servers.set(dropURL, strings.Repeat("#\n", 8<<20+1)) + }, + "the list is longer than 16 MiB", + }, + } { + t.Run(tc.name, func(t *testing.T) { + t.Parallel() + + synctest.Test(t, func(t *testing.T) { + var log bytes.Buffer + + servers := &standIn{lists: map[string]string{dropURL: drop}} + queue := newQueue() + p := params(dropURL) + p.ProcessLog = slog.New(slog.NewJSONHandler(&log, nil)) + p.Alerts = queue + lists := start(t, servers, p) + kept := lists.Snapshot() + + tc.fail(servers) + + // Each failure is tried again once refresh has passed since it. + for range 2 { + time.Sleep(refresh) + synctest.Wait() + } + + wantFetches(t, servers, 3) + wantListedBy(t, lists, "203.0.113.9", dropURL) + + want := kept[0] + want.Tried = time.Now() + + if got := lists.Snapshot(); !reflect.DeepEqual(got, []reputation.List{want}) { + t.Errorf("lists %+v, want the first copy, last tried now, %+v", got, want) + } + + if lists.Failures(dropURL) != 2 { + t.Errorf("%d failures, want 2", lists.Failures(dropURL)) + } + + // One alert for the first failure; the cooldown holds back the + // second. + wantAlert(t, queue, alerts.Alert{ + Time: time.Now().Add(-refresh), + Event: alerts.EventSourceFailure, + Reason: "fetching a list failed", + Detail: map[string]any{"source": dropURL, "error": tc.error}, + }) + + if !strings.Contains(log.String(), `"msg":"fetching a list failed",`+ + `"url":"`+dropURL+`","error":"`+tc.error) { + t.Errorf("logged\n%s\nwant the failures", log.String()) + } + }) + }) + } +} + +func TestFetchNotDoneWithinAMinuteFails(t *testing.T) { + t.Parallel() + + synctest.Test(t, func(t *testing.T) { + servers := &standIn{lists: map[string]string{}, hanging: true} + lists := start(t, servers, params(dropURL)) + + time.Sleep(time.Minute - time.Nanosecond) + synctest.Wait() + + if lists.Failures(dropURL) != 0 { + t.Errorf("%d failures before a minute, want none", lists.Failures(dropURL)) + } + + time.Sleep(time.Nanosecond) + synctest.Wait() + + if lists.Failures(dropURL) != 1 { + t.Errorf("%d failures after a minute, want 1", lists.Failures(dropURL)) + } + }) +} + +func TestKeptCopyIsFetchedAgainOnceRefreshHasPassedSinceItWasFetched(t *testing.T) { + t.Parallel() + + synctest.Test(t, func(t *testing.T) { + servers := &standIn{lists: map[string]string{ + dropURL: "203.0.113.10\n", torURL: "198.51.100.10\n", + }} + lists := reputation.New(params(dropURL, torURL)) + lists.SetTransport(servers) + + // drop.txt was fetched an hour ago, and tor.txt a refresh ago, as + // reputation.json says at start. + err := lists.Load([]reputation.List{ + {URL: dropURL, Fetched: time.Now().Add(-time.Hour), Lines: []string{"203.0.113.9"}}, + {URL: torURL, Fetched: time.Now().Add(-refresh), Lines: []string{"198.51.100.9"}}, + }) + if err != nil { + t.Fatalf("load: %v", err) + } + + run(t, lists) + + wantFetches(t, servers, 1) + wantListedBy(t, lists, "203.0.113.9", dropURL) + wantListedBy(t, lists, "198.51.100.10", torURL) + + time.Sleep(refresh - time.Hour - time.Nanosecond) + wantFetches(t, servers, 1) + + time.Sleep(time.Nanosecond) + wantFetches(t, servers, 2) + wantListedBy(t, lists, "203.0.113.10", dropURL) + }) +} + +func TestRestartWaitsRefreshAfterTheLastTryEvenOneThatFailed(t *testing.T) { + t.Parallel() + + synctest.Test(t, func(t *testing.T) { + // drop.txt is fetched, and a refresh later the fetch downloads it + // whole but fails on a line that does not read. + servers := &standIn{lists: map[string]string{dropURL: "198.51.100.1\n"}} + lists := start(t, servers, params(dropURL)) + + servers.set(dropURL, "198.51.100.2\n\n") + time.Sleep(refresh) + wantFetches(t, servers, 2) + + // Restarted with what reputation.json keeps, it waits a refresh + // after the failed try, as it does while it runs. + restarted := &standIn{lists: map[string]string{dropURL: "198.51.100.2\n"}} + again := reputation.New(params(dropURL)) + again.SetTransport(restarted) + + err := again.Load(lists.Snapshot()) + if err != nil { + t.Fatalf("load: %v", err) + } + + run(t, again) + wantFetches(t, restarted, 0) + wantListedBy(t, again, "198.51.100.1", dropURL) + + time.Sleep(refresh - time.Nanosecond) + wantFetches(t, restarted, 0) + + time.Sleep(time.Nanosecond) + wantFetches(t, restarted, 1) + wantListedBy(t, again, "198.51.100.2", dropURL) + }) +} + +func TestFetchCutOffAsItStopsIsNoFailureButARestartWaitsForIt(t *testing.T) { + t.Parallel() + + synctest.Test(t, func(t *testing.T) { + // Stopped 30 seconds into the fetch of drop.txt, before tor.txt's. + servers := &standIn{lists: map[string]string{}, hanging: true} + queue := newQueue() + p := params(dropURL, torURL) + p.Alerts = queue + + lists := reputation.New(p) + lists.SetTransport(servers) + + ctx, stop := context.WithCancel(t.Context()) + stopped := make(chan struct{}) + + go func() { + lists.Run(ctx) + close(stopped) + }() + + time.Sleep(30 * time.Second) + stop() + <-stopped + + wantFetches(t, servers, 1) + + if lists.Failures(dropURL) != 0 || len(waiting(queue)) != 0 { + t.Errorf("%d failures and alerts %+v, want none", lists.Failures(dropURL), + waiting(queue)) + } + + // Restarted an hour later with what reputation.json keeps, it fetches + // tor.txt, never tried, at once, and drop.txt a refresh after its + // cut-off try. + time.Sleep(time.Hour) + + restarted := &standIn{lists: map[string]string{ + dropURL: "203.0.113.7\n", torURL: "198.51.100.7\n", + }} + again := reputation.New(params(dropURL, torURL)) + again.SetTransport(restarted) + + err := again.Load(lists.Snapshot()) + if err != nil { + t.Fatalf("load: %v", err) + } + + run(t, again) + wantFetches(t, restarted, 1) + wantListedBy(t, again, "198.51.100.7", torURL) + + time.Sleep(refresh - time.Hour - time.Nanosecond) + wantFetches(t, restarted, 1) + + time.Sleep(time.Nanosecond) + wantFetches(t, restarted, 2) + wantListedBy(t, again, "203.0.113.7", dropURL) + }) +} + +func TestASNLimitPercentFileGivesEachASNumberItsLowestPercentage(t *testing.T) { + t.Parallel() + + synctest.Test(t, func(t *testing.T) { + servers := &standIn{lists: map[string]string{ + asnURL: "# hosting networks\nAS14061:50 ; DigitalOcean\nas16276:25\n\n" + + "AS14061:10\nAS14061:30\n", + }} + p := params() + p.ASNLimitPercentURL = asnURL + lists := start(t, servers, p) + + for asn, want := range map[string]int64{"AS14061": 10, "AS16276": 25} { + percent, listed := lists.ASNLimitPercent(asn) + if !listed || percent != want { + t.Errorf("%s has %d (listed %t), want %d", asn, percent, listed, want) + } + } + + if _, listed := lists.ASNLimitPercent("AS64496"); listed { + t.Error("AS64496 is listed") + } + + // A line that does not read fails the fetch. + servers.set(asnURL, "AS14061:50\nAS16276\n") + time.Sleep(refresh) + synctest.Wait() + + if lists.Failures(asnURL) != 1 { + t.Errorf("%d failures, want 1", lists.Failures(asnURL)) + } + }) +} + +func TestLoadDropsCopiesOfListsNotNamedAndRefusesOnesThatDoNotRead(t *testing.T) { + t.Parallel() + + fetched := time.Date(2026, 10, 6, 0, 0, 0, 0, time.UTC) + kept := reputation.List{ + URL: dropURL, Tried: fetched, Fetched: fetched, Lines: []string{"203.0.113.9"}, + } + lists := reputation.New(params(dropURL)) + + err := lists.Load([]reputation.List{kept, { + URL: torURL, Tried: fetched, Fetched: fetched, Lines: []string{"198.51.100.9"}, + }}) + if err != nil { + t.Fatalf("load: %v", err) + } + + if got := lists.Snapshot(); !reflect.DeepEqual(got, []reputation.List{kept}) { + t.Errorf("copies %+v, want only %+v", got, kept) + } + + err = lists.Load([]reputation.List{{URL: dropURL, Fetched: fetched, Lines: []string{ + "; DROP", "203.0.113.300", + }}}) + + const want = "the copy of " + dropURL + + ": line 2 is not an address or a netblock, such as 192.0.2.0/24" + if err == nil || err.Error() != want { + t.Errorf("error %v, want %s", err, want) + } + + if got := lists.Snapshot(); !reflect.DeepEqual(got, []reputation.List{kept}) { + t.Errorf("copies %+v after the error, want %+v still", got, kept) + } +} + +// standIn is a stand-in for the servers the lists are fetched from. It +// notes the URL of each fetch. +type standIn struct { + mu sync.Mutex + // lists are what it answers with, by URL; it answers a URL it has no + // list for with 503, and none at all while hanging. + lists map[string]string + hanging bool + fetches []string +} + +// RoundTrip has the stand-in answer req, in place of the network. A fetch +// abandoned before the stand-in answers fails, as over the network. +func (s *standIn) RoundTrip(req *http.Request) (*http.Response, error) { + s.mu.Lock() + s.fetches = append(s.fetches, req.URL.String()) + list, found := s.lists[req.URL.String()] + hanging := s.hanging + s.mu.Unlock() + + if hanging { + <-req.Context().Done() + + return nil, req.Context().Err() + } + + status := http.StatusOK + if !found { + status = http.StatusServiceUnavailable + } + + return &http.Response{ + StatusCode: status, + Status: fmt.Sprintf("%d %s", status, http.StatusText(status)), + Header: http.Header{}, + Body: io.NopCloser(strings.NewReader(list)), + Request: req, + }, nil +} + +// set has the stand-in answer listURL with list, or with 503 for "". +func (s *standIn) set(listURL, list string) { + s.mu.Lock() + defer s.mu.Unlock() + + if list == "" { + delete(s.lists, listURL) + + return + } + + s.lists[listURL] = list +} + +// params returns the Params of the blocklists at urls, refreshed every +// refresh, by the bubble's clock, with alerts to a queue that sends none. +func params(urls ...string) reputation.Params { + return reputation.Params{ + BlocklistURLs: urls, + Refresh: refresh, + Now: time.Now, + ProcessLog: slog.New(slog.DiscardHandler), + Alerts: newQueue(), + } +} + +// newQueue returns a queue of alerts to a webhook that is never sent +// them, with a cooldown of cooldown. +func newQueue() *alerts.Queue { + return alerts.New(alerts.Params{ + WebhookURL: &url.URL{Scheme: "https", Host: "alerts.example"}, + Events: alerts.Events(), + Cooldown: cooldown, + Now: time.Now, + }) +} + +// start returns the lists of p, fetched through servers by Run, which runs +// until the test ends, once Run has fetched those due at start. +func start(t *testing.T, servers *standIn, p reputation.Params) *reputation.Lists { + t.Helper() + + lists := reputation.New(p) + lists.SetTransport(servers) + run(t, lists) + + return lists +} + +// run runs lists' Run until the test ends, and waits until it has fetched +// the lists due. +func run(t *testing.T, lists *reputation.Lists) { + t.Helper() + + ctx, stop := context.WithCancel(t.Context()) + stopped := make(chan struct{}) + + go func() { + lists.Run(ctx) + close(stopped) + }() + + t.Cleanup(func() { + stop() + <-stopped + }) + + synctest.Wait() +} + +// wantFetches waits until Run has made the fetches due, and checks how +// many the servers have had. +func wantFetches(t *testing.T, servers *standIn, want int) { + t.Helper() + + synctest.Wait() + + servers.mu.Lock() + got := len(servers.fetches) + servers.mu.Unlock() + + if got != want { + t.Errorf("%d fetches, want %d", got, want) + } +} + +// wantListedBy checks the URLs of the blocklists lists says list addr. +func wantListedBy(t *testing.T, lists *reputation.Lists, addr string, want ...string) { + t.Helper() + + got := lists.ListedBy(netip.MustParseAddr(addr)) + if !slices.Equal(got, want) { + t.Errorf("%s is listed by %v, want %v", addr, got, want) + } +} + +// waiting returns the alerts waiting in queue. +func waiting(queue *alerts.Queue) []alerts.Alert { + return queue.Snapshot().Waiting[alerts.DestinationWebhook] +} + +// wantAlert checks that want is the one alert waiting in queue, and that +// the cooldown has held back one repeat of it. +func wantAlert(t *testing.T, queue *alerts.Queue, want alerts.Alert) { + t.Helper() + + got := waiting(queue) + if len(got) != 1 || !reflect.DeepEqual(got[0], want) || queue.Suppressed() != 1 { + t.Errorf("alerts waiting %+v, %d held back, want only %+v and 1", got, + queue.Suppressed(), want) + } +} diff --git a/internal/requestlog/requestlog.go b/internal/requestlog/requestlog.go index c583281..69c29d2 100644 --- a/internal/requestlog/requestlog.go +++ b/internal/requestlog/requestlog.go @@ -35,7 +35,8 @@ const ( // rule. ActionRuleBlocked = "rule_blocked" // ActionDenied is a request refused because its client is in - // SWWAF_DENY_NETS. + // SWWAF_DENY_NETS, or in a blocklist while SWWAF_BLOCKLIST_ACTION is + // deny. ActionDenied = "denied" // ActionCountryDenied is a request refused for its client's country. ActionCountryDenied = "country_denied" @@ -137,6 +138,8 @@ type Line struct { // Counts names its count: minute, hour or day for a rate limit, and // minute_bytes, hour_bytes or day_bytes for a byte limit. LimitHit string `json:"limit_hit,omitempty"` + // Reputation are the URLs of the blocklists that list the client. + Reputation []string `json:"reputation,omitempty"` // Offence is the offence the request was held as, OffenceLimit. Offence string `json:"offence,omitempty"` // BanExpires is when the ban the request made, or was refused under, diff --git a/internal/smallwebwaf/smallwebwaf.go b/internal/smallwebwaf/smallwebwaf.go index be9c446..ee9901d 100644 --- a/internal/smallwebwaf/smallwebwaf.go +++ b/internal/smallwebwaf/smallwebwaf.go @@ -200,6 +200,7 @@ func loadStateFiles( Ledger: server.Ledger, Limiter: server.Limiter, GeoJS: server.GeoJS, + Lists: server.Lists, Alerts: alertQueue, Anomalies: server.Anomalies, Now: now, @@ -267,8 +268,9 @@ func startSending( // serve serves requests on listener, writes the state files as they are // due, takes in an admin's edits of them, reads the rule files again as -// they change, and the lookup database when it is replaced, and sends the -// alerts, until ctx is done. Then it gives the requests in progress +// they change, and the lookup database when it is replaced, fetches the +// lists the settings name by URL as they are due, and sends the alerts, +// until ctx is done. Then it gives the requests in progress // shutdownTimeout to finish, and writes every state file, alerts.json with // the alerts still waiting. func serve( @@ -293,6 +295,7 @@ func serve( server.LookupFile.Watch(writing) } }) + listsFetched := inBackground(func() { server.Lists.Run(writing) }) alertsSent := inBackground(func() { alertQueue.Run(writing) }) select { @@ -336,6 +339,7 @@ func serve( <-watched <-rulesWatched <-lookupFileWatched + <-listsFetched <-alertsSent err = files.WriteAll() diff --git a/internal/smallwebwaf/smallwebwaf_test.go b/internal/smallwebwaf/smallwebwaf_test.go index fb708a3..3af0fdd 100644 --- a/internal/smallwebwaf/smallwebwaf_test.go +++ b/internal/smallwebwaf/smallwebwaf_test.go @@ -38,6 +38,7 @@ const ( stateWriteDelay = "SWWAF_STATE_WRITE_DELAY" stateCounterInterval = "SWWAF_STATE_COUNTER_INTERVAL" rateLimitPerDay = "SWWAF_RATE_LIMIT_PER_DAY" + rateLimitExemptNets = "SWWAF_RATE_LIMIT_EXEMPT_NETS" rulesDir = "SWWAF_RULES_DIR" lookupSource = "SWWAF_LOOKUP_SOURCE" lookupDBPath = "SWWAF_LOOKUP_DB_PATH" @@ -522,7 +523,7 @@ func TestLookupDatabaseReplacedWhileRunningTakesEffect(t *testing.T) { // The requests sent until a replacement takes effect, and those // for the metrics, must not break a rate limit, whose ban would // refuse them too. - "SWWAF_RATE_LIMIT_EXEMPT_NETS": placed + "," + localhost, + rateLimitExemptNets: placed + "," + localhost, } began := time.Now() // Each replacement is written beside the file and renamed over it, as @@ -573,6 +574,111 @@ func TestLookupDatabaseReplacedWhileRunningTakesEffect(t *testing.T) { } } +func TestBlocklistTriesAndCopiesKeptInReputationJSONAcrossRestarts(t *testing.T) { + t.Parallel() + + const ( + token = "0123456789abcdef0123456789abcdef" + torPath = "/tor.txt" + dropPath = "/drop.txt" + // copyright is the DROP list's date and copyright line. + copyright = "; Spamhaus DROP List 2026/10/07 - (c) 2026 The Spamhaus Project SLL" + ) + + failing, torFetches := new(atomic.Bool), new(atomic.Int32) + lists := map[string]string{ + torPath: "198.51.100.0/24\n", + dropPath: copyright + "\n" + placed + " ; SBL1\n", + } + server := httptest.NewServer(http.HandlerFunc( + func(w http.ResponseWriter, r *http.Request) { + if r.URL.Path == torPath { + torFetches.Add(1) + } + + if failing.Load() { + w.WriteHeader(http.StatusServiceUnavailable) + + return + } + + _, _ = io.WriteString(w, lists[r.URL.Path]) + })) + t.Cleanup(server.Close) + + torURL, dropURL := server.URL+torPath, server.URL+dropPath + dir := t.TempDir() + env := map[string]string{ + listenAddr: localhost + ":0", + upstreamURL: startApp(t), + stateDir: dir, + rulesDir: t.TempDir(), + trustedProxies: localhost + "/32", + metricsToken: token, + instanceName: instance, + "SWWAF_BLOCKLIST_URLS": torURL, + // The requests sent until the list takes effect, and those for the + // metrics, must not break a rate limit, whose ban would refuse them + // too. + rateLimitExemptNets: placed + "," + localhost, + } + failures := `smallwebwaf_reputation_failures_total{instance="fsn1app1/gitea",` + + `source="` + torURL + `"} ` + + // tor.txt cannot be fetched at first, which is counted. + failing.Store(true) + runUntilStopped(t, env, func(url string) { + metricsWith(t, url+"_smallwebwaf/metrics", token, failures+"1") + }) + + // Restarted with drop.txt named after it and the server answering, tor.txt + // waits SWWAF_BLOCKLIST_REFRESH after its failed try, kept in reputation.json, + // while drop.txt, never tried, is fetched at once. Lists are fetched in the + // order named, so once drop.txt refuses the client, tor.txt has had its turn. + failing.Store(false) + + env["SWWAF_BLOCKLIST_URLS"] = torURL + "," + dropURL + out := runUntilStopped(t, env, func(url string) { + for statusFrom(t, url, placed) != http.StatusForbidden { + time.Sleep(pollInterval) + } + }) + wantDeniedByList(t, out.line(t, "action", "denied"), dropURL) + + if fetches := torFetches.Load(); fetches != 1 { + t.Errorf("tor.txt fetched %d times, want once, before the restart", fetches) + } + + // After another restart, with the server failing, the copy of drop.txt + // kept in reputation.json, its copyright line included, refuses the + // client from the first request. + failing.Store(true) + + out = runUntilStopped(t, env, func(url string) { + wantStatus(t, url, placed, http.StatusForbidden) + }) + wantDeniedByList(t, out.line(t, "type", "request"), dropURL) + + path := filepath.Join(dir, "reputation.json") + + kept, err := os.ReadFile(path) //nolint:gosec // a file in the test's directory + if err != nil || !strings.Contains(string(kept), `"`+copyright+`"`) { + t.Errorf("reputation.json holds\n%s\nwant the copy with %q (%v)", kept, copyright, + err) + } +} + +// wantDeniedByList checks that the request log line is of a request the +// blocklist at listURL refused. +func wantDeniedByList(t *testing.T, line map[string]any, listURL string) { + t.Helper() + + reputation, _ := line["reputation"].([]any) + if line["action"] != "denied" || len(reputation) != 1 || reputation[0] != listURL { + t.Errorf("request log line %v, want one denied for %s", line, listURL) + } +} + func TestLookupDatabaseThatCannotBeReadStopsTheStart(t *testing.T) { t.Parallel() @@ -1014,7 +1120,7 @@ func wantStartingLine(t *testing.T, line map[string]any, appURL, dir string) { "SWWAF_REQUEST_MAX_BYTES": "100M", "SWWAF_RESPONSE_MAX_BYTES": "5G", "SWWAF_ALLOW_NETS": "", - "SWWAF_RATE_LIMIT_EXEMPT_NETS": "", + rateLimitExemptNets: "", "SWWAF_DENY_NETS": "", "SWWAF_RATE_LIMIT_PER_MINUTE": "1000", "SWWAF_RATE_LIMIT_PER_HOUR": "10000", diff --git a/internal/state/state.go b/internal/state/state.go index 6daf5c6..9ba75d4 100644 --- a/internal/state/state.go +++ b/internal/state/state.go @@ -1,9 +1,10 @@ // Package state keeps smallwebwaf's state in JSON files in // SWWAF_STATE_DIR, as the "Persistent state" section of SPEC.md describes: // bans.json holds the bans, clients.json each client's counters and -// history, lookups.json GeoJS's answers, and alerts.json the cooldowns, -// the hour under way, the alerts waiting for each destination and the -// anomaly counters. Load +// history, lookups.json GeoJS's answers, reputation.json the last try and +// last good copy of each list fetched from a URL, and alerts.json the +// cooldowns, the hour under way, the alerts waiting for each destination +// and the anomaly counters. Load // reads them at start, Watch takes in an admin's edit of one while // smallwebwaf runs, and Run and WriteAll write them. The disk is read and // written outside the parts' locks, which are held only to take a @@ -35,6 +36,7 @@ import ( "sneak.berlin/go/smallwebwaf/internal/lookup" "sneak.berlin/go/smallwebwaf/internal/metrics" "sneak.berlin/go/smallwebwaf/internal/ratelimit" + "sneak.berlin/go/smallwebwaf/internal/reputation" ) // version is the version of the files' format, the only one read. @@ -46,10 +48,11 @@ const fileMode = 0o600 // The state files' names. const ( - bansJSON = "bans.json" - clientsJSON = "clients.json" - lookupsJSON = "lookups.json" - alertsJSON = "alerts.json" + bansJSON = "bans.json" + clientsJSON = "clients.json" + lookupsJSON = "lookups.json" + reputationJSON = "reputation.json" + alertsJSON = "alerts.json" ) var ( @@ -73,12 +76,13 @@ type Params struct { // is (SWWAF_STATE_COUNTER_INTERVAL). WriteDelay time.Duration CounterInterval time.Duration - // Ledger, Limiter, GeoJS, Alerts and Anomalies hold the state. Alerts - // also receive a file_error alert for an edit set aside, and for a - // write that fails while smallwebwaf runs. + // Ledger, Limiter, GeoJS, Lists, Alerts and Anomalies hold the state. + // Alerts also receive a file_error alert for an edit set aside, and for + // a write that fails while smallwebwaf runs. Ledger *bans.Ledger Limiter *ratelimit.Limiter GeoJS *lookup.GeoJS + Lists *reputation.Lists Alerts *alerts.Queue Anomalies *anomaly.Counters // Now tells the time by which the counters' buckets run out, normally @@ -138,6 +142,13 @@ type lookupsFile struct { Lookups []lookup.Answer `json:"lookups"` } +// reputationFile is reputation.json, indented for an admin to read and +// edit, so that each line of a list's copy is on a line of its own. +type reputationFile struct { + Version int `json:"version"` + Lists []reputation.List `json:"lists"` +} + // alertsFile is alerts.json, indented for an admin to read and edit. // //nolint:tagliatelle // the state files use snake_case, as the request log does @@ -159,10 +170,10 @@ type stateFile interface { } // Load checks that files can be written in Dir, and reads the state files -// in it into the ledger, the limiter and GeoJS. A missing file is empty -// state, as on a first start. A file that does not parse, has an unknown -// version, or has an entry without a field it needs, is an error that -// names the file and, where the JSON decoder tells it, the line and +// in it into the parts of Params that hold the state. A missing file is +// empty state, as on a first start. A file that does not parse, has an +// unknown version, or has an entry without a field it needs, is an error +// that names the file and, where the JSON decoder tells it, the line and // column, or else the entry. func Load(params Params) (*Files, error) { err := checkWritable(params.Dir) @@ -175,16 +186,17 @@ func Load(params Params) (*Files, error) { bansRead, bansErr := f.read(bansJSON) clientsRead, clientsErr := f.read(clientsJSON) lookupsRead, lookupsErr := f.read(lookupsJSON) + reputationRead, reputationErr := f.read(reputationJSON) alertsRead, alertsErr := f.read(alertsJSON) - err = errors.Join(bansErr, clientsErr, lookupsErr, alertsErr) + err = errors.Join(bansErr, clientsErr, lookupsErr, reputationErr, alertsErr) if err != nil { return nil, err } params.ProcessLog.Info("read the state files", "directory", params.Dir, "bans", bansRead, "clients", clientsRead, "lookups", lookupsRead, - "alerts_waiting", alertsRead) + "lists", reputationRead, "alerts_waiting", alertsRead) return f, nil } @@ -213,7 +225,9 @@ func (f *Files) Run(ctx context.Context) { f.logFailure(bansJSON, f.writeFile(bansJSON)) case <-interval.C: - for _, name := range []string{bansJSON, clientsJSON, lookupsJSON, alertsJSON} { + for _, name := range []string{ + bansJSON, clientsJSON, lookupsJSON, reputationJSON, alertsJSON, + } { f.logFailure(name, f.writeFile(name)) } } @@ -224,7 +238,7 @@ func (f *Files) Run(ctx context.Context) { // fails does not keep the others from being written. func (f *Files) WriteAll() error { return errors.Join(f.writeFile(bansJSON), f.writeFile(clientsJSON), - f.writeFile(lookupsJSON), f.writeFile(alertsJSON)) + f.writeFile(lookupsJSON), f.writeFile(reputationJSON), f.writeFile(alertsJSON)) } // Watch watches Dir until ctx is done, and takes in an admin's edit of a @@ -259,7 +273,7 @@ func (f *Files) Watch(ctx context.Context) { return case event := <-watcher.Events: switch name := filepath.Base(event.Name); name { - case bansJSON, clientsJSON, lookupsJSON, alertsJSON: + case bansJSON, clientsJSON, lookupsJSON, reputationJSON, alertsJSON: f.fileChanged(name) } case err = <-watcher.Errors: @@ -405,33 +419,27 @@ func (f *Files) takeIn(name string, data []byte, edit bool) (int, error) { f.params.GeoJS.Load(file.Lookups) entries = len(file.Lookups) - case alertsJSON: - // waiting was a list, of the alerts waiting for the webhook, before - // alerts went to Slack and ntfy too. - var written struct { - Waiting json.RawMessage `json:"waiting"` - } - - if json.Unmarshal(data, &written) == nil && - bytes.HasPrefix(written.Waiting, []byte("[")) { - return 0, fmt.Errorf("%s: %w", path, errWaitingList) - } - - var file alertsFile + case reputationJSON: + var file reputationFile err := parse(path, data, &file) if err != nil { return 0, err } - f.params.Alerts.Load(alerts.State{ - Cooldowns: file.Cooldowns, Hour: file.Hour, Waiting: file.Waiting, - }) - f.params.Anomalies.Load(file.AnomalyCounters, f.params.Now()) - - for _, waiting := range file.Waiting { - entries += len(waiting) + err = f.params.Lists.Load(file.Lists) + if err != nil { + return 0, fmt.Errorf("%s: %w", path, err) } + + entries = len(file.Lists) + case alertsJSON: + waiting, err := f.takeInAlerts(path, data) + if err != nil { + return 0, err + } + + entries = waiting } f.sums[name] = sha256.Sum256(data) @@ -439,6 +447,41 @@ func (f *Files) takeIn(name string, data []byte, edit bool) (int, error) { return entries, nil } +// takeInAlerts parses data, what alerts.json, at path, holds, puts it +// into the alerts and the anomaly counters, in place of what they held, +// and returns how many alerts wait in it, as takeIn describes. +func (f *Files) takeInAlerts(path string, data []byte) (int, error) { + // waiting was a list, of the alerts waiting for the webhook, before + // alerts went to Slack and ntfy too. + var written struct { + Waiting json.RawMessage `json:"waiting"` + } + + if json.Unmarshal(data, &written) == nil && + bytes.HasPrefix(written.Waiting, []byte("[")) { + return 0, fmt.Errorf("%s: %w", path, errWaitingList) + } + + var file alertsFile + + err := parse(path, data, &file) + if err != nil { + return 0, err + } + + f.params.Alerts.Load(alerts.State{ + Cooldowns: file.Cooldowns, Hour: file.Hour, Waiting: file.Waiting, + }) + f.params.Anomalies.Load(file.AnomalyCounters, f.params.Now()) + + entries := 0 + for _, waiting := range file.Waiting { + entries += len(waiting) + } + + return entries, nil +} + // writeFile writes the state file name from what smallwebwaf holds. An // edit made since smallwebwaf last read or wrote the file is taken in // first, so that it is not overwritten, or set aside if it does not @@ -514,34 +557,38 @@ func (f *Files) setAside(name string, parseErr error) error { func (f *Files) encode(name string) ([]byte, error) { switch name { case bansJSON: - file := bansFile{Version: version, Bans: BanEntries(f.params.Ledger.Snapshot())} - - data, err := json.MarshalIndent(file, "", " ") - if err != nil { - return nil, err - } - - return append(data, '\n'), nil + return encodeIndented(bansFile{ + Version: version, Bans: BanEntries(f.params.Ledger.Snapshot()), + }) case clientsJSON: return encodeOnePerLine("clients", f.params.Limiter.Snapshot()) case lookupsJSON: return encodeOnePerLine("lookups", f.params.GeoJS.Snapshot()) + case reputationJSON: + return encodeIndented(reputationFile{ + Version: version, Lists: f.params.Lists.Snapshot(), + }) default: // alerts.json held := f.params.Alerts.Snapshot() - file := alertsFile{ + + return encodeIndented(alertsFile{ Version: version, Cooldowns: held.Cooldowns, Hour: held.Hour, Waiting: held.Waiting, AnomalyCounters: f.params.Anomalies.Snapshot(), - } - - data, err := json.MarshalIndent(file, "", " ") - if err != nil { - return nil, err - } - - return append(data, '\n'), nil + }) } } +// encodeIndented encodes file, a state file's struct, indented for an +// admin to read and edit. +func encodeIndented(file any) ([]byte, error) { + data, err := json.MarshalIndent(file, "", " ") + if err != nil { + return nil, err + } + + return append(data, '\n'), nil +} + // BanEntries returns held as bans.json lists them, an empty list for // none. func BanEntries(held []bans.Ban) []BanEntry { @@ -679,6 +726,27 @@ func (f *lookupsFile) check(data []byte) error { return nil } +// check refuses a list without its URL, which would name no list, or the +// time it was last tried, which would have it fetched at once, and a copy +// of it without the time it was fetched, or without its lines, which hold +// the list. +func (f *reputationFile) check([]byte) error { + for i, kept := range f.Lists { + switch { + case kept.URL == "": + return missing(i, "url") + case kept.Tried.IsZero(): + return missing(i, "tried") + case kept.Fetched.IsZero() && kept.Lines != nil: + return missing(i, "fetched") + case kept.Lines == nil && !kept.Fetched.IsZero(): + return missing(i, "lines") + } + } + + return nil +} + // check refuses a cooldown without its event or when its alert was sent, // which would hold back no repeat, alerts waiting for a destination with // another name than webhook, slack or ntfy, most likely misspelt, an diff --git a/internal/state/state_test.go b/internal/state/state_test.go index 62ad793..1b0e604 100644 --- a/internal/state/state_test.go +++ b/internal/state/state_test.go @@ -27,15 +27,20 @@ import ( "sneak.berlin/go/smallwebwaf/internal/lookup" "sneak.berlin/go/smallwebwaf/internal/metrics" "sneak.berlin/go/smallwebwaf/internal/ratelimit" + "sneak.berlin/go/smallwebwaf/internal/reputation" "sneak.berlin/go/smallwebwaf/internal/state" ) const ( // The state files. - bansJSON = "bans.json" - clientsJSON = "clients.json" - lookupsJSON = "lookups.json" - alertsJSON = "alerts.json" + bansJSON = "bans.json" + clientsJSON = "clients.json" + lookupsJSON = "lookups.json" + reputationJSON = "reputation.json" + alertsJSON = "alerts.json" + // blocklistURL and torURL are the blocklists the tests' lists name. + blocklistURL = "https://lists.example/drop.txt" + torURL = "https://lists.example/tor.txt" // The AS number and AS name the tests' clients are looked up in. asn = "AS64496" asName = "Example Net" @@ -204,6 +209,29 @@ const filledAlertsJSON = `{ } ` +// filledReputationJSON is reputation.json holding the blocklists' last +// tries and the copy of one, with its comment line, as fill puts them in. +const filledReputationJSON = `{ + "version": 1, + "lists": [ + { + "url": "https://lists.example/drop.txt", + "tried": "2026-10-06T00:00:00Z", + "fetched": "2026-10-05T23:00:00Z", + "lines": [ + "; Spamhaus DROP List 2026/10/05 - (c) 2026 The Spamhaus Project SLL", + "203.0.113.0/24 ; SBL1", + "2001:db8::/32 ; SBL2" + ] + }, + { + "url": "https://lists.example/tor.txt", + "tried": "2026-10-06T00:00:00Z" + } + ] +} +` + func TestFilesWrittenAndReadBack(t *testing.T) { t.Parallel() @@ -230,6 +258,11 @@ func TestFilesWrittenAndReadBack(t *testing.T) { wantEqual(t, clientsJSON, after.Limiter.Snapshot(), before.Limiter.Snapshot()) wantEqual(t, lookupsJSON, after.GeoJS.Snapshot(), before.GeoJS.Snapshot()) + if got, want := after.Lists.Snapshot(), before.Lists.Snapshot(); !reflect.DeepEqual( + got, want) { + t.Errorf("%s read back\n%+v\nwant\n%+v", reputationJSON, got, want) + } + if got, want := after.Alerts.Snapshot(), before.Alerts.Snapshot(); !reflect.DeepEqual( got, want) { t.Errorf("%s read back\n%+v\nwant\n%+v", alertsJSON, got, want) @@ -238,12 +271,32 @@ func TestFilesWrittenAndReadBack(t *testing.T) { wantEqual(t, alertsJSON, after.Anomalies.Snapshot(), before.Anomalies.Snapshot()) // Each one-per-line file lists its entries by client, and nothing - // but the four files is left in the directory. + // but the five files is left in the directory. wantEntries(t, filepath.Join(dir, clientsJSON), "clients", "192.0.2.1/32", "203.0.113.9/32", "2001:db8::/64") wantEntries(t, filepath.Join(dir, lookupsJSON), "lookups", "192.0.2.1/32", "203.0.113.9/32") - wantFiles(t, dir, alertsJSON, bansJSON, clientsJSON, lookupsJSON) + wantFiles(t, dir, alertsJSON, bansJSON, clientsJSON, lookupsJSON, reputationJSON) +} + +func TestReputationJSONKeepsEachCopyWholeOneLineToALine(t *testing.T) { + t.Parallel() + + dir := t.TempDir() + params := newParams(dir) + fill(params) + + files := load(t, params) + + err := files.WriteAll() + if err != nil { + t.Fatalf("write: %v", err) + } + + got := readFile(t, filepath.Join(dir, reputationJSON)) + if got != filledReputationJSON { + t.Errorf("reputation.json\n%s\nwant\n%s", got, filledReputationJSON) + } } func TestAlertsJSONIsIndentedWithTheCooldownsTheHourAndTheAlertsWaiting(t *testing.T) { @@ -368,8 +421,9 @@ func TestMissingFilesAreEmptyState(t *testing.T) { held := params.Alerts.Snapshot() if len(params.Ledger.Snapshot()) != 0 || len(params.Limiter.Snapshot()) != 0 || - len(params.GeoJS.Snapshot()) != 0 || len(held.Cooldowns) != 0 || - len(held.Waiting[alerts.DestinationWebhook]) != 0 || held.Hour.Sent != 0 { + len(params.GeoJS.Snapshot()) != 0 || len(params.Lists.Snapshot()) != 0 || + len(held.Cooldowns) != 0 || len(held.Waiting[alerts.DestinationWebhook]) != 0 || + held.Hour.Sent != 0 { t.Error("state from no files") } } @@ -428,6 +482,14 @@ func TestFileThatDoesNotParseStopsTheStart(t *testing.T) { `: anomaly_counters entry 2's scope "nett" is not client, net, asn, total ` + `or watch`, }, + { + "a copy of a list with a line that does not read", reputationJSON, + `{"version": 1, "lists": [{"url": "` + blocklistURL + `", ` + + `"tried": "2026-10-06T00:00:00Z", "fetched": "2026-10-06T00:00:00Z", ` + + `"lines": ["; DROP", "203.0.113.300"]}]}`, + `: the copy of ` + blocklistURL + `: line 2 is not an address or a netblock, ` + + `such as 192.0.2.0/24`, + }, } { t.Run(tc.name, func(t *testing.T) { t.Parallel() @@ -523,6 +585,51 @@ func TestEntryWithoutAFieldItNeedsStopsTheStart(t *testing.T) { } } +func TestReputationJSONEntryWithoutAFieldItNeedsStopsTheStart(t *testing.T) { + t.Parallel() + + const ( + drop = `"url": "` + blocklistURL + `", ` + tried = `"tried": "2026-10-06T00:00:00Z", ` + fetched = `"fetched": "2026-10-06T00:00:00Z"` + ) + + for _, tc := range []struct { + name, content string + // want is what the error says after the file's path. + want string + }{ + { + "a list without its URL", + `{"version": 1, "lists": [{` + tried + fetched + `, "lines": []}]}`, + `: entry 1 has no "url"`, + }, + { + "a list without the time it was last tried", + `{"version": 1, "lists": [{` + drop + fetched + `, "lines": []}]}`, + `: entry 1 has no "tried"`, + }, + { + "a copy of a list without the time it was fetched", + `{"version": 1, "lists": [{` + drop + tried + `"lines": []}]}`, + `: entry 1 has no "fetched"`, + }, + { + // An empty list has no lines, which is not having none. + "a copy of a list without its lines", + `{"version": 1, "lists": [{` + drop + tried + fetched + `, "lines": []}, ` + + `{"url": "` + torURL + `", ` + tried + fetched + `}]}`, + `: entry 2 has no "lines"`, + }, + } { + t.Run(tc.name, func(t *testing.T) { + t.Parallel() + + wantRefused(t, reputationJSON, tc.content, tc.want) + }) + } +} + func TestAlertsJSONEntryWithoutAFieldItNeedsStopsTheStart(t *testing.T) { t.Parallel() @@ -626,7 +733,9 @@ func TestBanWithAnotherCauseStopsTheStart(t *testing.T) { func TestUnknownVersionStopsTheStart(t *testing.T) { t.Parallel() - for _, file := range []string{bansJSON, clientsJSON, lookupsJSON, alertsJSON} { + for _, file := range []string{ + bansJSON, clientsJSON, lookupsJSON, reputationJSON, alertsJSON, + } { for _, content := range []string{`{"version": 2}`, `{}`} { t.Run(file+" "+content, func(t *testing.T) { t.Parallel() @@ -728,8 +837,9 @@ func TestEveryFileWrittenEveryCounterInterval(t *testing.T) { time.Sleep(time.Nanosecond) synctest.Wait() - wantFiles(t, dir, alertsJSON, bansJSON, clientsJSON, lookupsJSON) - removeFiles(t, dir, alertsJSON, bansJSON, clientsJSON, lookupsJSON) + wantFiles(t, dir, alertsJSON, bansJSON, clientsJSON, lookupsJSON, reputationJSON) + removeFiles(t, dir, alertsJSON, bansJSON, clientsJSON, lookupsJSON, + reputationJSON) } }) } @@ -948,7 +1058,7 @@ func TestFileThatCannotBeReadIsNotWrittenOver(t *testing.T) { t.Errorf("bans.json is now %v (%v), want the socket", info, err) } - wantFiles(t, dir, alertsJSON, bansJSON, clientsJSON, lookupsJSON) + wantFiles(t, dir, alertsJSON, bansJSON, clientsJSON, lookupsJSON, reputationJSON) wantWriteFailed(t, params, bansJSON) } @@ -1021,6 +1131,17 @@ func TestEditOfEachFileTakenIn(t *testing.T) { wantEqual(t, lookupsJSON, params.GeoJS.Snapshot(), []lookup.Answer{{Client: client, Country: "FR", Answered: midnight()}}) + edit(t, dir, reputationJSON, `{"version": 1, "lists": [{"url": "`+blocklistURL+`", `+ + `"tried": "2026-10-06T00:00:00Z", "fetched": "2026-10-06T00:00:00Z", `+ + `"lines": ["198.51.100.7"]}]}`) + wantTakenIn(t, lines, dir, reputationJSON) + + listedBy := params.Lists.ListedBy(client.Addr()) + if !slices.Equal(listedBy, []string{blocklistURL}) { + t.Errorf("%s taken in lists %s on %v, want on the blocklist", reputationJSON, + client.Addr(), listedBy) + } + // A netblock with bits past its length is read as the netblock it is // in. edit(t, dir, alertsJSON, `{"version": 1, "cooldowns": [{"event": "ban", `+ @@ -1248,7 +1369,7 @@ func TestBrokenEditSetAsideAtTheNextWrite(t *testing.T) { edit(t, dir, bansJSON, broken) edit(t, dir, clientsJSON, `{"version": 1, "clients": []}`) wantTakenIn(t, lines, dir, clientsJSON) - wantFiles(t, dir, alertsJSON, bansJSON, clientsJSON, lookupsJSON) + wantFiles(t, dir, alertsJSON, bansJSON, clientsJSON, lookupsJSON, reputationJSON) // The next write sets it aside, logged with where the error is, and // writes bans.json again from what smallwebwaf still holds. @@ -1272,7 +1393,8 @@ func TestBrokenEditSetAsideAtTheNextWrite(t *testing.T) { t.Errorf("alerts waiting %+v, want a file_error alert for %s", waiting, path+".bad") } - wantFiles(t, dir, alertsJSON, bansJSON, bansJSON+".bad", clientsJSON, lookupsJSON) + wantFiles(t, dir, alertsJSON, bansJSON, bansJSON+".bad", clientsJSON, lookupsJSON, + reputationJSON) if got := readFile(t, path+".bad"); got != broken { t.Errorf("bans.json.bad holds\n%s\nwant the edit", got) @@ -1401,9 +1523,10 @@ func midnight() time.Time { } // newParams returns Params for the state files in dir, with parts that -// hold nothing yet. GeoJS is never asked, and the alerts, at most two an -// hour, are never sent. The anomaly counters count the scopes fill -// counts, with thresholds fill does not reach. +// hold nothing yet. GeoJS is never asked, the lists, two blocklists, are +// never fetched, and the alerts, at most two an hour, are +// never sent. The anomaly counters count the scopes fill counts, with +// thresholds fill does not reach. func newParams(dir string) state.Params { discard := slog.New(slog.DiscardHandler) m := metrics.New(1, "app") @@ -1431,6 +1554,10 @@ func newParams(dir string) state.Params { GeoJS: lookup.New(lookup.Params{ Now: midnight, ProcessLog: discard, Metrics: m, }), + Lists: reputation.New(reputation.Params{ + BlocklistURLs: []string{blocklistURL, torURL}, Refresh: 24 * time.Hour, + Now: midnight, ProcessLog: discard, Alerts: queue, + }), Alerts: queue, Anomalies: anomaly.New(anomaly.Params{ Net: anomaly.Thresholds{RequestsPerMinute: 1000}, @@ -1455,8 +1582,9 @@ func office() netip.Prefix { // fill puts a permanent ban an admin made, a ban for a broken limit and // one for a clear sign of attack, clients with counts and histories, -// GeoJS answers, and alerts and anomaly counters, as filledAlertsJSON -// holds them, into the parts of params. +// GeoJS answers, the blocklists' last tries and the copy of one, as +// filledReputationJSON holds them, and alerts and anomaly counters, as +// filledAlertsJSON holds them, into the parts of params. func fill(params state.Params) { now := midnight() client := netip.MustParsePrefix("203.0.113.9/32") @@ -1489,6 +1617,20 @@ func fill(params state.Params) { }, }) + // The copy of drop.txt was fetched an hour ago, and the fetches of it + // and of tor.txt tried since failed. + err := params.Lists.Load([]reputation.List{{ + URL: blocklistURL, Tried: now, Fetched: now.Add(-time.Hour), + Lines: []string{ + "; Spamhaus DROP List 2026/10/05 - (c) 2026 The Spamhaus Project SLL", + "203.0.113.0/24 ; SBL1", + "2001:db8::/32 ; SBL2", + }, + }, {URL: torURL, Tried: now}}) + if err != nil { + panic(err) // the copy reads + } + // An alert waiting, a repeat of it the cooldown holds back, another // alert waiting, and one past the two an hour, for the hour's summary. ban := alerts.Alert{