Files
smallwebwaf/SPEC.md
T
sneak 8fc7539f1a SPEC and README: review fixes, country lists, GeoJS lookups (closes #6)
Address the review of the spec update and fold in issues 8 and 11.
Default ban rules are anchored at the site root. `bans.json` is bounded
by `MAX_BANS` and rewritten only when a ban is made, lifted or made
permanent. `clients.json` keeps one client per line, holds at most 20000
clients and is written every 15 minutes. Anomaly counters and alert
state go to a new `alerts.json`, the AbuseIPDB count to
`reputation.json`. 401 no longer counts toward the error burst. New
settings: `DENIED_COUNTRIES`, `EXCLUSIVELY_ALLOWED_COUNTRIES`,
`LOOKUP_SOURCE` (the IPinfo file or GeoJS), `LOOKUP_TIMEOUT`,
`CLIENT_REQUEST_HEADER_MAX_BYTES` and `CLIENT_IDLE_TIMEOUT`.

Model: opus-5-5
2026-09-23 13:37:34 +00:00

62 KiB

smallwebwaf SPEC (draft): protective reverse-proxy sidecar

Status: third draft, with the owner's rulings to date applied. Nothing has been built yet. EVALUATION.md beside this file explains why no existing tool was chosen.

Purpose

One small container that sits between traefik and one application container. The traefik router for the public hostname points at the sidecar; the sidecar forwards to the application. The sidecar limits request and byte rates per client, bounds the size and duration of every request and response, detects attacks, bans abusers (briefly at first, for seven days on a clear sign of attack, permanently when they keep at it), consults IP reputation sources, looks up the AS number and country of each client, can refuse whole countries, lowers its limits for listed AS numbers and countries, and sends alerts.

It is meant to protect an app on the open internet from the first request with one setting, UPSTREAM_URL: every other setting has a default chosen for a service facing the internet in 2026. Everything is configured by environment variables, apart from the attack detection rules, which are read from a directory of hand-editable text files.

Non-goals

  • TLS termination, certificates, hostname routing: traefik's job.
  • More than one upstream application per sidecar. Run one sidecar per app.
  • Browser challenges (captcha, proof of work). If wanted, chain Anubis between the sidecar and the app.
  • Defence against traffic floods that saturate the host's network link. That needs help upstream of the host.
  • A web UI or a configuration file. Settings are environment variables. Apart from its own state files and the lookup database, the only files read are the rule files, which hold one regex per line and nothing more elaborate.
  • Sharing bans between sidecars in the first version: each sidecar keeps its own. Running one CrowdSec engine per host, which every sidecar would report its bans to, is reconsidered once real ban volumes are known. Reading a CrowdSec decision list is supported (see Reputation).

Architecture

  • One statically linked Go binary in one container, running as a non-root user, read-only root filesystem, one writable volume for state. The image declares that volume (/data), so docker supplies an anonymous one when none is mounted, and a sidecar started with only UPSTREAM_URL has somewhere to write.
  • Three listeners, of which only the first is routed by traefik:
    • the proxy listener;
    • a metrics listener serving Prometheus metrics, which can be switched off;
    • an admin listener for health and ban management.
  • Components inside the process:
    • Client identification: works out the real client IP from X-Forwarded-For, trusting only the proxy netblocks in TRUSTED_PROXIES, by default the private address ranges.
    • Lookup: AS number and country, from one of two sources the operator chooses: a database file the operator supplies, held in memory, or the GeoJS web service, whose answers are kept for seven days.
    • Reputation: static blocklists fetched on a schedule; DNSBL and reputation API queries made in the background and cached.
    • Counters: per-client request and byte counts per minute, hour and day, held in memory and written out with the rest of the state, so they survive a restart.
    • Attack detection: regex rules read from a directory of plain text rule files and read again whenever the files change (see "Rule files"), Coraza with the OWASP Core Rule Set, and simple signals (requests for listed trap paths, bursts of error responses).
    • Ban ledger: active and past bans with their notes, and each client's history, held in memory and kept in JSON files that the sidecar watches for an admin's edits (see "Bans" and "Persistent state"). No database of any kind is used.
    • Request log: one JSON object per request on stdout, optionally also sent to a remote syslog server (see "Request log").
    • Metrics: Prometheus counters and gauges on their own listener (see "Metrics endpoint").
    • Alerting: a queue with de-duplication feeding webhook, Slack and ntfy senders.
    • Proxy: the standard library's net/http/httputil.ReverseProxy, streaming in both directions within the size and time limits, with WebSocket upgrade support.
  • Proposed libraries (to be checked against the owner's Go dependency defaults before any are added):
    • standard library for the proxy, HTTP clients, DNS, logging (log/slog), state files (encoding/json, os.Rename);
    • github.com/corazawaf/coraza/v3 and github.com/corazawaf/coraza-coreruleset for attack detection;
    • github.com/oschwald/maxminddb-golang/v2, the current major version, to read the lookup database;
    • github.com/fsnotify/fsnotify to notice files that are edited or added while the sidecar runs;
    • github.com/prometheus/client_golang for metrics;
    • remote log sending is written in this project: the standard library's log/syslog is frozen and writes only the older syslog format, not the current one (RFC 5424). RELP is not in the first release; it is revisited against the available packages afterwards.

Data flow for one request

Steps run in this order; the first step that produces a final answer ends processing. The size and time limits (see "Configuration surface") apply to the whole exchange.

  • Identify the client.
    • If the TCP peer is inside TRUSTED_PROXIES, walk X-Forwarded-For from the right and take the first address not inside TRUSTED_PROXIES. Otherwise use the TCP peer address and ignore the header.
    • IPv6 clients are grouped by prefix (IPV6_GROUP_PREFIX, default 64) for counting and banning, because one abuser usually controls a whole /64. A client is therefore one IPv4 address or one IPv6 group, a /64 by default.
  • Static lists.
    • In ALLOW_NETS: skip every check below and forward. Still counted for anomaly alerts.
    • In DENY_NETS: refuse.
  • Ban ledger. An active ban on the client's netblock: refuse with BAN_RESPONSE. If a clear sign of attack caused the ban, the ban becomes permanent (see "Bans").
  • Look up the AS number and country, unless LOOKUP_SOURCE is off. With GeoJS, a client whose answer is not yet kept waits for it, up to LOOKUP_TIMEOUT. A client the lookup cannot place has an unknown country and AS number: the exclusive country list refuses it, the biased thresholds give it UNKNOWN_LIMIT_PERCENT, and nothing else treats it differently.
  • Country lists. A client whose country is in DENIED_COUNTRIES, or, when EXCLUSIVELY_ALLOWED_COUNTRIES is set, is not in it, is refused with BAN_RESPONSE. Every step up to here needs only the client's address, so the request body has not been read yet, and the refusal skips everything below, the rule files and the Core Rule Set included. It is counted in the metrics, but it is not an offence and makes no ban.
  • Reputation.
    • Address inside a fetched blocklist: apply BLOCKLIST_ACTION.
    • Cached DNSBL or reputation API result: apply REPUTATION_ACTION.
    • No cached result: queue a background query and carry on. A first request is never delayed by a reputation query.
  • Work out the client's limit percentage: the lowest of the percentages that apply (AS number, country, reputation), or 100 if none applies.
  • Request rate limits. Unless the client is in RATE_LIMIT_EXEMPT_NETS, check the minute, hour and day counters against each limit times the percentage. A request over any of them breaks that limit: it is refused with BAN_RESPONSE and the client is banned (see "Bans").
  • Rule files. The request is checked against the rules loaded from RULES_DIR, in file name order then line order. Each rule that matches takes its action: only log, refuse with 403, or refuse and ban. Matching stops at the first rule that refuses or bans.
  • Core Rule Set inspection of the request (headers, URL, and body up to WAF_BODY_LIMIT). In block mode, the default, a match at or over the anomaly threshold is refused with 403; in detect mode it is only logged and alerted.
  • Forward to UPSTREAM_URL, streaming. Add X-Forwarded-For and, if enabled, X-Client-ASN and X-Client-Country for the app's own logs.
  • After the response.
    • Add response bytes (and request body bytes) to the byte counters. A client whose byte total passes a byte limit times the percentage has broken that limit and is banned. The response in progress is not cut off.
    • Count the client's requests answered with 403 or 404, whether the app gave that answer or the sidecar refused the request after a rule file or Core Rule Set match. More than ERROR_BURST_THRESHOLD of them within a minute breaks a limit and bans the client.
    • Update the client's history and the anomaly counters, and evaluate alert thresholds.

Counting method

  • Each window (minute, hour, day) uses two adjacent fixed buckets per client, with the previous bucket weighted by how much of it still overlaps the sliding window. This costs a few integers per client per window and avoids the burst at bucket boundaries that a single fixed bucket allows.
  • Counters are read and updated in memory; a request never waits on the disk. They are written to clients.json every STATE_COUNTER_INTERVAL and at shutdown and loaded again at start (see "Persistent state"), so a restart does not hand every client a fresh allowance. Buckets whose time has passed are discarded on load.
  • Memory is bounded by MAX_TRACKED_CLIENTS and MAX_BANS. When the table of clients is full, the least recently seen client is dropped first, with its history and its lookup answer. When MAX_BANS bans are held, the ban that has gone longest without a request from its netblock is dropped first, whether it is past, active or permanent (see "Bans").

Bans

A ban refuses every request from a netblock, answering with BAN_RESPONSE (403 by default), until the ban ends. A permanent ban does not run out: it ends only when an admin lifts it, or when a new ban is made while MAX_BANS bans are held and it is the ban that has gone longest without a request. A scanner whose ban was dropped that way is refused and banned again by its next probe.

The netblock a ban covers is the client: its IPv4 address, or its IPv6 group of IPV6_GROUP_PREFIX (a /64 by default). BAN_SCOPE_V4_PREFIX can widen an IPv4 ban to the surrounding netblock. At the defaults a ban covers one address or one /64, so a permanent ban does not reach neighbours who did nothing.

An offence is a request the sidecar holds against the client: one that carries a clear sign of attack, breaks a limit, or is refused by a rule file or the Core Rule Set. Offences are counted by kind in the client's history. Two kinds lead to a ban, each by its own rule:

  • A clear sign of attack: a request that matches a rule file rule whose action is ban, or asks for a path in TRAP_PATHS. Examples are a probe for a .env file or a .git directory, and a known scanner's user agent.
    • The first one bans the netblock for ATTACK_BAN_DURATION (default 7d).
    • Any further request from the netblock while that ban lasts shows it is malicious: the ban becomes permanent.
    • Once that ban has run out, the netblock is served like any other, but its next clear sign of attack bans it permanently at once.
  • A broken limit: a request over a request limit, a response that takes the client's byte total past a byte limit, or more error responses than ERROR_BURST_THRESHOLD within a minute.
    • The first such ban, or one that comes more than LIMIT_BAN_REPEAT_WINDOW (default 24h) after the last such ban ended, lasts LIMIT_BAN_DURATION (default 1h).
    • Breaking a limit again within that window makes the new ban three times as long as the last one: one hour, then 3, 9, 27 and 81 hours.
    • A ban that would be longer than MAX_BAN_DURATION (default 7d) is permanent instead; with the defaults, that is the sixth ban in a row.

A Core Rule Set match refuses only the request it matched, without a ban, because the Core Rule Set has false positives; a block rule does the same. Those refusals count toward the error burst like any other 403, so a client that keeps setting them off is banned under the second rule. Requests refused under a ban are not counted against limits, but they are counted in the client's history and in the ban's notes.

Every ban carries notes with what an admin needs to decide whether to lift it (see "Persistent state"). An admin can add or lift a ban at any time by editing bans.json, or through the admin listener when ADMIN_TOKEN is set. A ban lifted before it ends is kept, marked as lifted, and does not count toward a longer ban; deleting its entry from bans.json forgets it entirely.

Clients of the AS numbers and countries listed in the biased thresholds, and clients listed by a reputation source whose action is limit:<percent>, get lower limits (see "Biased thresholds"), so the same rules ban them after fewer requests. For them the result is always a ban, first a temporary one and then, for repeated abuse, a permanent one.

Some refusals make no ban at all. DENY_NETS, a blocklist or reputation hit whose action is deny, and the country lists (DENIED_COUNTRIES, EXCLUSIVELY_ALLOWED_COUNTRIES) refuse each request they cover and record nothing in bans.json.

Configuration surface

Conventions: lists are comma separated; netblocks are CIDR (a bare address means /32 or /128); durations use Go syntax plus d for days (90s, 15m, 24h, 7d); byte sizes accept K, M, G suffixes. Countries are two-letter ISO codes in either case (de and DE are the same); a code that is not a country code, such as nk (North Korea is kp), stops the start with a message naming it.

  • Only UPSTREAM_URL is required. Every other setting has a default chosen for a service facing the internet in 2026, or stays off until the operator supplies or chooses what it needs: an alert destination, an account key, a lookup source.
  • Any limit or threshold can be switched off with the value off.
  • A list set to an empty value is an empty list, and replaces the default.
  • Every variable may instead be given as NAME_FILE pointing at a file holding the value, for secrets and long lists.
  • Settings, including NAME_FILE files, are read once at start; changing one means restarting the container. The files the sidecar watches while it runs are its state files, its rule files and the lookup database.
  • Invalid configuration stops the process at start with a message naming the variable. At start the effective configuration is logged with secrets masked.

The settings, by group:

  • Core
    • UPSTREAM_URL (required): the application, for example http://gitea:3000.
    • LISTEN_ADDR (default :8080): proxy listener.
    • ADMIN_LISTEN_ADDR (default 127.0.0.1:9090): admin listener.
    • ADMIN_TOKEN: bearer token required for the ban management endpoints. Unset by default, which switches those endpoints off; bans are then managed by editing bans.json.
    • INSTANCE_NAME (default: the host name in UPSTREAM_URL, for example gitea): included in every log line, metric and alert. Set it, for example to fsn1app1/gitea, when several sidecars report to one place.
    • MODE (default enforce): enforce, or observe to log and alert on every decision while refusing nothing.
    • TRUSTED_PROXIES (default 10.0.0.0/8,172.16.0.0/12,192.168.0.0/16, the private address ranges): netblocks whose X-Forwarded-For is believed, normally where traefik reaches the sidecar from. The header is used only when the TCP peer is inside this list; otherwise the peer's own address is the client, so a client that connects directly cannot claim another address. A list given replaces the default; set but empty, it trusts nothing.
    • IPV6_GROUP_PREFIX (default 64).
    • MAX_TRACKED_CLIENTS (default 20000): clients held in memory and in clients.json (see "Persistent state").
    • MAX_BANS (default 5000): bans held in memory and in bans.json, past, active and permanent (see "Bans").
  • Persistent state
    • STATE_DIR (default /data): the JSON state files. The image declares /data as a volume, so it is writable even when nothing is mounted; a named volume, as in the examples below, keeps the state in a place that outlives the container.
    • STATE_WRITE_DELAY (default 10s): after a ban is made, lifted or made permanent, bans.json is written this long later with every change made in between, so a burst of changes becomes one write.
    • STATE_COUNTER_INTERVAL (default 15m): how often clients.json, reputation.json and alerts.json are written, and bans.json when only the counts in its notes changed.
  • Logging
    • The request log on stdout is always on and has no switch.
    • LOG_LEVEL (default info): for the process's own messages (start-up, fetch failures, state writes), which are JSON lines on stdout too, marked "type":"process". It does not filter the request log.
    • LOG_REQUEST_HEADERS (default accept,accept-language,accept-encoding,content-type,origin,range): extra request headers to record. Authorization, Cookie and Set-Cookie values are never logged, only whether they were present.
    • LOG_REMOTE_URL: when set, every log line is also sent to this endpoint. Forms: syslog+udp://host:514, syslog+tcp://host:514, syslog+tls://host:6514.
    • LOG_REMOTE_TLS_CA_FILE: optional CA certificate for the +tls forms.
    • LOG_REMOTE_BUFFER (default 10000): lines held in memory while the endpoint is unreachable; when full the oldest are dropped and counted.
    • LOG_REMOTE_FACILITY (default local0), LOG_REMOTE_APP_NAME (default INSTANCE_NAME): syslog header fields.
  • Metrics
    • METRICS_ENABLED (default true).
    • METRICS_LISTEN_ADDR (default :9100): its own listener, so it can be bound to a monitoring network without exposing the admin endpoints.
    • METRICS_PATH (default /metrics).
    • METRICS_TOP_N (default 50): how many AS numbers and countries get their own series; the rest are summed as other.
    • METRICS_TOKEN: optional bearer token; unset means no authentication, which is the usual arrangement for a scraper on a private network.
  • Static lists
    • ALLOW_NETS: bypass everything (monitoring, the owner's own networks).
    • RATE_LIMIT_EXEMPT_NETS: bypass request and byte limits only; the error burst threshold, attack detection and bans still apply.
    • DENY_NETS: always refused.
  • Request rate limits, per client (R1, R2). Breaking one bans the client (see "Bans"), so the defaults sit several times above what one busy person produces: a browser loading a heavy page makes a few hundred requests, a git clone a handful, and several people often share one address.
    • RATE_LIMIT_PER_MINUTE (default 1000), RATE_LIMIT_PER_HOUR (default 10000), RATE_LIMIT_PER_DAY (default 50000).
    • RATE_LIMIT_EXEMPT_PATHS: path prefixes not counted (static assets, health checks).
  • Byte limits, per client. A response's bytes are counted when it ends, so every default sits above the largest response allowed (CLIENT_RESPONSE_MAX_BYTES, 5 GB) and no single download breaks one.
    • BYTES_LIMIT_PER_MINUTE (default 10G), BYTES_LIMIT_PER_HOUR (default 20G), BYTES_LIMIT_PER_DAY (default 50G).
    • BYTES_COUNT (default both): response, request or both.
  • Size and time limits, per request, in both directions. The client-facing limits apply between the client and the sidecar, the app-facing ones between the sidecar and the app. Bodies stream straight through, so a request body reaches the app while the client is still sending it, and of two matching limits the lower one acts first.
    • CLIENT_REQUEST_TIMEOUT (default 60s): how long a client may take to send its whole request, headers and body.
    • CLIENT_REQUEST_MAX_BYTES (default 100M): the largest request body a client may send.
    • CLIENT_REQUEST_HEADER_MAX_BYTES (default 32K): the largest request line and headers a client may send. Over it, the sidecar answers 431 and closes the connection, and nothing reaches the app.
    • CLIENT_IDLE_TIMEOUT (default 120s): how long a kept-open connection may wait for its next request before the sidecar closes it. It is longer than the 90 seconds after which traefik, by default, closes a connection it is not using, so traefik closes first and never sends a request on a connection the sidecar is closing.
    • CLIENT_RESPONSE_TIMEOUT (default 30m): how long the sidecar may take to deliver one response to the client, from the end of the request to the last byte.
    • CLIENT_RESPONSE_MAX_BYTES (default 5G): the largest response body sent to a client.
    • UPSTREAM_REQUEST_TIMEOUT (default 60s): how long the sidecar may take to connect to the app and send it the whole request.
    • UPSTREAM_REQUEST_MAX_BYTES (default 100M): the largest request body sent to the app.
    • UPSTREAM_RESPONSE_TIMEOUT (default 30m): how long the app may take to send one whole response, from the end of the request to the last byte.
    • UPSTREAM_RESPONSE_MAX_BYTES (default 5G): the largest response body taken from the app.
    • When a limit is passed before the response has started, the sidecar answers itself: 413 for a request body that is too large, 408 for a client that is too slow, 502 for a response that is too large, 504 for an app that is too slow. A request that announces a body larger than its limit is refused before anything reaches the app. Once the response has started it can only be cut off, and the connection is closed.
    • A WebSocket connection leaves these limits behind once it is upgraded: it stays open until either side closes it.
  • Lookup of AS number and country (R7). Off by default: the file needs an account only the operator can hold, and GeoJS is told every visitor's address.
    • LOOKUP_SOURCE (default off): off, file or geojs.
    • LOOKUP_DB_PATH: the file for file, the IPinfo Lite database in its .mmdb form (ipinfo_lite.mmdb), one file that carries both country and AS number; the sidecar reads its asn, as_name and country_code fields. The operator downloads it with a free IPinfo account and mounts it read-only. The sidecar never fetches it and holds no account or token for it. A file that is missing or unreadable at start stops the start.
    • Refreshing the file is the operator's business; IPinfo updates it daily. The sidecar notices when the file is replaced and reads it again; a replacement it cannot read is ignored, logged and sent as one file_error alert, and the file it has stays in use. Mount the directory that holds the file rather than the file itself: docker does not show a single mounted file being replaced on the host.
    • geojs asks the free GeoJS web service, which needs no account or key, about several addresses in one request (https://get.geojs.io/v1/ip/geo.json?ip=a,b,c), and reads each answer's country_code, asn and organization_name (the AS name). A country_code of null, which GeoJS gives for some ranges, and an asn of 64512, which it gives when it knows none, count as unknown. GeoJS is told the address of every new visitor.
    • Each GeoJS answer is kept for 7 days with the client's record, in memory and in clients.json, so it survives a restart; after 7 days the client's next request asks again. A client dropped from the table loses its answer.
    • LOOKUP_TIMEOUT (default 1s): how long a client with no kept answer waits for one; GeoJS normally answers in a fraction of that. At most one request to GeoJS is under way at a time, the addresses that arrive meanwhile are asked about together in the next one, and a request that takes longer than LOOKUP_TIMEOUT is abandoned. A client whose answer has not come in time counts as unknown until it comes: its later requests do not wait, and its address is asked about again in the background.
    • GeoJS publishes no rate limit, but its terms forbid "an excessive amount of API requests", judged by GeoJS alone, and let it block a caller. While GeoJS is slow, down or refusing the sidecar, clients with a kept answer are unaffected and new clients count as unknown, so EXCLUSIVELY_ALLOWED_COUNTRIES, when set, refuses them. The sidecar keeps asking with backoff and sends one source_failure alert per cooldown. A service that cannot accept this uses the file.
    • One source at a time: LOOKUP_SOURCE=file without LOOKUP_DB_PATH, or LOOKUP_DB_PATH with any other LOOKUP_SOURCE, stops the start with a message naming both. So does a setting that needs lookups while LOOKUP_SOURCE is off: the country lists and the biased thresholds below, ADD_LOOKUP_HEADERS, and the per-AS-number anomaly thresholds.
    • ADD_LOOKUP_HEADERS (default false): pass X-Client-ASN and X-Client-Country to the app.
  • Country lists. Both are empty by default, since they need a lookup source. Clients in ALLOW_NETS are not checked.
    • DENIED_COUNTRIES: for example cn,ru,kp,ir,ua,by. Every request from a listed country is refused.
    • EXCLUSIVELY_ALLOWED_COUNTRIES: for example us,de. Every request from any other country is refused, and so is every request from a client the lookup cannot place, such as an address missing from the database or one GeoJS has not answered for in time. A list that let those through would let every new client in whenever GeoJS stops answering.
    • Both may be set. DENIED_COUNTRIES then adds nothing, since the exclusive list already refuses every other country, and a code on both lists stops the start with a message naming it.
    • A refused request is answered with BAN_RESPONSE, as a banned client is, as soon as the client's address has been looked up, before its body is read. It skips the reputation checks, the limits, the rule files and the Core Rule Set, is logged with the action country_denied, and is counted in the metrics. It is a refusal, not a ban: it is not an offence and makes no ban record.
  • Biased thresholds (R8). The lists are empty by default: they need a lookup source, which only the operator can choose.
    • ASN_LIMIT_PERCENT: for example AS14061:50,AS16276:50,AS45102:25. Clients in a listed AS number get that percentage of every request and byte limit, so the rules in "Bans" ban them after fewer requests than others. 0 is a zero allowance: the client's first request breaks a limit and bans it.
    • COUNTRY_LIMIT_PERCENT: same form with ISO country codes, for example CN:25,RU:50. Each client from a listed country gets that percentage of the normal per-client limits.
    • No budget is shared by a whole country or AS number: one abuser could use it up and lock out everyone else there, a denial of service nobody chose and the one this tool exists to prevent. Refusing a whole country is left to the operator's choice, through the country lists.
    • When more than one percentage applies to a client (its AS number, its country, a reputation hit), the lowest applies.
    • ASN_BYTES_PERCENT, COUNTRY_BYTES_PERCENT: optional overrides applied to byte limits only, when the byte percentage should differ from the request percentage.
    • UNKNOWN_LIMIT_PERCENT (default 100): for clients the lookup cannot place.
    • ASN_LIMIT_PERCENT_URL: optional URL of a text file of AS:percent lines, so one abuse-source list can be shared by every sidecar in the fleet. It is fetched and refreshed like the blocklists.
  • Attack detection (R9)
    • RULES_DIR (default /etc/smallwebwaf/rules.d): directory of rule files, read at start and again whenever a file in it changes; format under "Rule files".
    • RULES_ENABLED (default true): false skips rule files entirely.
    • WAF_MODE (default block): off, detect (log and alert only), or block.
    • WAF_PARANOIA_LEVEL (default 1), WAF_ANOMALY_THRESHOLD (default 5): the Core Rule Set's own two tuning values, at the Core Rule Set's own defaults.
    • WAF_DISABLED_RULES (default 920420,920440): rule ids to switch off when an app trips a false positive. The two in the default refuse requests by content type and by file extension; in front of a code forge they would refuse git's clone and push over HTTP and the display of source files such as .sh or .sql. A list given replaces the default, so include them in it.
    • WAF_EXEMPT_PATHS: path prefixes not inspected.
    • WAF_BODY_LIMIT (default 128K): bodies are inspected up to this size and streamed beyond it without buffering, so large uploads and git pushes are not held in memory.
    • TRAP_PATHS: paths the app never serves and only scanners ask for, for example /wp-login.php,/xmlrpc.php in front of gitea. A request for one is a clear sign of attack. This is the env-var short form of a path rule with the ban action, for deployments that mount no rule files.
    • ERROR_BURST_THRESHOLD (default 30): responses with status 403 or 404 per client per minute; more than this breaks a limit (see "Bans"). Scanners walking lists of paths cause hundreds a minute; a person rarely causes more than a few. 401 is not counted: git and container registry clients send their first request without credentials and are answered 401 at the start of every push, every fetch from a private repository and every image pull.
  • Bans (R5), following the rules under "Bans"
    • ATTACK_BAN_DURATION (default 7d): the ban for a first clear sign of attack.
    • LIMIT_BAN_DURATION (default 1h): the ban for a first broken limit.
    • LIMIT_BAN_REPEAT_WINDOW (default 24h): breaking a limit again within this time after a ban for a broken limit ended makes the next ban three times as long.
    • MAX_BAN_DURATION (default 7d): a ban that would be longer is permanent instead.
    • BAN_RESPONSE (default 403): 403, 429, or close to drop the connection without an answer. 403 makes a ban easy to recognise when debugging; close tells the client nothing.
    • BAN_SCOPE_V4_PREFIX (default 32): widen to for example 24 to ban the surrounding netblock.
  • Reputation (R6). No source is on by default. A DNS blocklist would be told the address of every visitor. AbuseIPDB and CrowdSec need an account or an engine of the operator's own. A downloaded list reveals nothing about visitors, but each sidecar fetches its own copy, and the list most fit to be a default, Spamhaus DROP, may be fetched at most once a day: a default would break that on any host that runs several sidecars.
    • BLOCKLIST_URLS: text files of addresses and netblocks, one per line; anything after a ; or # on a line is ignored. For example the Spamhaus DROP list, https://www.spamhaus.org/drop/drop.txt; Spamhaus asks that it be fetched at most once a day and that products using it credit The Spamhaus Project. Refreshed every BLOCKLIST_REFRESH (default 24h); the last good copy is kept on failure and across restarts.
    • BLOCKLIST_ACTION (default deny): deny, limit:<percent>, or log. deny refuses every request from a listed address, which suits lists of networks that send nothing legitimate, such as DROP. limit:<percent> gives listed clients that percentage of every limit, so they are banned after fewer requests; it suits a list of addresses shared with ordinary visitors, such as Tor exits.
    • DNSBL_ZONES: for example dnsbl.dronebl.org. Zones meant for mail (lists of residential ranges) will block ordinary visitors and should not be used; Spamhaus zones need their keyed query service, given as a zone name containing the key.
    • DNSBL_RESOLVER: optional resolver address, since public resolvers are refused by several list operators.
    • ABUSEIPDB_KEY, ABUSEIPDB_MIN_SCORE (default 75), ABUSEIPDB_DAILY_BUDGET (default 900; the free tier allows 1000 checks a day). Only clients that have already committed one offence are queried, so the budget is spent on suspects.
    • CROWDSEC_LAPI_URL, CROWDSEC_LAPI_KEY: optional. If a CrowdSec engine exists on the host, pull its decision list on a schedule and treat listed addresses as banned. This is how a fleet-wide blocklist can arrive without the sidecar depending on CrowdSec. The list is kept like a fetched blocklist, and a listed client's request makes a ban with the cause crowdsec that lasts as long as CrowdSec's decision, so a long list does not fill bans.json.
    • REPUTATION_ACTION (default limit:25): deny, limit:<percent>, or log, for DNSBL and API hits. Such verdicts are less certain than a blocklist, so by default a listed client gets a quarter of every limit and is banned after a quarter of the requests.
    • REPUTATION_CACHE_TTL (default 24h), REPUTATION_TIMEOUT (default 2s).
  • Alerting (R3)
    • ALERT_WEBHOOK_URL: JSON POST; schema below. ALERT_WEBHOOK_HEADERS: optional Name:value pairs for authentication.
    • ALERT_SLACK_WEBHOOK_URL: Slack incoming webhook, formatted message.
    • ALERT_NTFY_URL (full topic URL), ALERT_NTFY_TOKEN: title, priority and tags set from the event type.
    • ALERT_EVENTS (default ban,permanent_ban,waf_block,anomaly,reputation_hit,source_failure,file_error): which event types are sent. source_failure is a reputation source or GeoJS failing or refusing the sidecar. file_error is a rule file or state file edited while running that does not parse, a replacement lookup database that cannot be read, or a state file that cannot be written.
    • ALERT_COOLDOWN (default 15m): the same event type for the same client or netblock is not repeated within this time; a count of suppressed repeats is included in the next one.
    • ALERT_MAX_PER_HOUR (default 60): beyond this, alerts are rolled into one summary per hour so a wide attack cannot flood the channel.
  • Anomaly thresholds, alert only, nothing is blocked (R4). None is set by default, and a threshold left unset sends no alert: these only alert, an alert needs a destination only the operator can supply, and what counts as unusual depends on each service's normal traffic, which the metrics show.
    • Per client: ANOMALY_CLIENT_REQUESTS_PER_MINUTE, ANOMALY_CLIENT_REQUESTS_PER_HOUR, ANOMALY_CLIENT_BYTES_PER_MINUTE, ANOMALY_CLIENT_BYTES_PER_HOUR.
    • Per surrounding netblock (ANOMALY_NET_V4_PREFIX default 24, ANOMALY_NET_V6_PREFIX default 48): ANOMALY_NET_REQUESTS_PER_MINUTE, ..._PER_HOUR, ANOMALY_NET_BYTES_PER_MINUTE, ..._PER_HOUR.
    • Per AS number, which needs a lookup source: ANOMALY_ASN_REQUESTS_PER_MINUTE, ..._PER_HOUR, ANOMALY_ASN_BYTES_PER_MINUTE, ..._PER_HOUR.
    • Whole service: ANOMALY_TOTAL_REQUESTS_PER_MINUTE, ..._PER_HOUR, ANOMALY_TOTAL_BYTES_PER_MINUTE, ..._PER_HOUR.
    • Named netblocks: WATCH_NETS, for example office=203.0.113.0/24,scraper-x=198.51.100.0/22, with WATCH_REQUESTS_PER_MINUTE, ..._PER_HOUR, WATCH_BYTES_PER_MINUTE, ..._PER_HOUR applied to each named block as a whole.
    • Anomaly counting includes clients in ALLOW_NETS and RATE_LIMIT_EXEMPT_NETS, since an exempt client misbehaving is worth knowing about.

Rule files

A required feature: the sidecar reads every *.rules file in RULES_DIR and checks each request against them. The files are plain text meant to be edited by hand, so that a new scanning pattern seen in the request log can be turned into a rule in one line. The sidecar watches the directory: a file edited, added or removed there takes effect while the sidecar runs.

  • One rule per line, four fields separated by spaces or tabs; the fourth field runs to the end of the line:

    <id>  <target>  <action>  <regex>
    
  • Blank lines and lines starting with # are ignored.

  • id: a short name of letters, digits, - and _, unique across all files. It appears in the request log, metrics, alerts and ban notes.

  • target: what the regex is matched against.

    • path: the URL path as received, before any decoding.
    • query: the raw query string.
    • uri: path and query together, both as received and once percent-decoded, so an encoded probe cannot slip past.
    • method, host, user_agent, referer.
    • header:<Name>: any one request header.
    • Request bodies are not available to rule files; body inspection is the Core Rule Set's job.
  • action:

    • log: note the match in the request log and do nothing else.
    • block: refuse the request with 403. The client is not banned for it, but the refusal counts toward the error burst (see "Bans").
    • ban: the request is a clear sign of attack. Refuse it and ban the client's netblock for seven days (ATTACK_BAN_DURATION); any further request during those days, or a later clear sign of attack, makes the ban permanent (see "Bans"). Use it only for requests no real visitor sends, and anchor a path at the site root (^/): a file of the same name deeper in a site can be ordinary content, such as a file in a gitea repository.
  • regex: Go regular expression syntax (RE2). It has no backreferences or lookaround, and in exchange matching time is linear in the input, so no rule can be made to stall the proxy. (?i) at the front makes a rule case-insensitive. A rule matches if the regex matches anywhere in the target; anchor with ^ and $ when that is not wanted.

  • Files are read in name order (00-default.rules before 50-gitea.rules), rules in line order.

  • Rules are compiled when their file is read and held in memory. At start, a line that does not parse, a regex that does not compile, or a duplicate id stops the process with a message naming the file and line. While running, the same faults leave the rules as they were, including the earlier version of that file, and the log and one file_error alert name the file and line; once the file is fixed, it is read again. A missing or empty directory is not an error: the log says that no rules were loaded.

  • The image ships a default file in RULES_DIR. Mounting a directory over it replaces the defaults; mounting single files into it adds to them. Docker does not show a single mounted file being replaced on the host, which is how many editors save, so rules meant to be edited while the sidecar runs belong in a mounted directory, with a copy of the default file if the defaults are to stay.

  • In MODE=observe every action is logged as what would have happened and nothing is refused.

  • Clients in ALLOW_NETS are not checked.

Example file:

# 00-default.rules: probes no real visitor sends

# id             target      action  regex
env-file         path        ban     (?i)^/\.env(\.[a-z]+)?$
git-dir          path        ban     ^/\.git/(config|HEAD|index)$
php-shell        path        ban     (?i)^/(shell|c99|r57|wso|alfa)\.php$
scanner-agent    user_agent  ban     (?i)\b(sqlmap|nikto|nuclei|masscan|zgrab|wpscan)\b
path-traversal   uri         block   (\.\./){2,}
empty-agent      user_agent  log     ^$

The image ships one default file of this kind. It is kept short and limited to patterns that are wrong for every app, and its path rules are anchored at the site root: in front of gitea, /.env is a probe, while /<owner>/<repo>/src/branch/main/.env.example is a file in a repository that any visitor or search crawler may open. Anything app-specific belongs in a file the deployer mounts: a request for /wp-login.php, for example, is a clear sign of attack in front of gitea and an ordinary login in front of WordPress.

# 50-gitea.rules: WordPress probes, which gitea never serves
wp-probe         path        ban     (?i)^/(wp-login\.php|xmlrpc\.php|wp-admin/)

Persistent state

All state lives in memory, and the files in STATE_DIR hold a copy of all of it, so an orderly stop loses nothing. The one exception is log lines still waiting to be sent to LOG_REMOTE_URL, which stdout has already carried. No database is used, and nothing is read from disk while serving a request. The files are meant for people as well: an admin can read or edit them at any time, and the running sidecar takes the edit in.

  • Files, each holding one kind of state:
    • bans.json: active and past bans, up to MAX_BANS. Per entry: the netblock, start, expiry (null for permanent), what caused it (attack, limit, admin or crowdsec), a short reason, when it was lifted if an admin lifted it, and a notes field (below).
    • clients.json: per client, its minute, hour and day counters and its history since it was first seen: first and last time seen, the AS number, AS name and country last looked up and when (with GeoJS, the answer kept for 7 days), total requests and bytes in each direction, how many requests were forwarded and how many refused, responses by status class, and offences by kind. Clients that were never banned are kept too, so every client's history survives a restart; GET /clients/<ip> shows it.
    • reputation.json: the last good copy of each list fetched from a URL, cached DNSBL and reputation API verdicts, each with the time it was fetched, and the AbuseIPDB checks spent today, so a restart does not reset the daily budget.
    • alerts.json: the anomaly counters for surrounding netblocks, AS numbers, named netblocks and the whole service; for each event type and client or netblock, when it was last sent and how many repeats the cooldown has held back since; the alerts sent this hour; and alerts still waiting to be sent.
  • Ban notes. The notes on a ban hold what an admin needs to decide whether to lift it, drawn from what the sidecar already knows; nothing is looked up to fill them:
    • the AS number, AS name and country (empty when lookups are off);
    • what was broken: the rule ids and target that matched, or the limit, its window, the count reached and the client's limit percentage with what set it; and any reputation sources that listed the client;
    • the requests that caused the ban, up to the last ten: time, method, host, path with its query string, status and user agent, each text cut to 256 bytes;
    • how many requests counted toward the ban, and the time span over which they came;
    • the netblock's total requests since it was first seen, and the requests refused under this ban so far, kept up to date while the ban lasts;
    • how many earlier bans of each kind the netblock has had.
  • Size and disk writes. Each file is rewritten whole, so each has a bound:
    • clients.json takes about 1 KiB per client. Clients are dropped only when the table is full, so on a public service the file grows to the default MAX_TRACKED_CLIENTS of 20,000, about 20 MiB. Written every 15 minutes, that is under 2 GiB of disk writes a day.
    • bans.json takes about 2 KiB per ban and at most about 8 KiB, since the texts in the notes are cut short. At the default MAX_BANS of 5,000 it is about 10 MiB, and never more than about 40 MiB. It is written when a ban is made, lifted or made permanent, at most once every 10 seconds, and otherwise with the 15-minute write, so its writes follow the bans made: with a full file, a hundred new bans a day come to about 1 GiB of disk writes.
    • reputation.json and alerts.json are usually a few MiB or less.
    • Writing a file of these sizes takes well under a second on an ordinary disk, in the background, so rewriting whole files needs nothing cleverer.
  • Format: indented JSON with a top-level version number, entries sorted by client address, times in RFC 3339 UTC, durations and sizes as plain numbers with the unit in the field name. The aim is that a person can open bans.json in an editor, find an address, and remove or add an entry. clients.json puts each client on one line instead, which halves its size and lets grep show everything about one client.
  • Writing:
    • serialise from a snapshot taken under the lock, so requests are not held up while the file is written;
    • write to a temporary file in the same directory, sync it, rename it over the real name, sync the directory. A crash at any point leaves either the old complete file or the new complete file, never a partial one;
    • bans.json is written STATE_WRITE_DELAY after a ban is made, lifted or made permanent. Changes that only update the counts in its notes wait for the write every STATE_COUNTER_INTERVAL, as all of clients.json, reputation.json and alerts.json do. All four are written on orderly shutdown (SIGTERM);
    • before writing a file, the sidecar checks whether it changed on disk since the sidecar last read or wrote it; if it did, the sidecar takes that edit in first (below), so an admin's edit is never overwritten.
  • Reading at start:
    • Active bans go into an in-memory prefix lookup; everything else into maps. Counter buckets and cache entries whose time has passed are discarded.
    • A missing file means empty state and is normal on first run.
    • A file that does not parse, or has an unknown version, stops the process with a message naming the file and position. Starting with empty state would silently forgive every repeat offender, and since writes are atomic a broken file can only come from a hand edit, which the editor should hear about.
  • Edits while running:
    • The sidecar watches STATE_DIR and notices a file that is edited, replaced or added. It tells its own writes from an admin's by comparing the file with what it last wrote.
    • An edit that parses is taken in at once: what the file says replaces what the sidecar held for that file. Removing a ban's entry from bans.json lifts the ban; adding an entry bans. Changes the sidecar made after the admin opened the file, such as a new ban, are lost when the admin saves over them; most editors warn when a file changed on disk while it was open.
    • An edit that does not parse does not stop the running sidecar. It keeps the state it has, renames the edited file to <name>.bad (for example bans.json.bad) so the edit is kept, writes the file again from memory, logs the file and the position of the error, and sends them in one file_error alert. The admin fixes the .bad file and moves it back.
  • STATE_DIR not writable at start: the process exits. A write that fails while running: state stays correct in memory, the failure is logged, counted in metrics, sent as a file_error alert once per cooldown, and retried at the next write.
  • What a hard kill can lose: ban changes from the last STATE_WRITE_DELAY (10 seconds), and everything else from the last STATE_COUNTER_INTERVAL (15 minutes): counts and history, the counts in ban notes, GeoJS answers, reputation verdicts and the AbuseIPDB count, anomaly counters and alert cooldowns. Losing the volume loses ban history, not service.

Request log

One JSON object per line on stdout for every request, including refused ones. stdout is always on. When LOG_REMOTE_URL is set the same lines are also sent to the remote endpoint, so a deployment can stop depending on docker's log handling while docker logs keeps working.

  • Standard web log fields: time (RFC 3339 with milliseconds), instance, client_ip, method, scheme, host, path, query, protocol, status, request_bytes, response_bytes, referer, user_agent.
  • Request detail: request_id (generated if traefik did not supply one, and passed to the app), peer_ip (the TCP peer, normally traefik), forwarded_for (the header as received), client_group (the /64 or configured prefix used for counting), asn, as_name, country, content_type, content_length, the headers named in LOG_REQUEST_HEADERS, has_authorization and has_cookie as booleans, websocket when the connection was upgraded.
  • Response detail: response_content_type, upstream_status (differs from status when the sidecar answered itself), cache_control, location on redirects, aborted when the client went away early.
  • Decision: action (forward, banned, denied, country_denied, rate_limited, rule_blocked, waf_blocked, too_large, timed_out, upstream_error), would_action in observe mode, limit_percent and which rule set it, counts (the client's minute, hour and day request and byte totals after this request), limit_hit (which window), rule_ids (rule file rules that matched), waf_rule_ids, waf_score, reputation (sources that listed the client), offence when one was recorded, ban_expires.
  • Timings in milliseconds: duration_total, duration_checks (everything the sidecar did before forwarding), duration_waf, duration_upstream_connect, duration_upstream_first_byte, duration_upstream_total.
  • Bodies are never logged. Query strings are logged as received; an app that carries secrets in query strings needs that fixed in the app.
  • The process's own messages share the stream as JSON lines with "type":"process"; request lines carry "type":"request".
  • Remote sending:
    • syslog forms send each line as the message of an RFC 5424 record, with octet-counted framing on TCP and TLS;
    • sending happens on its own goroutine from a bounded buffer (LOG_REMOTE_BUFFER). An unreachable or slow endpoint never delays a request and never stops stdout; it reconnects with backoff, drops the oldest lines when the buffer is full, and counts the drops in metrics. UDP gives no delivery signal at all and is offered only for compatibility.

Metrics endpoint

Prometheus text format on METRICS_LISTEN_ADDR at METRICS_PATH, on by default, switched off with METRICS_ENABLED=false. It has its own listener so it can be reached by a scraper without exposing ban management.

  • Traffic: requests and bytes in and out, by status class and action; request duration and upstream duration histograms; requests in flight.
  • Limits and bans: limit hits by window and kind (requests or bytes), size and time limit hits by limit, offences by kind, bans created by cause, permanent bans, active bans (gauge); requests refused by the country lists, by country.
  • Attack detection: rule file matches by rule id and action, and the number of rules loaded; Core Rule Set matches by mode and rule id (label limited to the rules that actually fired).
  • Lookup and reputation: requests and bytes by AS number and by country, limited to the METRICS_TOP_N (default 50) busiest of each with the rest summed as other, so the label set stays bounded; reputation queries, hits, failures and remaining daily budget by source; GeoJS requests and failures, and clients that counted as unknown because it did not answer in time; age of each blocklist and lookup database.
  • Housekeeping: tracked clients (gauge), state file writes, write failures, last successful write time and size per file; files read again after an edit, and edits set aside because they did not parse; alerts sent, failed and suppressed by destination; remote log lines sent, dropped and buffer depth; the standard Go runtime and process metrics.
  • No metric carries a client IP address as a label; per-address questions are answered by the request log and GET /clients/<ip>.

Admin listener

  • GET /healthz: for the container health check.
  • GET /bans, POST /bans (client or netblock, duration, reason), DELETE /bans/<client>: need ADMIN_TOKEN, and are switched off while it is unset. Editing bans.json does the same without a token.
  • GET /clients/<ip>: current counters, history, lookup result, reputation, offences, and bans with their notes; for answering "why was this address refused" and "what has this netblock been doing".

Alert webhook schema

One JSON object per alert:

  • instance, time, event (one of the ALERT_EVENTS values)
  • client, netblock, asn, as_name, country
  • reason: short human-readable sentence
  • detail: event-specific fields, for example window, count, limit, limit_percent, rule_ids, path, ban_expires, for a ban its notes, and for file_error the file and the position of the error
  • suppressed_repeats: number of identical alerts held back by the cooldown

Deployment as a sidecar

docker-compose shape, using gitea as the example. The app carries no traefik labels and publishes no ports; the sidecar carries the labels and points at the app's hostname. UPSTREAM_URL is the only setting it needs:

services:
    gitea:
        image: gitea/gitea:1
        # no traefik labels, no published ports
        networks: [internal]

    gitea-guard:
        image: <registry>/<image>:<pinned digest>
        read_only: true
        user: "65532:65532"
        environment:
            UPSTREAM_URL: http://gitea:3000
        volumes:
            - gitea-guard-state:/data
        networks: [internal, traefik]
        labels:
            traefik.enable: "true"
            traefik.http.routers.gitea.rule: Host(`git.example.invalid`)
            traefik.http.services.gitea.loadbalancer.server.port: "8080"

volumes:
    gitea-guard-state:

networks:
    internal:
    traefik:
        external: true
  • An app deployed by upaas takes the same shape, without the app's part of this file. upaas deploys the app with no traefik labels; the sidecar runs beside it from its own docker-compose file, carries the traefik labels, and points UPSTREAM_URL at the app's hostname. The sidecar must be able to reach that hostname, for example over a docker network the two share.
  • The application container leaves the traefik network, so the sidecar cannot be bypassed.
  • Traefik routes only port 8080. The metrics port is reached by the scraper over a docker network it shares with the sidecar; the admin port stays on loopback inside the container and is used through docker exec.
  • The state volume holds the JSON state files, a few tens of MiB at most with the defaults (see "Persistent state"); it needs no backup beyond whatever the host already does.
  • SSH access to gitea does not pass through the sidecar and is not protected by it.
  • Rollout per service: point traefik at the sidecar with only UPSTREAM_URL set. It protects the app from the first request, and nothing needs tuning first. Afterwards the request log and the ban notes say why each client was refused. A real visitor refused by mistake is let through with an exclusion (WAF_DISABLED_RULES, WAF_EXEMPT_PATHS, RATE_LIMIT_EXEMPT_PATHS) or an exemption (RATE_LIMIT_EXEMPT_NETS, ALLOW_NETS), and its ban is lifted in bans.json. Additions to consider at any time: an alert destination, a lookup source with country lists or biased thresholds, reputation sources, a remote log endpoint, and an app-specific rule file.
  • Notes specific to gitea:
    • A clone is a response, and fits the defaults of 30 minutes and 5 GB for all but the largest repositories and slowest links. A push is a request: at the defaults, one larger than 100 MB or taking more than 60 seconds to upload is cut off, so a gitea that takes large pushes needs CLIENT_REQUEST_MAX_BYTES, UPSTREAM_REQUEST_MAX_BYTES, CLIENT_REQUEST_TIMEOUT and UPSTREAM_REQUEST_TIMEOUT raised to fit. WAF_BODY_LIMIT keeps pack uploads out of memory.
    • The default WAF_DISABLED_RULES lets git's clone and push over HTTP, and the display of source files, through the Core Rule Set. Pages where people post code (issues, pull requests, the web editor) can still trip other rules; such a match refuses only that request, and the request log names the rule in waf_rule_ids for WAF_DISABLED_RULES.
    • Archive download and blame or history pages are what scrapers hammer; request limits do most of the work there.

Failure behaviour

  • Lookup database: a configured file that is missing or unreadable at start stops the start. A replacement that cannot be read while running is ignored: the file already loaded stays in use, the problem is logged, one file_error alert.
  • GeoJS slow, down or refusing the sidecar: clients with a kept answer are unaffected, new clients count as unknown, the sidecar keeps asking with backoff, one source_failure alert per cooldown.
  • Reputation source down or over quota: no verdict, service continues, one source_failure alert per cooldown.
  • Alert destination down: retried with backoff from a bounded queue, oldest dropped first, drops counted in metrics.
  • Upstream down: 502 from the sidecar, not counted as client offences.
  • Attack detection engine error on a request: request is forwarded, error logged and counted.
  • Remote log endpoint down: stdout continues, lines are buffered then dropped oldest first, drops counted in metrics.
  • State file write fails while running: memory stays authoritative, logged, counted, one file_error alert per cooldown, retried.
  • A state file or rule file edited while running that does not parse: the sidecar keeps running on what it has. The state file is set aside as <name>.bad and written again from memory; the rule file is left as it is. Logged, one file_error alert.
  • In short: the only things that stop the process happen at start: invalid configuration, a configured lookup database that is missing or unreadable, a rule file or state file that does not parse, and an unwritable STATE_DIR. Once running, a broken helper or a broken edit never takes the protected service down. The nearest it comes is GeoJS failing while EXCLUSIVELY_ALLOWED_COUNTRIES is set: new clients then cannot be placed, and that list refuses them, as it refuses every client it cannot place.

Risks the design has to handle

  • Forged X-Forwarded-For: handled by believing the header only from a peer inside TRUSTED_PROXIES and by walking it from the right. The default trusts every private address, which in the intended deployment means traefik; where other containers can reach the sidecar directly, setting TRUSTED_PROXIES to traefik's own network closes that gap.
  • Many clients behind one address (mobile carriers, offices, Tor): they share its limits and its bans. The default limits sit several times above what one busy person produces, RATE_LIMIT_EXEMPT_NETS and ALLOW_NETS take known shared addresses out, and the ban notes show an admin what happened before a ban is lifted.
  • A ban rule that matches a real visitor bans it for seven days, and the visitor's next request during those days makes the ban permanent. So the default rule file keeps to requests no real visitor sends, with its paths anchored at the site root, where no app serves .env files or web shells. A visitor banned by mistake is let back in by lifting the ban in bans.json.
  • IPv6 address rotation inside a /64: handled by grouping.
  • Widely distributed scrapers using thousands of addresses at low rates each: per-client limits do not see them. The AS number and netblock anomaly alerts reveal them once their thresholds are set, and a low ASN_LIMIT_PERCENT for that AS number is the response: it lowers the limits of each of its clients. There is no shared budget for a whole AS number or country, which one abuser could use up and so lock out everyone else there.
  • Core Rule Set false positives against real apps (gitea's editor, API payloads): a match refuses only that request and bans no one by itself; the request log names the rule, and exclusions by rule id and path fix it.
  • GeoJS as the lookup source: every new visitor's address goes to a third party, and a swarm of fresh addresses, when lookups peak, is when GeoJS may slow down or block the sidecar. Keeping answers for 7 days and asking about many addresses in one request keep the number of requests low; the file source has neither risk.
  • Slow-request attacks: CLIENT_REQUEST_TIMEOUT bounds how long a request may take to arrive, CLIENT_REQUEST_HEADER_MAX_BYTES how large its headers may be, and CLIENT_IDLE_TIMEOUT closes a kept-open connection that sends nothing more.
  • Alert floods: cooldown and hourly cap.

Build order

  • First: proxy with its size and time limits, client identification, static lists, three-window request limits and the bans they lead to, the ban ledger and the JSON state files with edits taken in while running, exemptions, observe mode, the full request log on stdout, the metrics endpoint, health.
  • Second: rule files, admin endpoints, alerting to all three destinations, remote log sending.
  • Third: AS number and country lookup from the file or GeoJS, the country lists, biased thresholds, byte limits, anomaly thresholds.
  • Fourth: blocklists, DNSBL, AbuseIPDB, optional CrowdSec decision feed.
  • Fifth: attack detection with Coraza and the Core Rule Set, trap paths, error bursts.
  • Each stage is usable on its own; the first three already cover the traffic problem the fleet has today.