# smallwebwaf `smallwebwaf` is a simple, fast, logging web application firewall for people who host their own services. It is one small container that sits between your reverse proxy (traefik) and one application: traefik points at `smallwebwaf`, and `smallwebwaf` points at the app. It needs one setting, the address of the app, and protects the app from the first request with defaults chosen for a service on the open internet. It keeps its state in memory and in JSON files you can read and edit, and writes a detailed JSON log line for every request. Status: design stage. This repository currently holds the documents only; no code has been written. The design is in [`SPEC.md`](SPEC.md), and the survey of existing tools that led to it is in [`EVALUATION.md`](EVALUATION.md). ## Why Small self-hosted sites now receive a great deal of traffic nobody asked for: scrapers that ignore `robots.txt` and crawl every commit of every repository on a public git server, vulnerability scanners walking through lists of WordPress and `.env` paths, and credential-guessing bots. Most of it comes from a small number of hosting networks and countries. A single-person operation has no abuse desk and no CDN contract; it needs something small that can be put in front of one service and left alone. The existing tools each solve part of this. Rule-based firewalls catch attack payloads but do not limit request rates. Rate limiters count requests but cannot tell a residential visitor from a rented server farm. The products that do most of it want several containers, a database and a web console. None of them can say "clients from these networks are banned after half as many requests as anyone else", which is the most useful thing to be able to say when nearly all abuse comes from a known list of AS numbers. [`EVALUATION.md`](EVALUATION.md) goes through the candidates one by one. `smallwebwaf` is meant to fill that gap: - protect a service from misbehaving scrapers and scanners with per-client request and byte limits over a minute, an hour and a day; - lower those limits for the countries and AS numbers that abuse commonly comes from, so their clients are banned after fewer requests than others; - ban abusers: briefly at first, longer each time they come back, and permanently when they keep at it; a scanner's first probe bans it for seven days; - log everything in a form that is easy to search and ship elsewhere; - stay small enough to understand: one binary, one container, environment variables, no database, one required setting. ## Proposed features - Reverse proxy for one upstream application, streaming in both directions, with WebSocket support. One `smallwebwaf` per app. - Internet-ready out of the box: only `UPSTREAM_URL` must be set, and every other setting has a default chosen for a service facing the internet in 2026. - Real client address worked out from `X-Forwarded-For`, trusting only the proxy networks you list, by default the private address ranges. IPv6 clients are counted by /64 by default. - Size and time limits on requests and responses, both between the client and `smallwebwaf` and between `smallwebwaf` and the app: by default a request may take 60 seconds and 100 MB, a response 30 minutes and 5 GB. - Rate limits per client on requests per minute, per hour and per day, and on bytes per minute, per hour and per day, on by default and set well above what real visitors need. - Netblocks that bypass rate limiting, netblocks that bypass everything, and netblocks that are always refused. - AS number and country lookup for every client, off until you choose a source: the IPinfo Lite database file, which you download and mount, or the free GeoJS web service (see "Country and AS number lookup" below). - Country lists: `DENIED_COUNTRIES` refuses every request from the countries listed, `EXCLUSIVELY_ALLOWED_COUNTRIES` every request from anywhere else. Such a request gets the answer a banned client gets as soon as the client's address has been looked up, before its body is read and without the rule files or the Core Rule Set looking at it, and no ban is made. - Biased limits: listed AS numbers and countries get a percentage of every limit, for example 50 percent for common abuse-source networks, so their clients are banned after fewer requests. Zero percent is a zero allowance: the first request breaks the limit and bans the client. - Attack detection: - a directory of plain text rule files, one regex per line, for catching scanning and penetration probes; easy to edit by hand, and picked up while running; - the OWASP Core Rule Set, run by the Coraza engine, refusing the requests it flags; - trap paths, and a ban for a client that the rule files or the Core Rule Set refuse again and again. - Bans: - a clear sign of attack, such as a probe for a `.env` file or a scanner's user agent, bans for seven days on the first request, and any further request during those days makes the ban permanent; - breaking a limit bans for an hour; breaking one again within a day of a ban ending triples the length, and a ban that would last longer than seven days is permanent instead; - every ban carries notes on why it was made, to help decide whether to lift it. - IP reputation: downloadable blocklists, DNS blocklists, AbuseIPDB, and an optional feed of decisions from a CrowdSec engine. Lookups happen in the background and never delay a request. None is on until you add it. - Alerts on attacks and bans to a generic webhook, Slack or ntfy, with a cooldown and an hourly cap so a wide attack cannot flood the channel. - Anomaly alerts when requests or bytes per minute or hour cross a threshold you set, for a single client, its surrounding netblock, an AS number, a named netblock or the whole service. - Observe mode: log and alert on every decision while refusing nothing. - Request log: one JSON object per request on stdout with the usual web log fields, the decision taken and why, AS number and country, and timings. Optionally also sent to a remote syslog server. - Prometheus metrics on their own port. - State (bans with their notes, each client's counters and history, the GeoJS answers, the reputation cache, the alerting state) held in memory and kept in readable JSON files, written regularly and at every stop, so a restart loses nothing. Edit a file, or add a rule file, and the running `smallwebwaf` picks up the change. Nothing is read from disk while serving a request. - A small admin endpoint for health checks, listing, adding and lifting bans, and asking why a given address was refused. Not planned: TLS termination, routing for several apps, browser challenges (captcha or proof of work), a web console, or defence against floods large enough to fill the host's network link. ## How it works, in short For each request `smallwebwaf`: - works out who the client really is; - lets it straight through if it is on the bypass list, refuses it if it is on the deny list or currently banned; - looks up its AS number and country, and refuses it if that country is denied, or is not among the only ones allowed; - checks for a cached reputation verdict; - picks the client's limit percentage from those; - checks the minute, hour and day request counters against the limits, and bans the client if it breaks one; - checks the request against the rule files and the Core Rule Set, and bans the client at once for a clear sign of attack; - forwards it to the app and streams the response back, within the size and time limits; - counts the bytes and any refusal by the rule files or the Core Rule Set, bans the client if it broke a limit, updates its history, sends any alerts that are due, and writes the log line. A minimal deployment beside an app in docker-compose. `UPSTREAM_URL` is the only setting: ```yaml services: app: image: example/app networks: [internal] waf: image: /smallwebwaf: environment: UPSTREAM_URL: http://app:3000 volumes: [waf-state:/data] networks: [internal, traefik] labels: traefik.enable: "true" traefik.http.routers.app.rule: Host(`app.example.invalid`) traefik.http.services.app.loadbalancer.server.port: "8080" volumes: waf-state: networks: internal: traefik: external: true ``` The volume keeps bans and client history in a place you choose; without it `smallwebwaf` still starts, on a volume docker creates for it. A rule file is one rule per line: a name, what to match against, what to do, and a regex. ``` env-file path ban (?i)^/\.env(\.[a-z]+)?$ scanner-agent user_agent ban (?i)\b(sqlmap|nikto|nuclei|wpscan)\b ``` [`SPEC.md`](SPEC.md) has the full design: every environment variable, the ban rules, the rule file format, the state files, the log fields, the metrics, failure behaviour and the build order. ## Country and AS number lookup AS number and country lookups, and the country lists and biased limits that use them, are off until you choose one of two sources with `LOOKUP_SOURCE`. `LOOKUP_SOURCE=file` reads the free IPinfo Lite database (`ipinfo_lite.mmdb`). You download it with your own IPinfo account, mount the directory that holds it into the container, point `LOOKUP_DB_PATH` at the file and refresh it when you choose; `smallwebwaf` never downloads it itself, and reads it again when you replace it. It has to be the directory rather than the file itself: docker does not show a single mounted file being replaced, so a refresh would go unseen. IPinfo releases it under the Creative Commons Attribution-ShareAlike 4.0 International License and asks for attribution, in its own words on https://ipinfo.io/lite: "The attribution requirements can be met by giving our service credit as your data source. Simply place a link to IPinfo on the website, application, or social media account that uses our data." Its example of such a credit is a link mentioning "IP address data is powered by IPinfo". A service that uses the database through `smallwebwaf` should carry that link. `LOOKUP_SOURCE=geojs` asks the free GeoJS web service instead, with no account and no file. Every new visitor's address is sent to GeoJS. Each answer is kept for seven days, across restarts, and many addresses are asked about in one request. GeoJS publishes no rate limit but may block a caller it thinks asks too much; while it is not answering, new visitors count as coming from an unknown country, which `EXCLUSIVELY_ALLOWED_COUNTRIES` refuses. Neither source can place a private address, so a client on one, such as a visitor on your local network, another container or your monitoring, has no country: `EXCLUSIVELY_ALLOWED_COUNTRIES` refuses it unless you list it in `ALLOW_NETS`. Such addresses are never sent to GeoJS. ## Documents - [`SPEC.md`](SPEC.md): the design. - [`EVALUATION.md`](EVALUATION.md): what already exists, what each tool covers and misses, and why none was adopted.