sneak 43c63faa04 SPEC and README: owner rulings, issues 2 to 5, internet-ready defaults (closes #6)
Rewrite `SPEC.md` and `README.md` to the owner's rulings. The seven answers
to the spec's questions now stand where their topics live, and the questions
section is gone. Bans follow the owner's model: seven days for a clear sign of
attack and permanent on any further request, an hour for a broken limit,
tripled on a repeat within a day, permanent past seven days. The state files
hold all state, are watched, and take in an admin's edits while running. Every
ban carries notes, and each client's history survives a restart. Size and time
limits apply in both directions. Only `UPSTREAM_URL` is required: the Core Rule
Set refuses what it flags and every limit has a default.

Model: opus-5-5
2026-09-23 12:46:05 +00:00

smallwebwaf

smallwebwaf is a simple, fast, logging web application firewall for people who host their own services. It is one small container that sits between your reverse proxy (traefik) and one application: traefik points at smallwebwaf, and smallwebwaf points at the app. It needs one setting, the address of the app, and protects the app from the first request with defaults chosen for a service on the open internet. It keeps its state in memory and in JSON files you can read and edit, and writes a detailed JSON log line for every request.

Status: design stage. This repository currently holds the documents only; no code has been written. The design is in SPEC.md, and the survey of existing tools that led to it is in EVALUATION.md.

Why

Small self-hosted sites now receive a great deal of traffic nobody asked for: scrapers that ignore robots.txt and crawl every commit of every repository on a public git server, vulnerability scanners walking through lists of WordPress and .env paths, and credential-guessing bots. Most of it comes from a small number of hosting networks and countries. A single-person operation has no abuse desk and no CDN contract; it needs something small that can be put in front of one service and left alone.

The existing tools each solve part of this. Rule-based firewalls catch attack payloads but do not limit request rates. Rate limiters count requests but cannot tell a residential visitor from a rented server farm. The products that do most of it want several containers, a database and a web console. None of them can say "clients from these networks are banned after half as many requests as anyone else", which is the most useful thing to be able to say when nearly all abuse comes from a known list of AS numbers. EVALUATION.md goes through the candidates one by one.

smallwebwaf is meant to fill that gap:

  • protect a service from misbehaving scrapers and scanners with per-client request and byte limits over a minute, an hour and a day;
  • lower those limits for the countries and AS numbers that abuse commonly comes from, so their clients are banned after fewer requests than others;
  • ban abusers: briefly at first, longer each time they come back, and permanently when they keep at it; a scanner's first probe bans it for seven days;
  • log everything in a form that is easy to search and ship elsewhere;
  • stay small enough to understand: one binary, one container, environment variables, no database, one required setting.

Proposed features

  • Reverse proxy for one upstream application, streaming in both directions, with WebSocket support. One smallwebwaf per app.
  • Internet-ready out of the box: only UPSTREAM_URL must be set, and every other setting has a default chosen for a service facing the internet in 2026.
  • Real client address worked out from X-Forwarded-For, trusting only the proxy networks you list, by default the private address ranges. IPv6 clients are counted by /64.
  • Size and time limits on requests and responses, both between the client and smallwebwaf and between smallwebwaf and the app: by default a request may take 60 seconds and 100 MB, a response 30 minutes and 5 GB.
  • Rate limits per client on requests per minute, per hour and per day, and on bytes per minute, per hour and per day, on by default and set well above what real visitors need.
  • Netblocks that bypass rate limiting, netblocks that bypass everything, and netblocks that are always refused.
  • AS number and country lookup for every client from the IPinfo Lite database file, which you download and mount (see "Lookup database" below).
  • Biased limits: listed AS numbers and countries get a percentage of every limit, for example 50 percent for common abuse-source networks, so their clients are banned after fewer requests. Zero percent is a zero allowance: the first request breaks the limit and bans the client.
  • Attack detection:
    • a directory of plain text rule files, one regex per line, for catching scanning and penetration probes; easy to edit by hand, and picked up while running;
    • the OWASP Core Rule Set, run by the Coraza engine, refusing the requests it flags;
    • trap paths and bursts of error responses.
  • Bans:
    • a clear sign of attack, such as a probe for a .env file or a scanner's user agent, bans for seven days on the first request, and any further request during those days makes the ban permanent;
    • breaking a limit bans for an hour; breaking one again within a day of a ban ending triples the length, and a ban that would last longer than seven days is permanent instead;
    • every ban carries notes on why it was made, to help decide whether to lift it.
  • IP reputation: downloadable blocklists, DNS blocklists, AbuseIPDB, and an optional feed of decisions from a CrowdSec engine. Lookups happen in the background and never delay a request. None is on until you add it.
  • Alerts on attacks and bans to a generic webhook, Slack or ntfy, with a cooldown and an hourly cap so a wide attack cannot flood the channel.
  • Anomaly alerts when requests or bytes per minute or hour cross a threshold you set, for a single client, its surrounding netblock, an AS number, a named netblock or the whole service.
  • Observe mode: log and alert on every decision while refusing nothing.
  • Request log: one JSON object per request on stdout with the usual web log fields, the decision taken and why, AS number and country, and timings. Optionally also sent to a remote syslog server.
  • Prometheus metrics on their own port.
  • State (bans with their notes, each client's counters and history, the reputation cache) held in memory and kept in readable JSON files that always hold a copy of it, so a restart loses nothing. Edit a file, or add a rule file, and the running smallwebwaf picks up the change. Nothing is read from disk while serving a request.
  • A small admin endpoint for health checks, listing, adding and lifting bans, and asking why a given address was refused.

Not planned: TLS termination, routing for several apps, browser challenges (captcha or proof of work), a web console, or defence against floods large enough to fill the host's network link.

How it works, in short

For each request smallwebwaf:

  • works out who the client really is;
  • lets it straight through if it is on the bypass list, refuses it if it is on the deny list or currently banned;
  • looks up its AS number and country, and any cached reputation verdict;
  • picks the client's limit percentage from those;
  • checks the minute, hour and day request counters against the limits, and bans the client if it breaks one;
  • checks the request against the rule files and the Core Rule Set, and bans the client at once for a clear sign of attack;
  • forwards it to the app and streams the response back, within the size and time limits;
  • counts the bytes and any error response, bans the client if it broke a limit, updates its history, sends any alerts that are due, and writes the log line.

A minimal deployment beside an app in docker-compose. UPSTREAM_URL is the only setting:

services:
    app:
        image: example/app
        networks: [internal]

    waf:
        image: <registry>/smallwebwaf:<pinned digest>
        environment:
            UPSTREAM_URL: http://app:3000
        volumes: [waf-state:/data]
        networks: [internal, traefik]
        labels:
            traefik.enable: "true"
            traefik.http.routers.app.rule: Host(`app.example.invalid`)
            traefik.http.services.app.loadbalancer.server.port: "8080"

volumes:
    waf-state:

networks:
    internal:
    traefik:
        external: true

The volume keeps bans and client history in a place you choose; without it smallwebwaf still starts, on a volume docker creates for it.

A rule file is one rule per line: a name, what to match against, what to do, and a regex.

env-file       path        ban  (?i)/\.env(\.[a-z]+)?$
scanner-agent  user_agent  ban  (?i)\b(sqlmap|nikto|nuclei|wpscan)\b

SPEC.md has the full design: every environment variable, the ban rules, the rule file format, the state files, the log fields, the metrics, failure behaviour and the build order.

Lookup database

AS number and country lookups, and the biased limits that use them, need the free IPinfo Lite database (ipinfo_lite.mmdb). You download it with your own IPinfo account, mount it into the container, point LOOKUP_DB_PATH at it and refresh it when you choose; smallwebwaf never downloads it itself. IPinfo releases it under the Creative Commons Attribution-ShareAlike 4.0 International License and asks for attribution, in its own words on https://ipinfo.io/lite: "The attribution requirements can be met by giving our service credit as your data source. Simply place a link to IPinfo on the website, application, or social media account that uses our data." Its example of such a credit is a link mentioning "IP address data is powered by IPinfo". A service that uses the database through smallwebwaf should carry that link.

Documents

  • SPEC.md: the design.
  • EVALUATION.md: what already exists, what each tool covers and misses, and why none was adopted.
S
Description
No description provided
Readme
78 KiB
Languages
Markdown 100%