Author SHA1 Message Date
clawbot ba5102092f Stamp the tag or short commit in a plain docker build (closes #166)
check / check (push) Successful in 5m29s
.dockerignore now lets .git into the build context, without
.git/config, which can hold a remote URL with a credential. ARG VERSION
has no default: given none, the build stage takes the version from
git describe --tags --always, and fails if the context carries .git
and no version comes out. pixad now logs its name, version and
architecture as its first log line, through the existing
Logger.Identify, which nothing called.

Model: opus-5-5
2026-10-02 02:41:23 +00:00
33 changed files with 284 additions and 2159 deletions
+1 -4
View File
@@ -1,7 +1,4 @@
# .git is sent without its config. Without a VERSION build argument the
# stage that compiles runs `git describe --tags --always` on .git, which
# does not need .git/config; that file can hold a credential, such as a
# password in a remote URL or the token the CI checkout step stores there.
# .git is sent so the build can stamp the version, without its config.
.git/config
.gitignore
.DS_Store
+5 -11
View File
@@ -18,14 +18,9 @@ COPY . .
# Tells script/lint it is inside a container, so it runs the linter.
ENV container=docker
# Run formatting check and linter. script/cibuild and script/docker pass
# a new CHECK_EPOCH on every run, and each check step names it in its
# command, so a new value reruns the step instead of reusing a cached
# success that checked nothing. A plain `docker build .` leaves it empty
# and reuses the check steps only for an identical build context.
ARG CHECK_EPOCH
RUN echo "check epoch: ${CHECK_EPOCH}" && make fmt-check
RUN echo "check epoch: ${CHECK_EPOCH}" && make lint
# Run formatting check and linter
RUN make fmt-check
RUN make lint
# Build stage
# golang:1.25.4-alpine, 2026-02-25
@@ -44,9 +39,8 @@ RUN script/bootstrap
# Copy source code
COPY . .
# Run tests; a new CHECK_EPOCH reruns them, as in the lint stage.
ARG CHECK_EPOCH
RUN echo "check epoch: ${CHECK_EPOCH}" && make test
# Run tests
RUN make test
# VERSION is declared here, not earlier: a new value reruns only the
# build, not script/bootstrap or the tests. Given none, the version is
+22 -153
View File
@@ -71,17 +71,13 @@ prevent abuse, and allowlisted source hosts for open access.
### Storage
- **Source content**:
`<state_dir>/cache/sources/<ab>/<cd>/<sha256 of source content>`
`<statedir>/cache/src-content/<ab>/<cd>/<sha256 of source content>`
- **Source metadata**:
`<state_dir>/cache/metadata/<hostname>/<sha256 of path and query>.json`
(host, path and query, content hash, upstream status and headers, fetch time)
- **Database**: `<state_dir>/state.sqlite3` (SQLite)
- **Transformed images**:
`<state_dir>/cache/variants/<ab>/<cd>/<sha256 of host, path, query, size, format, quality and fit>`,
each with a `.meta` file beside it holding its content type
`<ab>` and `<cd>` are the first and second pairs of characters of the file's
name.
`<statedir>/cache/src-metadata/<hostname>/<sha256 of path>.json`
(fetch time, original headers, request, content hash)
- **Database**: `<statedir>/state.sqlite3` (SQLite)
- **Output documents**:
`<statedir>/cache/dst-content/<ab>/<cd>/<sha256 of output content>`
Multiple source paths may reference the same content blob; the
database tracks references rather than using filesystem refcounting.
@@ -92,86 +88,17 @@ the metadata file stored beside it.
### Routes
pixa answers these routes; any other path answers 404. A path in this list asked
with a method the list does not give answers 405, except `/static/<file>`, which
answers any method as it answers `GET`. A browser's CORS preflight request
(`OPTIONS` with `Origin` and `Access-Control-Request-Method` headers) to any
path under `/v1/` answers 200, in maintenance mode too.
- `GET /` — the login page, or the URL generator page with a login session
(see Encrypted URLs). Needs: nothing. Answers: 200.
- `POST /` — log in with the signing key typed into the login page. Needs: the
login page's form (below). Answers: 303 to `/` with a login session cookie
that lasts 30 days for the right key; 200 with the login page and an error for
a wrong key; 429 over the login limit (below).
- `POST /generate` — make an encrypted URL from the generator page's form.
Needs: a login session and the generator page's form (below); without a login
session it answers 303 to `/`. Answers: 200 with the page showing the URL; 400
with the page naming a field that is not valid; 500 when the URL cannot be
made.
- `GET /logout` — end the login session. Needs: nothing. Answers: 303 to `/`.
- `GET` or `HEAD` `/v1/image/<host>/<path>/<size>.<format>` — an image, fetched,
resized and converted (below). Needs: a signature, unless the host is
allowlisted (see Source Hosts). Answers: 200; 304 when `If-None-Match` matches
the image's `ETag`; 400 for a URL or parameter that is not valid; 401 for a
missing or wrong signature, a missing `exp` or an `exp` in the past; 403 when
the upstream host, or a host it redirects to, is `localhost`, ends in
`.localhost` or `.local`, or has an address in a blocked network (see
`blocked_networks`); 502 when the upstream answered with an error status, and
for 5 minutes after that for the same source URL; 503 when pixa is busy or in
maintenance mode; 500 for any other failure.
- `GET /v1/e/<token>/<name>` — an image through an encrypted URL (see Encrypted
URLs). Needs: nothing but the URL. Answers: 200; 400 for a token that does not
decrypt, or that asks for a size or fit that is not valid; 410 once it has
expired; 504 when the upstream has not sent its response headers within
`upstream_fetch_timeout`, but 500 when that time runs out while the image
itself is still arriving; 403, 502, 503 and 500 as for `/v1/image/`.
- `GET /robots.txt` — asks every crawler to stay away (`Disallow: /`). Needs:
nothing. Answers: 200.
- `GET /.well-known/healthcheck.json` — JSON with `status` (`ok`), `now`,
`uptime_seconds`, `uptime_human`, `version`, `appname` and
`maintenance_mode`. Needs: nothing. Answers: 200, always.
- `GET /static/<file>` — the script the login and generator pages load. Needs:
nothing. Answers: 200, or 404 for a file that does not exist.
- `GET /metrics` — Prometheus metrics (see Architecture). Needs: HTTP basic
authentication with `metrics.username` and `metrics.password`. Answers: 200;
401 without them; 404 when they are not set, as the route then does not exist.
Both `POST` routes accept only a form that pixa's own page served: the page puts
a token in the form and sets a cookie to match, and a request without both is
refused with 403, so another site cannot submit the form from a visitor's
browser. The login and generator pages are meant to be opened over HTTPS: while
`debug` is off, a form sent from a page opened over plain HTTP is refused with
403, and while it is on, so is one sent from a page opened over HTTPS. Plain
HTTP is for development on the browser's own machine: the login session cookie
is always marked `Secure`, and over plain HTTP a browser keeps such a cookie
only for its own machine (`localhost`), if at all. A form is also refused with
403 when the page's host is not the `Host` header pixa receives, so a reverse
proxy in front of pixa must pass that header on unchanged. A form body over
1 MiB is refused with 413. The image routes answer the errors listed for them
with JSON holding `error`, `status` and `timestamp`.
An image URL has this form:
```
/v1/image/<host>/<path>/<size>.<format>?sig=<signature>&exp=<expiration>&q=<quality>&fit=<fit>
/v1/image/<host>/<path>/<size>.<format>?sig=<signature>&exp=<expiration>
```
Images are only fetched from origins using TLS with valid certificates, unless
`allow_http` is set: then pixa fetches every image over plain HTTP, which is for
testing only.
Images are only fetched from origins using TLS with valid certificates.
A request whose query string cannot be decoded, or gives any parameter more
than once, is refused with 400.
- `<format>`: one of `orig` (or `original`), `jpeg` (or `jpg`), `png`, `webp`,
`avif`, `gif`
- `<format>`: one of `orig`, `png`, `jpeg`, `webp`
- `<size>`: `orig` or `<width>x<height>` (e.g. `800x600`)
- `sig` and `exp`: the signature and its expiry, needed unless the host is
allowlisted (see Signature Specification)
- `q` and `fit`: the output quality and how the image is fitted to `<size>`,
both optional (values under Signature Specification). Both are part of what
is cached, so each value of either is a separate cached image.
An image is served with `Cache-Control: public, max-age=<seconds>, immutable`.
When the URL has an expiry (an `exp`, or the TTL of an encrypted URL),
@@ -180,14 +107,6 @@ or proxy cache keeps the image after pixa would refuse the URL. A URL with no
expiry gets one year. `immutable` only stops a client revalidating while its
copy is fresh.
When several requests for the same image, size, format, quality and fit miss
the cache at once, they share one upstream fetch (or one read of the cached
source) and one transcode: the first request does the work, and the others wait
for its image or its error, holding no upstream connection or processing slot
of their own. A waiting request stops waiting when its own client goes away.
The work goes on for the others even if the first request's client goes away,
until that request's `downstream_timeout` ends.
The login form (`POST /`) is limited to 5 attempts per minute per client
address, counting an IPv6 client by its /64; an attempt over the limit is
refused with 429 and a `Retry-After` header. Behind a reverse proxy the client
@@ -205,39 +124,6 @@ its own `X-Forwarded-For`, whether it connects directly or through the proxy,
because its own address is trusted too. Setting `trusted_proxies` to only the
address pixa sees for requests that come through the proxy closes this.
### Encrypted URLs
An encrypted URL is an image URL made on pixa's own web page by someone who
knows the signing key. It works for any upstream host, allowlisted or not,
without a signature, and whoever gets it can neither read the source URL from it
nor change what it asks for.
1. Open `/` in a browser over HTTPS (or over plain HTTP while `debug` is on, see
Routes) and log in with the signing key (`signing_key`). The login session
lasts 30 days, or until `/logout`.
2. On the generator page, give the source image's URL, the width and height, the
format, quality and fit, and how long the URL lasts, then submit the form
(`POST /generate`). Width and height both empty or `0` keep the original
size; if only one of them is empty or `0`, that side is scaled to keep the
image's proportions.
3. The page shows the URL, `https://<host>/v1/e/<token>/img.<format>`, and when
it expires. `<host>` is the host the page was opened on, and the URL starts
with `http` instead while `debug` is on. The name after the token is ignored
and only gives the URL a file extension, `jpg` for `orig`.
The token holds the source's host, path and query and the size, format,
quality, fit and expiry, encrypted with a key derived from `signing_key`. The
source URL's scheme is not kept: the image is fetched like any other (see
Routes), and the blocked networks still apply.
How long the URL lasts is chosen on the page, from 1 minute to 1 year, or
never. The expiry is fixed in the token when the URL is made and cannot be
changed or revoked afterwards. Until then the image is served with a `max-age`
that ends at the expiry (see Routes); after it the URL answers 410
`URL has expired`. A URL made to last forever stops working only when
`signing_key` changes: changing it makes every encrypted URL already handed out
answer 400, and ends every login session.
### Image Metadata
pixa decodes and re-encodes every image it serves, and removes all metadata from
@@ -279,8 +165,7 @@ Where:
- `query` — source query string, empty string if none
- `width` — requested width in pixels, `0` for original
- `height` — requested height in pixels, `0` for original
- `format` — output format, one of those listed under Routes, with `original`
signed as `orig` and `jpg` as `jpeg`
- `format` — output format (jpeg, png, webp, avif, gif, orig)
- `expiration` — the URL's `exp` query parameter, the Unix timestamp when
the signature expires; a request whose `exp` is not a whole number, an
empty `exp=` included, is refused with 400
@@ -336,17 +221,6 @@ startup naming it, as an unknown config key does. The one other accepted
name is `PIXA_CONFIG_PATH`, the config file's path (like `--config`). The
variables set by the file's `env:` section are checked the same way.
pixa reads at most one config file: the one given with `--config` (or `-c`),
otherwise the one `PIXA_CONFIG_PATH` names, otherwise the first of these that
pixa finds: `/etc/pixa/config.yml`, `/etc/pixa/config.yaml`,
`~/.config/pixa/config.yml`, `~/.config/pixa/config.yaml`, then `config.yml`
and `config.yaml` in the working directory. A named file that does not exist,
cannot be read or does not parse aborts startup. Of the files pixa looks for on
its own, one it finds but cannot read or parse aborts startup; one it cannot
find, for any reason, is passed over without a message, even when the file is
there in a directory pixa may not enter. With no file, pixa uses the environment
and the defaults.
| Variable | Config key | Meaning |
| ------------------------------------ | ------------------------------- | ---------------------------------------------------------------------------- |
| `PIXA_SIGNING_KEY` | `signing_key` | Required: secret for signed and encrypted URLs and login, 32+ characters |
@@ -364,7 +238,7 @@ and the defaults.
| `PIXA_UPSTREAM_FETCH_TIMEOUT` | `upstream_fetch_timeout` | Time allowed for one fetch from an upstream host; default `30s` |
| `PIXA_UPSTREAM_MAX_RESPONSE_SIZE` | `upstream_max_response_size` | Largest upstream response accepted, in bytes; default 50 MiB |
| `PIXA_DOWNSTREAM_TIMEOUT` | `downstream_timeout` | Time allowed for answering one client request; default `60s` |
| `PIXA_ACCESS_CONTROL_ALLOW_ORIGIN` | `access_control_allow_origin` | CORS origin allowed to read image responses: `*` or one origin; default `*` |
| `PIXA_ACCESS_CONTROL_ALLOW_ORIGIN` | `access_control_allow_origin` | CORS origin allowed to read responses: `*` or one origin; default `*` |
| `PIXA_METRICS_USERNAME` | `metrics.username` | Username for `/metrics`, which is served only when both are set |
| `PIXA_METRICS_PASSWORD` | `metrics.password` | Password for `/metrics`; set together with the username |
| `PIXA_SENTRY_DSN` | `sentry_dsn` | Sentry DSN for error reporting; empty disables it |
@@ -373,9 +247,8 @@ and the defaults.
Key settings in more detail:
- `access_control_allow_origin` — the origin a browser lets read the responses
of the image routes, `/v1/image/` and `/v1/e/`, sent as the CORS
`Access-Control-Allow-Origin` header; no other route sends it. `*`, the
- `access_control_allow_origin` — the origin a browser lets read pixa's
responses, sent as the CORS `Access-Control-Allow-Origin` header: `*`, the
default, is any site; otherwise one `http` or `https` origin such as
`https://example.com`, whose host is a lowercase host name (letters,
digits, hyphens and dots, with a letter in its last part) or an IP address
@@ -432,12 +305,12 @@ Key settings in more detail:
seconds for one to free up; if none does, and `downstream_timeout` has not
ended first, it is answered 503 the same way
- `maintenance_mode` — while `true`, the image routes (`/v1/image/` and
`/v1/e/`) answer every request for an image with 503, a `Retry-After` header
and a JSON error body. The health check (`/.well-known/healthcheck.json`)
still answers 200 and reports `"maintenance_mode": true`. It stays 200
because the image's Docker `HEALTHCHECK` requests it: a 503 there would make
the container unhealthy, and upaas marks a deploy failed when its container
is unhealthy. The login and URL generator pages and `/metrics` keep working
`/v1/e/`) answer every request with 503, a `Retry-After` header and a JSON
error body. The health check (`/.well-known/healthcheck.json`) still answers
200 and reports `"maintenance_mode": true`. It stays 200 because the image's
Docker `HEALTHCHECK` requests it: a 503 there would make the container
unhealthy, and upaas marks a deploy failed when its container is unhealthy.
The login and URL generator pages and `/metrics` keep working
See `config.example.yml` for all options with defaults.
@@ -448,10 +321,7 @@ See `config.example.yml` for all options with defaults.
- **Image processing**: govips (CGO wrapper for libvips)
- **Database**: SQLite via modernc.org/sqlite
- **Static assets**: embedded via `//go:embed`
- **Metrics**: Prometheus, at `/metrics`: generic HTTP request metrics
(duration, response size, requests in flight) and the Go runtime and process
metrics; requests are measured and `/metrics` is served only when
`metrics.username` and `metrics.password` are set
- **Metrics**: Prometheus
- **Logging**: stdlib slog
## Entrypoints
@@ -474,9 +344,8 @@ them. We provide:
- `script/check` — run test, lint, and fmt-check
- `script/docker` — build the Docker image tagged via `script/projectname`
- `script/docker-smoke` — build the image, start it, wait for it to be healthy
- `script/cibuild` — CI entrypoint: `docker build .` with a new
`CHECK_EPOCH` on every run, so the Dockerfile's checks run instead of
coming from the build cache, and a green run implies a green repo
- `script/cibuild` — CI entrypoint: `docker build .` (the Dockerfile
runs the checks, so a green build implies a green repo)
- `script/precommit` — pre-commit checks (`go mod tidy` guard, then
`script/check`)
- `script/install-precommit` — install the git pre-commit hook that
+2 -90
View File
@@ -29,96 +29,6 @@ P2: security: referer blacklist
# Completed Steps
- 2026-10-04 the metrics basic auth, CORS preflight, request logging and
metrics recording have tests (closes #79): `MetricsAuth` on its own answers
401 with a challenge without credentials or with a wrong username or password
and lets the configured ones through; a preflight request gets `*` for any
origin when `access_control_allow_origin` is `*` and no
`Access-Control-Allow-Origin` from another origin than the configured one; a
`POST /` carrying the signing key leaves no trace of it in the request log
line, and the login handler's own log lines leave out the submitted key; the
metrics middleware on its own records a request it served, and the router
records nothing while no metrics username is set. Not tested: that the router
puts the basic auth in front of `/metrics` and records requests when a
metrics username is set. Only one test per package can set up `/metrics`, and
in `internal/server` that is `TestMaintenanceModeKeepsOtherRoutes`, which
needs the owner's approval to change; #180 holds it. Tests only; the basic
auth library already compares the password in constant time.
- 2026-10-04 routes, encrypted URLs and config file documented (closes #75):
"Routes" in `README.md` lists every route with its method, purpose, what it
needs and the status codes it answers with, and says `q` and `fit` are part
of what is cached; "Encrypted URLs" covers logging in, making one on the
generator page, how long it lasts and the 410 once it has expired;
"Configuration" gives the order in which pixa looks for its config file;
`config.example.yml` lists `db_url` and `env` and gives every key's default;
`scripts/manual-test.sh` is left to #97.
- 2026-10-04 shutdown stops cache eviction in progress (closes #102):
`StartEviction` runs the eviction goroutine with its own context, which
`StopEviction` cancels, so a pass in progress stops at its next database
call, file, row or eviction candidate instead of running to completion, and
no pass starts after it, so a stop logs at most one warning;
`StopEviction` takes a context and, when that context ends before the
goroutine exits, stops waiting and returns its error; the handlers' stop hook
passes fx's stop context, so an eviction still running when fx's stop
deadline ends fails the stop and makes the exit code 1.
- 2026-10-04 dead code in `internal/imgcache` is gone (closes #73): `Purge`,
which only returned an error and which nothing called, is no longer part of
the `ImageCache` interface or `Service`; the `SignatureValidator`,
`Allowlist` and `Storage` interfaces, which nothing implemented or used, are
deleted. Nothing else changes.
- 2026-10-04 upstream host semaphores and variant `.meta` files no longer
outlive their use (closes #87): the fetcher counts the fetches holding or
waiting for a slot of each upstream host's semaphore and removes the host's
semaphore once none is left, so fetches from many hosts no longer leave one
semaphore each until restart; `VariantStorage.Delete` removes the variant's
`.meta` file along with it, a missing `.meta` file not being an error, and
`DeleteWithMeta`, which eviction called for that, is gone.
- 2026-10-04 `README.md` matches the code (closes #74): "Storage" names the
cache directories pixa uses (`cache/sources`, `cache/metadata`,
`cache/variants`) and how files are named in each, and the comments in
`001_schema.sql` name the same paths; the routes and the signature section
list the same output formats, `jpg` and `original` included; the TLS
sentence names `allow_http` as its exception; "Metrics" says only generic
HTTP and Go runtime metrics exist, measured and served only when the metrics
username and password are set.
- 2026-10-03 shutdown sets the exit code and waits for image processing
(closes #86): fx alone handles SIGINT and SIGTERM, and the server's own
signal handler is gone; fx's `Run` in `cmd/pixad` exits with the shutdown's
code: 0 for a signal, 1 when the HTTP server cannot listen or the app fails
to start or to stop; the server's stop hook, which fx waits for, stops the
HTTP server, waits for the images still being processed, both within 5
seconds, then flushes Sentry; images still being processed after that are
logged with their count and make the exit code 1; a Sentry DSN that cannot be
used fails startup, so the stop hooks of what had already started run,
instead of exiting the process from a goroutine; the eviction loop is left to
#102.
- 2026-10-03 every `script/cibuild` and `script/docker` run executes the checks
(closes #101): the `Dockerfile` declares `CHECK_EPOCH` above `make fmt-check`
and `make lint` in the lint stage and above `make test` in the build stage,
and each of those steps names it in its command; both scripts pass a new value
on every run, so Docker runs the checks instead of reusing cached results,
while the `script/bootstrap` steps stay cached; a plain `docker build .` still
works, leaves it empty, and reuses the check steps only for an identical build
context; the `script/cibuild` comment and `README.md` no longer say that any
successful build implies a green repo.
- 2026-09-29 share concurrent misses (closes #65): requests that miss the same
variant at once (the same cache key, so quality and fit included) share one
upstream fetch or cached source read and one transcode through
`golang.org/x/sync/singleflight`; the first request's processing ignores its
cancellation but keeps its deadline, and the others wait for its image or
error holding no upstream connection or processing slot, and stop waiting
when their own context ends; the request doing the processing waits for it
even then, up to its deadline; a request whose context has already ended
starts nothing; each request counts one miss, and the processing counts its
fetch and transcode once; a panic while processing is reported to Sentry when
`sentry_dsn` is set and becomes an error for every waiting request instead of
stopping pixad; documented in `README.md`.
- 2026-09-29 only the image routes send CORS headers (closes #98): the CORS
middleware, with the `access_control_allow_origin` origin, moved from the
router root onto a `/v1` subrouter holding `/v1/image/` and `/v1/e/`, where it
still answers a preflight `OPTIONS` request; the login and URL generator
pages, `/metrics` and the other routes send no `Access-Control-Allow-Origin`;
documented in `README.md` and `config.example.yml`.
- 2026-10-02 a plain `docker build .` stamps the tag or short commit, not
`dev` (closes #166): `.dockerignore` lets `.git` into the build context,
without `.git/config`; with no `VERSION` build argument the `Dockerfile`
@@ -464,5 +374,7 @@ P2: security: referer blacklist
- integration tests for the image proxy flow
- load tests to verify the 1k to 5k req/s target
- P2: documentation
- configuration options
- API endpoints
- deployment guide
- example nginx or caddy reverse proxy config
-5
View File
@@ -4,8 +4,6 @@ package main
import (
"fmt"
"os"
"os/signal"
"syscall"
"github.com/spf13/cobra"
"go.uber.org/fx"
@@ -47,9 +45,6 @@ func run(_ *cobra.Command, _ []string) {
_ = os.Setenv("PIXA_CONFIG_PATH", configPath)
}
// A write to a closed stdout or stderr must not end the process.
signal.Ignore(syscall.SIGPIPE)
fx.New(
fx.Provide(
config.New,
+11 -30
View File
@@ -12,38 +12,27 @@
# Durations are Go duration strings such as 30s or 2m and must be
# positive; a bare number has no unit and aborts startup. Sizes are a
# whole number of bytes.
#
# A key left out takes the default its comment gives.
# Port to listen on (default: 8080)
# Server settings
port: 8080
# Debug logging and plain-HTTP local development (default: false)
debug: false
# While true, the image routes (/v1/image/ and /v1/e/) answer every request
# for an image with 503 and a Retry-After header. The health check keeps
# answering 200 and reports maintenance_mode as true. It stays 200 because
# the image's Docker HEALTHCHECK requests it: a 503 there would make the
# container unhealthy, and upaas marks a deploy failed when its container is
# unhealthy. (default: false)
# with 503 and a Retry-After header. The health check keeps answering 200 and
# reports maintenance_mode as true. It stays 200 because the image's Docker
# HEALTHCHECK requests it: a 503 there would make the container unhealthy, and
# upaas marks a deploy failed when its container is unhealthy.
maintenance_mode: false
# Data directory for SQLite database and cache files
# (default: /var/lib/pixa)
state_dir: ./data
# SQLite database URL (default:
# file:<state_dir>/state.sqlite3?_journal_mode=WAL). An empty value aborts
# startup; leave the key out to use the default.
# db_url: "file:./data/state.sqlite3?_journal_mode=WAL"
# Image proxy settings
# HMAC signing key for URL signatures (required, at least 32 characters)
# Generate with: openssl rand -base64 32
signing_key: "CHANGE_ME_generate_with_openssl_rand_base64_32"
# Hosts that don't require signatures (default: none)
# Hosts that don't require signatures
# Use "." prefix for wildcard subdomain matching (e.g., ".example.com" matches "cdn.example.com")
allowlist_hosts:
- s3.sneak.cloud
@@ -56,7 +45,7 @@ allowlist_hosts:
# SSRF protection. These are added to the always-enforced built-in ranges
# (loopback, RFC 1918 private, link-local, CGNAT, benchmark, NAT64, and
# similar), never replacing them. Each entry must be a valid CIDR in IPv4
# or IPv6 form; an invalid entry aborts startup. (default: none)
# or IPv6 form; an invalid entry aborts startup.
# blocked_networks:
# - 100.64.0.0/10
# - 2001:db8::/32
@@ -83,7 +72,6 @@ allowlist_hosts:
# - 2001:db8::/32
# Allow HTTP upstream (only for testing, always use HTTPS in production)
# (default: false)
allow_http: false
# Maximum concurrent connections per upstream host (default: 20)
@@ -115,9 +103,8 @@ upstream_max_response_size: 52428800
# longer than upstream_fetch_timeout plus 20 seconds.
downstream_timeout: 60s
# The origin a browser lets read the responses of the image routes,
# /v1/image/ and /v1/e/, sent as the CORS Access-Control-Allow-Origin
# header; no other route sends it. "*" (the default) is any site;
# The origin a browser lets read pixa's responses, sent as the CORS
# Access-Control-Allow-Origin header: "*" (the default) is any site;
# otherwise one http or https origin such as https://example.com, whose
# host is a lowercase host name (letters, digits, hyphens and dots, with a
# letter in its last part) or an IP address (IPv6 in brackets, in its
@@ -133,16 +120,10 @@ access_control_allow_origin: "*"
# with a minimum of 500 MiB.
# cache_max_bytes: 10737418240
# Sentry DSN for error reporting (default: empty, which turns it off)
# Sentry error reporting (optional)
sentry_dsn: ""
# Username and password for /metrics, set together (default: unset). Metrics
# are measured and /metrics is served only when both are set.
# Metrics endpoint authentication (optional)
# metrics:
# username: "admin"
# password: "secret"
# Environment variables set while this file loads, as described at the top
# (default: none)
# env:
# PIXA_DEBUG: "true"
+1 -1
View File
@@ -20,7 +20,6 @@ require (
github.com/spf13/cobra v1.10.2
go.uber.org/fx v1.24.0
golang.org/x/crypto v0.41.0
golang.org/x/sync v0.19.0
modernc.org/sqlite v1.42.2
)
@@ -135,6 +134,7 @@ require (
golang.org/x/image v0.34.0 // indirect
golang.org/x/net v0.43.0 // indirect
golang.org/x/oauth2 v0.30.0 // indirect
golang.org/x/sync v0.19.0 // indirect
golang.org/x/sys v0.36.0 // indirect
golang.org/x/term v0.34.0 // indirect
golang.org/x/text v0.32.0 // indirect
+3 -4
View File
@@ -2,7 +2,7 @@
-- Creates all tables for the pixa caching image proxy
-- Source content blobs
-- Files stored at: cache/sources/<ab>/<cd>/<sha256>
-- Files stored at: cache/src-content/<ab>/<cd>/<sha256>
-- last_accessed_at is NULL until the first LRU touch; eviction falls
-- back to fetched_at for rows that have never been touched.
CREATE TABLE IF NOT EXISTS source_content (
@@ -16,7 +16,7 @@ CREATE INDEX IF NOT EXISTS idx_source_content_last_accessed
ON source_content(last_accessed_at);
-- Source URL metadata - maps URLs to content hashes
-- JSON stored at: cache/metadata/<hostname>/<path_hash>.json
-- JSON stored at: cache/src-metadata/<hostname>/<path_hash>.json
CREATE TABLE IF NOT EXISTS source_metadata (
id INTEGER PRIMARY KEY AUTOINCREMENT,
source_host TEXT NOT NULL,
@@ -56,8 +56,7 @@ CREATE INDEX IF NOT EXISTS idx_variant_content_last_accessed
ON variant_content(last_accessed_at);
-- Output/transformed content blobs
-- Not written: transformed images are stored in cache/variants and
-- tracked in variant_content above.
-- Files stored at: cache/dst-content/<ab>/<cd>/<sha256>
CREATE TABLE IF NOT EXISTS output_content (
content_hash TEXT PRIMARY KEY,
content_type TEXT NOT NULL,
@@ -1,59 +0,0 @@
package handlers
import (
"bytes"
"log/slog"
"net/http"
"net/http/httptest"
"net/url"
"strings"
"testing"
"sneak.berlin/go/pixa/internal/config"
"sneak.berlin/go/pixa/internal/session"
)
// TestLoginLogLeavesOutSubmittedKey verifies that the log lines for a
// failed and for a successful login do not contain the submitted key.
func TestLoginLogLeavesOutSubmittedKey(t *testing.T) {
t.Parallel()
const wrongKey = "wrong-signing-key-fedcba9876543210"
var buf bytes.Buffer
sessMgr, err := session.NewManager(testSigningKey)
if err != nil {
t.Fatalf("session.NewManager() error = %v", err)
}
h := &Handlers{
log: slog.New(slog.NewJSONHandler(&buf, nil)),
config: &config.Config{SigningKey: testSigningKey},
sessMgr: sessMgr,
}
submittedKeys := []string{wrongKey, testSigningKey}
for _, key := range submittedKeys {
form := url.Values{loginKeyField: {key}}
req := httptest.NewRequestWithContext(
t.Context(), http.MethodPost, "/",
strings.NewReader(form.Encode()))
req.Header.Set("Content-Type", "application/x-www-form-urlencoded")
h.handleLoginPost(httptest.NewRecorder(), req)
}
for _, msg := range []string{"failed login attempt", "successful login"} {
if !strings.Contains(buf.String(), msg) {
t.Fatalf("log missing %q; got %q", msg, buf.String())
}
}
for _, key := range submittedKeys {
if strings.Contains(buf.String(), key) {
t.Errorf("log contains submitted key %q; got %q", key, buf.String())
}
}
}
+12 -11
View File
@@ -58,28 +58,29 @@ func New(lc fx.Lifecycle, params Params) (*Handlers, error) {
}
lc.Append(fx.Hook{
//nolint:contextcheck // the eviction loop outlives OnStart; OnStop cancels it
// The eviction goroutine must outlive OnStart, so it cannot
// inherit this hook's context. It makes its own instead, which
// leaves it uncancellable: an in-flight pass runs to completion
// during OnStop regardless of the shutdown deadline. Making the
// loop cancellable changes shutdown semantics and is tracked
// separately in issue #102, rather than being folded into the
// lint-conformance change that surfaced it.
//nolint:contextcheck // see issue #102
OnStart: func(_ context.Context) error {
return s.initImageService()
},
OnStop: func(ctx context.Context) error {
if s.imgCache == nil {
return nil
OnStop: func(_ context.Context) error {
if s.imgCache != nil {
s.imgCache.StopEviction()
}
return s.imgCache.StopEviction(ctx)
return nil
},
})
return s, nil
}
// WaitForProcessing waits until no image is being processed, or until ctx
// ends, and returns how many images were still being processed then.
func (s *Handlers) WaitForProcessing(ctx context.Context) int {
return s.imgSvc.WaitForProcessing(ctx)
}
// initImageService initializes the image cache and service.
func (s *Handlers) initImageService() error {
// Create the cache. cache_max_bytes: 0 disables the disk cache
+1 -17
View File
@@ -191,23 +191,7 @@ func fetchBody(t *testing.T, f *HTTPFetcher, path string) string {
// semLen reports how many per-host semaphore slots are currently held.
func semLen(f *HTTPFetcher, host string) int {
f.hostSemMu.Lock()
defer f.hostSemMu.Unlock()
sem, ok := f.hostSems[host]
if !ok {
return 0
}
return len(sem.slots)
}
// hostSemCount reports how many hosts have a semaphore in hostSems.
func hostSemCount(f *HTTPFetcher) int {
f.hostSemMu.Lock()
defer f.hostSemMu.Unlock()
return len(f.hostSems)
return len(f.getHostSemaphore(host))
}
func TestFetchRedirectToPrivateIPBlocked(t *testing.T) {
+8 -44
View File
@@ -160,13 +160,10 @@ func DefaultConfig() *Config {
// HTTPFetcher implements Fetcher with SSRF protection and connection limits
// per host and for all hosts together.
type HTTPFetcher struct {
client *http.Client
config *Config
// hostSems holds the semaphore of each host with a fetch holding or
// waiting for one of its slots; the entry is removed when the host's
// last such fetch gives its slot back or stops waiting.
hostSems map[string]*hostSemaphore
hostSemMu sync.Mutex // protects hostSems and each entry's count
client *http.Client
config *Config
hostSems map[string]chan struct{} // per-host semaphores
hostSemMu sync.Mutex // protects hostSems map
// allHostsSemaphore has one slot per connection allowed to all hosts
// together (config.MaxConnections).
allHostsSemaphore chan struct{}
@@ -174,14 +171,6 @@ type HTTPFetcher struct {
connectionWaitTimeout time.Duration
}
// hostSemaphore is one host's connection slots
// (config.MaxConnectionsPerHost) and the number of fetches holding or
// waiting for one of them.
type hostSemaphore struct {
slots chan struct{}
count int
}
// New creates a new HTTPFetcher with SSRF protection.
func New(config *Config) *HTTPFetcher {
if config == nil {
@@ -222,7 +211,7 @@ func New(config *Config) *HTTPFetcher {
return &HTTPFetcher{
client: client,
config: config,
hostSems: make(map[string]*hostSemaphore),
hostSems: make(map[string]chan struct{}),
allHostsSemaphore: make(chan struct{}, config.MaxConnections),
connectionWaitTimeout: ConnectionWaitTimeout,
}
@@ -318,8 +307,6 @@ func (f *HTTPFetcher) acquireConnection(
select {
case hostSem <- struct{}{}:
case <-ctx.Done():
f.putHostSemaphore(host)
return nil, ctx.Err()
}
@@ -327,55 +314,32 @@ func (f *HTTPFetcher) acquireConnection(
case f.allHostsSemaphore <- struct{}{}:
case <-time.After(f.connectionWaitTimeout):
<-hostSem
f.putHostSemaphore(host)
return nil, ErrTooManyConnections
case <-ctx.Done():
<-hostSem
f.putHostSemaphore(host)
return nil, ctx.Err()
}
return func() {
<-hostSem
f.putHostSemaphore(host)
<-f.allHostsSemaphore
}, nil
}
// getHostSemaphore returns the semaphore for a host, creating it if
// necessary, and counts the caller among the fetches using it. The caller
// calls putHostSemaphore once it holds no slot and waits for none.
// getHostSemaphore returns the semaphore for a host, creating it if necessary.
func (f *HTTPFetcher) getHostSemaphore(host string) chan struct{} {
f.hostSemMu.Lock()
defer f.hostSemMu.Unlock()
sem, ok := f.hostSems[host]
if !ok {
sem = &hostSemaphore{
slots: make(chan struct{}, f.config.MaxConnectionsPerHost),
}
sem = make(chan struct{}, f.config.MaxConnectionsPerHost)
f.hostSems[host] = sem
}
sem.count++
return sem.slots
}
// putHostSemaphore stops counting the caller among the fetches using the
// host's semaphore, and removes the semaphore when no fetch uses it.
func (f *HTTPFetcher) putHostSemaphore(host string) {
f.hostSemMu.Lock()
defer f.hostSemMu.Unlock()
sem := f.hostSems[host]
sem.count--
if sem.count == 0 {
delete(f.hostSems, host)
}
return sem
}
// buildResult validates the upstream response and assembles a FetchResult
@@ -5,7 +5,6 @@ import (
"errors"
"net"
"strconv"
"sync"
"testing"
"time"
)
@@ -120,98 +119,6 @@ func TestFetchFreesHostSlotWhenContextEndsWaitingForConnection(t *testing.T) {
}
}
// TestFetchRemovesIdleHostSemaphores checks that a host's semaphore is
// removed once no fetch holds or waits for one of its slots: after 100
// concurrent fetches from 50 hosts have all finished, no semaphore is left.
func TestFetchRemovesIdleHostSemaphores(t *testing.T) {
t.Parallel()
srv := startUpstream(t)
f, _ := newServerFetcher(t, srv, nil)
ctx := testContext(t)
var wg sync.WaitGroup
for i := range 100 {
wg.Go(func() {
res, err := f.Fetch(ctx, imageURLOnPort(1+i%50))
if err != nil {
t.Errorf("Fetch() error = %v", err)
return
}
_ = res.Content.Close()
})
}
wg.Wait()
if n := hostSemCount(f); n != 0 {
t.Errorf("%d host semaphores left after every fetch finished, want 0", n)
}
}
// TestFetchRemovesHostSemaphoreWhenNoConnection checks that a fetch that
// ends without a connection leaves no semaphore behind: when it is refused
// after waiting for a connection shared by all hosts, when its context ends
// while it waits for its host's slot, and when its context ends while it
// waits for a connection shared by all hosts, long before the 10 second
// wait timeout.
func TestFetchRemovesHostSemaphoreWhenNoConnection(t *testing.T) {
t.Parallel()
srv := startUpstream(t)
cfg := DefaultConfig()
cfg.MaxConnections = 1
cfg.MaxConnectionsPerHost = 1
f, _ := newServerFetcher(t, srv, cfg)
f.connectionWaitTimeout = 100 * time.Millisecond
open, err := f.Fetch(testContext(t), imageURLOnPort(81))
if err != nil {
t.Fatalf("first Fetch() error = %v", err)
}
_, err = f.Fetch(testContext(t), imageURLOnPort(82))
if !errors.Is(err, ErrTooManyConnections) {
t.Fatalf("Fetch() from another host: error = %v, "+
"want ErrTooManyConnections", err)
}
ctx, cancel := context.WithTimeout(t.Context(), 100*time.Millisecond)
defer cancel()
_, err = f.Fetch(ctx, imageURLOnPort(81))
if !errors.Is(err, context.DeadlineExceeded) {
t.Fatalf("Fetch() from the busy host: error = %v, "+
"want context.DeadlineExceeded", err)
}
// Back to the 10 second wait, so the next fetch's context ends first.
f.connectionWaitTimeout = ConnectionWaitTimeout
ctx, cancel = context.WithTimeout(t.Context(), 100*time.Millisecond)
defer cancel()
_, err = f.Fetch(ctx, imageURLOnPort(83))
if !errors.Is(err, context.DeadlineExceeded) {
t.Fatalf("Fetch() from a host with nothing open: error = %v, "+
"want context.DeadlineExceeded", err)
}
err = open.Content.Close()
if err != nil {
t.Fatalf("close first body: %v", err)
}
if n := hostSemCount(f); n != 0 {
t.Errorf("%d host semaphores left after every fetch finished, want 0", n)
}
}
// TestFetchReleasesConnectionOnError checks that a fetch that fails after
// taking its connection gives it back: with MaxConnections at 1, the slot
// must be free after the failure and the next fetch must succeed.
-25
View File
@@ -331,31 +331,6 @@ func FormatToMIME(format Format) string {
}
}
// WaitForProcessing waits until no image is being processed, or until ctx
// ends, and returns how many images were still being processed then. It
// waits by taking each slot in processingSemaphore as it frees up until it
// holds them all, or until ctx ends, then gives back the slots it took.
func (p *ImageProcessor) WaitForProcessing(ctx context.Context) int {
taken := 0
defer func() {
for range taken {
<-p.processingSemaphore
}
}()
for taken < cap(p.processingSemaphore) {
select {
case p.processingSemaphore <- struct{}{}:
taken++
case <-ctx.Done():
return len(p.processingSemaphore) - taken
}
}
return 0
}
// acquireSlot takes a slot in processingSemaphore, waiting at most
// processingWaitTimeout for one to free up, and returns the func that gives
// it back. A free slot is taken even when ctx has ended; only the wait for
@@ -300,69 +300,3 @@ func TestProcessReleasesSlotOnError(t *testing.T) {
})
}
}
// TestWaitForProcessing holds a processing slot with a Process call that
// cannot finish reading its input. WaitForProcessing must report that image
// when its context ends first, wait for it otherwise, return 0 once it has
// finished, and give back the slots it took while waiting.
func TestWaitForProcessing(t *testing.T) {
t.Parallel()
proc := New(Params{MaxConcurrentProcessing: 2})
gate := make(chan struct{})
entered := make(chan struct{}, 1)
results := make(chan error, 1)
openGate := sync.OnceFunc(func() { close(gate) })
t.Cleanup(openGate)
processInBackground(proc, &gatedReader{
data: bytes.NewReader(createTestJPEG(t, 10, 10)), gate: gate,
entered: entered, counter: &readingCounter{},
}, results)
waitForEntries(t, entered, 1)
ctx, cancel := context.WithTimeout(t.Context(), 100*time.Millisecond)
defer cancel()
stillProcessing := proc.WaitForProcessing(ctx)
t.Logf("WaitForProcessing() after its context ended: %d", stillProcessing)
if stillProcessing != 1 {
t.Errorf("WaitForProcessing() after its context ended = %d, want 1",
stillProcessing)
}
waited := make(chan int, 1)
go func() { waited <- proc.WaitForProcessing(t.Context()) }()
select {
case got := <-waited:
t.Fatalf("WaitForProcessing() = %d while an image was being processed",
got)
case <-time.After(100 * time.Millisecond):
}
openGate()
err := <-results
if err != nil {
t.Errorf("Process() error = %v, want nil", err)
}
select {
case got := <-waited:
if got != 0 {
t.Errorf("WaitForProcessing() once processing finished = %d, want 0",
got)
}
case <-time.After(5 * time.Second):
t.Fatal("WaitForProcessing() did not return once processing finished")
}
if held := len(proc.processingSemaphore); held != 0 {
t.Errorf("%d slots still held after WaitForProcessing() returned", held)
}
}
+18 -25
View File
@@ -11,6 +11,7 @@ import (
"io"
"log/slog"
"path/filepath"
"sync"
"time"
lru "github.com/hashicorp/golang-lru/v2"
@@ -68,11 +69,11 @@ type Cache struct {
// Eviction machinery. The channels are created in NewCache so
// stores can signal write pressure without racing StartEviction.
// evictionCancel, set by StartEviction, cancels the eviction
// goroutine's context.
evictionPressure chan struct{}
evictionStop chan struct{}
evictionDone chan struct{}
evictionCancel context.CancelFunc
evictionStarted bool
evictionStopOnce sync.Once
// metaCache holds the content types of the variants most recently
// stored or served, so a hit does not read the variant's .meta file.
@@ -111,6 +112,7 @@ func NewCache(db *sql.DB, config CacheConfig) (*Cache, error) {
log: log,
disabled: config.DisableDiskCache,
evictionPressure: make(chan struct{}, 1),
evictionStop: make(chan struct{}),
evictionDone: make(chan struct{}),
metaCache: metaCache,
contentLocks: newContentLock(),
@@ -469,8 +471,7 @@ func (c *Cache) Stats(ctx context.Context) (*CacheStats, error) {
return &stats, nil
}
// IncrementStats counts a cache hit or miss, and an upstream fetch that read
// fetchBytes bytes, as IncrementUpstreamFetch does.
// IncrementStats increments cache statistics.
func (c *Cache) IncrementStats(ctx context.Context, hit bool, fetchBytes int64) {
var err error
@@ -494,26 +495,18 @@ func (c *Cache) IncrementStats(ctx context.Context, hit bool, fetchBytes int64)
c.log.Warn("failed to count cache hit or miss", "hit", hit, "error", err)
}
c.IncrementUpstreamFetch(ctx, fetchBytes)
}
// IncrementUpstreamFetch counts one upstream fetch that read fetchBytes bytes.
// A fetch that read no bytes is not counted.
func (c *Cache) IncrementUpstreamFetch(ctx context.Context, fetchBytes int64) {
if fetchBytes <= 0 {
return
}
_, err := c.db.ExecContext(ctx, `
UPDATE cache_stats
SET upstream_fetch_count = upstream_fetch_count + 1,
upstream_fetch_bytes = upstream_fetch_bytes + ?,
last_updated_at = CURRENT_TIMESTAMP
WHERE id = 1
`, fetchBytes)
if err != nil {
c.log.Warn("failed to count upstream fetch",
"fetch_bytes", fetchBytes, "error", err)
if fetchBytes > 0 {
_, err = c.db.ExecContext(ctx, `
UPDATE cache_stats
SET upstream_fetch_count = upstream_fetch_count + 1,
upstream_fetch_bytes = upstream_fetch_bytes + ?,
last_updated_at = CURRENT_TIMESTAMP
WHERE id = 1
`, fetchBytes)
if err != nil {
c.log.Warn("failed to count upstream fetch",
"fetch_bytes", fetchBytes, "error", err)
}
}
}
@@ -1,509 +0,0 @@
package imgcache
import (
"bytes"
"context"
"errors"
"image/jpeg"
"io"
"io/fs"
"strings"
"sync"
"sync/atomic"
"testing"
"time"
"github.com/getsentry/sentry-go"
"sneak.berlin/go/pixa/internal/httpfetcher"
"sneak.berlin/go/pixa/internal/magic"
)
// arrivalWait is how long a test gives requests it has started to reach the
// point where they wait for a held fetch.
const arrivalWait = 100 * time.Millisecond
// heldFetcher counts the fetches it is asked for and holds each one until
// releaseFetches is called, so that a test can have requests arrive while a
// fetch is in progress. started receives once for every fetch.
type heldFetcher struct {
upstream httpfetcher.Fetcher
fetches atomic.Int32
started chan struct{}
release chan struct{}
releaseOnce sync.Once
}
func (f *heldFetcher) Fetch(
ctx context.Context, url string,
) (*httpfetcher.FetchResult, error) {
f.fetches.Add(1)
f.started <- struct{}{}
select {
case <-f.release:
case <-ctx.Done():
return nil, ctx.Err()
}
return f.upstream.Fetch(ctx, url)
}
// releaseFetches lets every held fetch, and every later one, go on.
func (f *heldFetcher) releaseFetches() {
f.releaseOnce.Do(func() { close(f.release) })
}
// setupHeldFetchService returns a test service whose fetches go through a
// heldFetcher. Its database is limited to one connection: each connection to
// an in-memory SQLite database opens a new, empty one, so requests running at
// once must share the connection that holds the schema.
func setupHeldFetchService(t *testing.T) (*Service, *TestFixtures, *heldFetcher) {
t.Helper()
svc, fixtures := SetupTestService(t)
svc.cache.db.SetMaxOpenConns(1)
fetcher := &heldFetcher{
upstream: svc.fetcher,
started: make(chan struct{}, 100),
release: make(chan struct{}),
}
svc.fetcher = fetcher
t.Cleanup(fetcher.releaseFetches)
return svc, fixtures, fetcher
}
// photoVariant asks for the test photo, 100x100, at 50x25 as a JPEG of the
// given quality and fit mode. Each call returns a new request, as Get writes
// to the request it is given.
func photoVariant(fixtures *TestFixtures, quality int, fit FitMode) *ImageRequest {
return &ImageRequest{
SourceHost: fixtures.GoodHost,
SourcePath: testPathPhoto,
Size: Size{Width: 50, Height: 25},
Format: FormatJPEG,
Quality: quality,
FitMode: fit,
}
}
// getResult is what one Get call returned, with the image read out.
type getResult struct {
image []byte
err error
}
// startGet calls Get in a goroutine of its own and delivers what it returned
// on the channel.
func startGet(
ctx context.Context, svc *Service, req *ImageRequest,
) <-chan getResult {
results := make(chan getResult, 1)
go func() {
resp, err := svc.Get(ctx, req)
if err != nil {
results <- getResult{err: err}
return
}
defer func() { _ = resp.Content.Close() }()
image, err := io.ReadAll(resp.Content)
results <- getResult{image: image, err: err}
}()
return results
}
// jpegSize returns the width and height of the JPEG image in data, or 0 and 0
// if data is not one.
func jpegSize(data []byte) (int, int) {
config, err := jpeg.DecodeConfig(bytes.NewReader(data))
if err != nil {
return 0, 0
}
return config.Width, config.Height
}
// TestService_Get_ConcurrentMissesShareOneFetch starts several requests for
// one uncached variant while the first one's fetch is held. Between them they
// must fetch the source once and transcode it once, every one must be answered
// with the same 50x25 JPEG, and each must count one miss.
func TestService_Get_ConcurrentMissesShareOneFetch(t *testing.T) {
t.Parallel()
svc, fixtures, fetcher := setupHeldFetchService(t)
const requests = 8
pending := make([]<-chan getResult, 0, requests)
for range requests {
pending = append(pending,
startGet(t.Context(), svc, photoVariant(fixtures, 85, FitCover)))
}
<-fetcher.started
time.Sleep(arrivalWait)
fetcher.releaseFetches()
var first []byte
for i, results := range pending {
got := <-results
if got.err != nil {
t.Fatalf("request %d: Get() error = %v", i, got.err)
}
if width, height := jpegSize(got.image); width != 50 || height != 25 {
t.Errorf("request %d: image is %dx%d, want a 50x25 JPEG", i, width, height)
}
if first == nil {
first = got.image
} else if !bytes.Equal(got.image, first) {
t.Errorf("request %d: image differs from request 0's", i)
}
}
if fetches := fetcher.fetches.Load(); fetches != 1 {
t.Errorf("%d requests made %d upstream fetches, want 1", requests, fetches)
}
// NewTestFS builds the same files the test service's fetcher serves.
testFS, _ := NewTestFS(t)
photo, err := fs.ReadFile(testFS, fixtures.GoodHostJPEG)
if err != nil {
t.Fatal(err)
}
want := cacheStatsCounters{0, requests, 1, int64(len(photo)), 1}
if got := readCacheStatsCounters(t, svc.cache); got != want {
t.Errorf("counters = %+v, want %+v", got, want)
}
}
// TestService_Get_ConcurrentVariantsStayApart requests three variants of the
// test photo at once that differ only in quality or fit. Each must be made by
// a fetch and a transcode of its own, and each answer must be its own variant.
func TestService_Get_ConcurrentVariantsStayApart(t *testing.T) {
t.Parallel()
svc, fixtures, fetcher := setupHeldFetchService(t)
cover := startGet(t.Context(), svc, photoVariant(fixtures, 85, FitCover))
lowQuality := startGet(t.Context(), svc, photoVariant(fixtures, 40, FitCover))
contain := startGet(t.Context(), svc, photoVariant(fixtures, 85, FitContain))
for range 3 {
select {
case <-fetcher.started:
case <-time.After(5 * time.Second):
t.Fatal("fewer fetches started than variants requested: " +
"variants differing in quality or fit were merged")
}
}
fetcher.releaseFetches()
images := make(map[string][]byte)
for _, variant := range []struct {
name string
results <-chan getResult
width, height int
}{
{"q=85 fit=cover", cover, 50, 25},
{"q=40 fit=cover", lowQuality, 50, 25},
{"q=85 fit=contain", contain, 25, 25},
} {
got := <-variant.results
if got.err != nil {
t.Fatalf("%s: Get() error = %v", variant.name, got.err)
}
width, height := jpegSize(got.image)
if width != variant.width || height != variant.height {
t.Errorf("%s: image is %dx%d, want a %dx%d JPEG", variant.name,
width, height, variant.width, variant.height)
}
images[variant.name] = got.image
}
if bytes.Equal(images["q=85 fit=cover"], images["q=40 fit=cover"]) {
t.Error("q=40 was answered with the q=85 image")
}
}
// TestService_Get_WaiterStopsWhenItsContextEnds has a second request for a
// variant join the first one's held fetch, then ends the second request's
// context. The second request must return at once with the context's error,
// while the fetch is still held, and the first must still be answered.
func TestService_Get_WaiterStopsWhenItsContextEnds(t *testing.T) {
t.Parallel()
svc, fixtures, fetcher := setupHeldFetchService(t)
first := startGet(t.Context(), svc, photoVariant(fixtures, 85, FitCover))
<-fetcher.started
waiterCtx, cancelWaiter := context.WithCancel(t.Context())
waiter := startGet(waiterCtx, svc, photoVariant(fixtures, 85, FitCover))
time.Sleep(arrivalWait)
cancelWaiter()
select {
case got := <-waiter:
if !errors.Is(got.err, context.Canceled) {
t.Errorf("waiting request: Get() error = %v, want %v",
got.err, context.Canceled)
}
case <-time.After(time.Second):
t.Fatal("waiting request did not return when its context ended")
}
fetcher.releaseFetches()
got := <-first
if got.err != nil {
t.Fatalf("first request: Get() error = %v", got.err)
}
if width, height := jpegSize(got.image); width != 50 || height != 25 {
t.Errorf("first request: image is %dx%d, want a 50x25 JPEG", width, height)
}
if fetches := fetcher.fetches.Load(); fetches != 1 {
t.Errorf("upstream fetches = %d, want 1", fetches)
}
}
// TestService_Get_FirstRequestLeavingKeepsTheWork ends the context of the
// request whose fetch is held, after a second request has joined it. The fetch
// and transcode must go on and answer the second request. The first request
// waits for its own work, as every request did before misses were shared, and
// is answered too.
func TestService_Get_FirstRequestLeavingKeepsTheWork(t *testing.T) {
t.Parallel()
svc, fixtures, fetcher := setupHeldFetchService(t)
firstCtx, cancelFirst := context.WithCancel(t.Context())
first := startGet(firstCtx, svc, photoVariant(fixtures, 85, FitCover))
<-fetcher.started
second := startGet(t.Context(), svc, photoVariant(fixtures, 85, FitCover))
time.Sleep(arrivalWait)
cancelFirst()
fetcher.releaseFetches()
for name, results := range map[string]<-chan getResult{
"first request": first, "second request": second,
} {
got := <-results
if got.err != nil {
t.Fatalf("%s: Get() error = %v", name, got.err)
}
if width, height := jpegSize(got.image); width != 50 || height != 25 {
t.Errorf("%s: image is %dx%d, want a 50x25 JPEG", name, width, height)
}
}
if fetches := fetcher.fetches.Load(); fetches != 1 {
t.Errorf("upstream fetches = %d, want 1", fetches)
}
}
// TestService_Get_ConcurrentMissesShareAFailure has several requests for an
// image that cannot be served arrive while its fetch is held: the one fetch
// answers all of them with its error. The request after them is answered from
// the negative cache when the failure is kept there, and fetches again when it
// is not.
func TestService_Get_ConcurrentMissesShareAFailure(t *testing.T) {
t.Parallel()
tests := []struct {
name string
path string
wantErr error // what every request at once gets
wantNextErr error // what the request after them gets
wantFetches int32 // fetches once the request after them is answered
}{
{"upstream answers 404, kept in the negative cache",
"/images/missing.jpg", httpfetcher.ErrUpstreamError,
ErrNegativeCached, 1},
{"source fails the magic byte check, not kept",
"/images/text.png", magic.ErrUnknownFormat, magic.ErrUnknownFormat, 2},
}
for _, tc := range tests {
t.Run(tc.name, func(t *testing.T) {
t.Parallel()
svc, fixtures, fetcher := setupHeldFetchService(t)
request := func() *ImageRequest {
req := photoVariant(fixtures, 85, FitCover)
req.SourcePath = tc.path
return req
}
pending := make([]<-chan getResult, 0, 4)
for range 4 {
pending = append(pending, startGet(t.Context(), svc, request()))
}
<-fetcher.started
time.Sleep(arrivalWait)
fetcher.releaseFetches()
for i, results := range pending {
if got := <-results; !errors.Is(got.err, tc.wantErr) {
t.Errorf("request %d: Get() error = %v, want %v", i, got.err, tc.wantErr)
}
}
_, err := svc.Get(t.Context(), request())
if !errors.Is(err, tc.wantNextErr) {
t.Errorf("next request: Get() error = %v, want %v", err, tc.wantNextErr)
}
if fetches := fetcher.fetches.Load(); fetches != tc.wantFetches {
t.Errorf("upstream fetches = %d, want %d", fetches, tc.wantFetches)
}
})
}
}
// TestService_Get_EndedRequestFetchesNothing checks that a request whose
// context has already ended when it misses the cache starts no fetch.
func TestService_Get_EndedRequestFetchesNothing(t *testing.T) {
t.Parallel()
svc, fixtures, fetcher := setupHeldFetchService(t)
ctx, cancel := context.WithCancel(t.Context())
cancel()
_, err := svc.Get(ctx, photoVariant(fixtures, 85, FitCover))
if !errors.Is(err, context.Canceled) {
t.Errorf("Get() error = %v, want %v", err, context.Canceled)
}
if fetches := fetcher.fetches.Load(); fetches != 0 {
t.Errorf("upstream fetches = %d, want 0", fetches)
}
}
// TestService_Get_ReturnsByItsDeadline gives a request a deadline and holds
// its fetch until the fetch's context ends, as the fetcher does while it waits
// for a free connection to the host. The request must return by its deadline
// with the deadline's error.
func TestService_Get_ReturnsByItsDeadline(t *testing.T) {
t.Parallel()
svc, fixtures, _ := setupHeldFetchService(t)
const timeout = 200 * time.Millisecond
ctx, cancel := context.WithTimeout(t.Context(), timeout)
defer cancel()
results := startGet(ctx, svc, photoVariant(fixtures, 85, FitCover))
select {
case got := <-results:
if !errors.Is(got.err, context.DeadlineExceeded) {
t.Errorf("Get() error = %v, want %v", got.err, context.DeadlineExceeded)
}
case <-time.After(timeout + time.Second):
t.Fatal("request did not return by its deadline")
}
}
// panickingFetcher panics on every fetch.
type panickingFetcher struct{}
func (panickingFetcher) Fetch(
context.Context, string,
) (*httpfetcher.FetchResult, error) {
panic("upstream fetcher panicked")
}
// TestService_Get_PanicBecomesAnError checks that a panic while a variant is
// being made reaches its request as an error naming the panic, instead of
// being raised again.
func TestService_Get_PanicBecomesAnError(t *testing.T) {
t.Parallel()
svc, fixtures := SetupTestService(t)
svc.fetcher = panickingFetcher{}
var err error
func() {
defer func() {
if recovered := recover(); recovered != nil {
t.Fatalf("Get() panicked: %v", recovered)
}
}()
_, err = svc.Get(t.Context(), photoVariant(fixtures, 85, FitCover))
}()
if err == nil || !strings.Contains(err.Error(), "upstream fetcher panicked") {
t.Errorf("Get() error = %v, want one naming the panic", err)
}
}
// TestService_Get_PanicIsReportedToSentry checks that a panic while a variant
// is being made is reported through the Sentry hub on the request's context,
// where the Sentry middleware puts one when sentry_dsn is set.
func TestService_Get_PanicIsReportedToSentry(t *testing.T) {
t.Parallel()
svc, fixtures := SetupTestService(t)
svc.fetcher = panickingFetcher{}
transport := &sentry.MockTransport{}
client, err := sentry.NewClient(sentry.ClientOptions{
Dsn: "https://abc123@sentry.example.com/42",
Transport: transport,
})
if err != nil {
t.Fatal(err)
}
ctx := sentry.SetHubOnContext(t.Context(),
sentry.NewHub(client, sentry.NewScope()))
func() {
defer func() {
if recovered := recover(); recovered != nil {
t.Fatalf("Get() panicked: %v", recovered)
}
}()
_, _ = svc.Get(ctx, photoVariant(fixtures, 85, FitCover))
}()
events := transport.Events()
if len(events) != 1 || events[0].Message != "upstream fetcher panicked" {
t.Errorf("Sentry events = %+v, want one for the panic", events)
}
}
+24 -65
View File
@@ -117,10 +117,7 @@ func (c *Cache) EvictToLimit(ctx context.Context) error {
// evictBatch fetches one batch of LRU candidates across variants and
// source blobs and evicts them oldest-first until excessBytes are
// freed or the batch is exhausted. It returns the bytes freed. A
// candidate that fails once ctx is cancelled (every one started after
// that fails at its first database call) ends the batch with ctx's
// error, without a warning.
// freed or the batch is exhausted. It returns the bytes freed.
func (c *Cache) evictBatch(ctx context.Context, excessBytes int64) (int64, error) {
candidates, err := c.evictionCandidates(ctx)
if err != nil {
@@ -136,10 +133,6 @@ func (c *Cache) evictBatch(ctx context.Context, excessBytes int64) (int64, error
err := c.evictCandidate(ctx, candidate)
if err != nil {
if ctx.Err() != nil {
return freed, ctx.Err()
}
c.log.Warn("failed to evict cache entry",
"cache_key", candidate.cacheKey,
"content_hash", candidate.contentHash,
@@ -290,7 +283,7 @@ func (c *Cache) evictVariant(ctx context.Context, cacheKey VariantKey) error {
c.metaCache.Remove(cacheKey)
err = c.variants.Delete(cacheKey)
err = c.variants.DeleteWithMeta(cacheKey)
if err != nil {
return err
}
@@ -431,44 +424,37 @@ func (c *Cache) notifyWritePressure() {
// startup and again on every periodic tick thereafter, and evicts to
// the configured limit on the given periodic interval and on
// write-pressure notifications. It is a no-op on a disabled cache or
// when already started. The goroutine outlives the caller, so it runs
// with its own context, which StopEviction cancels.
// when already started.
func (c *Cache) StartEviction(interval time.Duration) {
if c.disabled || c.evictionCancel != nil {
if c.disabled || c.evictionStarted {
return
}
ctx, cancel := context.WithCancel(context.Background())
c.evictionCancel = cancel
c.evictionStarted = true
go c.evictionLoop(ctx, interval)
go c.evictionLoop(interval)
}
// StopEviction cancels the background eviction goroutine, which
// interrupts a pass in progress, and waits for it to exit or for ctx to
// end, whichever comes first. In the second case it returns an error
// wrapping ctx's error. It is safe to call when eviction was never
// started, and safe to call more than once.
func (c *Cache) StopEviction(ctx context.Context) error {
if c.evictionCancel == nil {
return nil
// StopEviction stops the background eviction goroutine and waits for
// it to exit. It is safe to call when eviction was never started, and
// safe to call more than once.
func (c *Cache) StopEviction() {
if !c.evictionStarted {
return
}
c.evictionCancel()
select {
case <-c.evictionDone:
return nil
case <-ctx.Done():
return fmt.Errorf("cache eviction still running: %w", ctx.Err())
}
c.evictionStopOnce.Do(func() {
close(c.evictionStop)
<-c.evictionDone
})
}
// evictionLoop is the body of the background eviction goroutine. It
// returns when ctx is cancelled, and starts no pass after that.
func (c *Cache) evictionLoop(ctx context.Context, interval time.Duration) {
// evictionLoop is the body of the background eviction goroutine.
func (c *Cache) evictionLoop(interval time.Duration) {
defer close(c.evictionDone)
ctx := context.Background()
c.runReconciliationPass(ctx)
c.runEvictionPass(ctx)
@@ -477,7 +463,7 @@ func (c *Cache) evictionLoop(ctx context.Context, interval time.Duration) {
for {
select {
case <-ctx.Done():
case <-c.evictionStop:
return
case <-ticker.C:
// Reconciliation walks the cache directories, so it only
@@ -498,13 +484,8 @@ func (c *Cache) evictionLoop(ctx context.Context, interval time.Duration) {
}
// runEvictionPass runs one eviction pass, logging failures instead of
// propagating them (the loop must keep running). It does nothing once
// ctx is cancelled.
// propagating them (the loop must keep running).
func (c *Cache) runEvictionPass(ctx context.Context) {
if ctx.Err() != nil {
return
}
err := c.EvictToLimit(ctx)
if err != nil {
c.log.Warn("cache eviction pass failed", "error", err)
@@ -512,13 +493,8 @@ func (c *Cache) runEvictionPass(ctx context.Context) {
}
// runReconciliationPass runs one reconciliation pass, logging failures
// instead of propagating them (the loop must keep running). It does
// nothing once ctx is cancelled.
// instead of propagating them (the loop must keep running).
func (c *Cache) runReconciliationPass(ctx context.Context) {
if ctx.Err() != nil {
return
}
err := c.reconcileAccounting(ctx)
if err != nil {
c.log.Warn("cache accounting reconciliation failed", "error", err)
@@ -535,8 +511,7 @@ func (c *Cache) runReconciliationPass(ctx context.Context) {
// know (and rows whose files are gone), and sweeps stale temp files
// left behind by crashed writes. Running it periodically, not just
// once, bounds how long such drift can accumulate unaccounted for on a
// long-running process to one eviction interval. Once ctx is cancelled,
// it stops at the next file or row and returns ctx's error.
// long-running process to one eviction interval.
func (c *Cache) reconcileAccounting(ctx context.Context) error {
if c.disabled {
return nil
@@ -571,10 +546,6 @@ func (c *Cache) reconcileVariantFiles(ctx context.Context) error {
return filepath.WalkDir(
c.variants.baseDir,
func(path string, entry fs.DirEntry, err error) error {
if ctx.Err() != nil {
return ctx.Err()
}
if err != nil || entry.IsDir() {
return err
}
@@ -663,10 +634,6 @@ func (c *Cache) reconcileVariantRows(ctx context.Context) error {
}
for _, key := range keys {
if ctx.Err() != nil {
return ctx.Err()
}
if c.variants.Exists(key) {
continue
}
@@ -732,10 +699,6 @@ func (c *Cache) reconcileSourceFiles(ctx context.Context) error {
return filepath.WalkDir(
c.srcContent.baseDir,
func(path string, entry fs.DirEntry, err error) error {
if ctx.Err() != nil {
return ctx.Err()
}
if err != nil || entry.IsDir() {
return err
}
@@ -797,10 +760,6 @@ func (c *Cache) reconcileSourceRows(ctx context.Context) error {
}
for _, hash := range hashes {
if ctx.Err() != nil {
return ctx.Err()
}
if c.srcContent.Exists(hash) {
continue
}
+4 -243
View File
@@ -4,12 +4,9 @@ import (
"bytes"
"context"
"database/sql"
"errors"
"io/fs"
"log/slog"
"os"
"path/filepath"
"strings"
"testing"
"time"
@@ -660,7 +657,7 @@ func TestEvictionRunsUnderWritePressure(t *testing.T) {
// An interval far longer than the test ensures only write
// pressure can trigger eviction here.
cache.StartEviction(time.Hour)
defer func() { _ = cache.StopEviction(t.Context()) }()
defer cache.StopEviction()
keys := []VariantKey{
testVariantKeyOne, testVariantKeyTwo, testVariantKeyThree,
@@ -693,7 +690,7 @@ func TestEvictionRunsOnPeriodicSchedule(t *testing.T) {
// write-pressure notification fires and only the periodic ticker
// can trigger eviction.
cache.StartEviction(100 * time.Millisecond)
defer func() { _ = cache.StopEviction(t.Context()) }()
defer cache.StopEviction()
keys := []VariantKey{
testVariantKeyOne, testVariantKeyTwo, testVariantKeyThree,
@@ -754,7 +751,7 @@ func TestStartEvictionReconcilesAccountingWithDisk(t *testing.T) {
}
cache.StartEviction(time.Hour)
defer func() { _ = cache.StopEviction(t.Context()) }()
defer cache.StopEviction()
deadline := time.Now().Add(5 * time.Second)
@@ -811,7 +808,7 @@ func TestPeriodicReconciliationAdoptsFileThatAppearsAfterStartup(t *testing.T) {
const interval = 100 * time.Millisecond
cache.StartEviction(interval)
defer func() { _ = cache.StopEviction(t.Context()) }()
defer cache.StopEviction()
// Let startup reconciliation run and settle on an empty cache
// before introducing the untracked file, so the adoption we assert
@@ -865,242 +862,6 @@ func TestPeriodicReconciliationAdoptsFileThatAppearsAfterStartup(t *testing.T) {
}
}
// TestStopEvictionInterruptsPassInProgress holds the test database's
// only connection, so the startup reconciliation pass waits for it, and
// checks that StopEviction stops that pass instead of waiting for the
// connection to come free, and that the stop logs one warning: the
// interrupted reconciliation's, with no eviction pass started after it.
func TestStopEvictionInterruptsPassInProgress(t *testing.T) {
t.Parallel()
cache, _ := newEvictionTestCache(t, 1<<30)
var logBuf bytes.Buffer
cache.log = slog.New(slog.NewJSONHandler(&logBuf, nil))
conn, err := cache.db.Conn(t.Context())
if err != nil {
t.Fatalf("failed to take the database connection: %v", err)
}
defer func() { _ = conn.Close() }()
cache.StartEviction(time.Hour)
// The pass is in progress once it waits for the connection.
deadline := time.Now().Add(5 * time.Second)
for cache.db.Stats().WaitCount == 0 {
if time.Now().After(deadline) {
t.Fatal("the reconciliation pass never waited for the database")
}
time.Sleep(10 * time.Millisecond)
}
ctx, cancel := context.WithTimeout(t.Context(), 5*time.Second)
defer cancel()
err = cache.StopEviction(ctx)
t.Logf("StopEviction() error = %v", err)
if err != nil {
t.Fatalf("StopEviction() error = %v, want nil: the pass waiting for "+
"the database did not stop", err)
}
t.Logf("log output: %s", logBuf.String())
warnings := strings.Count(logBuf.String(), `"level":"WARN"`)
if warnings != 1 {
t.Errorf("the stop logged %d warnings, want 1", warnings)
}
}
// TestEvictToLimitStopsAtNextCandidateOnceCancelled cancels the context
// while the oldest of three source blobs is being evicted, and checks that
// EvictToLimit then returns context.Canceled without evicting the other
// two or logging a warning for either of them.
func TestEvictToLimitStopsAtNextCandidateOnceCancelled(t *testing.T) {
t.Parallel()
cache, _ := newEvictionTestCache(t, 1)
var logBuf bytes.Buffer
cache.log = slog.New(slog.NewJSONHandler(&logBuf, nil))
hashes := []ContentHash{
storeEvictionTestSource(t, cache, "cancel.example.com", "/a.jpg",
bytes.Repeat([]byte{0x61}, 1000)),
storeEvictionTestSource(t, cache, "cancel.example.com", "/b.jpg",
bytes.Repeat([]byte{0x62}, 1000)),
storeEvictionTestSource(t, cache, "cancel.example.com", "/c.jpg",
bytes.Repeat([]byte{0x63}, 1000)),
}
base := time.Now().Add(-time.Hour)
for i, hash := range hashes {
setSourceLastAccessed(t, cache, hash, base.Add(time.Duration(i)*time.Minute))
}
ctx, cancel := context.WithCancel(t.Context())
defer cancel()
cache.evictSourceBlobTestHook = func(ContentHash) { cancel() }
err := cache.EvictToLimit(ctx)
t.Logf("EvictToLimit() error = %v", err)
t.Logf("log output: %s", logBuf.String())
if !errors.Is(err, context.Canceled) {
t.Errorf("EvictToLimit() error = %v, want context.Canceled", err)
}
if cache.srcContent.Exists(hashes[0]) {
t.Errorf("source blob %s, evicted when the context was cancelled, "+
"is still on disk", hashes[0])
}
for _, hash := range hashes[1:] {
if !cache.srcContent.Exists(hash) {
t.Errorf("source blob %s was evicted after the context was cancelled", hash)
}
}
if strings.Contains(logBuf.String(), `"level":"WARN"`) {
t.Errorf("EvictToLimit logged a warning after the context was cancelled")
}
assertNoDanglingReferences(t, cache)
}
// TestStopEvictionReturnsWhenItsContextEnds pauses an eviction pass where
// cancellation cannot reach it, after a source blob's rows are deleted and
// before its file is removed, and checks that StopEviction returns its
// context's error when that context ends instead of waiting for the pass.
// Once the pass goes on, the goroutine exits and no row points at a
// missing file.
func TestStopEvictionReturnsWhenItsContextEnds(t *testing.T) {
t.Parallel()
cache, _ := newEvictionTestCache(t, 1)
paused := make(chan struct{})
resume := make(chan struct{})
cache.evictSourceBlobTestHook = func(ContentHash) {
close(paused)
<-resume
}
hash := storeEvictionTestSource(t, cache, "stop.example.com", "/a.jpg",
bytes.Repeat([]byte{0x61}, 1000))
cache.StartEviction(time.Hour)
select {
case <-paused:
case <-time.After(5 * time.Second):
t.Fatal("the eviction pass never reached the source blob")
}
ctx, cancel := context.WithTimeout(t.Context(), 50*time.Millisecond)
defer cancel()
err := cache.StopEviction(ctx)
t.Logf("StopEviction() error = %v", err)
if !errors.Is(err, context.DeadlineExceeded) {
t.Errorf("StopEviction() error = %v, want context.DeadlineExceeded", err)
}
close(resume)
err = cache.StopEviction(t.Context())
if err != nil {
t.Fatalf("second StopEviction() error = %v, want nil", err)
}
assertNoDanglingReferences(t, cache)
if cache.srcContent.Exists(hash) {
t.Errorf("source blob %s is still on disk after its rows were deleted", hash)
}
}
// TestReconciliationWalksStopOnceCancelled checks that both directory
// walks of a reconciliation pass return the context's error once it is
// cancelled, leaving in place a stale temp file they would otherwise
// remove.
func TestReconciliationWalksStopOnceCancelled(t *testing.T) {
t.Parallel()
cache, _ := newEvictionTestCache(t, 1<<30)
ctx, cancel := context.WithCancel(t.Context())
cancel()
staleTime := time.Now().Add(-2 * staleTempFileAge)
walks := map[string]func(context.Context) error{
cache.variants.baseDir: cache.reconcileVariantFiles,
cache.srcContent.baseDir: cache.reconcileSourceFiles,
}
for dir, walk := range walks {
tempFile := filepath.Join(dir, tempFilePrefix+"stale")
err := os.WriteFile(tempFile, []byte("partial"), 0o600)
if err != nil {
t.Fatalf("failed to write temp file: %v", err)
}
err = os.Chtimes(tempFile, staleTime, staleTime)
if err != nil {
t.Fatalf("failed to backdate temp file: %v", err)
}
err = walk(ctx)
t.Logf("walk of %s: error = %v", dir, err)
if !errors.Is(err, context.Canceled) {
t.Errorf("walk of %s: error = %v, want context.Canceled", dir, err)
}
_, err = os.Stat(tempFile)
if err != nil {
t.Errorf("walk of %s went on after cancellation: %v", dir, err)
}
}
}
// TestReconciliationPassLogsNoWarningOnceCancelled checks that a
// reconciliation pass run with an already cancelled context logs no
// warning, so a periodic tick the loop takes after a stop adds no
// warning to the one from the pass the stop interrupted.
func TestReconciliationPassLogsNoWarningOnceCancelled(t *testing.T) {
t.Parallel()
cache, _ := newEvictionTestCache(t, 1<<30)
var logBuf bytes.Buffer
cache.log = slog.New(slog.NewJSONHandler(&logBuf, nil))
ctx, cancel := context.WithCancel(t.Context())
cancel()
cache.runReconciliationPass(ctx)
t.Logf("log output: %s", logBuf.String())
if strings.Contains(logBuf.String(), `"level":"WARN"`) {
t.Errorf("runReconciliationPass logged a warning with a cancelled context")
}
}
// TestEvictSourceBlobExcludesConcurrentStoreOfIdenticalContent exercises
// the exact TOCTOU window between evictSourceBlob's row-deletion
// transaction commit and its content file unlink: a concurrent
+30
View File
@@ -5,6 +5,7 @@ import (
"context"
"errors"
"io"
"net/url"
"time"
)
@@ -153,6 +154,9 @@ type ImageCache interface {
// Warm pre-fetches and caches an image without returning it
Warm(ctx context.Context, req *ImageRequest) error
// Purge removes a cached image
Purge(ctx context.Context, req *ImageRequest) error
// Stats returns cache statistics
Stats(ctx context.Context) (*CacheStats, error)
}
@@ -172,3 +176,29 @@ type CacheStats struct {
// HitRate is HitCount / (HitCount + MissCount)
HitRate float64
}
// SignatureValidator validates request signatures
type SignatureValidator interface {
// Validate checks if the signature is valid for the request
Validate(req *ImageRequest) error
// Generate creates a signature for a request
Generate(req *ImageRequest) string
}
// Allowlist checks if a URL is allowlisted (no signature required)
type Allowlist interface {
// IsAllowlisted returns true if the URL doesn't require a signature
IsAllowlisted(u *url.URL) bool
}
// Storage handles persistent storage of cached content
type Storage interface {
// Store saves content and returns its hash
Store(ctx context.Context, content io.Reader) (hash string, err error)
// Load retrieves content by hash
Load(ctx context.Context, hash string) (io.ReadCloser, error)
// Delete removes content by hash
Delete(ctx context.Context, hash string) error
// Exists checks if content exists
Exists(ctx context.Context, hash string) (bool, error)
}
+36 -134
View File
@@ -8,12 +8,9 @@ import (
"io"
"log/slog"
"net/url"
"runtime/debug"
"time"
"github.com/dustin/go-humanize"
"github.com/getsentry/sentry-go"
"golang.org/x/sync/singleflight"
"sneak.berlin/go/pixa/internal/allowlist"
"sneak.berlin/go/pixa/internal/httpfetcher"
"sneak.berlin/go/pixa/internal/imageprocessor"
@@ -32,9 +29,6 @@ type Service struct {
log *slog.Logger
allowHTTP bool
maxResponseSize int64
// variantsInProgress lets the requests that miss the same variant at the
// same time share one fetch and one transcode.
variantsInProgress singleflight.Group
}
// ServiceConfig holds configuration for the image service.
@@ -56,10 +50,11 @@ type ServiceConfig struct {
Logger *slog.Logger
}
// Static errors for service construction.
// Static errors for service construction and unimplemented operations.
var (
errCacheRequired = errors.New("cache is required")
errSigningKeyRequired = errors.New("signing key is required")
errCacheRequired = errors.New("cache is required")
errSigningKeyRequired = errors.New("signing key is required")
errPurgeNotImplemented = errors.New("purge not implemented")
)
// NewService creates a new image service.
@@ -165,12 +160,14 @@ func (s *Service) Get(ctx context.Context, req *ImageRequest) (*ImageResponse, e
}
}
// Cache miss - get the variant, processed once for all the requests that
// miss it at the same time, then count this request's miss, also when it
// failed or the request context has ended meanwhile
response, err := s.processOrWait(ctx, req)
// Cache miss - process the cached source or fetch it, then count the
// miss with the bytes it fetched from upstream, also when it failed or
// the request context has ended meanwhile
cacheKey := CacheKey(req)
s.cache.IncrementStats(context.WithoutCancel(ctx), false, 0)
response, fetchedBytes, err := s.processFromSourceOrFetch(ctx, req, cacheKey)
s.cache.IncrementStats(context.WithoutCancel(ctx), false, fetchedBytes)
if err != nil {
return nil, err
@@ -188,17 +185,16 @@ func (s *Service) Warm(ctx context.Context, req *ImageRequest) error {
return err
}
// Purge removes a cached image. Purging is not implemented yet.
func (s *Service) Purge(_ context.Context, _ *ImageRequest) error {
return errPurgeNotImplemented
}
// Stats returns cache statistics.
func (s *Service) Stats(ctx context.Context) (*CacheStats, error) {
return s.cache.Stats(ctx)
}
// WaitForProcessing waits until no image is being processed, or until ctx
// ends, and returns how many images were still being processed then.
func (s *Service) WaitForProcessing(ctx context.Context) int {
return s.processor.WaitForProcessing(ctx)
}
// ValidateRequest validates the request signature if required.
func (s *Service) ValidateRequest(req *ImageRequest) error {
// Check if host is allowed (no signature required)
@@ -245,96 +241,6 @@ func (s *Service) GenerateSignedURL(
baseURL, path, sig, exp, req.Quality, req.FitMode), nil
}
// errPanicked is returned when processing a variant panicked.
var errPanicked = errors.New("panic while processing image")
// processOrWait returns the variant req asks for. The first of the requests
// that miss a variant at the same time processes it, and singleflight hands
// its result, or its error, to the others: they fetch nothing, read no source
// and take no upstream connection or processing slot. The processing ignores
// the first request's cancellation, so the others are still served if that
// client goes away, but keeps its deadline.
func (s *Service) processOrWait(
ctx context.Context, req *ImageRequest,
) (*ImageResponse, error) {
// A request that has already ended starts no processing
if ctx.Err() != nil {
return nil, ctx.Err()
}
cacheKey := CacheKey(req)
// Closed when this request's own function runs, which singleflight does
// only when no other request is processing the variant
processing := make(chan struct{})
results := s.variantsInProgress.DoChan(string(cacheKey),
func() (_ any, err error) {
close(processing)
// singleflight would raise a panic again in a goroutine of its
// own, where no handler recovers it, and stop pixad. It is
// reported through the Sentry hub that the Sentry middleware puts
// on the request's context when sentry_dsn is set.
defer func() {
recovered := recover()
if recovered != nil {
s.log.Error("panic while processing image",
"host", req.SourceHost, "path", req.SourcePath,
"panic", recovered, "stack", string(debug.Stack()))
if hub := sentry.GetHubFromContext(ctx); hub != nil {
hub.RecoverWithContext(ctx, recovered)
}
err = fmt.Errorf("%w: %v", errPanicked, recovered)
}
}()
processingCtx := context.WithoutCancel(ctx)
if deadline, ok := ctx.Deadline(); ok {
var cancel context.CancelFunc
processingCtx, cancel = context.WithDeadline(processingCtx, deadline)
defer cancel()
}
return s.processFromSourceOrFetch(processingCtx, req, cacheKey)
})
var result singleflight.Result
select {
case result = <-results:
case <-ctx.Done():
select {
case <-processing:
// This request is processing the variant: it waits for the
// result, as every request did before misses were shared
result = <-results
default:
// Another request is processing the variant, or this request's
// function has not started yet; the processing goes on without it
return nil, ctx.Err()
}
}
if result.Err != nil {
return nil, result.Err
}
variant, _ := result.Val.(*processedVariant)
return &ImageResponse{
Content: io.NopCloser(bytes.NewReader(variant.data)),
ContentLength: int64(len(variant.data)),
ContentType: variant.contentType,
FetchedBytes: variant.fetchedBytes,
ETag: formatETag(cacheKey),
}, nil
}
// loadCachedSource opens source content from cache, without reading it, and
// returns it with its size; nil if the cached data is unavailable, empty or
// exceeds maxResponseSize.
@@ -369,12 +275,13 @@ func (s *Service) loadCachedSource(
}
// processFromSourceOrFetch processes an image, using cached source content
// if available.
// if available. It also returns the number of bytes fetched from upstream,
// as fetchAndProcess does, or 0 when the cached source was used.
func (s *Service) processFromSourceOrFetch(
ctx context.Context,
req *ImageRequest,
cacheKey VariantKey,
) (*processedVariant, error) {
) (*ImageResponse, int64, error) {
// Check if we have cached source content
contentHash, _, err := s.cache.LookupSource(ctx, req)
if err != nil {
@@ -401,17 +308,19 @@ func (s *Service) processFromSourceOrFetch(
// Process using cached source; nothing was fetched from upstream. The
// image processor reads the source only once it has a processing slot,
// so a request waiting for one holds none of it in memory.
return s.processAndStore(ctx, req, cacheKey, source, sourceSize)
resp, err := s.processAndStore(ctx, req, cacheKey, source, sourceSize)
return resp, 0, err
}
// fetchAndProcess fetches from upstream, processes, and caches the result.
// It counts the fetch with the bytes read from upstream, including when
// It also returns the number of bytes read from upstream, including when
// reading the response or a later step fails.
func (s *Service) fetchAndProcess(
ctx context.Context,
req *ImageRequest,
cacheKey VariantKey,
) (*processedVariant, error) {
) (*ImageResponse, int64, error) {
// Fetch from upstream
sourceURL := req.SourceURL()
@@ -430,7 +339,7 @@ func (s *Service) fetchAndProcess(
}
}
return nil, fmt.Errorf("upstream fetch failed: %w", err)
return nil, 0, fmt.Errorf("upstream fetch failed: %w", err)
}
// Closing the body frees the upstream connection. It is closed only
@@ -443,11 +352,8 @@ func (s *Service) fetchAndProcess(
sourceData, err := io.ReadAll(fetchResult.Content)
fetchBytes := int64(len(sourceData))
// Counted also when the request context has ended meanwhile
s.cache.IncrementUpstreamFetch(context.WithoutCancel(ctx), fetchBytes)
if err != nil {
return nil, fmt.Errorf("failed to read upstream response: %w", err)
return nil, fetchBytes, fmt.Errorf("failed to read upstream response: %w", err)
}
// Calculate download bitrate
@@ -475,7 +381,7 @@ func (s *Service) fetchAndProcess(
// Validate magic bytes match content type
err = magic.ValidateMagicBytes(sourceData, fetchResult.ContentType)
if err != nil {
return nil, fmt.Errorf("content validation failed: %w", err)
return nil, fetchBytes, fmt.Errorf("content validation failed: %w", err)
}
// Store source content
@@ -485,17 +391,11 @@ func (s *Service) fetchAndProcess(
// Continue even if caching fails
}
return s.processAndStore(
resp, err := s.processAndStore(
ctx, req, cacheKey, bytes.NewReader(sourceData), fetchBytes,
)
}
// processedVariant is a variant as processAndStore made it. Each request that
// shared its processing serves it through a reader of its own.
type processedVariant struct {
data []byte
contentType string
fetchedBytes int64
return resp, fetchBytes, err
}
// processAndStore processes the image read from source and stores the
@@ -506,7 +406,7 @@ func (s *Service) processAndStore(
cacheKey VariantKey,
source io.Reader,
fetchBytes int64,
) (*processedVariant, error) {
) (*ImageResponse, error) {
// Process the image
processStart := time.Now()
@@ -570,10 +470,12 @@ func (s *Service) processAndStore(
// Continue even if caching fails
}
return &processedVariant{
data: processedData,
contentType: processResult.ContentType,
fetchedBytes: fetchBytes,
return &ImageResponse{
Content: io.NopCloser(bytes.NewReader(processedData)),
ContentLength: outputSize,
ContentType: processResult.ContentType,
FetchedBytes: fetchBytes,
ETag: formatETag(cacheKey),
}, nil
}
+13 -3
View File
@@ -564,8 +564,7 @@ func (s *VariantStorage) Exists(key VariantKey) bool {
return err == nil
}
// Delete removes the content at the given key together with its .meta
// sidecar file. A missing file is not an error.
// Delete removes content at the given key.
func (s *VariantStorage) Delete(key VariantKey) error {
path := s.keyToPath(key)
@@ -574,7 +573,18 @@ func (s *VariantStorage) Delete(key VariantKey) error {
return fmt.Errorf("failed to delete content: %w", err)
}
metaPath := path + ".meta"
return nil
}
// DeleteWithMeta removes the content at the given key together with
// its .meta sidecar file. A missing file is not an error.
func (s *VariantStorage) DeleteWithMeta(key VariantKey) error {
err := s.Delete(key)
if err != nil {
return err
}
metaPath := s.keyToPath(key) + ".meta"
err = os.Remove(metaPath)
if err != nil && !os.IsNotExist(err) {
@@ -438,73 +438,3 @@ func TestVariantStorage_StoreLogsFailedMetaWrite(t *testing.T) {
t.Errorf("log missing %s; got %q", want, logBuf.String())
}
}
// storeTestVariant stores one variant, with its .meta file, in a new
// VariantStorage and returns the storage and the variant's key.
func storeTestVariant(t *testing.T) (*VariantStorage, VariantKey) {
t.Helper()
storage, err := NewVariantStorage(t.TempDir(), slog.New(slog.DiscardHandler))
if err != nil {
t.Fatalf("NewVariantStorage() error = %v", err)
}
key := CacheKey(&ImageRequest{SourceHost: testHostCDN, SourcePath: testPathCat})
_, err = storage.Store(key, bytes.NewReader([]byte("variant data")), "image/webp")
if err != nil {
t.Fatalf("Store() error = %v", err)
}
return storage, key
}
// TestVariantStorage_DeleteRemovesMeta verifies that Delete removes the
// variant's .meta file along with the variant file.
func TestVariantStorage_DeleteRemovesMeta(t *testing.T) {
t.Parallel()
storage, key := storeTestVariant(t)
metaPath := storage.keyToPath(key) + ".meta"
_, err := os.Stat(metaPath)
if err != nil {
t.Fatalf("Store() wrote no .meta file: %v", err)
}
err = storage.Delete(key)
if err != nil {
t.Fatalf("Delete() error = %v", err)
}
if storage.Exists(key) {
t.Error("Exists() = true after delete, want false")
}
_, err = os.Stat(metaPath)
if !os.IsNotExist(err) {
t.Errorf(".meta file left after Delete() (stat err=%v)", err)
}
}
// TestVariantStorage_DeleteWithoutMeta verifies that Delete succeeds for a
// variant whose .meta file is missing.
func TestVariantStorage_DeleteWithoutMeta(t *testing.T) {
t.Parallel()
storage, key := storeTestVariant(t)
err := os.Remove(storage.keyToPath(key) + ".meta")
if err != nil {
t.Fatalf("removing .meta file: %v", err)
}
err = storage.Delete(key)
if err != nil {
t.Fatalf("Delete() error = %v, want nil", err)
}
if storage.Exists(key) {
t.Error("Exists() = true after delete, want false")
}
}
@@ -1,16 +1,11 @@
package middleware
import (
"bytes"
"log/slog"
"net/http"
"net/http/httptest"
"net/url"
"strings"
"testing"
"github.com/prometheus/client_golang/prometheus/promhttp"
"sneak.berlin/go/pixa/internal/config"
)
@@ -61,203 +56,6 @@ func TestCORSAnswersWithConfiguredOrigin(t *testing.T) {
}
}
// TestCORSAnswersPreflightWithConfiguredOrigin checks that the CORS
// middleware answers a preflight request, which the CORS library handles
// apart from other requests, the same way: "*" lets any origin read
// responses and a single origin lets only that origin read them.
func TestCORSAnswersPreflightWithConfiguredOrigin(t *testing.T) {
t.Parallel()
const appOrigin = "https://app.example.com"
cases := []struct {
configured string
requestOrigin string
want string
}{
{"*", "https://any.example.com", "*"},
{appOrigin, appOrigin, appOrigin},
{appOrigin, "https://other.example.com", ""},
}
testHandler := http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) {
w.WriteHeader(http.StatusOK)
})
for _, tc := range cases {
mw := &Middleware{
log: slog.Default(),
config: &config.Config{AccessControlAllowOrigin: tc.configured},
}
handler := mw.CORS()(testHandler)
// An OPTIONS request naming the method it asks about is the
// preflight a browser sends before some cross-origin requests.
req := httptest.NewRequestWithContext(
t.Context(), http.MethodOptions, "/v1/image/example.com/a.jpg/1x1.png", nil)
req.Header.Set("Origin", tc.requestOrigin)
req.Header.Set("Access-Control-Request-Method", http.MethodGet)
rec := httptest.NewRecorder()
handler.ServeHTTP(rec, req)
got := rec.Header().Get("Access-Control-Allow-Origin")
if got != tc.want {
t.Errorf("configured %q, preflight from %q: "+
"Access-Control-Allow-Origin = %q, want %q",
tc.configured, tc.requestOrigin, got, tc.want)
}
}
}
// TestMetricsAuthRequiresConfiguredCredentials checks that MetricsAuth on
// its own answers 401 with a challenge to a request without credentials or
// with a wrong username or password, and lets a request with the configured
// username and password through. That the router puts it in front of
// /metrics is not tested.
func TestMetricsAuthRequiresConfiguredCredentials(t *testing.T) {
t.Parallel()
const (
username = "metricsuser"
password = "metricspass"
challenge = `Basic realm="metrics"`
)
// An empty username stands for a request sent without credentials.
cases := []struct {
name string
username string
password string
wantReached bool
}{
{"no credentials", "", "", false},
{"wrong username", "someone", password, false},
{"wrong password", username, "wrongpass", false},
{"configured credentials", username, password, true},
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
t.Parallel()
mw := &Middleware{
log: slog.Default(),
config: &config.Config{
MetricsUsername: username,
MetricsPassword: password,
},
}
reached := false
handler := mw.MetricsAuth()(http.HandlerFunc(
func(http.ResponseWriter, *http.Request) {
reached = true
}))
req := httptest.NewRequestWithContext(
t.Context(), http.MethodGet, "/metrics", nil)
if tc.username != "" {
req.SetBasicAuth(tc.username, tc.password)
}
rec := httptest.NewRecorder()
handler.ServeHTTP(rec, req)
if reached != tc.wantReached {
t.Fatalf("request reached /metrics = %v, want %v",
reached, tc.wantReached)
}
if tc.wantReached {
return
}
if rec.Code != http.StatusUnauthorized {
t.Errorf("status = %d, want %d",
rec.Code, http.StatusUnauthorized)
}
if got := rec.Header().Get("WWW-Authenticate"); got != challenge {
t.Errorf("WWW-Authenticate = %q, want %q", got, challenge)
}
})
}
}
// TestMetricsRecordsServedRequest checks that the metrics middleware
// records a request it served, so /metrics reports it. It is the only test
// in this package that sets up the metrics middleware, which registers with
// the process-wide Prometheus registry and can do so only once.
func TestMetricsRecordsServedRequest(t *testing.T) {
t.Parallel()
// The line /metrics shows once one GET /test has been served.
const want = `http_request_duration_seconds_count{` +
`code="200",handler="/test",method="GET",service=""} 1`
mw := &Middleware{log: slog.Default(), config: &config.Config{}}
handler := mw.Metrics()(http.HandlerFunc(
func(w http.ResponseWriter, _ *http.Request) {
w.WriteHeader(http.StatusOK)
}))
handler.ServeHTTP(httptest.NewRecorder(), httptest.NewRequestWithContext(
t.Context(), http.MethodGet, "/test", nil))
rec := httptest.NewRecorder()
promhttp.Handler().ServeHTTP(rec, httptest.NewRequestWithContext(
t.Context(), http.MethodGet, "/metrics", nil))
if !strings.Contains(rec.Body.String(), want) {
t.Errorf("/metrics does not report the GET /test served; "+
"want the line %q in:\n%s", want, rec.Body.String())
}
}
// TestLoggingLeavesOutSubmittedSigningKey checks that a login, a POST /
// whose form carries the signing key, leaves no trace of the key in the
// request's log line.
func TestLoggingLeavesOutSubmittedSigningKey(t *testing.T) {
t.Parallel()
const signingKey = "test-signing-key-0123456789abcdef"
var buf bytes.Buffer
mw := newTestMiddleware(t, &buf)
// The handler reads the key from the form, as the login handler does.
handler := mw.Logging()(http.HandlerFunc(
func(w http.ResponseWriter, r *http.Request) {
if got := r.FormValue("key"); got != signingKey {
t.Errorf("key in form = %q, want %q", got, signingKey)
}
w.WriteHeader(http.StatusSeeOther)
}))
form := url.Values{"key": {signingKey}}
req := httptest.NewRequestWithContext(t.Context(), http.MethodPost, "/",
strings.NewReader(form.Encode()))
req.Header.Set("Content-Type", "application/x-www-form-urlencoded")
handler.ServeHTTP(httptest.NewRecorder(), req)
if !strings.Contains(buf.String(), `"method":"POST"`) {
t.Fatalf("no log line for the request; got %q", buf.String())
}
if strings.Contains(buf.String(), signingKey) {
t.Errorf("log output contains the signing key; got %q", buf.String())
}
}
func TestSecurityHeaders(t *testing.T) {
t.Parallel()
-64
View File
@@ -1,64 +0,0 @@
package server
import (
"net/http"
"net/http/httptest"
"testing"
)
// TestCORSOnlyOnImageRoutes verifies that the image routes answer with the
// configured access_control_allow_origin, a preflight request included, and
// that the login and URL generator pages send no Access-Control-Allow-Origin,
// so no other site can read them. /metrics is left out: its middleware
// registers with the process-wide Prometheus registry, which only one test
// in this package can do.
func TestCORSOnlyOnImageRoutes(t *testing.T) {
t.Parallel()
const appOrigin = "https://app.example.com"
s := newTestServer(t)
s.config.AccessControlAllowOrigin = appOrigin
s.SetupRoutes()
requests := []struct {
method string
path string
want string
}{
{http.MethodGet, unsignedImagePath, appOrigin},
{http.MethodHead, unsignedImagePath, appOrigin},
{http.MethodOptions, unsignedImagePath, appOrigin},
{http.MethodGet, encryptedImagePath, appOrigin},
{http.MethodGet, "/", ""},
{http.MethodOptions, "/", ""},
{http.MethodPost, "/generate", ""},
{http.MethodGet, "/logout", ""},
}
for _, tc := range requests {
t.Run(tc.method+" "+tc.path, func(t *testing.T) {
t.Parallel()
req := httptest.NewRequestWithContext(
t.Context(), tc.method, tc.path, nil)
req.Header.Set("Origin", appOrigin)
// An OPTIONS request naming the method it asks about is the
// preflight a browser sends before some cross-origin requests.
if tc.method == http.MethodOptions {
req.Header.Set("Access-Control-Request-Method", http.MethodGet)
}
rec := httptest.NewRecorder()
s.ServeHTTP(rec, req)
t.Logf("status %d", rec.Code)
got := rec.Header().Get("Access-Control-Allow-Origin")
if got != tc.want {
t.Errorf("Access-Control-Allow-Origin = %q, want %q",
got, tc.want)
}
})
}
}
+6 -8
View File
@@ -5,8 +5,6 @@ import (
"fmt"
"net/http"
"time"
"go.uber.org/fx"
)
// HTTP server configuration constants.
@@ -38,19 +36,19 @@ func (s *Server) newHTTPServer() *http.Server {
}
}
// serveUntilShutdown serves on s.httpServer until it is shut down. When it
// stops for any other reason, such as its port being in use, it asks fx to
// shut down with exit code 1.
func (s *Server) serveUntilShutdown() {
s.httpServer = s.newHTTPServer()
s.SetupRoutes()
s.log.Info("http begin listen", "listenaddr", s.httpServer.Addr)
err := s.httpServer.ListenAndServe()
if err != nil && !errors.Is(err, http.ErrServerClosed) {
s.log.Error("listen error", "error", err)
err = s.shutdowner.Shutdown(fx.ExitCode(1))
if err != nil {
s.log.Error("shutdown request failed", "error", err)
if s.cancelFunc != nil {
s.cancelFunc()
}
}
}
+2 -6
View File
@@ -13,10 +13,6 @@ import (
// unsignedImagePath is an image URL that carries no signature.
const unsignedImagePath = "/v1/image/cdn.example.com/cat.jpg/100x100.jpeg"
// encryptedImagePath is an encrypted image URL whose token cannot be
// decrypted.
const encryptedImagePath = "/v1/e/token/cat.jpg"
// TestMaintenanceModeRefusesImageRequests verifies that while maintenance
// mode is on, both image routes answer 503 Service Unavailable with a
// Retry-After header and the JSON error body the image handlers send.
@@ -32,7 +28,7 @@ func TestMaintenanceModeRefusesImageRequests(t *testing.T) {
}{
{http.MethodGet, unsignedImagePath},
{http.MethodHead, unsignedImagePath},
{http.MethodGet, encryptedImagePath},
{http.MethodGet, "/v1/e/token/cat.jpg"},
}
for _, tc := range requests {
@@ -99,7 +95,7 @@ func TestImageRequestsServedWithoutMaintenanceMode(t *testing.T) {
}{
{http.MethodGet, unsignedImagePath, http.StatusUnauthorized},
{http.MethodHead, unsignedImagePath, http.StatusUnauthorized},
{http.MethodGet, encryptedImagePath, http.StatusBadRequest},
{http.MethodGet, "/v1/e/token/cat.jpg", http.StatusBadRequest},
}
for _, tc := range requests {
-32
View File
@@ -1,32 +0,0 @@
package server
import (
"net/http"
"net/http/httptest"
"strings"
"testing"
"github.com/prometheus/client_golang/prometheus/promhttp"
)
// TestNoMetricsRecordedWithoutMetricsUsername checks that with no metrics
// username set the router records nothing about the requests it serves.
// /metrics is not served then, so the process-wide Prometheus registry is
// read directly. TestMaintenanceModeKeepsOtherRoutes records into the same
// registry, but never a GET /robots.txt.
func TestNoMetricsRecordedWithoutMetricsUsername(t *testing.T) {
t.Parallel()
s := newTestServer(t)
s.ServeHTTP(httptest.NewRecorder(), httptest.NewRequestWithContext(
t.Context(), http.MethodGet, "/robots.txt", nil))
rec := httptest.NewRecorder()
promhttp.Handler().ServeHTTP(rec, httptest.NewRequestWithContext(
t.Context(), http.MethodGet, "/metrics", nil))
if strings.Contains(rec.Body.String(), `handler="/robots.txt"`) {
t.Errorf("with no metrics username GET /robots.txt was recorded:\n%s",
rec.Body.String())
}
}
+15 -23
View File
@@ -38,6 +38,7 @@ func (s *Server) SetupRoutes() {
s.router.Use(s.mw.Metrics())
}
s.router.Use(s.mw.CORS())
s.router.Use(middleware.Timeout(s.config.DownstreamTimeout))
if s.sentryEnabled {
@@ -73,31 +74,22 @@ func (s *Server) SetupRoutes() {
s.router.Get("/logout", s.h.HandleLogout())
// Image routes, the only ones that send CORS headers, as pages on other
// sites read them. They are a subrouter rather than a group: a group's
// middleware runs only for a request that matches one of its routes,
// and a browser's preflight OPTIONS request matches none, so the CORS
// middleware could not answer it.
s.router.Route("/v1", func(r chi.Router) {
r.Use(s.mw.CORS())
// Image routes, refused while maintenance mode is on. Only these: the
// image's Docker HEALTHCHECK requests the health check, a 503 there
// would make the container unhealthy, and upaas marks a deploy failed
// when its container is unhealthy.
s.router.Group(func(r chi.Router) {
r.Use(s.refuseDuringMaintenance)
// Refused while maintenance mode is on. Only these: the image's
// Docker HEALTHCHECK requests the health check, a 503 there would
// make the container unhealthy, and upaas marks a deploy failed
// when its container is unhealthy.
r.Group(func(r chi.Router) {
r.Use(s.refuseDuringMaintenance)
// Main image proxy route
// /v1/image/<host>/<path>/<width>x<height>.<format>
r.Get("/v1/image/*", s.h.HandleImage())
r.Head("/v1/image/*", s.h.HandleImage())
// Main image proxy route
// /v1/image/<host>/<path>/<width>x<height>.<format>
r.Get("/image/*", s.h.HandleImage())
r.Head("/image/*", s.h.HandleImage())
// Encrypted image URL route
// The trailing filename (e.g., /img.jpg) is ignored but helps
// browsers with content type
r.Get("/e/{token}/*", s.h.HandleImageEnc())
})
// Encrypted image URL route
// The trailing filename (e.g., /img.jpg) is ignored but helps
// browsers with content type
r.Get("/v1/e/{token}/*", s.h.HandleImageEnc())
})
// Metrics endpoint with auth
+63 -52
View File
@@ -3,10 +3,12 @@ package server
import (
"context"
"errors"
"fmt"
"log/slog"
"net/http"
"os"
"os/signal"
"syscall"
"time"
"github.com/getsentry/sentry-go"
@@ -25,10 +27,6 @@ const (
SentryFlushTimeout = 2 * time.Second
)
// errStillProcessing is returned by the server's stop hook when images are
// still being processed once ShutdownTimeout has passed.
var errStillProcessing = errors.New("images still being processed at shutdown")
// Params defines dependencies for Server.
type Params struct {
fx.In
@@ -38,7 +36,6 @@ type Params struct {
Config *config.Config
Middleware *middleware.Middleware
Handlers *handlers.Handlers
Shutdowner fx.Shutdowner
}
// Server is the main HTTP server.
@@ -48,58 +45,59 @@ type Server struct {
globals *globals.Globals
mw *middleware.Middleware
h *handlers.Handlers
shutdowner fx.Shutdowner
startupTime time.Time
exitCode int
sentryEnabled bool
cancelFunc context.CancelFunc
httpServer *http.Server
router *chi.Mux
}
// New creates a new Server instance. Its start hook starts Sentry and the
// HTTP server; its stop hook, which fx runs on SIGINT, SIGTERM or a
// shutdown request, shuts them down.
// New creates a new Server instance.
func New(lc fx.Lifecycle, params Params) (*Server, error) {
s := &Server{
log: params.Logger.Get(),
config: params.Config,
globals: params.Globals,
mw: params.Middleware,
h: params.Handlers,
shutdowner: params.Shutdowner,
log: params.Logger.Get(),
config: params.Config,
globals: params.Globals,
mw: params.Middleware,
h: params.Handlers,
}
lc.Append(fx.Hook{
OnStart: func(_ context.Context) error {
OnStart: func(ctx context.Context) error {
s.startupTime = time.Now()
go s.Run(context.WithoutCancel(ctx))
err := s.enableSentry()
if err != nil {
return err
return nil
},
OnStop: func(_ context.Context) error {
if s.cancelFunc != nil {
s.cancelFunc()
}
s.SetupRoutes()
s.httpServer = s.newHTTPServer()
go s.serveUntilShutdown()
return nil
},
OnStop: s.cleanShutdown,
})
return s, nil
}
// Run starts the server.
func (s *Server) Run(ctx context.Context) {
s.enableSentry()
s.serve(ctx)
}
// MaintenanceMode returns whether maintenance mode is enabled.
func (s *Server) MaintenanceMode() bool {
return s.config.MaintenanceMode
}
func (s *Server) enableSentry() error {
func (s *Server) enableSentry() {
s.sentryEnabled = false
if s.config.SentryDSN == "" {
return nil
return
}
err := sentry.Init(sentry.ClientOptions{
@@ -107,42 +105,55 @@ func (s *Server) enableSentry() error {
Release: fmt.Sprintf("%s-%s", s.globals.Appname, s.globals.Version),
})
if err != nil {
return fmt.Errorf("sentry init failure: %w", err)
s.log.Error("sentry init failure", "error", err)
os.Exit(1)
}
s.log.Info("sentry error reporting activated")
s.sentryEnabled = true
return nil
}
// cleanShutdown stops the HTTP server, waits for the images still being
// processed, then flushes Sentry. The first two share ShutdownTimeout. It
// returns errStillProcessing when images are still being processed after
// that, as their work is abandoned.
func (s *Server) cleanShutdown(ctx context.Context) error {
s.log.Info("shutting down")
func (s *Server) serve(ctx context.Context) int {
ctx, cancelFunc := context.WithCancel(ctx)
s.cancelFunc = cancelFunc
ctxShutdown, shutdownCancel := context.WithTimeout(ctx, ShutdownTimeout)
go func() {
c := make(chan os.Signal, 1)
signal.Ignore(syscall.SIGPIPE)
signal.Notify(c, os.Interrupt, syscall.SIGTERM)
sig := <-c
s.log.Info("signal received", "signal", sig)
if s.cancelFunc != nil {
s.cancelFunc()
}
}()
go s.serveUntilShutdown()
<-ctx.Done()
s.cleanShutdown(ctx)
return s.exitCode
}
func (s *Server) cleanShutdown(ctx context.Context) {
s.exitCode = 0
ctxShutdown, shutdownCancel := context.WithTimeout(
context.WithoutCancel(ctx), ShutdownTimeout)
defer shutdownCancel()
err := s.httpServer.Shutdown(ctxShutdown)
if err != nil {
s.log.Error("server clean shutdown failed", "error", err)
if s.httpServer != nil {
err := s.httpServer.Shutdown(ctxShutdown)
if err != nil {
s.log.Error("server clean shutdown failed", "error", err)
}
}
stillProcessing := s.h.WaitForProcessing(ctxShutdown)
if s.sentryEnabled {
sentry.Flush(SentryFlushTimeout)
}
if stillProcessing > 0 {
s.log.Error("images still being processed at shutdown",
"count", stillProcessing)
return errStillProcessing
}
return nil
}
-96
View File
@@ -1,96 +0,0 @@
package server
import (
"log/slog"
"net"
"testing"
"time"
"go.uber.org/fx"
"go.uber.org/fx/fxtest"
"sneak.berlin/go/pixa/internal/config"
"sneak.berlin/go/pixa/internal/globals"
"sneak.berlin/go/pixa/internal/logger"
)
// shutdownRecorder is an fx.Shutdowner that sends the options of each
// shutdown request on requests.
type shutdownRecorder struct {
requests chan []fx.ShutdownOption
}
func (r shutdownRecorder) Shutdown(opts ...fx.ShutdownOption) error {
r.requests <- opts
return nil
}
// TestSentryInitFailureFailsStartup checks that a Sentry DSN that cannot be
// used makes the server's start hook fail, so fx stops what has already
// started, instead of the process exiting from a goroutine.
func TestSentryInitFailureFailsStartup(t *testing.T) {
t.Parallel()
lc := fxtest.NewLifecycle(t)
log, err := logger.New(lc, logger.Params{Globals: &globals.Globals{}})
if err != nil {
t.Fatalf("logger.New() error = %v", err)
}
_, err = New(lc, Params{
Logger: log,
Globals: &globals.Globals{Appname: "pixad"},
Config: &config.Config{SentryDSN: "not-a-dsn"},
})
if err != nil {
t.Fatalf("New() error = %v", err)
}
err = lc.Start(t.Context())
t.Logf("Start() error = %v", err)
if err == nil {
t.Fatal("Start() error = nil, want the Sentry initialization error")
}
}
// TestListenErrorRequestsShutdownWithExitCode1 occupies the server's port
// and checks that the listen error asks fx to shut down with exit code 1.
func TestListenErrorRequestsShutdownWithExitCode1(t *testing.T) {
t.Parallel()
busy, err := (&net.ListenConfig{}).Listen(t.Context(), "tcp", ":0")
if err != nil {
t.Fatalf("Listen() error = %v", err)
}
t.Cleanup(func() { _ = busy.Close() })
addr, ok := busy.Addr().(*net.TCPAddr)
if !ok {
t.Fatalf("listener address %v is not a TCP address", busy.Addr())
}
requests := make(chan []fx.ShutdownOption, 1)
s := &Server{
log: slog.New(slog.DiscardHandler),
config: &config.Config{Port: addr.Port},
shutdowner: shutdownRecorder{requests: requests},
}
s.httpServer = s.newHTTPServer()
go s.serveUntilShutdown()
select {
case opts := <-requests:
t.Logf("shutdown options = %v", opts)
if len(opts) != 1 || opts[0] != fx.ExitCode(1) {
t.Errorf("shutdown options = %v, want [fx.ExitCode(1)]", opts)
}
case <-time.After(5 * time.Second):
t.Fatal("no shutdown was requested after the listen error")
}
}
+4 -7
View File
@@ -1,18 +1,15 @@
#!/bin/sh
# script/cibuild: run the CI build. The Dockerfile runs the checks
# (make fmt-check, lint, test) as build steps. This script passes a new
# CHECK_EPOCH on every run, so Docker runs those steps instead of
# reusing cached results: a successful run means the checks ran and
# passed on this tree. Generic: needs no adaptation. The Gitea workflow
# runs this on push.
# (make fmt-check, lint, test), so a successful build implies a green
# repo. Generic: needs no adaptation. The Gitea workflow runs this on
# push.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
main() {
cd "$ROOT"
epoch="$(date +%s)$$"
docker build --build-arg CHECK_EPOCH="$epoch" .
docker build .
}
main "$@"
+3 -7
View File
@@ -1,9 +1,7 @@
#!/bin/sh
# script/docker: build the Docker image tagged with the project name.
# Identical in all repos; the tag comes from script/projectname. Like
# script/cibuild, it passes a new CHECK_EPOCH, so the build runs the
# checks instead of reusing cached results. Generic: needs no
# adaptation.
# Identical in all repos; the tag comes from script/projectname.
# Generic: needs no adaptation.
set -eu
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd -P)"
@@ -11,9 +9,7 @@ ROOT="$(cd "$SCRIPT_DIR/.." && pwd -P)"
main() {
cd "$ROOT"
epoch="$(date +%s)$$"
docker build --build-arg CHECK_EPOCH="$epoch" \
-t "$("$SCRIPT_DIR/projectname")" .
docker build -t "$("$SCRIPT_DIR/projectname")" .
}
main "$@"