4 Commits
Author SHA1 Message Date
clawbot 742b6fa4c1 Test the warnings a stop logs and stopping between eviction candidates
check / check (push) Failing after 3m0s
The stop-during-reconciliation test now also expects exactly one
warning; it fails while the loop still starts an eviction pass after a
cancelled reconciliation. A new test cancels the context while the
oldest of three source blobs is evicted and expects EvictToLimit to
return context.Canceled, leave the other two on disk and log no
warning for them.

Model: opus-5-5
2026-10-04 05:46:14 +00:00
clawbot c7f2c38733 Stop cache eviction in progress at shutdown (closes #102)
StartEviction runs the eviction goroutine with its own context, which
StopEviction cancels in place of the old stop channel, so a pass in
progress stops at its next database call, file, row or eviction
candidate instead of running to completion. StopEviction takes a
context: when it ends before the goroutine exits, StopEviction stops
waiting and returns an error wrapping it. The handlers' stop hook
passes fx's stop context, so an eviction still running at fx's stop
deadline fails the stop and the exit code is 1. The contextcheck
suppression on the start hook stays, with a one-line reason: the loop
outlives OnStart.

Model: opus-5-5
2026-10-04 05:42:52 +00:00
clawbot 5b17d1f555 Remove idle host semaphores and delete .meta with its variant (closes #87)
check / check (push) Successful in 10m10s
Each upstream host's semaphore now counts the fetches holding or waiting
for one of its slots, and is removed from hostSems when the last of them
gives its slot back or stops waiting, so a long-running pixad no longer
keeps one semaphore per host it ever fetched from.

VariantStorage.Delete removes the variant's .meta file too, a missing one
not being an error; DeleteWithMeta, which eviction called for that, is
gone.

Model: opus-5-5
2026-10-04 07:11:31 +02:00
clawbot 3a274aaa44 Make the README's storage, formats, TLS and metrics match the code (closes #74)
check / check (push) Successful in 10m10s
"Storage" names the cache directories pixa uses (cache/sources,
cache/metadata, cache/variants) and how files are named in each; the
schema comments name the same paths, and the output_content comment says
the table is not written. Routes and the signature section give one list
of output formats, jpg and original included, and say how those two are
signed. The TLS sentence names allow_http as its exception. "Metrics"
says what exists: generic HTTP and Go runtime metrics at /metrics,
measured and served only when the metrics username and password are set.

Model: opus-5-5
2026-10-04 06:24:36 +02:00
12 changed files with 553 additions and 78 deletions
+21 -10
View File
@@ -71,13 +71,17 @@ prevent abuse, and allowlisted source hosts for open access.
### Storage
- **Source content**:
`<statedir>/cache/src-content/<ab>/<cd>/<sha256 of source content>`
`<state_dir>/cache/sources/<ab>/<cd>/<sha256 of source content>`
- **Source metadata**:
`<statedir>/cache/src-metadata/<hostname>/<sha256 of path>.json`
(fetch time, original headers, request, content hash)
- **Database**: `<statedir>/state.sqlite3` (SQLite)
- **Output documents**:
`<statedir>/cache/dst-content/<ab>/<cd>/<sha256 of output content>`
`<state_dir>/cache/metadata/<hostname>/<sha256 of path and query>.json`
(host, path and query, content hash, upstream status and headers, fetch time)
- **Database**: `<state_dir>/state.sqlite3` (SQLite)
- **Transformed images**:
`<state_dir>/cache/variants/<ab>/<cd>/<sha256 of host, path, query, size, format, quality and fit>`,
each with a `.meta` file beside it holding its content type
`<ab>` and `<cd>` are the first and second pairs of characters of the file's
name.
Multiple source paths may reference the same content blob; the
database tracks references rather than using filesystem refcounting.
@@ -92,12 +96,15 @@ the metadata file stored beside it.
/v1/image/<host>/<path>/<size>.<format>?sig=<signature>&exp=<expiration>
```
Images are only fetched from origins using TLS with valid certificates.
Images are only fetched from origins using TLS with valid certificates, unless
`allow_http` is set: then pixa fetches every image over plain HTTP, which is for
testing only.
A request whose query string cannot be decoded, or gives any parameter more
than once, is refused with 400.
- `<format>`: one of `orig`, `png`, `jpeg`, `webp`
- `<format>`: one of `orig` (or `original`), `jpeg` (or `jpg`), `png`, `webp`,
`avif`, `gif`
- `<size>`: `orig` or `<width>x<height>` (e.g. `800x600`)
An image is served with `Cache-Control: public, max-age=<seconds>, immutable`.
@@ -173,7 +180,8 @@ Where:
- `query` — source query string, empty string if none
- `width` — requested width in pixels, `0` for original
- `height` — requested height in pixels, `0` for original
- `format` — output format (jpeg, png, webp, avif, gif, orig)
- `format` — output format, one of those listed under Routes, with `original`
signed as `orig` and `jpg` as `jpeg`
- `expiration` — the URL's `exp` query parameter, the Unix timestamp when
the signature expires; a request whose `exp` is not a whole number, an
empty `exp=` included, is refused with 400
@@ -330,7 +338,10 @@ See `config.example.yml` for all options with defaults.
- **Image processing**: govips (CGO wrapper for libvips)
- **Database**: SQLite via modernc.org/sqlite
- **Static assets**: embedded via `//go:embed`
- **Metrics**: Prometheus
- **Metrics**: Prometheus, at `/metrics`: generic HTTP request metrics
(duration, response size, requests in flight) and the Go runtime and process
metrics; requests are measured and `/metrics` is served only when
`metrics.username` and `metrics.password` are set
- **Logging**: stdlib slog
## Entrypoints
+23
View File
@@ -29,6 +29,29 @@ P2: security: referer blacklist
# Completed Steps
- 2026-10-04 shutdown stops cache eviction in progress (closes #102):
`StartEviction` runs the eviction goroutine with its own context, which
`StopEviction` cancels, so a pass in progress stops at its next database
call, file, row or eviction candidate instead of running to completion;
`StopEviction` takes a context and, when that context ends before the
goroutine exits, stops waiting and returns its error; the handlers' stop hook
passes fx's stop context, so an eviction still running when fx's stop
deadline ends fails the stop and makes the exit code 1.
- 2026-10-04 upstream host semaphores and variant `.meta` files no longer
outlive their use (closes #87): the fetcher counts the fetches holding or
waiting for a slot of each upstream host's semaphore and removes the host's
semaphore once none is left, so fetches from many hosts no longer leave one
semaphore each until restart; `VariantStorage.Delete` removes the variant's
`.meta` file along with it, a missing `.meta` file not being an error, and
`DeleteWithMeta`, which eviction called for that, is gone.
- 2026-10-04 `README.md` matches the code (closes #74): "Storage" names the
cache directories pixa uses (`cache/sources`, `cache/metadata`,
`cache/variants`) and how files are named in each, and the comments in
`001_schema.sql` name the same paths; the routes and the signature section
list the same output formats, `jpg` and `original` included; the TLS
sentence names `allow_http` as its exception; "Metrics" says only generic
HTTP and Go runtime metrics exist, measured and served only when the metrics
username and password are set.
- 2026-10-03 shutdown sets the exit code and waits for image processing
(closes #86): fx alone handles SIGINT and SIGTERM, and the server's own
signal handler is gone; fx's `Run` in `cmd/pixad` exits with the shutdown's
+4 -3
View File
@@ -2,7 +2,7 @@
-- Creates all tables for the pixa caching image proxy
-- Source content blobs
-- Files stored at: cache/src-content/<ab>/<cd>/<sha256>
-- Files stored at: cache/sources/<ab>/<cd>/<sha256>
-- last_accessed_at is NULL until the first LRU touch; eviction falls
-- back to fetched_at for rows that have never been touched.
CREATE TABLE IF NOT EXISTS source_content (
@@ -16,7 +16,7 @@ CREATE INDEX IF NOT EXISTS idx_source_content_last_accessed
ON source_content(last_accessed_at);
-- Source URL metadata - maps URLs to content hashes
-- JSON stored at: cache/src-metadata/<hostname>/<path_hash>.json
-- JSON stored at: cache/metadata/<hostname>/<path_hash>.json
CREATE TABLE IF NOT EXISTS source_metadata (
id INTEGER PRIMARY KEY AUTOINCREMENT,
source_host TEXT NOT NULL,
@@ -56,7 +56,8 @@ CREATE INDEX IF NOT EXISTS idx_variant_content_last_accessed
ON variant_content(last_accessed_at);
-- Output/transformed content blobs
-- Files stored at: cache/dst-content/<ab>/<cd>/<sha256>
-- Not written: transformed images are stored in cache/variants and
-- tracked in variant_content above.
CREATE TABLE IF NOT EXISTS output_content (
content_hash TEXT PRIMARY KEY,
content_type TEXT NOT NULL,
+5 -12
View File
@@ -58,23 +58,16 @@ func New(lc fx.Lifecycle, params Params) (*Handlers, error) {
}
lc.Append(fx.Hook{
// The eviction goroutine must outlive OnStart, so it cannot
// inherit this hook's context. It makes its own instead, which
// leaves it uncancellable: an in-flight pass runs to completion
// during OnStop regardless of the shutdown deadline. Making the
// loop cancellable changes shutdown semantics and is tracked
// separately in issue #102, rather than being folded into the
// lint-conformance change that surfaced it.
//nolint:contextcheck // see issue #102
//nolint:contextcheck // the eviction loop outlives OnStart; OnStop cancels it
OnStart: func(_ context.Context) error {
return s.initImageService()
},
OnStop: func(_ context.Context) error {
if s.imgCache != nil {
s.imgCache.StopEviction()
OnStop: func(ctx context.Context) error {
if s.imgCache == nil {
return nil
}
return nil
return s.imgCache.StopEviction(ctx)
},
})
+17 -1
View File
@@ -191,7 +191,23 @@ func fetchBody(t *testing.T, f *HTTPFetcher, path string) string {
// semLen reports how many per-host semaphore slots are currently held.
func semLen(f *HTTPFetcher, host string) int {
return len(f.getHostSemaphore(host))
f.hostSemMu.Lock()
defer f.hostSemMu.Unlock()
sem, ok := f.hostSems[host]
if !ok {
return 0
}
return len(sem.slots)
}
// hostSemCount reports how many hosts have a semaphore in hostSems.
func hostSemCount(f *HTTPFetcher) int {
f.hostSemMu.Lock()
defer f.hostSemMu.Unlock()
return len(f.hostSems)
}
func TestFetchRedirectToPrivateIPBlocked(t *testing.T) {
+44 -8
View File
@@ -160,10 +160,13 @@ func DefaultConfig() *Config {
// HTTPFetcher implements Fetcher with SSRF protection and connection limits
// per host and for all hosts together.
type HTTPFetcher struct {
client *http.Client
config *Config
hostSems map[string]chan struct{} // per-host semaphores
hostSemMu sync.Mutex // protects hostSems map
client *http.Client
config *Config
// hostSems holds the semaphore of each host with a fetch holding or
// waiting for one of its slots; the entry is removed when the host's
// last such fetch gives its slot back or stops waiting.
hostSems map[string]*hostSemaphore
hostSemMu sync.Mutex // protects hostSems and each entry's count
// allHostsSemaphore has one slot per connection allowed to all hosts
// together (config.MaxConnections).
allHostsSemaphore chan struct{}
@@ -171,6 +174,14 @@ type HTTPFetcher struct {
connectionWaitTimeout time.Duration
}
// hostSemaphore is one host's connection slots
// (config.MaxConnectionsPerHost) and the number of fetches holding or
// waiting for one of them.
type hostSemaphore struct {
slots chan struct{}
count int
}
// New creates a new HTTPFetcher with SSRF protection.
func New(config *Config) *HTTPFetcher {
if config == nil {
@@ -211,7 +222,7 @@ func New(config *Config) *HTTPFetcher {
return &HTTPFetcher{
client: client,
config: config,
hostSems: make(map[string]chan struct{}),
hostSems: make(map[string]*hostSemaphore),
allHostsSemaphore: make(chan struct{}, config.MaxConnections),
connectionWaitTimeout: ConnectionWaitTimeout,
}
@@ -307,6 +318,8 @@ func (f *HTTPFetcher) acquireConnection(
select {
case hostSem <- struct{}{}:
case <-ctx.Done():
f.putHostSemaphore(host)
return nil, ctx.Err()
}
@@ -314,32 +327,55 @@ func (f *HTTPFetcher) acquireConnection(
case f.allHostsSemaphore <- struct{}{}:
case <-time.After(f.connectionWaitTimeout):
<-hostSem
f.putHostSemaphore(host)
return nil, ErrTooManyConnections
case <-ctx.Done():
<-hostSem
f.putHostSemaphore(host)
return nil, ctx.Err()
}
return func() {
<-hostSem
f.putHostSemaphore(host)
<-f.allHostsSemaphore
}, nil
}
// getHostSemaphore returns the semaphore for a host, creating it if necessary.
// getHostSemaphore returns the semaphore for a host, creating it if
// necessary, and counts the caller among the fetches using it. The caller
// calls putHostSemaphore once it holds no slot and waits for none.
func (f *HTTPFetcher) getHostSemaphore(host string) chan struct{} {
f.hostSemMu.Lock()
defer f.hostSemMu.Unlock()
sem, ok := f.hostSems[host]
if !ok {
sem = make(chan struct{}, f.config.MaxConnectionsPerHost)
sem = &hostSemaphore{
slots: make(chan struct{}, f.config.MaxConnectionsPerHost),
}
f.hostSems[host] = sem
}
return sem
sem.count++
return sem.slots
}
// putHostSemaphore stops counting the caller among the fetches using the
// host's semaphore, and removes the semaphore when no fetch uses it.
func (f *HTTPFetcher) putHostSemaphore(host string) {
f.hostSemMu.Lock()
defer f.hostSemMu.Unlock()
sem := f.hostSems[host]
sem.count--
if sem.count == 0 {
delete(f.hostSems, host)
}
}
// buildResult validates the upstream response and assembles a FetchResult
@@ -5,6 +5,7 @@ import (
"errors"
"net"
"strconv"
"sync"
"testing"
"time"
)
@@ -119,6 +120,98 @@ func TestFetchFreesHostSlotWhenContextEndsWaitingForConnection(t *testing.T) {
}
}
// TestFetchRemovesIdleHostSemaphores checks that a host's semaphore is
// removed once no fetch holds or waits for one of its slots: after 100
// concurrent fetches from 50 hosts have all finished, no semaphore is left.
func TestFetchRemovesIdleHostSemaphores(t *testing.T) {
t.Parallel()
srv := startUpstream(t)
f, _ := newServerFetcher(t, srv, nil)
ctx := testContext(t)
var wg sync.WaitGroup
for i := range 100 {
wg.Go(func() {
res, err := f.Fetch(ctx, imageURLOnPort(1+i%50))
if err != nil {
t.Errorf("Fetch() error = %v", err)
return
}
_ = res.Content.Close()
})
}
wg.Wait()
if n := hostSemCount(f); n != 0 {
t.Errorf("%d host semaphores left after every fetch finished, want 0", n)
}
}
// TestFetchRemovesHostSemaphoreWhenNoConnection checks that a fetch that
// ends without a connection leaves no semaphore behind: when it is refused
// after waiting for a connection shared by all hosts, when its context ends
// while it waits for its host's slot, and when its context ends while it
// waits for a connection shared by all hosts, long before the 10 second
// wait timeout.
func TestFetchRemovesHostSemaphoreWhenNoConnection(t *testing.T) {
t.Parallel()
srv := startUpstream(t)
cfg := DefaultConfig()
cfg.MaxConnections = 1
cfg.MaxConnectionsPerHost = 1
f, _ := newServerFetcher(t, srv, cfg)
f.connectionWaitTimeout = 100 * time.Millisecond
open, err := f.Fetch(testContext(t), imageURLOnPort(81))
if err != nil {
t.Fatalf("first Fetch() error = %v", err)
}
_, err = f.Fetch(testContext(t), imageURLOnPort(82))
if !errors.Is(err, ErrTooManyConnections) {
t.Fatalf("Fetch() from another host: error = %v, "+
"want ErrTooManyConnections", err)
}
ctx, cancel := context.WithTimeout(t.Context(), 100*time.Millisecond)
defer cancel()
_, err = f.Fetch(ctx, imageURLOnPort(81))
if !errors.Is(err, context.DeadlineExceeded) {
t.Fatalf("Fetch() from the busy host: error = %v, "+
"want context.DeadlineExceeded", err)
}
// Back to the 10 second wait, so the next fetch's context ends first.
f.connectionWaitTimeout = ConnectionWaitTimeout
ctx, cancel = context.WithTimeout(t.Context(), 100*time.Millisecond)
defer cancel()
_, err = f.Fetch(ctx, imageURLOnPort(83))
if !errors.Is(err, context.DeadlineExceeded) {
t.Fatalf("Fetch() from a host with nothing open: error = %v, "+
"want context.DeadlineExceeded", err)
}
err = open.Content.Close()
if err != nil {
t.Fatalf("close first body: %v", err)
}
if n := hostSemCount(f); n != 0 {
t.Errorf("%d host semaphores left after every fetch finished, want 0", n)
}
}
// TestFetchReleasesConnectionOnError checks that a fetch that fails after
// taking its connection gives it back: with MaxConnections at 1, the slot
// must be free after the failure and the next fetch must succeed.
+3 -5
View File
@@ -11,7 +11,6 @@ import (
"io"
"log/slog"
"path/filepath"
"sync"
"time"
lru "github.com/hashicorp/golang-lru/v2"
@@ -69,11 +68,11 @@ type Cache struct {
// Eviction machinery. The channels are created in NewCache so
// stores can signal write pressure without racing StartEviction.
// evictionCancel, set by StartEviction, cancels the eviction
// goroutine's context.
evictionPressure chan struct{}
evictionStop chan struct{}
evictionDone chan struct{}
evictionStarted bool
evictionStopOnce sync.Once
evictionCancel context.CancelFunc
// metaCache holds the content types of the variants most recently
// stored or served, so a hit does not read the variant's .meta file.
@@ -112,7 +111,6 @@ func NewCache(db *sql.DB, config CacheConfig) (*Cache, error) {
log: log,
disabled: config.DisableDiskCache,
evictionPressure: make(chan struct{}, 1),
evictionStop: make(chan struct{}),
evictionDone: make(chan struct{}),
metaCache: metaCache,
contentLocks: newContentLock(),
+51 -22
View File
@@ -117,7 +117,8 @@ func (c *Cache) EvictToLimit(ctx context.Context) error {
// evictBatch fetches one batch of LRU candidates across variants and
// source blobs and evicts them oldest-first until excessBytes are
// freed or the batch is exhausted. It returns the bytes freed.
// freed or the batch is exhausted. It returns the bytes freed. Once ctx
// is cancelled, it stops at the next candidate and returns ctx's error.
func (c *Cache) evictBatch(ctx context.Context, excessBytes int64) (int64, error) {
candidates, err := c.evictionCandidates(ctx)
if err != nil {
@@ -131,6 +132,10 @@ func (c *Cache) evictBatch(ctx context.Context, excessBytes int64) (int64, error
break
}
if ctx.Err() != nil {
return freed, ctx.Err()
}
err := c.evictCandidate(ctx, candidate)
if err != nil {
c.log.Warn("failed to evict cache entry",
@@ -283,7 +288,7 @@ func (c *Cache) evictVariant(ctx context.Context, cacheKey VariantKey) error {
c.metaCache.Remove(cacheKey)
err = c.variants.DeleteWithMeta(cacheKey)
err = c.variants.Delete(cacheKey)
if err != nil {
return err
}
@@ -424,37 +429,44 @@ func (c *Cache) notifyWritePressure() {
// startup and again on every periodic tick thereafter, and evicts to
// the configured limit on the given periodic interval and on
// write-pressure notifications. It is a no-op on a disabled cache or
// when already started.
// when already started. The goroutine outlives the caller, so it runs
// with its own context, which StopEviction cancels.
func (c *Cache) StartEviction(interval time.Duration) {
if c.disabled || c.evictionStarted {
if c.disabled || c.evictionCancel != nil {
return
}
c.evictionStarted = true
ctx, cancel := context.WithCancel(context.Background())
c.evictionCancel = cancel
go c.evictionLoop(interval)
go c.evictionLoop(ctx, interval)
}
// StopEviction stops the background eviction goroutine and waits for
// it to exit. It is safe to call when eviction was never started, and
// safe to call more than once.
func (c *Cache) StopEviction() {
if !c.evictionStarted {
return
// StopEviction cancels the background eviction goroutine, which
// interrupts a pass in progress, and waits for it to exit or for ctx to
// end, whichever comes first. In the second case it returns an error
// wrapping ctx's error. It is safe to call when eviction was never
// started, and safe to call more than once.
func (c *Cache) StopEviction(ctx context.Context) error {
if c.evictionCancel == nil {
return nil
}
c.evictionStopOnce.Do(func() {
close(c.evictionStop)
<-c.evictionDone
})
c.evictionCancel()
select {
case <-c.evictionDone:
return nil
case <-ctx.Done():
return fmt.Errorf("cache eviction still running: %w", ctx.Err())
}
}
// evictionLoop is the body of the background eviction goroutine.
func (c *Cache) evictionLoop(interval time.Duration) {
// evictionLoop is the body of the background eviction goroutine. It
// returns when ctx is cancelled.
func (c *Cache) evictionLoop(ctx context.Context, interval time.Duration) {
defer close(c.evictionDone)
ctx := context.Background()
c.runReconciliationPass(ctx)
c.runEvictionPass(ctx)
@@ -463,7 +475,7 @@ func (c *Cache) evictionLoop(interval time.Duration) {
for {
select {
case <-c.evictionStop:
case <-ctx.Done():
return
case <-ticker.C:
// Reconciliation walks the cache directories, so it only
@@ -511,7 +523,8 @@ func (c *Cache) runReconciliationPass(ctx context.Context) {
// know (and rows whose files are gone), and sweeps stale temp files
// left behind by crashed writes. Running it periodically, not just
// once, bounds how long such drift can accumulate unaccounted for on a
// long-running process to one eviction interval.
// long-running process to one eviction interval. Once ctx is cancelled,
// it stops at the next file or row and returns ctx's error.
func (c *Cache) reconcileAccounting(ctx context.Context) error {
if c.disabled {
return nil
@@ -546,6 +559,10 @@ func (c *Cache) reconcileVariantFiles(ctx context.Context) error {
return filepath.WalkDir(
c.variants.baseDir,
func(path string, entry fs.DirEntry, err error) error {
if ctx.Err() != nil {
return ctx.Err()
}
if err != nil || entry.IsDir() {
return err
}
@@ -634,6 +651,10 @@ func (c *Cache) reconcileVariantRows(ctx context.Context) error {
}
for _, key := range keys {
if ctx.Err() != nil {
return ctx.Err()
}
if c.variants.Exists(key) {
continue
}
@@ -699,6 +720,10 @@ func (c *Cache) reconcileSourceFiles(ctx context.Context) error {
return filepath.WalkDir(
c.srcContent.baseDir,
func(path string, entry fs.DirEntry, err error) error {
if ctx.Err() != nil {
return ctx.Err()
}
if err != nil || entry.IsDir() {
return err
}
@@ -760,6 +785,10 @@ func (c *Cache) reconcileSourceRows(ctx context.Context) error {
}
for _, hash := range hashes {
if ctx.Err() != nil {
return ctx.Err()
}
if c.srcContent.Exists(hash) {
continue
}
+219 -4
View File
@@ -4,9 +4,12 @@ import (
"bytes"
"context"
"database/sql"
"errors"
"io/fs"
"log/slog"
"os"
"path/filepath"
"strings"
"testing"
"time"
@@ -657,7 +660,7 @@ func TestEvictionRunsUnderWritePressure(t *testing.T) {
// An interval far longer than the test ensures only write
// pressure can trigger eviction here.
cache.StartEviction(time.Hour)
defer cache.StopEviction()
defer func() { _ = cache.StopEviction(t.Context()) }()
keys := []VariantKey{
testVariantKeyOne, testVariantKeyTwo, testVariantKeyThree,
@@ -690,7 +693,7 @@ func TestEvictionRunsOnPeriodicSchedule(t *testing.T) {
// write-pressure notification fires and only the periodic ticker
// can trigger eviction.
cache.StartEviction(100 * time.Millisecond)
defer cache.StopEviction()
defer func() { _ = cache.StopEviction(t.Context()) }()
keys := []VariantKey{
testVariantKeyOne, testVariantKeyTwo, testVariantKeyThree,
@@ -751,7 +754,7 @@ func TestStartEvictionReconcilesAccountingWithDisk(t *testing.T) {
}
cache.StartEviction(time.Hour)
defer cache.StopEviction()
defer func() { _ = cache.StopEviction(t.Context()) }()
deadline := time.Now().Add(5 * time.Second)
@@ -808,7 +811,7 @@ func TestPeriodicReconciliationAdoptsFileThatAppearsAfterStartup(t *testing.T) {
const interval = 100 * time.Millisecond
cache.StartEviction(interval)
defer cache.StopEviction()
defer func() { _ = cache.StopEviction(t.Context()) }()
// Let startup reconciliation run and settle on an empty cache
// before introducing the untracked file, so the adoption we assert
@@ -862,6 +865,218 @@ func TestPeriodicReconciliationAdoptsFileThatAppearsAfterStartup(t *testing.T) {
}
}
// TestStopEvictionInterruptsPassInProgress holds the test database's
// only connection, so the startup reconciliation pass waits for it, and
// checks that StopEviction stops that pass instead of waiting for the
// connection to come free, and that the stop logs one warning: the
// interrupted reconciliation's, with no eviction pass started after it.
func TestStopEvictionInterruptsPassInProgress(t *testing.T) {
t.Parallel()
cache, _ := newEvictionTestCache(t, 1<<30)
var logBuf bytes.Buffer
cache.log = slog.New(slog.NewJSONHandler(&logBuf, nil))
conn, err := cache.db.Conn(t.Context())
if err != nil {
t.Fatalf("failed to take the database connection: %v", err)
}
defer func() { _ = conn.Close() }()
cache.StartEviction(time.Hour)
// The pass is in progress once it waits for the connection.
deadline := time.Now().Add(5 * time.Second)
for cache.db.Stats().WaitCount == 0 {
if time.Now().After(deadline) {
t.Fatal("the reconciliation pass never waited for the database")
}
time.Sleep(10 * time.Millisecond)
}
ctx, cancel := context.WithTimeout(t.Context(), 5*time.Second)
defer cancel()
err = cache.StopEviction(ctx)
t.Logf("StopEviction() error = %v", err)
if err != nil {
t.Fatalf("StopEviction() error = %v, want nil: the pass waiting for "+
"the database did not stop", err)
}
t.Logf("log output: %s", logBuf.String())
warnings := strings.Count(logBuf.String(), `"level":"WARN"`)
if warnings != 1 {
t.Errorf("the stop logged %d warnings, want 1", warnings)
}
}
// TestEvictToLimitStopsAtNextCandidateOnceCancelled cancels the context
// while the oldest of three source blobs is being evicted, and checks that
// EvictToLimit then returns context.Canceled without evicting the other
// two or logging a warning for either of them.
func TestEvictToLimitStopsAtNextCandidateOnceCancelled(t *testing.T) {
t.Parallel()
cache, _ := newEvictionTestCache(t, 1)
var logBuf bytes.Buffer
cache.log = slog.New(slog.NewJSONHandler(&logBuf, nil))
hashes := []ContentHash{
storeEvictionTestSource(t, cache, "cancel.example.com", "/a.jpg",
bytes.Repeat([]byte{0x61}, 1000)),
storeEvictionTestSource(t, cache, "cancel.example.com", "/b.jpg",
bytes.Repeat([]byte{0x62}, 1000)),
storeEvictionTestSource(t, cache, "cancel.example.com", "/c.jpg",
bytes.Repeat([]byte{0x63}, 1000)),
}
base := time.Now().Add(-time.Hour)
for i, hash := range hashes {
setSourceLastAccessed(t, cache, hash, base.Add(time.Duration(i)*time.Minute))
}
ctx, cancel := context.WithCancel(t.Context())
defer cancel()
cache.evictSourceBlobTestHook = func(ContentHash) { cancel() }
err := cache.EvictToLimit(ctx)
t.Logf("EvictToLimit() error = %v", err)
t.Logf("log output: %s", logBuf.String())
if !errors.Is(err, context.Canceled) {
t.Errorf("EvictToLimit() error = %v, want context.Canceled", err)
}
if cache.srcContent.Exists(hashes[0]) {
t.Errorf("source blob %s, evicted when the context was cancelled, "+
"is still on disk", hashes[0])
}
for _, hash := range hashes[1:] {
if !cache.srcContent.Exists(hash) {
t.Errorf("source blob %s was evicted after the context was cancelled", hash)
}
}
if strings.Contains(logBuf.String(), `"level":"WARN"`) {
t.Errorf("EvictToLimit logged a warning after the context was cancelled")
}
assertNoDanglingReferences(t, cache)
}
// TestStopEvictionReturnsWhenItsContextEnds pauses an eviction pass where
// cancellation cannot reach it, after a source blob's rows are deleted and
// before its file is removed, and checks that StopEviction returns its
// context's error when that context ends instead of waiting for the pass.
// Once the pass goes on, the goroutine exits and no row points at a
// missing file.
func TestStopEvictionReturnsWhenItsContextEnds(t *testing.T) {
t.Parallel()
cache, _ := newEvictionTestCache(t, 1)
paused := make(chan struct{})
resume := make(chan struct{})
cache.evictSourceBlobTestHook = func(ContentHash) {
close(paused)
<-resume
}
hash := storeEvictionTestSource(t, cache, "stop.example.com", "/a.jpg",
bytes.Repeat([]byte{0x61}, 1000))
cache.StartEviction(time.Hour)
select {
case <-paused:
case <-time.After(5 * time.Second):
t.Fatal("the eviction pass never reached the source blob")
}
ctx, cancel := context.WithTimeout(t.Context(), 50*time.Millisecond)
defer cancel()
err := cache.StopEviction(ctx)
t.Logf("StopEviction() error = %v", err)
if !errors.Is(err, context.DeadlineExceeded) {
t.Errorf("StopEviction() error = %v, want context.DeadlineExceeded", err)
}
close(resume)
err = cache.StopEviction(t.Context())
if err != nil {
t.Fatalf("second StopEviction() error = %v, want nil", err)
}
assertNoDanglingReferences(t, cache)
if cache.srcContent.Exists(hash) {
t.Errorf("source blob %s is still on disk after its rows were deleted", hash)
}
}
// TestReconciliationWalksStopOnceCancelled checks that both directory
// walks of a reconciliation pass return the context's error once it is
// cancelled, leaving in place a stale temp file they would otherwise
// remove.
func TestReconciliationWalksStopOnceCancelled(t *testing.T) {
t.Parallel()
cache, _ := newEvictionTestCache(t, 1<<30)
ctx, cancel := context.WithCancel(t.Context())
cancel()
staleTime := time.Now().Add(-2 * staleTempFileAge)
walks := map[string]func(context.Context) error{
cache.variants.baseDir: cache.reconcileVariantFiles,
cache.srcContent.baseDir: cache.reconcileSourceFiles,
}
for dir, walk := range walks {
tempFile := filepath.Join(dir, tempFilePrefix+"stale")
err := os.WriteFile(tempFile, []byte("partial"), 0o600)
if err != nil {
t.Fatalf("failed to write temp file: %v", err)
}
err = os.Chtimes(tempFile, staleTime, staleTime)
if err != nil {
t.Fatalf("failed to backdate temp file: %v", err)
}
err = walk(ctx)
t.Logf("walk of %s: error = %v", dir, err)
if !errors.Is(err, context.Canceled) {
t.Errorf("walk of %s: error = %v, want context.Canceled", dir, err)
}
_, err = os.Stat(tempFile)
if err != nil {
t.Errorf("walk of %s went on after cancellation: %v", dir, err)
}
}
}
// TestEvictSourceBlobExcludesConcurrentStoreOfIdenticalContent exercises
// the exact TOCTOU window between evictSourceBlob's row-deletion
// transaction commit and its content file unlink: a concurrent
+3 -13
View File
@@ -564,7 +564,8 @@ func (s *VariantStorage) Exists(key VariantKey) bool {
return err == nil
}
// Delete removes content at the given key.
// Delete removes the content at the given key together with its .meta
// sidecar file. A missing file is not an error.
func (s *VariantStorage) Delete(key VariantKey) error {
path := s.keyToPath(key)
@@ -573,18 +574,7 @@ func (s *VariantStorage) Delete(key VariantKey) error {
return fmt.Errorf("failed to delete content: %w", err)
}
return nil
}
// DeleteWithMeta removes the content at the given key together with
// its .meta sidecar file. A missing file is not an error.
func (s *VariantStorage) DeleteWithMeta(key VariantKey) error {
err := s.Delete(key)
if err != nil {
return err
}
metaPath := s.keyToPath(key) + ".meta"
metaPath := path + ".meta"
err = os.Remove(metaPath)
if err != nil && !os.IsNotExist(err) {
@@ -438,3 +438,73 @@ func TestVariantStorage_StoreLogsFailedMetaWrite(t *testing.T) {
t.Errorf("log missing %s; got %q", want, logBuf.String())
}
}
// storeTestVariant stores one variant, with its .meta file, in a new
// VariantStorage and returns the storage and the variant's key.
func storeTestVariant(t *testing.T) (*VariantStorage, VariantKey) {
t.Helper()
storage, err := NewVariantStorage(t.TempDir(), slog.New(slog.DiscardHandler))
if err != nil {
t.Fatalf("NewVariantStorage() error = %v", err)
}
key := CacheKey(&ImageRequest{SourceHost: testHostCDN, SourcePath: testPathCat})
_, err = storage.Store(key, bytes.NewReader([]byte("variant data")), "image/webp")
if err != nil {
t.Fatalf("Store() error = %v", err)
}
return storage, key
}
// TestVariantStorage_DeleteRemovesMeta verifies that Delete removes the
// variant's .meta file along with the variant file.
func TestVariantStorage_DeleteRemovesMeta(t *testing.T) {
t.Parallel()
storage, key := storeTestVariant(t)
metaPath := storage.keyToPath(key) + ".meta"
_, err := os.Stat(metaPath)
if err != nil {
t.Fatalf("Store() wrote no .meta file: %v", err)
}
err = storage.Delete(key)
if err != nil {
t.Fatalf("Delete() error = %v", err)
}
if storage.Exists(key) {
t.Error("Exists() = true after delete, want false")
}
_, err = os.Stat(metaPath)
if !os.IsNotExist(err) {
t.Errorf(".meta file left after Delete() (stat err=%v)", err)
}
}
// TestVariantStorage_DeleteWithoutMeta verifies that Delete succeeds for a
// variant whose .meta file is missing.
func TestVariantStorage_DeleteWithoutMeta(t *testing.T) {
t.Parallel()
storage, key := storeTestVariant(t)
err := os.Remove(storage.keyToPath(key) + ".meta")
if err != nil {
t.Fatalf("removing .meta file: %v", err)
}
err = storage.Delete(key)
if err != nil {
t.Fatalf("Delete() error = %v, want nil", err)
}
if storage.Exists(key) {
t.Error("Exists() = true after delete, want false")
}
}