Commit Graph
8 Commits
Author SHA1 Message Date
clawbot ff23551b07 Install Dockerfile build dependencies through script/bootstrap (closes #95)
check / check (push) Successful in 4m20s
The Dockerfile lint and build stages and Dockerfile.lint each carried
their own apk add list, a copy of what script/bootstrap installs. They
now copy script/, go.mod and go.sum and run script/bootstrap, so that
layer is reused until one of those changes. script/bootstrap gains a C
compiler check: the golang image has none, and cgo needs one.

The build adds -trimpath and -s -w; CGO_ENABLED=1 stays, as govips
links libvips. ARG VERSION moves to just above the build, so a new
version reruns neither script/bootstrap nor the tests.

Model: opus-5-5
2026-09-29 05:51:36 +00:00
clawbot 582ff66ba6 Describe what --health-interval does in docker-smoke on current Docker (closes #132)
check / check (push) Successful in 14s
The old comment said the flag stops the image's 30-second interval from
delaying the first probe past the wait. From Docker 25 on, the first
probe runs 5 seconds after start either way; the flag makes probes
after the 10-second start period come every second instead of every
30. Only the comment changes; TODO.md is left alone because the issue
limits the change to this script.

Model: opus-5-5
2026-09-28 15:48:53 +02:00
clawbot 2f7365cc9b Run all linting in Docker through script/lint (closes #104)
check / check (push) Successful in 12s
make lint calls script/lint, the only way golangci-lint is run. Inside a
container it runs the linter; anywhere else it builds Dockerfile.lint,
whose last step runs script/lint again. Both Dockerfiles set
container=docker to mark the container, since /.dockerenv is missing in
build steps and present on hosts that are themselves containers. The
Dockerfile lint stage runs make lint.

A new CACHEBUST build-arg on every run keeps the lint step from being
served from cache; a tmpfs mount keeps Go's and golangci-lint's caches
out of that step's layer, so runs do not pile up build cache.
script/bootstrap and the nix-shell package lists no longer carry
golangci-lint. golangci-lint config verify is not run: it fetches its
schema over an unpinned live HTTPS call.

Model: opus-4-8 (implementation); opus-5-5 (rework)
2026-09-28 14:27:35 +02:00
clawbot b7c1226c38 Add a Docker HEALTHCHECK and make docker-smoke (closes #111)
check / check (push) Successful in 11s
The runtime stage declares a HEALTHCHECK that probes
/.well-known/healthcheck.json with busybox wget. script/docker-smoke
(make docker-smoke) builds the image with script/docker, starts it with
a random PIXA_SIGNING_KEY, and passes only once Docker reports the
container healthy within 30 seconds; the container is removed on exit
and its log printed on failure. The Gitea workflow runs it after
script/cibuild; it is not part of make check.

It waits on Docker's health status instead of polling a published host
port because the Gitea job runs in its own container on its own
network, where such a port is not reachable at localhost.

Model: opus-5-5
2026-09-28 12:06:29 +02:00
clawbot b95ef1eb69 fix: script/test conditional-verbose-rerun with -cover (closes #59)
check / check (push) Failing after 2s
script/test now runs the suite quietly first (with -race and -cover, 30s timeout) and re-runs it with -v only when that run fails, then exits non-zero. This is the pattern REPO_POLICIES.md mandates; before, every green run printed full per-test output.

What a reader would trip over: the whole compound command is passed as one string to run_with_cgo_deps, so it behaves the same on the host path and under the nix-shell fallback. -cover is on the first run only; the verbose rerun exists for diagnostics.

Disclosure: no test was written for the wrapper script itself; the failure path was exercised by hand by author and reviewer.
Disclosure: the nix-shell fallback is kept; moving tests into Docker belongs to #101 and #104.

Model: opus-4-8 (implementation, review); fable-5-1 (landing message)
2026-09-21 19:43:21 +02:00
clawbot 04b5db6fbf next -> main (1.0.0 milestone) (#105)
check / check (push) Successful in 5s
Accumulating milestone branch. One squashed commit per closed issue; `next` is kept green and mergeable to `main` at any time without notice.

Landed so far:

- `chore: update golangci-lint to v2.12.2 with canonical config` (#54) — canonical v2-schema `.golangci.yml`, pins bumped in `Dockerfile` and `script/bootstrap`, tree at `0 issues.`. Three behaviour deltas are recorded in that PR's body: `Cache.StoreVariant` takes a context, `MetadataStorage.Store` no longer leaks temp files on failure, and the `signing_key` too-short error text gained a `value too short:` prefix.

Sequencing for the milestone is tracked in #103.

Reviewed-on: #105
Co-authored-by: clawbot <clawbot@noreply.example.org>
2026-09-21 09:31:54 +02:00
clawbotandsneak 63fbc98e63 feat: cache size management and LRU eviction (closes #51) (#55)
check / check (push) Has been cancelled
Implements #51 per the issue DoD and the owner direction comment (issuecomment-44068).

## Behavior

**Config: `cache_max_bytes`** (integrates with the #52/#53 validation framework)

- Strict int64 parsing via a new `getInt64`/`int64Val` getter in the existing strict-loader pattern; the key is registered in the known-keys list. A SET but invalid value — negative, float, null, non-numeric string, boolean, list — aborts startup with exit 1 naming the key and the offending value.
- Explicit values are used exactly as given, any non-negative amount, no floor. `cache_max_bytes: 0` is a valid value that disables the disk cache entirely.
- Omitted: after `state_dir` validation, the default resolves to `max(75% of free bytes on the filesystem containing &lt;state_dir&gt;/cache/, 500 MiB)`. The cache directory is created first and statfs runs on that actual path, so the measurement hits the right filesystem. The probe is injectable (`FreeSpaceProbeFunc`) so tests do not depend on the host disk. The effective limit (and disabled state) is logged at startup.

**Size accounting** (no directory scans on the hot path)

- Migration `002` adds a `variant_content` table — processed variants were previously untracked anywhere — and a `last_accessed_at` column on `source_content`, both indexed. Total usage is two SUM queries.
- Stores record accounting rows; cache hits touch the LRU timestamps (same cost class as the existing per-request stats UPDATEs). The variant accounting insert is best-effort with a warning: the reconciliation pass (below) adopts any file that missed its row, and this keeps the pre-migration inline test schema working.

**Eviction policy: global LRU across both content classes**

- Candidates are the least-recently-used entries from `variant_content` and `source_content` (batched, 100 per class per pass, merged oldest-first by `COALESCE(last_accessed_at, fetched_at)`), evicted until usage is at or below the limit.
- Why global LRU: recency of actual use is the best cheap predictor of future use for a CDN-style cache, and treating both classes in one ordering avoids pathologies of class-priority schemes (e.g. evicting every variant before any cold source blob, which would tank hit rate, or the reverse, which would hoard stale sources). Byte-for-byte, the coldest data goes first regardless of what kind it is. LFU-style schemes need more bookkeeping for marginal gain at this scale.
- Reference safety (the multi-reference DoD case): evicting a source blob deletes ALL `source_metadata` rows referencing it plus its `source_content` row in a single transaction BEFORE the file is unlinked. A blob referenced by multiple source paths is only ever removed together with all of its references, and DB rows never point at deleted files (the crash window leaves at worst an orphaned file, which reconciliation sweeps). The JSON metadata sidecars for removed rows are deleted as well.

**Triggers, off the request path**

- A background goroutine (started in the handlers OnStart hook, stopped in OnStop) runs an eviction pass on a periodic ticker (5 min) and on write pressure: every store sends a non-blocking notification on a capacity-1 channel. Requests never wait on eviction.
- On startup the goroutine first reconciles accounting with the disk (off the hot path): adopts untracked variant files (size/mtime from disk, content type from the `.meta` sidecar), drops accounting rows whose files are missing, removes source blob files the DB does not know (unreachable, since lookups go through `source_metadata`), removes rows whose files are gone, and sweeps `.tmp-*` files older than an hour.

**`cache_max_bytes: 0` disables the disk cache**

- No cache directories are created, lookups always miss, `StoreSource`/`StoreVariant` are no-ops, no evictor runs; every request fetches and processes uncached. Verified end-to-end (below).

## Notes for review

- At the `imgcache.CacheConfig` layer, disabling is an explicit `DisableDiskCache` flag rather than `MaxBytes == 0`, because existing test fixtures construct `CacheConfig` without `MaxBytes` and rely on the legacy "no limit" behavior; per repo rules those tests were not touched. The config layer maps `cache_max_bytes: 0` to the flag in `handlers`. `MaxBytes == 0` at that layer means "no limit enforced" and is unreachable from production config (the computed default is always at least 500 MiB).
- The negative cache stays active in disabled mode: it is DB-backed (TTL-expired rows in SQLite), not part of the disk cache this issue bounds, and it protects against hammering failing upstreams. Flagging explicitly since the direction said "no cache reads, no cache writes" — I read that as the disk cache; happy to disable it too if intended.
- One deviation from pure red/green: after the red commit I extended the new-test fixture helper (`newEvictionTestCache`) to pass `DisableDiskCache: maxBytes == 0`, mirroring the production mapping, when the flag design emerged. Assertions were not touched; no pre-existing tests were modified.
- Sidecar files (`.meta`, metadata JSON) are not counted in usage; they are bounded by entry counts and small (tens of bytes to ~1 KiB per entry) while content bytes dominate. Documented here for transparency.
- Discovered while working: `Cache.Stats` reads the never-populated `output_content`/`request_cache` tables, so `TotalItems`/`TotalSizeBytes` are always 0. Out of scope here; filing as a separate issue.

## Verification

- TDD: commit `3963ec3` adds the failing tests first (18 new tests covering strict parsing, default computation with injected probe including floor and 75% branches, explicit-no-floor, zero-disables, size accounting, dedup accounting, LRU order, multi-reference blob eviction with the no-dangling-references invariant, under-limit no-op, write-pressure trigger, periodic trigger, reconciliation); implementation follows in `8cb09b6`/`bdd86a4` until green.
- `make check` green (all tests, lint 0 issues, fmt-check) at HEAD.
- Pinned CI lint gate: `docker build --target lint .` green (golangci-lint v2.10.1).
- End-to-end with the built binary:
  - omitted key: startup logs `computed default cache size limit from free space` and `effective cache size limit` (75% of the test host's free space);
  - `cache_max_bytes: banana`: exit 1 with `config key "cache_max_bytes": value "banana" is not an integer`;
  - `cache_max_bytes: 0`: `cache_disabled=true` logged, two identical requests both fetch upstream (2 upstream fetches logged), 200 `image/jpeg` responses, no `cache/` directory created, only `state.sqlite3` in the state dir;
  - enabled: second request served from cache (1 upstream fetch), `variant_content` and `source_content` rows match the on-disk file sizes.

Co-authored-by: sneak <sneak@sneak.berlin>
Reviewed-on: #55
Co-authored-by: clawbot <clawbot@noreply.example.org>
Co-committed-by: clawbot <clawbot@noreply.example.org>
2026-08-09 13:22:51 +02:00
sneak 504afea4f8 scripts-to-rule-them-all (#45)
check / check (push) Successful in 4s
Reviewed-on: #45
Co-authored-by: sneak <sneak@sneak.berlin>
Co-committed-by: sneak <sneak@sneak.berlin>
2026-07-07 02:14:03 +02:00