10 Commits

Author SHA1 Message Date
2bfa11c10c Walk the path CI runs in lint-once, and pin the branch the container installs (closes #33)
All checks were successful
check / check (push) Successful in 35s
The header of test/packaging/lint-once.test.ts claimed a duplicate prettier
pass is caught wherever it is added. It was not: the walk started at
`make check`, which never reads `Dockerfile`, so appending
`RUN yarn run prettier --check .` to the image that `script/cibuild` builds
left the suite green — two prettier passes on the one path where it matters
most. The walk now also starts at `.gitea/workflows/check.yml` and follows
its `run:` steps into `script/cibuild` and from there into both images, so
the graph under test is the one CI executes rather than the one it was
assumed to execute. Reaching `script/cibuild` and `Dockerfile` is asserted,
and the test and build image is asserted to invoke prettier zero times.

The lockfile assertion was a substring check against the whole of
`script/bootstrap`. That script has two install sites, and the containers
take the second, because the pinned node image ships yarn; changing that
site to a bare `yarn install` kept the suite green while the container's
install stopped being pinned. `install_js_deps` is now resolved out of the
script and split at its `missing yarn` guard, and every `yarn install`
occurrence in each branch is required to carry `--frozen-lockfile`. That the
container runs `script/bootstrap` at all is asserted too, so the lockfile
assertions cannot end up describing a script the image never executes.

Prettier is counted per occurrence instead of per line:
`prettier --check . && prettier --check src` was one invocation by the old
count. The `continue` that followed a counted line also dropped every
script, make, yarn and docker edge sharing that line, so a subtree could be
hidden behind a single `&&`; edges are now extracted from every line.

Undercounting is what would make this file worthless, so every way of
reaching nothing is a thrown error rather than a quiet zero: an unknown
Makefile target, an unknown package.json script, a missing script file, a
node that resolves to no commands, and an unknown node kind. All five are
tested, as is a walk that legitimately counts zero, and the cycle guard.

Every assertion in the file was mutation-tested: changed to assert something
else, run, and confirmed to fail for its own named reason. The two mutations
above were reproduced and both now turn the suite red.

test/packaging/entrypoints.test.ts said `make check` runs test, lint and
fmt-check. Formatting has been part of the lint container since the
duplicate host pass was removed, so the comment now says what it does.
2026-08-10 13:41:12 +00:00
a73f0abbe8 Check formatting once per make check, in the container (closes #29)
All checks were successful
check / check (push) Successful in 59s
script/check ran script/test, script/lint and script/fmt-check. Since
linting moved into Docker, script/lint is a build of Dockerfile.lint,
which runs `prettier --check .` as a build step — so make check checked
formatting twice over the same tree: once in the container and once on
the host. script/precommit had the same pair.

Drop the script/fmt-check call from both. The container keeps the check,
because a successful Dockerfile.lint build is what CI treats as proof of
a clean tree, and it is the stronger of the two verdicts: its prettier is
digest-pinned and installed under --frozen-lockfile, while the host's is
whatever the working tree happens to have. The pre-commit hook is
unchanged in what it catches — script/lint still fails a badly formatted
tree, and therefore the commit.

script/fmt-check survives as a standalone entrypoint, as REPO_POLICIES.md
requires, for asking the formatting question by itself without docker.
Its verdict cannot drift from the container's: prettier is pinned to an
exact version, installed from yarn.lock in both places, and reads
.gitignore as its default ignore file, which is why .dockerignore keeps
.gitignore in the build context.

The count is asserted rather than promised. test/packaging/lint-once.test.ts
walks the invocation graph from each entrypoint — through the Makefile
shims, the script/ calls, the package.json scripts and the docker build
into Dockerfile.lint's RUN steps — and counts prettier invocations: one
per make check, one per script/precommit, and one each for make lint and
make fmt-check alone, so neither can become a no-op that satisfies the
count trivially. The walk also asserts which nodes it reached, so a
restructure that defeats the resolver fails the test instead of quietly
counting zero.

Observed: 2 prettier invocations per make check before, 1 after.
2026-08-10 13:05:03 +00:00
fed39d19cf Run all linting in Docker via Dockerfile.lint (closes #30)
All checks were successful
check / check (push) Successful in 1m2s
Linting now happens in one place only: a new root Dockerfile.lint copies
the repo into the digest-pinned node image already used by Dockerfile and
runs eslint and prettier as build steps, so a successful build is a clean
lint. script/lint is reduced to building it, which also works where the
docker daemon is remote and bind mounts are impossible. No host lint path
survives: the "lint" script is gone from package.json, so there is no
second, unpinned way to get a lint verdict.

Caching is waived for lint, because a lint build over an unchanged tree
returns success in well under a second having linted nothing. LINT_EPOCH
is the cache buster and it fails closed exactly as CHECK_EPOCH does: an
unset ARG is the empty string, which is a perfectly stable cache key, so
the guard rejects it and a bare `docker build -f Dockerfile.lint .` errors
out instead of serving a green it did not earn. Both linters sit below the
guard, so a fresh epoch forces them to execute while the bootstrap and
dependency layers above stay cached.

That makes script/lint a docker build, which nothing inside a container
may call. script/check calls script/lint, so the Dockerfile image can no
longer run make check: the lint stage and its COPY --from=lint ordering
hack are deleted, and the remaining stage runs make test and make build
under the existing CHECK_EPOCH guard. script/cibuild is now the composite
gate and builds the lint image first, so a lint failure is reported before
the slower suite runs.

The .dockerignore exclusions are unchanged and still apply to the lint
build, including the .claude/ exclusion (eslint's flat config does not
ignore dot-directories, so a nested worktree in the context would be
linted) and the deliberate exception that keeps .gitignore in the context
for prettier. A new test asserts no per-Dockerfile ignore file shadows the
root one for either image, and test/packaging/lint-docker.test.ts asserts
the whole shape: the docker-only lint path, the digest pin, manifests
copied before sources, the fail-closed guard with both linters below it,
the absence of a lint stage or make check in Dockerfile, and the build
order in script/cibuild.
2026-08-10 12:39:23 +00:00
156fe871e8 Make the image build multi-stage and cache-proof (closes #4)
All checks were successful
check / check (push) Successful in 23s
The image build reported a green it had not earned. `script/cibuild` is a
bare `docker build .`, and with `COPY . .` followed by `RUN make check`,
an unchanged tree served that layer from cache: the suite never ran and
the build still exited 0, while the script's header comment asserted the
opposite.

CHECK_EPOCH, passed by `script/cibuild` and `script/docker`, changes the
cache key of the check and build layers on every invocation. It is
guarded, because an unset ARG is the empty string and therefore a stable
key: without the guard a plain `docker build .` — the command the policy
names, and the one anyone debugging types — would still get the false
green. A missing argument is now a hard failure rather than a silent
degradation to the behaviour the epoch was added to prevent.

The Dockerfile is now two stages: `fmt-check` and `lint` run first, and
the check stage takes a `COPY --from=lint` dependency on them, so a
formatting mistake fails the build in seconds instead of racing the
suite to the finish. Both stages stay pinned to the same digest.

The remaining fixes are one-liners that had made the target unusable:
`script/projectname` still printed the pre-rename name, so `make docker`
tagged its image after a name this project dropped in May;
`script/bootstrap` installed without fetching apt's package lists, which
cannot work on a Debian base; and `.dockerignore` had drifted far enough
from `.gitignore` to ship a ~100 MB compiled binary and any agent
worktree under `.claude/` into the build context. The second of those is
a correctness problem, not a size one — vitest globs a copied worktree's
tests alongside the real ones and runs the suite twice over. `.gitignore`
itself stays in the context, because prettier reads it as a default
ignore file and dropping it would change what `make fmt-check` sees.
2026-08-09 14:38:32 +00:00
118f8e22c3 Add failing tests for the project name and build context
script/projectname still prints the pre-rename name, so make docker tags
its image quack, and .dockerignore has drifted from .gitignore: a compiled
bin/quak, the vitest and tsc caches, the CLI's runtime directory and the
local worktree directory all reach the build context. Neither failure is
visible in a build that exits 0.

Red until the fixes land.
2026-08-09 14:27:50 +00:00
69bd6d1539 Compile bin/ alongside src/ so the package can be built (closes #3)
All checks were successful
check / check (push) Successful in 5s
rootDir was ./src while include also matched bin/**/*, which is TS6059: tsc
refuses to emit at all when a compiled file sits outside rootDir. rootDir is
now the repository root, which is the smallest change that makes the two
agree and leaves the source layout the README documents alone. Output keeps
the shape of the source tree, so main and types move to dist/src/index.js and
dist/src/index.d.ts while bin.quak stays at dist/bin/quak.js. The alternative,
moving the CLI body into src/ behind a shim in bin/, would hold main at
dist/index.js at the cost of churning the CLI and contradicting the layout
diagram in the README.

Clearing TS6059 exposed two type errors that had never been reached, because
the config error aborts before checking: StateAddress was read as a namespace
member off the default import, and the secretstream pull was called without
the additional-data argument, which libsodium does not make optional. The type
is now taken from the module's named export and the pull passes null for ad,
matching the null already passed on the push side in encryptBlob. Neither
changes what runs.

noEmitOnError stops a failed build from leaving output behind. It emitted
despite the error before, which is how a stale bin/quak.js came to sit next to
bin/quak.ts in a working tree, where eslint then read it and failed make check
on a generated file.

script/build compiles and then checks that the files package.json advertises
are among the ones the compiler wrote, since tsc knows nothing about the
manifest and a green build could still ship a package whose main resolves to
nothing. It also sets the executable bit on the bin entries, which tsc does
not carry over from the source even though it does copy the shebang. The
Makefile target is now a shim over it, as the other targets are, and
package.json's build script points at it so yarn build gets the same checks.

The Dockerfile runs make build after make check, so a branch that does not
compile cannot reach main. What make check itself runs is unchanged.

package.json gains a quak script, so the yarn quak commands the README's
Getting Started block has always listed resolve to the built CLI.
2026-08-09 10:06:31 +00:00
d79ed83f4d Add failing tests for the package entrypoint contract
The manifests disagreed and nothing noticed. tsconfig.json set rootDir to
./src while include also matched bin/**/*, which is TS6059, so no build had
succeeded; package.json meanwhile advertised main, types and a bin that a
successful build would have to produce. make check runs test, lint and
fmt-check, so neither half was ever exercised.

These tests read tsconfig.json and package.json and assert the contract
between them without invoking a compiler, which keeps them in the fast unit
suite: every include pattern must root under rootDir, and main, types and
bin.quak must equal the paths tsc will emit for src/index.ts and bin/quak.ts.
They also require a quak script pointing at the built CLI, which the README's
Getting Started block has always told the reader to run.

Three of them fail at this commit.
2026-08-09 10:01:32 +00:00
348f23bac9 Narrow the replay errno set to failures that prove no connection existed
All checks were successful
check / check (push) Successful in 4s
`CONNECT_CODES` drives `isSafeToReplay`, which is the only thing standing
between a transport failure and a replayed `POST /users/two-factor/verify`.
It included `EHOSTUNREACH`, `ENETUNREACH` and `ENETDOWN` on the stated
grounds that those errnos can only be reported before any request byte was
written. That is not true on Linux: an ICMP destination-unreachable
delivered on an already-established connection sets the socket error and
the next read or write returns `EHOSTUNREACH` or `ENETUNREACH`, and a local
interface going down after the request was fully written surfaces as
`ENETDOWN` the same way. In each case the server may already have received
and acted on the request -- exactly the ambiguity the rule exists to
exclude, on the paths that consume a second-factor attempt or register a
thumbnail.

The three are dropped from `CONNECT_CODES` and stay in `TRANSPORT_CODES`,
so they remain retryable for the idempotent calls; only replay eligibility
narrows. What is left -- `ENOTFOUND`, `EAI_AGAIN`, `ECONNREFUSED` -- means
no TCP connection to the server ever existed, so no request byte can have
been transmitted.

The justification is corrected everywhere it was stated: the comment on
`CONNECT_CODES`, the one on `isSafeToReplay`, the `postJSON` call site, the
README's idempotency section and the `client.test.ts` docblock. All of them
now describe what the narrowed set actually establishes rather than
claiming a proof it did not support.

The narrowing is enforced by the suite rather than asserted in a comment:
the three errnos join `ECONNRESET`/`EPIPE`/`ETIMEDOUT` in the
`isSafeToReplay`-returns-false test, with companion `isRetryable` assertions
so a future edit cannot make them non-retryable by accident. Putting the
three back into `CONNECT_CODES` turns that test red (1 failure, verified).
2026-08-09 05:41:59 +00:00
f3cf4af833 Retry transient network failures with exponential backoff (closes #2)
All checks were successful
check / check (push) Successful in 20s
No retry on 4xx, backoff on 5xx and transport failures, and a deadline on
every request. Before this, one transient 503 or TCP reset failed a file
for good, and a CDN connection that went quiet after accepting the request
blocked `quak backup` forever, because there was no timeout anywhere.

src/retry.ts holds the policy: a classifier that decides whether another
attempt could produce a different answer, and a loop that acts on it with
exponential backoff and full jitter. Retried: 5xx, 408, 429, transport
failures (the errno is read out of the cause chain, which is where Node's
fetch puts it), deadline aborts, and truncated transfers. Not retried:
every other 4xx, and anything unrecognised — a wrongly retried permanent
failure delays every remaining file, while a wrongly abandoned transient
one costs a single file the next run picks up. Attempt count, delays,
sleep and jitter source are all configurable through ApiClientOptions;
sleep being injectable is what lets the suite exercise the policy without
waiting.

Truncation needed a type before it could be classified. streamDecrypt
threw plain Errors whose messages began "download: stream truncated", and
classifying on message text would mean the next reword silently turned
every truncated download into a permanent failure. It now throws
TruncatedStreamError, which lives in src/errors.ts alongside ApiError so
the classifier can recognise both without importing the modules that
import it; api/client.ts re-exports ApiError, so it stays one class and
every existing import path still resolves.

Downloads retry the request, the stream consumption and the decryption
together. Only the first of those happens inside ApiClient: a socket
reset after the headers arrived throws in streamDecrypt, and retrying the
request alone would never see it. The client's own retry is switched off
for those two calls so the budgets do not multiply into sixteen requests
per file, and the atomic write stays outside the loop so a download that
took three attempts still performs one write and one rename.

Non-idempotent requests are not blindly replayed. postJSON and putJSON
reach create-session, two-factor/verify — which burns one of a few
second-factor attempts — and files/thumbnail, so they retry only when the
connection was never established and the server provably never saw the
request. putFile is exempt and retries fully: a presigned PUT stores one
whole object at one key, with no partial state to damage. It now throws
ApiError with the status, as do the two null-body paths, which previously
threw bare Errors that nothing could classify.

Timeouts come from AbortSignal.timeout(), renewed per attempt: 30s for
JSON and upload calls, 10 minutes for file bodies, since a value short
enough to keep a hung API call from stalling a backup would cancel a
legitimate multi-gigabyte download. The download deadline is enforced
over the body rather than only the headers, by racing each read against
the signal, so the guarantee does not depend on the fetch implementation
tearing down a stream it already handed over.

listMissingThumbnails now separates a genuine 404 from an exhausted
retry. Its bare catch reported both as missing, which after this change
would have let a few minutes of 500s talk fix-missing-thumbnails into
regenerating and re-uploading thumbnails that were fine. runBackup and
runMetadataBackup are untouched: the retry sits below them and their
per-file resilience is unchanged.
2026-08-09 05:21:31 +00:00
0cbe338b58 Add failing tests for the download retry policy
Tests only; the retry module they import does not exist yet, so the
branch is red at this commit.

New test/retry/retry.test.ts documents the classifier and the backoff:
which errors are worth another attempt, which are not, and how the delay
before each retry is derived. It asserts on the arguments handed to an
injected sleep function rather than on elapsed time, so the suite never
waits and the numbers are exact.

test/api/client.test.ts gains the request-count contract for each of the
six call sites, the deadline behaviour, the ApiError typing that the
presigned PUT and the null-body paths need in order to be classified at
all, and the replay rule for the two non-idempotent methods.

test/download/download.test.ts gains the case that motivates the whole
design: a socket reset after the response headers arrived, which happens
below ApiClient and can only be caught by retrying the request, the
stream consumption and the decryption together. It also pins that a
retried download stages exactly one temp file, and that the two retry
layers do not compose into a multiplied request budget. The existing
truncation tests now assert on the error type rather than its wording,
since that type is what the classifier reads.

test/thumbnails/thumbnails.test.ts separates a genuine 404 from an
exhausted retry, so a failing server can no longer make
fix-missing-thumbnails re-upload thumbnails that already exist.
2026-08-09 05:11:07 +00:00
34 changed files with 3488 additions and 161 deletions

View File

@@ -1,8 +1,50 @@
# Mirrors .gitignore, with one deliberate exception: .gitignore itself stays
# in the build context, because prettier 3 reads it as a default ignore file
# and dropping it would change what `make fmt-check` sees inside the image.
# VCS
.git .git
# OS
.DS_Store
Thumbs.db
# Editors
*.swp
*.swo
*~
*.bak
.idea/
.vscode/
*.sublime-*
# Node
node_modules node_modules
# TypeScript / build artifacts
dist dist
build build
*.tsbuildinfo
coverage coverage
.DS_Store .nyc_output/
# Vitest
.vitest-cache/
# Environment / secrets
.env .env
.env.* .env.*
*.pem
*.key
# Compiled binary (built by make build-bin); around 100 MB
bin/quak
# quak runtime data (in case anyone runs the CLI from inside the repo)
.quak/
# Local per-developer tool state, including agent worktrees. Correctness,
# not context size: a worktree copied in here has its own test/ tree, which
# vitest globs alongside the real one, so the containerised suite runs N+1
# times over and still reports success.
.claude/

3
.gitignore vendored
View File

@@ -36,5 +36,6 @@ bin/quak
# quak runtime data (in case anyone runs the CLI from inside the repo) # quak runtime data (in case anyone runs the CLI from inside the repo)
.quak/ .quak/
# Local Claude Code settings (per-developer) # Local per-developer tool settings and scratch state, including the
# worktrees agents check out under this directory
.claude/ .claude/

View File

@@ -1,10 +1,28 @@
# node 22-alpine, 2026-02-22 # Test and build image: the suite, then the compile.
FROM node@sha256:e4bf2a82ad0a4037d28035ae71529873c069b13eb0455466ae0bc13363826e34 #
# Linting deliberately does not happen here. `script/lint` is a build of
# Dockerfile.lint, and `script/check` calls `script/lint`, so running
# `make check` in this image would mean running `docker build` inside a
# container. Lint runs exactly once, in Dockerfile.lint; script/cibuild
# builds that first and this second.
# node 22.22.0 on Alpine 3.23.3 (node:22-alpine), 2026-08-09
FROM node@sha256:e4bf2a82ad0a4037d28035ae71529873c069b13eb0455466ae0bc13363826e34 AS check
WORKDIR /app WORKDIR /app
COPY script/ script/ COPY script/ script/
COPY package.json yarn.lock ./ COPY package.json yarn.lock ./
RUN script/bootstrap RUN script/bootstrap
COPY . . COPY . .
RUN make check # CHECK_EPOCH is a cache buster: without it Docker serves the test layer from
# cache on an unchanged tree, the suite never executes, and the build still
# exits 0. The guard makes an absent argument a hard failure — an unset ARG
# is the empty string, which is a perfectly stable cache key, so a plain
# `docker build .` would otherwise still get the false green. Fail closed.
ARG CHECK_EPOCH
RUN [ -n "$CHECK_EPOCH" ] || exit 1
RUN make test
ARG CHECK_EPOCH
RUN [ -n "$CHECK_EPOCH" ] || exit 1
RUN make build

35
Dockerfile.lint Normal file
View File

@@ -0,0 +1,35 @@
# Lint image: every lint run happens here, and nowhere else. The repo is
# COPYed into a digest-pinned image and the linters run as build steps, so a
# successful build IS a clean lint. `script/lint` does nothing but build this
# file, which also works where the docker daemon is remote and bind mounts are
# impossible. Nothing that runs inside a container may call `script/lint`:
# that is why Dockerfile no longer runs `make check`.
# node 22.22.0 on Alpine 3.23.3 (node:22-alpine), 2026-08-09
FROM node@sha256:e4bf2a82ad0a4037d28035ae71529873c069b13eb0455466ae0bc13363826e34 AS lint
WORKDIR /app
# Manifests before sources, so the dependency install layer stays cached
# until package.json or yarn.lock changes. script/bootstrap ends in
# `yarn install --frozen-lockfile`; the lint steps below are deliberately
# not cached.
COPY script/ script/
COPY package.json yarn.lock ./
RUN script/bootstrap
COPY . .
# LINT_EPOCH is a cache buster, with the same fail-closed contract as
# CHECK_EPOCH in Dockerfile. No lint cache is wanted: on an unchanged tree
# Docker serves the linter layers in well under a second, having linted
# nothing, and the build still exits 0. The guard makes an absent argument a
# hard failure — an unset ARG is the empty string, which is a perfectly
# stable cache key, so a plain `docker build -f Dockerfile.lint .` would
# otherwise get exactly that false green. Every layer below this one is a
# child of the guard, so a fresh epoch forces all of them to execute.
ARG LINT_EPOCH
RUN [ -n "$LINT_EPOCH" ] || exit 1
# The linters are invoked directly rather than through `make lint`, because
# `make lint` is the build of this file.
RUN yarn run eslint .
RUN yarn run prettier --check .

View File

@@ -24,7 +24,7 @@ check:
@script/check @script/check
build: build:
@$(YARN) tsc @script/build
build-bin: build-bin:
nix-shell -p bun --run "bun build bin/quak.ts --compile --outfile bin/quak" nix-shell -p bun --run "bun build bin/quak.ts --compile --outfile bin/quak"

204
README.md
View File

@@ -81,25 +81,76 @@ alpine. We provide:
`script/bootstrap`, then `script/install-precommit` `script/bootstrap`, then `script/install-precommit`
- `script/projectname` — output the project name (our own extension); used by - `script/projectname` — output the project name (our own extension); used by
`script/docker` for the image tag `script/docker` for the image tag
- `script/build` — compile the TypeScript sources into `dist/`, then verify that
the entrypoints `package.json` declares (`main`, `types`, `bin`) are among the
files the compiler wrote, and make the CLI executable (our own extension)
- `script/test` — run the test suite (vitest, hard-capped at 30s where `timeout` - `script/test` — run the test suite (vitest, hard-capped at 30s where `timeout`
is available, verbose rerun on failure) is available, verbose rerun on failure)
- `script/lint` — run eslint and a prettier check - `script/lint` — run eslint and a prettier check, by building
`Dockerfile.lint`; requires docker (see Linting below)
- `script/fmt` — format all files with prettier (writes) - `script/fmt` — format all files with prettier (writes)
- `script/fmt-check` — check formatting (read-only) - `script/fmt-check` — check formatting on the host (read-only); standalone, and
- `script/check` — run all checks: `test`, `lint`, `fmt-check` (our own not called by `script/check` or `script/precommit`, because `script/lint`
extension) already checks formatting in the container (see Linting below)
- `script/docker` — build the Docker image, tagged via `script/projectname` - `script/check` — run all checks: `test`, `lint` (our own extension)
(byte-identical across repos) - `script/docker` — build the test and build image, tagged via
- `script/cibuild` — cd to the repo root and `docker build .` (what CI runs; the `script/projectname`
image build runs `make check`) - `script/cibuild` — cd to the repo root and build both images (what CI runs):
`script/lint` first, then the `Dockerfile` image, which runs `make test` and
`make build`
- `script/precommit` — run by the git pre-commit hook (our own extension); runs - `script/precommit` — run by the git pre-commit hook (our own extension); runs
`script/lint` and `script/fmt-check` but deliberately not the tests, so the `script/lint`, which checks both lint and formatting, but deliberately not the
TDD red-phase commit can land tests, so the TDD red-phase commit can land
- `script/install-precommit` — installs the git pre-commit hook (our own - `script/install-precommit` — installs the git pre-commit hook (our own
extension); `make hooks` shims to it extension); `make hooks` shims to it
`make hooks` installs the pre-commit hook that runs `script/precommit`. `make hooks` installs the pre-commit hook that runs `script/precommit`.
### Linting
Linting runs in a container, one way, everywhere. `script/lint` builds
`Dockerfile.lint`, which copies the repo into a digest-pinned node image and
runs eslint and prettier as build steps, so a successful build is a clean lint.
There is no host lint path: docker is required to lint, and that also works
where the docker daemon is remote and bind mounts are impossible.
The formatting check is part of that, not a step beside it. `script/check` and
`script/precommit` therefore call `script/lint` and stop; neither calls
`script/fmt-check` as well, which would run prettier a second time over the same
tree for the same verdict — and the weaker of the two, since the host's prettier
is whatever the working tree has installed. So `make check` and the pre-commit
hook both still fail on a badly formatted tree, and prettier runs exactly once
in each. `test/packaging/lint-once.test.ts` asserts that count by walking the
invocation graph, so a second pass cannot creep back in unnoticed.
`script/fmt-check` remains as a standalone entrypoint for asking the formatting
question on its own, without docker and without the rest of lint. Its verdict
cannot drift from the container's: prettier is pinned to an exact version,
installed from `yarn.lock` under `--frozen-lockfile` in both places, and reads
`.gitignore` as its default ignore file — which is why `.dockerignore`
deliberately keeps `.gitignore` in the build context.
Lint happens in exactly one place, which constrains the rest of the build.
`script/check` calls `script/lint`, so `make check` cannot run inside a
container without asking for docker inside docker. The image built from
`Dockerfile` therefore runs `make test` and `make build` and does not lint;
`script/cibuild` builds `Dockerfile.lint` first and that image second, so CI
gets both verdicts.
### Build epochs
`script/lint` passes `--build-arg LINT_EPOCH="$(date +%s)"`, and `script/docker`
and `script/cibuild` pass `--build-arg CHECK_EPOCH="$(date +%s)"`. Both
Dockerfiles refuse to build without their argument. This is deliberate: on an
unchanged tree Docker would otherwise serve the linter and test layers from
cache, so nothing would run and the build would still exit 0 — a lint build over
an untouched tree returns success in well under a second, having linted nothing.
A changing epoch invalidates every layer below the guard on every invocation
while leaving the dependency layers above them cached, and the missing-argument
guard means a bare `docker build .` fails loudly instead of quietly reporting a
green it did not earn: an unset build argument is the empty string, which is a
perfectly stable cache key.
## Rationale ## Rationale
Ente is one of very few photo services with a credible end-to-end encryption Ente is one of very few photo services with a credible end-to-end encryption
@@ -133,8 +184,10 @@ All work on quak is test-driven. No exceptions.
3. Subsequent commits add the implementation and any refactors needed to make 3. Subsequent commits add the implementation and any refactors needed to make
the tests pass. the tests pass.
4. A feature branch can only be merged into `main` when `make check` is green. 4. A feature branch can only be merged into `main` when `make check` is green.
`main` is always green. The Dockerfile runs `make check`, so a red branch `main` is always green. CI runs `script/cibuild`, which lints via
cannot pass CI. `Dockerfile.lint` and then runs `make test` and `make build` in the
`Dockerfile` image, so neither a red branch nor one that does not compile can
pass CI.
5. Tests are the canonical API documentation for this library. Every test file 5. Tests are the canonical API documentation for this library. Every test file
is commented thoroughly enough that a reader who has never seen quak can is commented thoroughly enough that a reader who has never seen quak can
learn how to use it from the tests alone. Comments explain why a behavior learn how to use it from the tests alone. Comments explain why a behavior
@@ -148,10 +201,11 @@ All work on quak is test-driven. No exceptions.
history must still show tests landing before (or with) the matching history must still show tests landing before (or with) the matching
implementation. implementation.
8. The pre-commit hook installed by `make hooks` runs `script/precommit`, which 8. The pre-commit hook installed by `make hooks` runs `script/precommit`, which
runs the lint and format checks but not the full `make check`. This is runs `script/lint` — eslint and the prettier check, in the container — but
deliberate so the TDD red-phase commit (failing tests, no implementation yet) not the tests, and so not the full `make check`. This is deliberate so the
can land. The full `make check` runs as part of `docker build .`, which is TDD red-phase commit (failing tests, no implementation yet) can land. The
what CI executes, so a red branch still cannot reach `main`. suite runs as part of the image build, which is what CI executes via
`script/cibuild`, so a red branch still cannot reach `main`.
## Design ## Design
@@ -169,6 +223,8 @@ quak/
model/ decrypted Collection, File, Metadata types + decrypt fns model/ decrypted Collection, File, Metadata types + decrypt fns
download/ streaming file/thumbnail download + decryption download/ streaming file/thumbnail download + decryption
backup.ts resilient full-account backup with dedup backup.ts resilient full-account backup with dedup
errors.ts error types shared across layers
retry.ts retry classifier + exponential backoff with jitter
thumbnails.ts detect + regenerate missing thumbnails thumbnails.ts detect + regenerate missing thumbnails
client.ts high-level Client class assembled from the above client.ts high-level Client class assembled from the above
index.ts public library exports index.ts public library exports
@@ -176,11 +232,18 @@ quak/
quak.ts CLI entrypoint (commander.js) quak.ts CLI entrypoint (commander.js)
test/ unit + integration tests (vitest) test/ unit + integration tests (vitest)
Makefile Makefile
Dockerfile Dockerfile test suite and compile
Dockerfile.lint eslint and prettier, as build steps
package.json package.json
tsconfig.json tsconfig.json
``` ```
`make build` compiles that tree into `dist/`, preserving its shape: the library
lands in `dist/src/` and the CLI in `dist/bin/quak.js`, which is what
`package.json` points `main`, `types` and `bin` at. The compiler's `rootDir` is
the repository root rather than `src/`, because `bin/` is compiled too and
`rootDir` has to contain everything that is compiled.
### Cryptography ### Cryptography
All cryptography is done by `libsodium-wrappers-sumo` (the "sumo" build is All cryptography is done by `libsodium-wrappers-sumo` (the "sumo" build is
@@ -250,6 +313,93 @@ Endpoints used:
- `POST /files/upload-url`: mint a presigned upload URL (for thumbnail repair). - `POST /files/upload-url`: mint a presigned upload URL (for thumbnail repair).
- `PUT /files/thumbnail`: register an uploaded thumbnail's object key. - `PUT /files/thumbnail`: register an uploaded thumbnail's object key.
### Retries and timeouts
Every request in the library goes through one policy, in `src/retry.ts`. A
request is repeated only when repeating it could produce a different answer:
- `ApiError` with a 5xx status: retried. So are `408` and `429`, the two 4xx
codes that are statements about timing rather than about the request.
- Every other 4xx: not retried. A 404 in particular is an answer, and
`listMissingThumbnails` depends on getting it promptly and once.
- Transport failures — a `fetch` rejection, `ECONNRESET`, `ETIMEDOUT`, a DNS or
TLS failure — and deadline aborts: retried. The errno is looked for in the
error's `cause` chain, because that is where Node's `fetch` puts it.
- A truncated download: retried.
- Anything else, including a secretstream authentication failure that is not
truncation: not retried. The default answer is no. For a backup tool, retrying
a permanent failure spends round trips and delays every remaining file, while
declining to retry a transient one costs a single file that the next run picks
up.
Backoff is exponential with full jitter: the delay before retry _n_ is
`random() * min(maxDelayMs, baseDelayMs * 2 ** (n - 1))`. The exponential term
is the ceiling and the wait is drawn below it, so a client that lost many
parallel downloads to one CDN blip does not send them all again at the same
instant. Defaults, configurable through `ApiClientOptions.retry`:
| Option | Default | Meaning |
| ------------- | ------- | ----------------------------------- |
| `attempts` | `4` | total calls, not retries |
| `baseDelayMs` | `500` | ceiling for the first retry's delay |
| `maxDelayMs` | `10000` | upper bound on that ceiling |
With those defaults a file that is going to fail gives up after at most three
and a half seconds of waiting. `sleep` and `random` are injectable through the
same option, which is how the test suite exercises the whole policy without
waiting.
Two deadlines, applied with `AbortSignal.timeout()` and renewed for each
attempt:
| Option | Default | Applies to |
| ------------------- | -------- | ------------------------------------------- |
| `requestTimeoutMs` | `30000` | `getJSON`, `postJSON`, `putJSON`, `putFile` |
| `downloadTimeoutMs` | `600000` | file and thumbnail body transfers |
They are separate because one number cannot serve both: a value short enough to
keep a hung API call from stalling a backup would cancel a legitimate
multi-gigabyte download. The download deadline covers the body, not just the
headers — `getFileStream` returns as soon as headers arrive, so a deadline that
only guarded the initial request would leave the same hang one layer down.
**Non-idempotent requests are not blindly replayed.** `postJSON` and `putJSON`
reach `/users/srp/create-session`, `/users/two-factor/verify` — which consumes
one of a small number of second-factor attempts — and `/files/thumbnail`. They
are retried only on the three failures that establish no TCP connection to the
server ever existed, so no request byte can have been transmitted: `ENOTFOUND`
and `EAI_AGAIN` (name resolution produced no address) and `ECONNREFUSED` (the
peer refused the connection). A 5xx, a mid-flight reset and a deadline are all
left to the caller, because each of them can happen after the server has already
acted. The routing errnos `EHOSTUNREACH`, `ENETUNREACH` and `ENETDOWN` are
excluded for the same reason, despite looking like connect-time failures: on
Linux an ICMP unreachable arriving mid-flight, or a local interface going down
after the request was written, delivers them on an already-established socket.
They stay retryable for the idempotent calls. `putFile` is exempt: a presigned
PUT stores one whole object at one key in one request, so replaying it has no
partial state to damage.
A download is retried as a whole — request, stream consumption, and decryption —
because a socket reset after the response headers have arrived surfaces in the
download layer rather than in `ApiClient`, and that is the common failure for
multi-megabyte photos over a CDN. The secretstream pull state is not resumable
and these endpoints have no Range support, so a retry starts the file over. The
atomic write stays outside the retry, so a download that needed three attempts
still performs exactly one write and one rename. `runBackup` and
`runMetadataBackup` are unchanged: the retry sits below them, and a file that
fails after exhausting it is still logged, counted, and stepped over.
One imprecision is deliberate and worth knowing about. When a body ends part-way
through a secretstream chunk, Poly1305 fails and carries no framing signal, so a
cut connection and genuinely corrupt bytes are indistinguishable. quak reports
that as truncation, which means it is retried. For a body of more than one chunk
the distinction is real — a chunk that failed while the stream carried on past
it stays an authentication failure and is not retried — but for a single-chunk
body, which is most thumbnails and every small file, a wrong key, server-side
corruption and a mid-chunk cutoff all present alike and all get retried. The
cost is bounded by the attempt count, and it buys never silently keeping a
truncated file.
### Session handling ### Session handling
The `Client` class holds the auth token, master key, secret key, and public key The `Client` class holds the auth token, master key, secret key, and public key
@@ -308,10 +458,10 @@ code is non-zero if any files failed.
## TODO ## TODO
- [ ] Retry policy: no retry on 4xx, exponential backoff on 5xx and network - [x] Retry policy: no retry on 4xx, exponential backoff on 5xx and network
errors errors
- [ ] Update the API reference section below to match the current implementation - [ ] Update the API reference section below to match the current implementation
- [ ] `make docker` green - [x] `make docker` green
- [ ] Tag `v1.0.0` - [ ] Tag `v1.0.0`
Future (desktop client, separate repo): Future (desktop client, separate repo):
@@ -333,7 +483,10 @@ are correct.
The key types and their actual signatures can be found in: The key types and their actual signatures can be found in:
- `src/client.ts`: `Client`, `LoginOptions`, `ClientSnapshot` - `src/client.ts`: `Client`, `LoginOptions`, `ClientSnapshot`
- `src/api/client.ts`: `ApiClient`, `ApiClientOptions`, `ApiError` - `src/api/client.ts`: `ApiClient`, `ApiClientOptions`, `ApiError`,
`StreamOptions`
- `src/errors.ts`: `ApiError`, `TruncatedStreamError`
- `src/retry.ts`: `withRetry`, `isRetryable`, `isSafeToReplay`, `RetryOptions`
- `src/auth/types.ts`: `KeyAttributes`, `SRPAttributes`, - `src/auth/types.ts`: `KeyAttributes`, `SRPAttributes`,
`AuthorizationResponse`, `LoginChallenge` `AuthorizationResponse`, `LoginChallenge`
- `src/model/types.ts`: `Collection`, `EnteFile`, `FileMetadata`, `FileBlob`, - `src/model/types.ts`: `Collection`, `EnteFile`, `FileMetadata`, `FileBlob`,
@@ -367,9 +520,14 @@ documents:
implementation. Tests are the canonical API documentation and must be implementation. Tests are the canonical API documentation and must be
commented thoroughly. `main` is always green. commented thoroughly. `main` is always green.
- **Required checks before every commit:** `make lint` (eslint + prettier check) - **Required checks before every commit:** `make lint` must pass — that is
and `make fmt-check` must pass. The pre-commit hook enforces this. eslint plus the prettier check, and it builds `Dockerfile.lint`, so it needs
`make check` (which also runs tests) must pass before merging to `main`. docker. The pre-commit hook enforces exactly that. `make check` (which also
runs the tests) must pass before merging to `main`. `make fmt-check` is
available for a host-side formatting check on its own, but it is not a
separate requirement: `make lint` already covers it, and running both would
check formatting twice. Never invoke eslint or prettier directly; linting runs
in the container only.
- **Formatting:** prettier with 4-space indents and `proseWrap: always` for - **Formatting:** prettier with 4-space indents and `proseWrap: always` for
markdown. Use `make fmt` to format. Use `yarn` not `npm`. markdown. Use `make fmt` to format. Use `yarn` not `npm`.

71
TODO.md
View File

@@ -14,12 +14,73 @@ pre-1.0
# Next Step # Next Step
Implement the download retry policy from the README TODO: no retry on 4xx, Update the README API reference section to match the current implementation.
exponential backoff on 5xx and network errors. Apply it to file and thumbnail
downloads, cover it with mock-server tests, and update the README TODO checkbox.
# Completed Steps # Completed Steps
- 2026-08-10: Made `lint-once.test.ts` enforce what its header claims. It walked
`make check` only, so it never read `Dockerfile` — the image CI builds through
`script/cibuild` — and a second `prettier --check .` could be added there with
the suite staying green. The walk now also starts at
`.gitea/workflows/check.yml` and follows its `run:` steps, so the graph under
test is the one CI executes rather than the one someone assumed it executes.
The lockfile assertion was a substring check against the whole of
`script/bootstrap`, which has two install sites and so reported the branch the
containers never take; the two branches are now resolved separately and every
`yarn install` in each is required to be `--frozen-lockfile`. Prettier is
counted per occurrence instead of per line, so two invocations chained with
`&&` no longer read as one, and edges are followed on counted lines instead of
being skipped. Every way for the walk to reach nothing — an unknown target, an
unknown script, a missing file, a node with no commands, an unknown node kind
— is a thrown error rather than a quiet zero. Every assertion in the file was
mutation-tested individually.
- 2026-08-10: Stopped `make check` running `prettier --check .` twice. Since
linting moved into Docker, the duplicate was one container pass and one host
pass of the same check: `script/lint` builds `Dockerfile.lint`, which runs
prettier as a build step, and `script/check` then called `script/fmt-check` as
well. The host call is gone from `script/check` and from `script/precommit`;
the container keeps checking formatting, because a successful
`Dockerfile.lint` build is what CI treats as proof of a clean tree, and it is
also what still fails the pre-commit hook on a badly formatted tree.
`script/fmt-check` survives as a standalone entrypoint, whose verdict cannot
drift from the container's. A test walks the invocation graph from each
entrypoint — through the Makefile shims, the `script/` calls and the
`docker build` — and asserts the prettier count, so the duplication cannot
come back unnoticed.
- 2026-08-10: Moved all linting into Docker. `script/lint` builds a new root
`Dockerfile.lint`, which copies the repo into the digest-pinned node image and
runs eslint and prettier as build steps, so a successful build is a clean
lint; no host lint path remains and `yarn lint` is gone from `package.json`. A
fail-closed `LINT_EPOCH` guard stops Docker serving the linter layers from
cache, which is how a lint build returns success in under a second having
linted nothing. The lint stage inside `Dockerfile` and its `COPY --from=lint`
ordering hack are gone: that image now runs `make test` and `make build` only,
because `script/check` calls `script/lint` and running it in a container would
mean docker inside docker. `script/cibuild` builds the lint image first, then
the test and build image.
- 2026-08-09: Made `make docker` green and policy-conformant. Multi-stage
Dockerfile: a lint stage runs `make fmt-check` and `make lint`, and the check
stage takes a `COPY --from=lint` dependency on it before running `make check`
and `make build`. `CHECK_EPOCH` and a fail-closed guard stop Docker serving
those two layers from cache, which is what let a build report success without
running the suite. `script/projectname` says `quak`, so the image is tagged
`quak`; `script/bootstrap` updates apt lists before installing, so a Debian
base works; `.dockerignore` no longer ships the compiled binary, the caches or
agent worktrees into the build context, and keeps `.gitignore` in it for
prettier.
- 2026-08-09: Fixed the TypeScript build. `rootDir` is the repo root, so `bin/`
compiles alongside `src/` instead of failing with TS6059; output is
`dist/src/` and `dist/bin/`, which is where `main`, `types` and `bin.quak` now
point. `script/build` verifies the declared entrypoints exist after the
compiler runs and makes the CLI executable, the Dockerfile runs `make build`
as well as `make check`, and a `quak` script makes the README's
`yarn quak <command>` examples work.
- 2026-08-09: Retry policy: no retry on 4xx (except `408` and `429`),
exponential backoff with full jitter on 5xx, transport failures and truncated
transfers, under per-attempt deadlines that cover the response body as well as
the request. Downloads retry request, stream consumption and decryption as one
unit; `postJSON` and `putJSON` are replayed only when the connection was never
established.
- 2026-08-09: Downloads verify the secretstream terminated on `TAG_FINAL` and - 2026-08-09: Downloads verify the secretstream terminated on `TAG_FINAL` and
write output atomically: a truncated body is rejected instead of landing on write output atomically: a truncated body is rejected instead of landing on
disk as a short file, and plaintext is staged in a sibling temp file and disk as a short file, and plaintext is staged in a sibling temp file and
@@ -45,10 +106,6 @@ downloads, cover it with mock-server tests, and update the README TODO checkbox.
# Future Steps # Future Steps
- Retry policy: no retry on 4xx, exponential backoff on 5xx and network errors
(the Next Step).
- Update the README API reference section to match the current implementation.
- Make `make docker` green.
- Tag v1.0.0. - Tag v1.0.0.
- Future desktop client, separate repo: - Future desktop client, separate repo:
- Electron app skeleton consuming this library. - Electron app skeleton consuming this library.

View File

@@ -10,8 +10,8 @@
"url": "https://git.eeqj.de/sneak/quak.git" "url": "https://git.eeqj.de/sneak/quak.git"
}, },
"type": "module", "type": "module",
"main": "./dist/index.js", "main": "./dist/src/index.js",
"types": "./dist/index.d.ts", "types": "./dist/src/index.d.ts",
"bin": { "bin": {
"quak": "./dist/bin/quak.js" "quak": "./dist/bin/quak.js"
}, },
@@ -21,9 +21,9 @@
"LICENSE" "LICENSE"
], ],
"scripts": { "scripts": {
"build": "tsc", "build": "script/build",
"quak": "node ./dist/bin/quak.js",
"test": "vitest run", "test": "vitest run",
"lint": "eslint .",
"fmt": "prettier --write .", "fmt": "prettier --write .",
"fmt-check": "prettier --check ." "fmt-check": "prettier --check ."
}, },

View File

@@ -19,6 +19,7 @@ YARN_VERSION="1.22.22"
PKGMGR="" PKGMGR=""
SUDO="" SUDO=""
APT_UPDATED=""
detect_pkgmgr() { detect_pkgmgr() {
[ -n "$PKGMGR" ] && return 0 [ -n "$PKGMGR" ] && return 0
@@ -42,12 +43,24 @@ detect_pkgmgr() {
fi fi
} }
# A fresh Debian image ships no package lists at all, so apt-get install
# fails with "E: Unable to locate package make" until they are fetched.
# Done once per run, since the lists do not go stale mid-bootstrap.
apt_update_once() {
[ -n "$APT_UPDATED" ] && return 0
$SUDO env DEBIAN_FRONTEND=noninteractive apt-get update
APT_UPDATED="yes"
}
# pkg_install <nix-attr> <apt-pkg> <brew-formula> <apk-pkg> # pkg_install <nix-attr> <apt-pkg> <brew-formula> <apk-pkg>
pkg_install() { pkg_install() {
detect_pkgmgr detect_pkgmgr
case "$PKGMGR" in case "$PKGMGR" in
nix) nix-env -iA "nixpkgs.$1" ;; nix) nix-env -iA "nixpkgs.$1" ;;
apt) $SUDO env DEBIAN_FRONTEND=noninteractive apt-get install -y "$2" ;; apt)
apt_update_once
$SUDO env DEBIAN_FRONTEND=noninteractive apt-get install -y "$2"
;;
brew) brew install "$3" ;; brew) brew install "$3" ;;
apk) apk add --no-cache "$4" ;; apk) apk add --no-cache "$4" ;;
esac esac

54
script/build Executable file
View File

@@ -0,0 +1,54 @@
#!/bin/sh
# script/build: compile the TypeScript sources into dist/, then verify that
# the artifacts package.json advertises are among the files the compiler
# actually wrote. tsc reports success by exit status alone and knows nothing
# about the manifest, so without this step a green build can still ship a
# package whose main, types or bin resolve to nothing. Our own extension to
# scripts-to-rule-them-all.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
# Reads package.json, requires every declared entrypoint to exist, requires
# each bin entry to have kept its shebang, and makes the bin entries
# executable: tsc copies the shebang through but not the mode bits, and an
# installed CLI has to be runnable.
verify_entrypoints() {
node -e '
const { readFileSync, statSync, chmodSync } = require("node:fs");
const pkg = JSON.parse(readFileSync("package.json", "utf-8"));
const bins = Object.values(pkg.bin ?? {});
const fail = (message) => {
console.error("build: " + message);
process.exit(1);
};
for (const declared of [pkg.main, pkg.types, ...bins]) {
if (!declared) continue;
try {
statSync(declared);
} catch {
fail("package.json declares " + declared + ", which the build did not produce");
}
console.log("build: verified " + declared);
}
for (const bin of bins) {
const firstLine = readFileSync(bin, "utf-8").split("\n")[0];
if (!firstLine.startsWith("#!")) {
fail(bin + " lost its shebang, so it cannot be executed directly");
}
chmodSync(bin, 0o755);
console.log("build: " + bin + " is executable (" + firstLine + ")");
}
'
}
main() {
cd "$ROOT"
yarn run tsc
verify_entrypoints
}
main "$@"

View File

@@ -1,6 +1,19 @@
#!/bin/sh #!/bin/sh
# script/check: run all checks (test, lint, fmt-check). Our own # script/check: run all checks (test, lint). Our own extension to
# extension to scripts-to-rule-them-all. Must not modify any files. # scripts-to-rule-them-all. Must not modify any files.
#
# The formatting check is part of lint, not a step of its own:
# script/lint builds Dockerfile.lint, which runs eslint AND
# `prettier --check .` as build steps. Calling script/fmt-check here as
# well would run prettier a second time over the same tree for the same
# verdict — the weaker of the two, since the host toolchain is whatever
# the working tree happens to have installed while the container's is
# digest-pinned. script/fmt-check remains a standalone entrypoint for
# asking the formatting question by itself.
#
# script/lint builds Dockerfile.lint, so this script requires docker and
# must never be run from inside a container: that is why the Dockerfile
# image runs script/test and script/build rather than this.
set -eu set -eu
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd -P)" SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd -P)"
@@ -8,7 +21,6 @@ SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd -P)"
main() { main() {
"$SCRIPT_DIR/test" "$SCRIPT_DIR/test"
"$SCRIPT_DIR/lint" "$SCRIPT_DIR/lint"
"$SCRIPT_DIR/fmt-check"
} }
main "$@" main "$@"

View File

@@ -1,13 +1,24 @@
#!/bin/sh #!/bin/sh
# script/cibuild: run the CI build. The Dockerfile runs script/check, so # script/cibuild: run the CI build, which is both images in a defined order.
# a successful build implies all checks pass. #
# First script/lint, which builds Dockerfile.lint and is the one and only
# place linting happens — it goes first so a lint failure is reported before
# the slower suite runs. Then the Dockerfile image, which runs script/test
# and script/build. CHECK_EPOCH and LINT_EPOCH differ on every invocation, so
# neither the linters nor the suite can be served from Docker's cache: a
# green build here means the checks ran now, not that a previous run was
# remembered. The layers below the epochs (bootstrap, yarn install) are
# unaffected and stay cached. A build that omits the arguments fails by
# design.
set -eu set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)" SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd -P)"
ROOT="$(cd "$SCRIPT_DIR/.." && pwd -P)"
main() { main() {
cd "$ROOT" cd "$ROOT"
docker build . "$SCRIPT_DIR/lint"
docker build --build-arg CHECK_EPOCH="$(date +%s)" .
} }
main "$@" main "$@"

View File

@@ -1,6 +1,10 @@
#!/bin/sh #!/bin/sh
# script/docker: build the Docker image tagged with the project name. # script/docker: build the Docker image tagged with the project name.
# Identical in all repos; the tag comes from script/projectname. # Identical in all repos; the tag comes from script/projectname.
# CHECK_EPOCH is passed for the same reason script/cibuild passes it: the
# Dockerfile refuses to build without it, so that no path to an image can
# quietly serve the test and build layers from cache. This builds the test
# and build image only; linting is a separate image, built by script/lint.
set -eu set -eu
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd -P)" SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd -P)"
@@ -8,7 +12,8 @@ ROOT="$(cd "$SCRIPT_DIR/.." && pwd -P)"
main() { main() {
cd "$ROOT" cd "$ROOT"
docker build -t "$("$SCRIPT_DIR/projectname")" . docker build --build-arg CHECK_EPOCH="$(date +%s)" \
-t "$("$SCRIPT_DIR/projectname")" .
} }
main "$@" main "$@"

View File

@@ -1,13 +1,24 @@
#!/bin/sh #!/bin/sh
# script/lint: run the linter (eslint plus a prettier check). # script/lint: run the linters. eslint and prettier are never run against
# the working tree from here: linting runs via docker only, one way,
# everywhere — script/lint builds Dockerfile.lint, which COPYs the repo into
# the pinned node image and runs the linters as build steps. That works even
# when the docker daemon is remote and bind mounts are impossible.
#
# LINT_EPOCH is passed on every invocation because no lint cache is wanted:
# on an unchanged tree Docker would otherwise serve the linter layers, having
# linted nothing, and still exit 0. Dockerfile.lint refuses to build without
# the argument, so no path to a lint result can quietly come from cache.
#
# Nothing that runs inside a container may call this script; see the header
# of Dockerfile.
set -eu set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)" ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
main() { main() {
cd "$ROOT" cd "$ROOT"
yarn run eslint . docker build --build-arg LINT_EPOCH="$(date +%s)" -f Dockerfile.lint .
yarn run prettier --check .
} }
main "$@" main "$@"

View File

@@ -2,17 +2,25 @@
# script/precommit: run by the git pre-commit hook; fails the commit if # script/precommit: run by the git pre-commit hook; fails the commit if
# checks fail. Our own extension to scripts-to-rule-them-all. # checks fail. Our own extension to scripts-to-rule-them-all.
# #
# Runs lint and fmt-check but deliberately NOT the tests, so the TDD # Runs lint but deliberately NOT the tests, so the TDD red-phase commit
# red-phase commit (failing tests, no implementation yet) can land. CI # (failing tests, no implementation yet) can land. CI runs
# runs make check via docker build, which catches any branch that # script/cibuild, which builds both images and so catches any branch
# ships red. # that ships red.
#
# The formatting check is still enforced here, because script/lint is a
# build of Dockerfile.lint and that runs `prettier --check .` as a build
# step: a badly formatted tree fails this hook, and therefore the
# commit. Calling script/fmt-check as well would only run prettier a
# second time over the same tree for the same verdict.
#
# script/lint is a docker build (Dockerfile.lint); docker is required to
# commit, which is the point of linting one way, everywhere.
set -eu set -eu
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd -P)" SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd -P)"
main() { main() {
"$SCRIPT_DIR/lint" "$SCRIPT_DIR/lint"
"$SCRIPT_DIR/fmt-check"
} }
main "$@" main "$@"

View File

@@ -6,7 +6,7 @@
set -eu set -eu
main() { main() {
echo "quack" echo "quak"
} }
main "$@" main "$@"

View File

@@ -1,8 +1,33 @@
import { ApiError } from "../errors.js";
import {
isSafeToReplay,
resolveRetryOptions,
withRetry,
type ResolvedRetryOptions,
type RetryOptions,
} from "../retry.js";
// `ApiError` is defined in `src/errors.ts` so that the retry classifier can
// recognise it without importing this module, which imports the classifier.
// It is re-exported here because this is where callers have always imported it
// from, and it must remain one class: a second copy would make `instanceof`
// fail in the classifier and every 5xx would look permanent.
export { ApiError };
const DEFAULT_API_ORIGIN = "https://api.ente.io"; const DEFAULT_API_ORIGIN = "https://api.ente.io";
const DEFAULT_FILES_ORIGIN = "https://files.ente.io"; const DEFAULT_FILES_ORIGIN = "https://files.ente.io";
const DEFAULT_THUMBS_ORIGIN = "https://thumbnails.ente.io"; const DEFAULT_THUMBS_ORIGIN = "https://thumbnails.ente.io";
const CLIENT_PACKAGE = "berlin.sneak.quak"; const CLIENT_PACKAGE = "berlin.sneak.quak";
// Two deadlines rather than one, because a single number cannot serve both
// jobs. Thirty seconds is generous for a JSON call and short enough that a
// hung API connection cannot stall a backup for long. A file body is a
// different shape of problem: the deadline has to cover the whole transfer,
// which for a large video on a slow link is minutes, so a value sane for JSON
// would cancel legitimate downloads.
export const DEFAULT_REQUEST_TIMEOUT_MS = 30_000;
export const DEFAULT_DOWNLOAD_TIMEOUT_MS = 600_000;
export interface ApiClientOptions { export interface ApiClientOptions {
apiOrigin?: string; apiOrigin?: string;
filesOrigin?: string; filesOrigin?: string;
@@ -10,26 +35,68 @@ export interface ApiClientOptions {
authToken?: string; authToken?: string;
fetch?: typeof globalThis.fetch; fetch?: typeof globalThis.fetch;
userAgent?: string; userAgent?: string;
retry?: RetryOptions;
requestTimeoutMs?: number;
downloadTimeoutMs?: number;
} }
export class ApiError extends Error { export interface StreamOptions {
readonly status: number; // Opt out of this client's own retry. Exactly one caller wants that: the
readonly code?: string; // download layer, which retries the request, the stream consumption and
readonly requestID?: string; // the decryption as one unit. Leaving both layers enabled would multiply
readonly body?: unknown; // the budgets — four attempts each becoming sixteen requests per file.
constructor( retry?: boolean;
message: string,
status: number,
opts?: { code?: string; requestID?: string; body?: unknown },
) {
super(message);
this.name = "ApiError";
this.status = status;
this.code = opts?.code;
this.requestID = opts?.requestID;
this.body = opts?.body;
} }
// Enforce a deadline over a response body, not merely over its headers.
//
// `getFileStream` returns as soon as headers arrive; the bytes are pulled
// later, in the download layer. Whether the signal passed to `fetch` also
// tears down the body afterwards is up to the fetch implementation, so this
// wrapper makes it a property of quak instead: every read races the signal,
// and an abort errors the stream with the abort reason — which the retry
// classifier recognises.
const deadlineStream = (
body: ReadableStream<Uint8Array>,
signal: AbortSignal,
): ReadableStream<Uint8Array> => {
const reader = body.getReader();
let rejectOnAbort: (reason: unknown) => void = () => undefined;
const aborted = new Promise<never>((_resolve, reject) => {
rejectOnAbort = reject;
});
// The abort may fire when nothing is awaiting `aborted` — after the body
// has been read in full, say. Without this, that rejection would surface
// as an unhandled rejection and take the process down.
void aborted.catch(() => undefined);
const onAbort = (): void => rejectOnAbort(signal.reason);
if (signal.aborted) onAbort();
else signal.addEventListener("abort", onAbort, { once: true });
const release = (): void => signal.removeEventListener("abort", onAbort);
return new ReadableStream<Uint8Array>({
async pull(controller) {
try {
const next = await Promise.race([reader.read(), aborted]);
if (next.done) {
release();
controller.close();
return;
} }
controller.enqueue(next.value);
} catch (err) {
release();
await reader.cancel(err).catch(() => undefined);
throw err;
}
},
async cancel(reason) {
release();
await reader.cancel(reason);
},
});
};
export class ApiClient { export class ApiClient {
private readonly apiOrigin: string; private readonly apiOrigin: string;
@@ -37,6 +104,9 @@ export class ApiClient {
private readonly filesOrigin: string; private readonly filesOrigin: string;
private readonly thumbsOrigin: string; private readonly thumbsOrigin: string;
private readonly _fetch: typeof globalThis.fetch; private readonly _fetch: typeof globalThis.fetch;
private readonly retry: ResolvedRetryOptions;
private readonly requestTimeoutMs: number;
private readonly downloadTimeoutMs: number;
private token: string | undefined; private token: string | undefined;
constructor(opts?: ApiClientOptions) { constructor(opts?: ApiClientOptions) {
@@ -53,6 +123,11 @@ export class ApiClient {
opts?.thumbsOrigin ?? DEFAULT_THUMBS_ORIGIN opts?.thumbsOrigin ?? DEFAULT_THUMBS_ORIGIN
).replace(/\/+$/, ""); ).replace(/\/+$/, "");
this._fetch = opts?.fetch ?? globalThis.fetch; this._fetch = opts?.fetch ?? globalThis.fetch;
this.retry = resolveRetryOptions(opts?.retry);
this.requestTimeoutMs =
opts?.requestTimeoutMs ?? DEFAULT_REQUEST_TIMEOUT_MS;
this.downloadTimeoutMs =
opts?.downloadTimeoutMs ?? DEFAULT_DOWNLOAD_TIMEOUT_MS;
this.token = opts?.authToken; this.token = opts?.authToken;
} }
@@ -64,6 +139,13 @@ export class ApiClient {
this.token = undefined; this.token = undefined;
} }
// The policy this client was configured with, so that a caller wrapping a
// whole operation in its own `withRetry` — the download layer — runs under
// the same settings rather than under the library defaults.
getRetryOptions(): ResolvedRetryOptions {
return this.retry;
}
private headers(extra?: Record<string, string>): Record<string, string> { private headers(extra?: Record<string, string>): Record<string, string> {
const h: Record<string, string> = { const h: Record<string, string> = {
"X-Client-Package": CLIENT_PACKAGE, "X-Client-Package": CLIENT_PACKAGE,
@@ -128,38 +210,54 @@ export class ApiClient {
} }
} }
} }
// A GET changes nothing, so it is retried under the full policy.
return withRetry(async () => {
const resp = await this._fetch(url.href, { const resp = await this._fetch(url.href, {
method: "GET", method: "GET",
headers: this.headers(), headers: this.headers(),
signal: AbortSignal.timeout(this.requestTimeoutMs),
}); });
await this.throwIfError(resp); await this.throwIfError(resp);
return (await resp.json()) as T; return (await resp.json()) as T;
}, this.retry);
} }
async postJSON<T>(path: string, body: unknown): Promise<T> { async postJSON<T>(path: string, body: unknown): Promise<T> {
const url = `${this.apiOrigin}${path}`; const url = `${this.apiOrigin}${path}`;
// Idempotency: this reaches `/users/srp/create-session`,
// `/users/two-factor/verify` and `/users/ott`, all of which change
// server state — verifying a second factor consumes one of a small
// number of attempts. So a POST is replayed only on a failure that
// establishes no TCP connection to the server ever existed: DNS
// produced no address, or the peer refused the connection. A 5xx, a
// mid-flight reset, a routing errno (which Linux also delivers on an
// established socket) and a timeout are all left to the caller,
// because each of them can occur after the server has already acted.
return withRetry(
async () => {
const resp = await this._fetch(url, { const resp = await this._fetch(url, {
method: "POST", method: "POST",
headers: this.headers({ "Content-Type": "application/json" }), headers: this.headers({
"Content-Type": "application/json",
}),
body: JSON.stringify(body), body: JSON.stringify(body),
signal: AbortSignal.timeout(this.requestTimeoutMs),
}); });
await this.throwIfError(resp); await this.throwIfError(resp);
return (await resp.json()) as T; return (await resp.json()) as T;
},
{ ...this.retry, isRetryable: isSafeToReplay },
);
} }
async getFileStream(fileID: number): Promise<ReadableStream<Uint8Array>> { async getFileStream(
fileID: number,
opts?: StreamOptions,
): Promise<ReadableStream<Uint8Array>> {
const url = this.isCustomOrigin const url = this.isCustomOrigin
? `${this.apiOrigin}/files/download/${fileID}` ? `${this.apiOrigin}/files/download/${fileID}`
: `${this.filesOrigin}/?fileID=${fileID}`; : `${this.filesOrigin}/?fileID=${fileID}`;
const resp = await this._fetch(url, { return this.streamRequest(url, opts);
method: "GET",
headers: this.headers(),
});
await this.throwIfError(resp);
if (!resp.body) {
throw new Error("response body is null");
}
return resp.body;
} }
async getUploadURL( async getUploadURL(
@@ -173,6 +271,11 @@ export class ApiClient {
} }
async putFile(presignedURL: string, data: Uint8Array): Promise<void> { async putFile(presignedURL: string, data: Uint8Array): Promise<void> {
// Idempotent despite being a write: a presigned PUT stores one whole
// object at one key in one request, so replaying it either overwrites
// the same bytes or lands them for the first time. There is no partial
// state to protect, hence the full policy rather than the POST rule.
await withRetry(async () => {
const resp = await this._fetch(presignedURL, { const resp = await this._fetch(presignedURL, {
method: "PUT", method: "PUT",
headers: { headers: {
@@ -180,21 +283,40 @@ export class ApiClient {
"Content-Length": String(data.length), "Content-Length": String(data.length),
}, },
body: data, body: data,
signal: AbortSignal.timeout(this.requestTimeoutMs),
}); });
if (!resp.ok) { if (!resp.ok) {
throw new Error(`PUT to presigned URL failed: HTTP ${resp.status}`); // An ApiError, not a bare Error: without the status on the
// error the upload path cannot be classified at all, and a
// 503 from S3 would be indistinguishable from a bug.
throw new ApiError(
`PUT to presigned URL failed: HTTP ${resp.status}`,
resp.status,
);
} }
}, this.retry);
} }
async putJSON<T>(path: string, body: unknown): Promise<T> { async putJSON<T>(path: string, body: unknown): Promise<T> {
const url = `${this.apiOrigin}${path}`; const url = `${this.apiOrigin}${path}`;
// Same idempotency rule as `postJSON`, for the same reason: this
// reaches `/files/thumbnail`, which registers an uploaded thumbnail
// against a file.
return withRetry(
async () => {
const resp = await this._fetch(url, { const resp = await this._fetch(url, {
method: "PUT", method: "PUT",
headers: this.headers({ "Content-Type": "application/json" }), headers: this.headers({
"Content-Type": "application/json",
}),
body: JSON.stringify(body), body: JSON.stringify(body),
signal: AbortSignal.timeout(this.requestTimeoutMs),
}); });
await this.throwIfError(resp); await this.throwIfError(resp);
return (await resp.json()) as T; return (await resp.json()) as T;
},
{ ...this.retry, isRetryable: isSafeToReplay },
);
} }
async updateThumbnail( async updateThumbnail(
@@ -210,18 +332,36 @@ export class ApiClient {
async getThumbnailStream( async getThumbnailStream(
fileID: number, fileID: number,
opts?: StreamOptions,
): Promise<ReadableStream<Uint8Array>> { ): Promise<ReadableStream<Uint8Array>> {
const url = this.isCustomOrigin const url = this.isCustomOrigin
? `${this.apiOrigin}/files/preview/${fileID}` ? `${this.apiOrigin}/files/preview/${fileID}`
: `${this.thumbsOrigin}/?fileID=${fileID}`; : `${this.thumbsOrigin}/?fileID=${fileID}`;
return this.streamRequest(url, opts);
}
private async streamRequest(
url: string,
opts?: StreamOptions,
): Promise<ReadableStream<Uint8Array>> {
const once = async (): Promise<ReadableStream<Uint8Array>> => {
// A fresh deadline per attempt, so a retry gets the whole budget
// rather than the remainder of the one that just expired.
const signal = AbortSignal.timeout(this.downloadTimeoutMs);
const resp = await this._fetch(url, { const resp = await this._fetch(url, {
method: "GET", method: "GET",
headers: this.headers(), headers: this.headers(),
signal,
}); });
await this.throwIfError(resp); await this.throwIfError(resp);
if (!resp.body) { if (!resp.body) {
throw new Error("response body is null"); // Carries the status, and is not retryable: a response that
// arrived without a body is malformed, and asking again
// produces the same malformed response.
throw new ApiError("response body is null", resp.status);
} }
return resp.body; return deadlineStream(resp.body, signal);
};
return opts?.retry === false ? once() : withRetry(once, this.retry);
} }
} }

View File

@@ -1,4 +1,4 @@
import sodium from "libsodium-wrappers-sumo"; import sodium, { type StateAddress } from "libsodium-wrappers-sumo";
// Plaintext chunk size used by Ente for file content streams. Hard-coded by // Plaintext chunk size used by Ente for file content streams. Hard-coded by
// the server; clients must match. // the server; clients must match.
@@ -47,8 +47,10 @@ export const encryptBlob = (
}; };
// Opaque handle to libsodium's secretstream pull state. Threaded through // Opaque handle to libsodium's secretstream pull state. Threaded through
// successive pullStreamChunk calls. // successive pullStreamChunk calls. The type comes from the named export;
export type StreamPullState = sodium.StateAddress; // the default import is the module's value side and has no type namespace
// under it.
export type StreamPullState = StateAddress;
// Initialise a pull stream from the per-file decryption header and the // Initialise a pull stream from the per-file decryption header and the
// per-file key. // per-file key.
@@ -84,9 +86,13 @@ export const pullStreamChunk = (
state: StreamPullState, state: StreamPullState,
ciphertext: Uint8Array, ciphertext: Uint8Array,
): { plaintext: Uint8Array; tag: number } => { ): { plaintext: Uint8Array; tag: number } => {
// The additional-data argument is not optional in libsodium's signature.
// null is "no additional data", matching the null passed on the push side
// in encryptBlob; Ente's file streams carry none.
const result = sodium.crypto_secretstream_xchacha20poly1305_pull( const result = sodium.crypto_secretstream_xchacha20poly1305_pull(
state, state,
ciphertext, ciphertext,
null,
); );
if (result === false) { if (result === false) {
throw new Error("secretstream chunk authentication failed"); throw new Error("secretstream chunk authentication failed");

View File

@@ -9,6 +9,8 @@ import {
STREAM_CHUNK_SIZE, STREAM_CHUNK_SIZE,
streamTagFinal, streamTagFinal,
} from "../crypto/index.js"; } from "../crypto/index.js";
import { TruncatedStreamError } from "../errors.js";
import { withRetry } from "../retry.js";
import type { ApiClient } from "../api/client.js"; import type { ApiClient } from "../api/client.js";
import type { EnteFile } from "../model/types.js"; import type { EnteFile } from "../model/types.js";
@@ -65,7 +67,7 @@ const streamDecrypt = async (
try { try {
pulled = pullStreamChunk(state, buffer); pulled = pullStreamChunk(state, buffer);
} catch (err) { } catch (err) {
throw new Error( throw new TruncatedStreamError(
`download: stream truncated: response body ended with ${buffer.length} trailing bytes that did not authenticate as a final chunk (transfer stopped mid-chunk, or the data is corrupt)`, `download: stream truncated: response body ended with ${buffer.length} trailing bytes that did not authenticate as a final chunk (transfer stopped mid-chunk, or the data is corrupt)`,
{ cause: err }, { cause: err },
); );
@@ -85,13 +87,13 @@ const streamDecrypt = async (
// Returning a short plaintext here would put a corrupt file on disk that // Returning a short plaintext here would put a corrupt file on disk that
// later backup runs would treat as complete. // later backup runs would treat as complete.
if (chunksPulled === 0) { if (chunksPulled === 0) {
throw new Error( throw new TruncatedStreamError(
"download: stream truncated: response body contained no secretstream chunks", "download: stream truncated: response body contained no secretstream chunks",
); );
} }
const tagFinal = streamTagFinal(); const tagFinal = streamTagFinal();
if (lastTag !== tagFinal) { if (lastTag !== tagFinal) {
throw new Error( throw new TruncatedStreamError(
`download: stream truncated: last chunk tag ${lastTag}, expected TAG_FINAL (${tagFinal})`, `download: stream truncated: last chunk tag ${lastTag}, expected TAG_FINAL (${tagFinal})`,
); );
} }
@@ -128,15 +130,49 @@ const writeAtomic = async (
} }
}; };
// Fetch a stream and decrypt it, retrying the whole sequence.
//
// The request is only the first third of a download. `getXStream` returns as
// soon as headers arrive, and the bytes are pulled here, so a socket reset
// mid-body — the dominant failure mode for multi-megabyte photos over a CDN —
// throws in `streamDecrypt` and never reaches `ApiClient` at all. Retrying the
// request alone would miss it entirely.
//
// The client's own retry is therefore switched off for these two calls: with
// both layers active the budgets would multiply, and the library default of
// four attempts would mean sixteen requests for one file. The policy comes
// from the client so a caller that configured one gets it here too.
//
// A retry starts the file over from byte zero: the secretstream pull state is
// not resumable and there is no Range support on these endpoints.
const fetchAndDecrypt = async (
api: ApiClient,
openStream: () => Promise<ReadableStream<Uint8Array>>,
header: Uint8Array,
key: Uint8Array,
): Promise<Uint8Array> =>
withRetry(async () => {
const stream = await openStream();
return streamDecrypt(stream, header, key);
}, api.getRetryOptions());
export const downloadFile = async ( export const downloadFile = async (
api: ApiClient, api: ApiClient,
file: EnteFile, file: EnteFile,
outPath?: string, outPath?: string,
): Promise<DownloadResult> => { ): Promise<DownloadResult> => {
const resolvedPath = outPath ?? file.metadata.title; const resolvedPath = outPath ?? file.metadata.title;
const stream = await api.getFileStream(file.id);
const header = fromBase64(file.file.decryptionHeader); const header = fromBase64(file.file.decryptionHeader);
const plaintext = await streamDecrypt(stream, header, file.key); const plaintext = await fetchAndDecrypt(
api,
() => api.getFileStream(file.id, { retry: false }),
header,
file.key,
);
// Outside the retry, deliberately: only the attempt that produced a
// complete, authenticated plaintext gets to stage a temporary file, so a
// download that needed three tries still performs exactly one write and
// one rename.
await writeAtomic(resolvedPath, plaintext); await writeAtomic(resolvedPath, plaintext);
return { path: resolvedPath, bytesWritten: plaintext.length }; return { path: resolvedPath, bytesWritten: plaintext.length };
}; };
@@ -147,9 +183,13 @@ export const downloadThumbnail = async (
outPath?: string, outPath?: string,
): Promise<DownloadResult> => { ): Promise<DownloadResult> => {
const resolvedPath = outPath ?? `thumb_${file.metadata.title}`; const resolvedPath = outPath ?? `thumb_${file.metadata.title}`;
const stream = await api.getThumbnailStream(file.id);
const header = fromBase64(file.thumbnail.decryptionHeader); const header = fromBase64(file.thumbnail.decryptionHeader);
const plaintext = await streamDecrypt(stream, header, file.key); const plaintext = await fetchAndDecrypt(
api,
() => api.getThumbnailStream(file.id, { retry: false }),
header,
file.key,
);
await writeAtomic(resolvedPath, plaintext); await writeAtomic(resolvedPath, plaintext);
return { path: resolvedPath, bytesWritten: plaintext.length }; return { path: resolvedPath, bytesWritten: plaintext.length };
}; };

45
src/errors.ts Normal file
View File

@@ -0,0 +1,45 @@
// Error types that more than one layer of quak needs to recognise.
//
// They live here rather than beside the code that throws them so that the
// retry classifier can identify them without importing the HTTP client or the
// download layer — both of which import the classifier. `ApiError` is
// re-exported from `src/api/client.ts`, which is where callers have always
// imported it from and where it still belongs conceptually.
export class ApiError extends Error {
readonly status: number;
readonly code?: string;
readonly requestID?: string;
readonly body?: unknown;
constructor(
message: string,
status: number,
opts?: { code?: string; requestID?: string; body?: unknown },
) {
super(message);
this.name = "ApiError";
this.status = status;
this.code = opts?.code;
this.requestID = opts?.requestID;
this.body = opts?.body;
}
}
// A response body that ended before the secretstream did.
//
// This is a type rather than a message prefix because it is a decision, not a
// diagnostic: the retry policy asks "was this a short transfer?" and acts on
// the answer. Matching on the wording of an error message would make the next
// person to reword a diagnostic silently turn every truncated download into a
// permanent failure, and the failure would look like a corrupt file rather
// than like a bug.
//
// `cause` carries the underlying authentication failure on the one path where
// there is one — a body that stopped part-way through a chunk, which Poly1305
// cannot distinguish from corruption.
export class TruncatedStreamError extends Error {
constructor(message: string, opts?: ErrorOptions) {
super(message, opts);
this.name = "TruncatedStreamError";
}
}

View File

@@ -1,7 +1,25 @@
export const VERSION = "0.0.0"; export const VERSION = "0.0.0";
export { Client, type LoginOptions, type ClientSnapshot } from "./client.js"; export { Client, type LoginOptions, type ClientSnapshot } from "./client.js";
export { ApiClient, ApiError, type ApiClientOptions } from "./api/client.js"; export {
ApiClient,
ApiError,
DEFAULT_DOWNLOAD_TIMEOUT_MS,
DEFAULT_REQUEST_TIMEOUT_MS,
type ApiClientOptions,
type StreamOptions,
} from "./api/client.js";
export { TruncatedStreamError } from "./errors.js";
export {
DEFAULT_RETRY_OPTIONS,
isRetryable,
isSafeToReplay,
resolveRetryOptions,
withRetry,
type ResolvedRetryOptions,
type RetryOptions,
type WithRetryOptions,
} from "./retry.js";
export { unwrapAuth, type UnwrapResult } from "./auth/unwrap.js"; export { unwrapAuth, type UnwrapResult } from "./auth/unwrap.js";
export { export {
beginLogin, beginLogin,

201
src/retry.ts Normal file
View File

@@ -0,0 +1,201 @@
// The retry policy shared by every network operation in quak.
//
// Two independent pieces: a classifier that decides whether an error is worth
// another attempt, and a loop that acts on that decision with exponential
// backoff. Keeping them apart is what lets the non-idempotent call sites reuse
// the loop under a stricter question (see `isSafeToReplay`).
//
// The classifier's default answer is no. For a backup tool, retrying a
// permanent failure spends round trips and delays every remaining file, while
// declining to retry a transient one costs a single file that the next run
// picks up anyway. When the evidence is ambiguous, fail fast.
import { ApiError, TruncatedStreamError } from "./errors.js";
export interface RetryOptions {
// Total calls, not retries: `attempts: 1` disables retrying.
attempts?: number;
// Ceiling for the first retry's delay; doubles with each retry.
baseDelayMs?: number;
// Upper bound on that ceiling, so a long outage settles into a steady
// poll instead of growing without limit.
maxDelayMs?: number;
// Injected so tests exercise the whole policy without waiting.
sleep?: (ms: number) => Promise<void>;
// Injected so the jitter is reproducible under test.
random?: () => number;
}
export type ResolvedRetryOptions = Required<RetryOptions>;
export const DEFAULT_RETRY_OPTIONS: ResolvedRetryOptions = {
attempts: 4,
baseDelayMs: 500,
maxDelayMs: 10_000,
sleep: (ms: number): Promise<void> =>
new Promise((resolve) => setTimeout(resolve, ms)),
random: Math.random,
};
export const resolveRetryOptions = (
opts?: RetryOptions,
): ResolvedRetryOptions => ({
attempts: opts?.attempts ?? DEFAULT_RETRY_OPTIONS.attempts,
baseDelayMs: opts?.baseDelayMs ?? DEFAULT_RETRY_OPTIONS.baseDelayMs,
maxDelayMs: opts?.maxDelayMs ?? DEFAULT_RETRY_OPTIONS.maxDelayMs,
sleep: opts?.sleep ?? DEFAULT_RETRY_OPTIONS.sleep,
random: opts?.random ?? DEFAULT_RETRY_OPTIONS.random,
});
// Transport failures: the request did not complete, for reasons below HTTP.
const TRANSPORT_CODES = new Set([
"ECONNRESET",
"ECONNABORTED",
"ETIMEDOUT",
"EPIPE",
"ENOTFOUND",
"EAI_AGAIN",
"ECONNREFUSED",
"EHOSTUNREACH",
"ENETUNREACH",
"ENETRESET",
"ENETDOWN",
]);
// The subset of the above that can only be reported before a TCP connection
// exists, and therefore before any request byte could have been written: name
// resolution produced no address (`ENOTFOUND`, `EAI_AGAIN`) or the peer
// refused the connection with an RST to the SYN (`ECONNREFUSED`).
//
// The routing errnos — `EHOSTUNREACH`, `ENETUNREACH`, `ENETDOWN` — are
// deliberately absent even though they look like connect-time failures. On
// Linux they are also delivered on an already-established socket: an ICMP
// destination-unreachable arriving mid-flight sets the socket error and the
// next read or write returns it, and a local interface going down after the
// request was fully written surfaces the same way. In those cases the server
// may already have received and acted on the request, which is exactly the
// ambiguity this set exists to exclude. They stay in `TRANSPORT_CODES`, so
// they remain retryable for idempotent calls; only replay eligibility is
// narrowed. See `isSafeToReplay`.
const CONNECT_CODES = new Set(["ENOTFOUND", "EAI_AGAIN", "ECONNREFUSED"]);
// `cause` is an arbitrary user-settable property and nothing prevents it from
// forming a cycle, so the walk is bounded. Hanging the process would be a
// worse outcome than any misclassification.
const MAX_CAUSE_DEPTH = 8;
// Collect every `code` in an error's cause chain. undici does not put the
// errno on the error it throws — it hangs the underlying socket error off
// `cause`, sometimes more than one level down — so a classifier that only read
// the top-level error would see a bare `Error` and call every dropped
// connection permanent.
const causeCodes = (err: unknown): string[] => {
const codes: string[] = [];
let current: unknown = err;
for (let depth = 0; depth < MAX_CAUSE_DEPTH; depth++) {
if (current === null || typeof current !== "object") break;
const { code, cause } = current as { code?: unknown; cause?: unknown };
if (typeof code === "string") codes.push(code);
if (cause === current) break;
current = cause;
}
return codes;
};
const isAbort = (err: unknown): boolean => {
if (err === null || typeof err !== "object") return false;
const { name } = err as { name?: unknown };
return name === "AbortError" || name === "TimeoutError";
};
// Is another attempt capable of producing a different answer?
//
// A note on the truncation case, recorded on issue #2. `TruncatedStreamError`
// is retried, and for a body of more than one chunk that is exactly right: a
// chunk that failed to authenticate while the stream carried on past it is
// corruption, stays an ordinary authentication failure, and is not retried.
//
// For a single-chunk body — most thumbnails, every small file — the split is
// not achievable. A wrong file key, server-side corruption and a connection
// cut mid-chunk are cryptographically identical: Poly1305 fails and carries no
// framing signal. All three are reported as truncation and therefore retried.
// That imprecision is deliberate and bounded by the attempt count: one wasted
// round trip on a genuinely corrupt file is a fair price for never silently
// keeping a truncated one, and the alternative — treating a single-chunk
// authentication failure as permanent — would reintroduce exactly the
// silent-corruption risk the truncation check exists to remove.
export const isRetryable = (err: unknown): boolean => {
if (err instanceof ApiError) {
// 408 and 429 are the two 4xx codes that are statements about timing
// rather than about the request, and backoff is the right answer to
// both. Every other 4xx will answer the same way however often it is
// asked.
if (err.status === 408 || err.status === 429) return true;
return err.status >= 500 && err.status <= 599;
}
if (err instanceof TruncatedStreamError) return true;
// quak only ever aborts a request on its own deadline, so an abort means
// this attempt ran out of time — which a later one may not.
if (isAbort(err)) return true;
// Node's fetch rejects with `TypeError: fetch failed` for everything below
// HTTP: DNS failure, refused connection, TLS error, reset socket. Nothing
// on the object separates it from a TypeError thrown by a bug, so this is
// deliberately literal. Demanding a recognised `cause` instead would
// classify real network failures as permanent and fail backups that should
// have succeeded; the cost of the imprecision is bounded by the attempt
// count.
if (err instanceof TypeError) return true;
return causeCodes(err).some((code) => TRANSPORT_CODES.has(code));
};
// Could the first attempt already have taken effect on the server?
//
// `isRetryable` is the wrong question for a request that changes state.
// quak's non-idempotent calls are `/users/srp/create-session`,
// `/users/two-factor/verify` — which consumes one of a small number of 2FA
// attempts — and `/files/thumbnail`. They are replayed only on the failures in
// `CONNECT_CODES`, which establish that no TCP connection to the server ever
// existed: there was no address to connect to, or the peer refused the
// connection outright. A request byte cannot have been transmitted, so the
// server cannot have acted.
//
// Everything else is ambiguous. A 5xx proves the server did process the
// request. A reset or a broken pipe can arrive after it was fully sent and
// acted on. A routing errno can be delivered on an established socket. A
// deadline says nothing at all about the server's state.
export const isSafeToReplay = (err: unknown): boolean =>
isRetryable(err) && causeCodes(err).some((code) => CONNECT_CODES.has(code));
export interface WithRetryOptions extends RetryOptions {
isRetryable?: (err: unknown) => boolean;
}
// Exponential backoff with full jitter: the exponential term is the ceiling,
// and the actual wait is drawn uniformly below it. Full jitter, rather than a
// fixed delay plus noise, is what stops a client that lost a hundred parallel
// downloads to one CDN blip from re-sending all hundred at the same instant.
const backoffMs = (retryNumber: number, policy: ResolvedRetryOptions): number =>
policy.random() *
Math.min(policy.maxDelayMs, policy.baseDelayMs * 2 ** (retryNumber - 1));
export const withRetry = async <T>(
fn: () => Promise<T>,
opts?: WithRetryOptions,
): Promise<T> => {
const policy = resolveRetryOptions(opts);
const retryable = opts?.isRetryable ?? isRetryable;
for (let attempt = 1; ; attempt++) {
try {
return await fn();
} catch (err) {
if (attempt >= policy.attempts || !retryable(err)) {
// The error escapes unwrapped, and it is the one from the
// final attempt: callers classify what they catch, and
// `listMissingThumbnails` in particular needs the `ApiError`
// and its status intact.
throw err;
}
await policy.sleep(backoffMs(attempt, policy));
}
}
};

View File

@@ -2,6 +2,7 @@ import { createHash } from "node:crypto";
import { readFileSync } from "node:fs"; import { readFileSync } from "node:fs";
import * as jpeg from "jpeg-js"; import * as jpeg from "jpeg-js";
import type { Client } from "./client.js"; import type { Client } from "./client.js";
import { ApiError } from "./api/client.js";
import { encryptBlob, toBase64 } from "./crypto/index.js"; import { encryptBlob, toBase64 } from "./crypto/index.js";
import { downloadFile } from "./download/index.js"; import { downloadFile } from "./download/index.js";
import type { EnteFile } from "./model/types.js"; import type { EnteFile } from "./model/types.js";
@@ -62,13 +63,31 @@ export const listMissingThumbnails = async (
reason: "empty thumbnail (0 bytes)", reason: "empty thumbnail (0 bytes)",
}); });
} }
} catch { } catch (err) {
// A 404 is the server stating the thumbnail is not there:
// that, and an empty body, are the only two answers that mean
// "missing". Anything else reaching this point is a failure
// that already exhausted its retries — a failing server, a
// dropped connection, a deadline — and says nothing about
// whether the thumbnail exists.
//
// The distinction is what stops `helper
// fix-missing-thumbnails` from downloading originals,
// regenerating thumbnails and uploading them over thumbnails
// that were fine all along, because the CDN was briefly
// returning 500s while this ran.
if (err instanceof ApiError && err.status === 404) {
missing.push({ missing.push({
fileID: file.id, fileID: file.id,
title: file.metadata.title, title: file.metadata.title,
collection: col.name, collection: col.name,
reason: "thumbnail fetch failed", reason: "thumbnail not found (HTTP 404)",
}); });
} else {
log(
`[${col.name}] Could not check ${file.metadata.title}: ${err instanceof Error ? err.message : String(err)} (not reported as missing)`,
);
}
} }
} }
} }

View File

@@ -29,12 +29,23 @@
* return a `ReadableStream<Uint8Array>` from the appropriate CDN * return a `ReadableStream<Uint8Array>` from the appropriate CDN
* (or the self-hosted fallback path). * (or the self-hosted fallback path).
* *
* - Retries and timeouts. Every request is issued under a deadline and,
* where it is safe to do so, retried with exponential backoff. The
* policy is `src/retry.ts`; what the last section of this file
* documents is which requests get it and which deliberately do not.
*
* All tests inject a fake `fetch` via the constructor so nothing touches * All tests inject a fake `fetch` via the constructor so nothing touches
* the network. The fake records every call for assertion. * the network. The fake records every call for assertion.
*/ */
import { describe, expect, it } from "vitest"; import { describe, expect, it } from "vitest";
import { ApiClient, ApiError } from "../../src/api/client.js"; import {
ApiClient,
ApiError,
DEFAULT_DOWNLOAD_TIMEOUT_MS,
DEFAULT_REQUEST_TIMEOUT_MS,
} from "../../src/api/client.js";
import type { RetryOptions } from "../../src/retry.js";
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
// Test helpers // Test helpers
@@ -102,6 +113,92 @@ const recordingFetch = (
return { fetch: fake as typeof globalThis.fetch, calls }; return { fetch: fake as typeof globalThis.fetch, calls };
}; };
/**
* One scripted outcome for a single `fetch` call:
*
* - a `Response`, returned as-is;
* - an `Error`, thrown — this is how `fetch` reports a network failure;
* - `HANG`, a request that never answers until its own deadline aborts it.
*
* `HANG` is what makes the timeout tests honest. A fake that resolved after a
* delay would be testing the clock; this one resolves *only* when the signal
* quak attached fires. If no signal is attached, or the signal is not wired to
* the body, the promise never settles and the test fails on its own timeout
* rather than passing by accident.
*/
const HANG = Symbol("hang until aborted");
type FetchStep = Response | Error | typeof HANG;
const scriptedFetch = (
...steps: FetchStep[]
): {
fetch: typeof globalThis.fetch;
calls: { url: string; init: RequestInit | undefined }[];
} => {
const calls: { url: string; init: RequestInit | undefined }[] = [];
let i = 0;
const fake = async (
input: RequestInfo | URL,
init?: RequestInit,
): Promise<Response> => {
const url =
typeof input === "string"
? input
: input instanceof URL
? input.href
: input.url;
calls.push({ url, init });
const step = steps[i++];
if (step === undefined) {
throw new Error(`scriptedFetch: no step for call #${i - 1}`);
}
if (step === HANG) {
return new Promise<Response>((_resolve, reject) => {
const signal = init?.signal;
if (!signal) return;
if (signal.aborted) {
reject(signal.reason as Error);
return;
}
signal.addEventListener(
"abort",
() => reject(signal.reason as Error),
{ once: true },
);
});
}
if (step instanceof Error) throw step;
return step;
};
return { fetch: fake as typeof globalThis.fetch, calls };
};
/** An error shaped like a Node transport failure: the errno is on `.code`. */
const errnoError = (code: string, message = code): Error =>
Object.assign(new Error(message), { code });
/**
* A retry policy with the waiting removed. Backoff arithmetic is covered in
* `test/retry/retry.test.ts`; what the tests below are about is *how many
* requests* each call site issues, so they inject a `sleep` that returns
* immediately. Nothing in this file waits.
*/
const noWait: RetryOptions = {
sleep: () => Promise.resolve(),
random: () => 0,
};
/** Drain a stream and return the bytes, so body-level failures surface. */
const readAll = async (stream: ReadableStream<Uint8Array>): Promise<number> => {
const reader = stream.getReader();
let total = 0;
for (;;) {
const { done, value } = await reader.read();
if (value) total += value.length;
if (done) return total;
}
};
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
// Tests // Tests
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
@@ -383,3 +480,446 @@ describe("ApiClient.getFileStream / getThumbnailStream", () => {
expect(headers.get("X-Auth-Token")).toBe("tk"); expect(headers.get("X-Auth-Token")).toBe("tk");
}); });
}); });
// ---------------------------------------------------------------------------
// Retries
//
// The rule the whole section turns on: a request is repeated only when
// repeating it could produce a different answer, and only when repeating it
// cannot do harm. Those are two separate questions and the second one is why
// `postJSON` and `putJSON` behave differently from everything else here.
//
// Every assertion below counts requests. None of them measures how long
// anything took: the retry policy's `sleep` is injected and returns
// immediately, so a machine under load and an idle one produce identical
// results.
// ---------------------------------------------------------------------------
describe("ApiClient retries", () => {
it("issues exactly one request for a 404", async () => {
// A 404 is an answer, not a failure to get one. Repeating it wastes
// a round trip and — for `listMissingThumbnails`, which reads a 404
// as "this thumbnail really is missing" — delays a correct result.
// The script holds five responses; only the first may be consumed.
const { fetch, calls } = scriptedFetch(
textResponse("Not Found", 404),
textResponse("Not Found", 404),
textResponse("Not Found", 404),
textResponse("Not Found", 404),
textResponse("Not Found", 404),
);
const client = new ApiClient({ fetch, retry: noWait });
await expect(client.getJSON("/missing")).rejects.toBeInstanceOf(
ApiError,
);
expect(calls).toHaveLength(1);
});
it("retries a 500 up to the configured attempt count, then throws", async () => {
// `attempts` is a total, not a number of retries: three attempts mean
// three requests. The script offers five responses so that a client
// which ignored the limit would be visible as a count of 4 or 5
// rather than as a crash.
const { fetch, calls } = scriptedFetch(
textResponse("boom", 500),
textResponse("boom", 500),
textResponse("boom", 500),
textResponse("boom", 500),
textResponse("boom", 500),
);
const client = new ApiClient({
fetch,
retry: { ...noWait, attempts: 3 },
});
const err: unknown = await client
.getJSON("/flaky")
.catch((e: unknown) => e);
expect(err).toBeInstanceOf(ApiError);
expect((err as ApiError).status).toBe(500);
expect(calls).toHaveLength(3);
});
it("returns the first successful response after a 503", async () => {
const { fetch, calls } = scriptedFetch(
textResponse("unavailable", 503),
jsonResponse({ ok: true }),
);
const client = new ApiClient({ fetch, retry: noWait });
await expect(
client.getJSON<{ ok: boolean }>("/health"),
).resolves.toEqual({ ok: true });
expect(calls).toHaveLength(2);
});
it("retries 408 and 429", async () => {
// The two 4xx codes that are about timing rather than about the
// request. Backoff is precisely the right response to both.
for (const status of [408, 429]) {
const { fetch, calls } = scriptedFetch(
textResponse("wait", status),
jsonResponse({ ok: true }),
);
const client = new ApiClient({ fetch, retry: noWait });
await expect(
client.getJSON("/rate-limited"),
).resolves.toBeDefined();
expect(calls).toHaveLength(2);
}
});
it("retries a fetch rejection and succeeds on a later attempt", async () => {
// A dropped or refused connection is the failure this policy exists
// for: the request never got an answer, so asking again is free of
// consequence and likely to work.
const { fetch, calls } = scriptedFetch(
new TypeError("fetch failed"),
errnoError("ECONNRESET", "read ECONNRESET"),
jsonResponse({ collections: [] }),
);
const client = new ApiClient({ fetch, retry: noWait });
await expect(client.getJSON("/collections/v2")).resolves.toEqual({
collections: [],
});
expect(calls).toHaveLength(3);
});
it("applies the policy to file and thumbnail streams", async () => {
// Both CDN endpoints go through the same wrapper. A 503 from
// files.ente.io during a large backup is common enough that not
// retrying it would fail files for no reason.
for (const get of ["getFileStream", "getThumbnailStream"] as const) {
const { fetch, calls } = scriptedFetch(
textResponse("unavailable", 503),
streamResponse(new Uint8Array([1, 2, 3])),
);
const client = new ApiClient({ fetch, retry: noWait });
const stream = await client[get](7);
expect(await readAll(stream)).toBe(3);
expect(calls).toHaveLength(2);
}
});
it("does not retry when the caller opts out", async () => {
// `{ retry: false }` exists for one caller: the download layer, which
// wraps request *and* body consumption *and* decryption in a single
// retry of its own. Without the opt-out the two budgets would
// multiply — four attempts each becoming sixteen requests for one
// file.
const { fetch, calls } = scriptedFetch(
textResponse("unavailable", 503),
streamResponse(new Uint8Array([1, 2, 3])),
);
const client = new ApiClient({ fetch, retry: noWait });
await expect(
client.getFileStream(7, { retry: false }),
).rejects.toBeInstanceOf(ApiError);
expect(calls).toHaveLength(1);
});
it("exposes its resolved retry policy to the download layer", async () => {
// The download layer runs its own `withRetry` and must run it under
// the same policy the client was configured with, not under the
// library defaults.
const { fetch } = scriptedFetch(jsonResponse({}));
const client = new ApiClient({
fetch,
retry: { ...noWait, attempts: 9, baseDelayMs: 7, maxDelayMs: 11 },
});
const policy = client.getRetryOptions();
expect(policy.attempts).toBe(9);
expect(policy.baseDelayMs).toBe(7);
expect(policy.maxDelayMs).toBe(11);
});
});
describe("ApiClient timeouts", () => {
it("ships bounded default deadlines", () => {
// Asserted here so the README and the code cannot drift. Two numbers
// rather than one, because a deadline that is sane for a JSON call is
// nowhere near enough for a multi-gigabyte body, and a deadline long
// enough for that body would let a hung API call stall a backup for
// ten minutes.
expect(DEFAULT_REQUEST_TIMEOUT_MS).toBe(30_000);
expect(DEFAULT_DOWNLOAD_TIMEOUT_MS).toBe(600_000);
});
it("attaches an abort signal to every request", async () => {
const { fetch, calls } = recordingFetch(
jsonResponse({}),
jsonResponse({}),
new Response(null, { status: 200 }),
jsonResponse({}),
);
const client = new ApiClient({ fetch, retry: noWait });
await client.getJSON("/a");
await client.postJSON("/b", {});
await client.putFile("https://s3.example/x", new Uint8Array([1]));
await client.putJSON("/c", {});
for (const call of calls) {
expect(call.init?.signal).toBeInstanceOf(AbortSignal);
}
});
it("gives up on a request that never answers, and retries it", async () => {
// Before this policy existed there was no timeout anywhere in quak: a
// CDN connection that accepted the request and then went quiet would
// hang `quak backup` forever. `HANG` reproduces exactly that — the
// fake never answers, so the only thing that can end the call is the
// deadline quak attached.
const { fetch, calls } = scriptedFetch(HANG, HANG, HANG);
const client = new ApiClient({
fetch,
requestTimeoutMs: 20,
retry: { ...noWait, attempts: 3 },
});
await expect(client.getJSON("/black-hole")).rejects.toThrow();
// A timeout is retryable, so all three attempts were spent...
expect(calls).toHaveLength(3);
// ...each under its own fresh deadline, not one shared one that
// expired during the first attempt.
const signals = calls.map((c) => c.init?.signal);
expect(new Set(signals).size).toBe(3);
}, 5000);
it("recovers when a later attempt answers in time", async () => {
const { fetch, calls } = scriptedFetch(HANG, jsonResponse({ ok: 1 }));
const client = new ApiClient({
fetch,
requestTimeoutMs: 20,
retry: noWait,
});
await expect(client.getJSON("/slow-then-fast")).resolves.toEqual({
ok: 1,
});
expect(calls).toHaveLength(2);
}, 5000);
it("aborts a body that stalls after the headers arrived", async () => {
// The failure mode that a naive timeout misses. `getFileStream`
// returns as soon as headers arrive; the bytes are pulled later, in
// the download layer. A deadline that only guarded the initial fetch
// would leave the identical hang one layer down — which is where
// multi-megabyte photo downloads actually stall.
//
// This fake resolves its headers immediately and then serves a body
// that never produces a chunk and never observes the signal, so the
// only thing that can unblock the read is quak's own enforcement of
// the deadline over the stream it hands out.
const stalling = new Response(
new ReadableStream<Uint8Array>({
pull: () => new Promise<void>(() => {}),
}),
{ status: 200 },
);
const { fetch } = scriptedFetch(stalling);
const client = new ApiClient({
fetch,
downloadTimeoutMs: 20,
retry: { ...noWait, attempts: 1 },
});
const stream = await client.getFileStream(42);
const err: unknown = await readAll(stream).catch((e: unknown) => e);
expect(err).toBeInstanceOf(Error);
expect((err as Error).name).toBe("TimeoutError");
}, 5000);
it("lets a body that arrives in time through untouched", async () => {
// The counterpart to the previous test: enforcing the deadline over
// the stream must not corrupt or truncate a body that is simply being
// read normally.
const payload = new Uint8Array([9, 8, 7, 6, 5]);
const { fetch } = scriptedFetch(streamResponse(payload));
const client = new ApiClient({ fetch, retry: noWait });
const stream = await client.getFileStream(42);
const reader = stream.getReader();
const chunks: Uint8Array[] = [];
for (;;) {
const { done, value } = await reader.read();
if (done) break;
chunks.push(value);
}
const joined = new Uint8Array(chunks.reduce((n, c) => n + c.length, 0));
let offset = 0;
for (const c of chunks) {
joined.set(c, offset);
offset += c.length;
}
expect(joined).toEqual(payload);
});
});
describe("ApiClient error typing", () => {
it("throws ApiError with the status when a presigned PUT fails", async () => {
// `putFile` used to throw a bare Error with the status baked into a
// string. Nothing downstream could classify it, so a 500 from S3 was
// indistinguishable from a bug and could never be retried.
const { fetch, calls } = scriptedFetch(
new Response("Forbidden", { status: 403 }),
);
const client = new ApiClient({ fetch, retry: noWait });
const err: unknown = await client
.putFile("https://s3.example/obj", new Uint8Array([1, 2]))
.catch((e: unknown) => e);
expect(err).toBeInstanceOf(ApiError);
expect((err as ApiError).status).toBe(403);
expect(calls).toHaveLength(1);
});
it("retries a presigned PUT on a 5xx", async () => {
// A presigned PUT writes the whole object at one key in one request,
// so repeating it either overwrites the same bytes or lands them for
// the first time. There is no partial state to protect.
const { fetch, calls } = scriptedFetch(
new Response("slow down", { status: 503 }),
new Response(null, { status: 200 }),
);
const client = new ApiClient({ fetch, retry: noWait });
await client.putFile("https://s3.example/obj", new Uint8Array([1, 2]));
expect(calls).toHaveLength(2);
});
it("throws ApiError when a download response has no body", async () => {
// Also previously a bare Error. It carries the response status so a
// caller can see what arrived — and it is *not* retried: a 200 with
// no body is a malformed response, and asking again produces the same
// malformed response.
for (const get of ["getFileStream", "getThumbnailStream"] as const) {
const { fetch, calls } = scriptedFetch(
new Response(null, { status: 200 }),
new Response(null, { status: 200 }),
);
const client = new ApiClient({ fetch, retry: noWait });
const err: unknown = await client[get](5).catch((e: unknown) => e);
expect(err).toBeInstanceOf(ApiError);
expect((err as ApiError).status).toBe(200);
expect(calls).toHaveLength(1);
}
});
});
describe("ApiClient non-idempotent requests", () => {
/**
* `postJSON` and `putJSON` carry quak's only requests that change server
* state: `/users/srp/create-session`, `/users/two-factor/verify` — which
* consumes one of a small number of 2FA attempts — and `/files/thumbnail`.
*
* They are retried only on a failure that establishes no TCP connection to
* the server ever existed — DNS produced no address, or the peer refused
* the connection — so no request byte can have been transmitted.
* Everything else is ambiguous: a 5xx proves the server did process the
* request, and a reset or a timeout can arrive after it did.
* Replaying under that ambiguity can burn a 2FA attempt or register a
* thumbnail twice, and neither is worth the round trip it saves.
*/
it("does not replay a POST after a 5xx", async () => {
const { fetch, calls } = scriptedFetch(
textResponse("boom", 500),
jsonResponse({ sessionID: "second" }),
);
const client = new ApiClient({ fetch, retry: noWait });
await expect(
client.postJSON("/users/two-factor/verify", { code: "123456" }),
).rejects.toBeInstanceOf(ApiError);
expect(calls).toHaveLength(1);
});
it("does not replay a POST after a mid-flight connection reset", async () => {
// A reset can happen after the request was fully sent and acted on.
// It is retryable in general — `getJSON` retries it — but it is not
// replay-safe.
const { fetch, calls } = scriptedFetch(
errnoError("ECONNRESET", "socket hang up"),
jsonResponse({ sessionID: "second" }),
);
const client = new ApiClient({ fetch, retry: noWait });
await expect(
client.postJSON("/users/srp/create-session", {}),
).rejects.toThrow(/ECONNRESET|socket hang up/);
expect(calls).toHaveLength(1);
});
it("does not replay a POST after a timeout", async () => {
// A deadline says nothing about whether the server acted.
const { fetch, calls } = scriptedFetch(HANG, jsonResponse({}));
const client = new ApiClient({
fetch,
requestTimeoutMs: 20,
retry: noWait,
});
await expect(client.postJSON("/users/ott", {})).rejects.toThrow();
expect(calls).toHaveLength(1);
}, 5000);
it("replays a POST when the connection was never established", async () => {
// A refused connection or a DNS failure happens before any request
// byte is written, so the server cannot have seen it. This is the one
// case where replaying is provably harmless.
const { fetch, calls } = scriptedFetch(
new TypeError("fetch failed", {
cause: errnoError("ECONNREFUSED", "connect ECONNREFUSED"),
}),
errnoError("EAI_AGAIN", "getaddrinfo EAI_AGAIN api.ente.io"),
jsonResponse({ sessionID: "third" }),
);
const client = new ApiClient({ fetch, retry: noWait });
await expect(
client.postJSON<{ sessionID: string }>(
"/users/srp/create-session",
{},
),
).resolves.toEqual({ sessionID: "third" });
expect(calls).toHaveLength(3);
});
it("applies the same rule to PUT", async () => {
// `/files/thumbnail` is reached through `putJSON`.
const failing = scriptedFetch(
textResponse("boom", 503),
jsonResponse({}),
);
const failingClient = new ApiClient({
fetch: failing.fetch,
retry: noWait,
});
await expect(
failingClient.updateThumbnail(1, "key", "header"),
).rejects.toBeInstanceOf(ApiError);
expect(failing.calls).toHaveLength(1);
const refused = scriptedFetch(
errnoError("ECONNREFUSED", "connect ECONNREFUSED"),
jsonResponse({}),
);
const refusedClient = new ApiClient({
fetch: refused.fetch,
retry: noWait,
});
await refusedClient.updateThumbnail(1, "key", "header");
expect(refused.calls).toHaveLength(2);
});
});

View File

@@ -411,11 +411,22 @@ describe("quak backup", () => {
// File 101 (sunset.jpg) will return HTTP 500. The other two // File 101 (sunset.jpg) will return HTTP 500. The other two
// files must still download. The result must report the failure // files must still download. The result must report the failure
// without throwing. // without throwing.
//
// A 500 is retryable, so this file now costs several requests before
// it is given up on — that is the point of the retry policy, and
// `runBackup`'s own resilience is unchanged by it: the retry lives
// strictly below this loop, and an exhausted file is still logged,
// counted, and stepped over rather than aborting the run. The
// injected `sleep` is what keeps the suite from actually waiting out
// the backoff.
const outDir = join(testDir, "partial-failure"); const outDir = join(testDir, "partial-failure");
const client = await Client.login({ const client = await Client.login({
email: TEST_EMAIL, email: TEST_EMAIL,
password: TEST_PASSWORD, password: TEST_PASSWORD,
apiOptions: { fetch: buildMockFetch(mock, { failFileID: 101 }) }, apiOptions: {
fetch: buildMockFetch(mock, { failFileID: 101 }),
retry: { sleep: () => Promise.resolve(), random: () => 0 },
},
}); });
const result = await runBackup(client, outDir); const result = await runBackup(client, outDir);

View File

@@ -34,6 +34,15 @@
* forever after. The staging file and the rename are observed directly (see * forever after. The staging file and the rename are observed directly (see
* the `rename` hook below), not inferred from an empty directory. * the `rename` hook below), not inferred from an empty directory.
* *
* 3. **A failed transfer is retried as a whole.** A download is a request, a
* stream consumption, and a decryption, and only the first of those three
* happens inside `ApiClient`. A socket reset after the response headers
* have arrived therefore surfaces here, in the download layer — and that is
* the dominant failure mode for multi-megabyte photos over a CDN. So the
* entire sequence is retried as one unit, not just the request. The
* secretstream pull state is not resumable and there is no Range support,
* so a retry starts the file over from byte zero.
*
* These tests build synthetic encrypted files using sodium's push API, * These tests build synthetic encrypted files using sodium's push API,
* serve them from a mock fetch, and verify the decrypted output on disk. * serve them from a mock fetch, and verify the decrypted output on disk.
*/ */
@@ -61,6 +70,8 @@ import {
} from "vitest"; } from "vitest";
import { init, toBase64, STREAM_CHUNK_SIZE } from "../../src/crypto/index.js"; import { init, toBase64, STREAM_CHUNK_SIZE } from "../../src/crypto/index.js";
import { ApiClient } from "../../src/api/client.js"; import { ApiClient } from "../../src/api/client.js";
import { ApiError, TruncatedStreamError } from "../../src/errors.js";
import type { RetryOptions } from "../../src/retry.js";
import { downloadFile, downloadThumbnail } from "../../src/download/index.js"; import { downloadFile, downloadThumbnail } from "../../src/download/index.js";
import type { EnteFile, FileMetadata } from "../../src/model/types.js"; import type { EnteFile, FileMetadata } from "../../src/model/types.js";
@@ -326,18 +337,92 @@ beforeAll(() => {
multiChunk = encryptMultiChunkBody(multiChunkKey, 1, 1024); multiChunk = encryptMultiChunkBody(multiChunkKey, 1, 1024);
}); });
/** An error shaped like a Node transport failure: the errno is on `.code`. */
const errnoError = (code: string, message = code): Error =>
Object.assign(new Error(message), { code });
/**
* A retry policy with the waiting removed, used by every fixture in this
* file. Backoff arithmetic belongs to `test/retry/retry.test.ts`; here the
* only interesting quantity is how many requests a download issued, so the
* injected `sleep` returns immediately and nothing in this file waits.
*/
const noWait: RetryOptions = {
sleep: () => Promise.resolve(),
random: () => 0,
};
/**
* One scripted outcome for a single request to the CDN.
*
* - `body` — a complete response body.
* - `status` — an HTTP error response.
* - `reset` — a response whose headers arrive, whose body delivers `bytes`,
* and which then dies with a socket reset. This is the failure that
* motivates retrying the download rather than the request: by the time it
* happens `ApiClient` has already returned successfully.
*/
type BodyStep =
| { kind: "body"; bytes: Uint8Array }
| { kind: "status"; status: number }
| { kind: "reset"; bytes: Uint8Array };
/**
* A fetch that serves one scripted step per call and counts the calls. It
* deliberately refuses to serve more requests than it was given steps for, so
* a retry loop that ran away is a test failure rather than a silent success.
*/
const scriptedCdnFetch = (
...steps: BodyStep[]
): { fetch: typeof globalThis.fetch; requests: () => number } => {
let calls = 0;
const fake = async (): Promise<Response> => {
const step = steps[calls++];
if (step === undefined) {
throw new Error(`scriptedCdnFetch: no step for request #${calls}`);
}
if (step.kind === "status") {
return new Response("error", { status: step.status });
}
if (step.kind === "body") {
return new Response(step.bytes, { status: 200 });
}
const bytes = step.bytes;
return new Response(
new ReadableStream<Uint8Array>({
start(controller) {
controller.enqueue(bytes);
controller.error(
errnoError("ECONNRESET", "aborted by peer"),
);
},
}),
{ status: 200 },
);
};
return { fetch: fake as typeof globalThis.fetch, requests: () => calls };
};
/** /**
* Build an EnteFile plus ApiClient whose file *and* thumbnail streams both * Build an EnteFile plus ApiClient whose file *and* thumbnail streams both
* serve `body` under `header`. The download path under test is otherwise * serve `body` under `header`. The download path under test is otherwise
* identical for the two, so every truncation/atomicity case below runs * identical for the two, so every truncation/atomicity case below runs
* against both entry points from a single fixture. * against both entry points from a single fixture.
*
* The default policy here is a single attempt. The failure-contract tests are
* about what the caller and the filesystem are left with, not about how many
* times quak asked; pinning attempts to one keeps them saying exactly that,
* and keeps them from re-decrypting a 4 MiB fixture four times over. The
* retry counts have their own tests at the bottom of this file, which set the
* attempt count explicitly.
*/ */
const fixtureFor = ( const fixtureFor = (
key: Uint8Array, key: Uint8Array,
header: Uint8Array, header: Uint8Array,
body: Uint8Array, body: Uint8Array,
retry: RetryOptions = { ...noWait, attempts: 1 },
): { api: ApiClient; file: EnteFile } => ({ ): { api: ApiClient; file: EnteFile } => ({
api: new ApiClient({ fetch: mockFetchForBody(body) }), api: new ApiClient({ fetch: mockFetchForBody(body), retry }),
file: buildMockEnteFile(key, header, header), file: buildMockEnteFile(key, header, header),
}); });
@@ -502,9 +587,17 @@ describe.each(entryPoints)(
); );
const outPath = join(freshDir(), "truncated.bin"); const outPath = join(freshDir(), "truncated.bin");
await expect(download(api, file, outPath)).rejects.toThrow( const err: unknown = await download(api, file, outPath).catch(
/truncated/i, (e: unknown) => e,
); );
// The type, not the wording, is the contract. The retry policy
// classifies truncation as worth another attempt, and it decides
// that with `instanceof`: matching on message text would make
// rewording a diagnostic silently turn every truncated download
// into a permanent failure.
expect(err).toBeInstanceOf(TruncatedStreamError);
expect((err as Error).message).toMatch(/truncated/i);
}); });
it("rejects a body whose final chunk arrived only in part", async () => { it("rejects a body whose final chunk arrived only in part", async () => {
@@ -537,7 +630,7 @@ describe.each(entryPoints)(
(e: unknown) => e, (e: unknown) => e,
); );
expect(err).toBeInstanceOf(Error); expect(err).toBeInstanceOf(TruncatedStreamError);
expect((err as Error).message).toMatch(/truncated/i); expect((err as Error).message).toMatch(/truncated/i);
expect((err as Error).cause).toBeInstanceOf(Error); expect((err as Error).cause).toBeInstanceOf(Error);
expect(((err as Error).cause as Error).message).toMatch( expect(((err as Error).cause as Error).message).toMatch(
@@ -564,8 +657,8 @@ describe.each(entryPoints)(
); );
const outPath = join(freshDir(), "empty.bin"); const outPath = join(freshDir(), "empty.bin");
await expect(download(api, file, outPath)).rejects.toThrow( await expect(download(api, file, outPath)).rejects.toBeInstanceOf(
/truncated/i, TruncatedStreamError,
); );
}); });
@@ -588,8 +681,8 @@ describe.each(entryPoints)(
const dir = freshDir(); const dir = freshDir();
const outPath = join(dir, "absent.bin"); const outPath = join(dir, "absent.bin");
await expect(download(api, file, outPath)).rejects.toThrow( await expect(download(api, file, outPath)).rejects.toBeInstanceOf(
/truncated/i, TruncatedStreamError,
); );
expect(existsSync(outPath)).toBe(false); expect(existsSync(outPath)).toBe(false);
@@ -621,10 +714,18 @@ describe.each(entryPoints)(
const dir = freshDir(); const dir = freshDir();
const outPath = join(dir, "corrupt.bin"); const outPath = join(dir, "corrupt.bin");
await expect(download(api, file, outPath)).rejects.toThrow( const err: unknown = await download(api, file, outPath).catch(
/authentication failed/i, (e: unknown) => e,
); );
expect(err).toBeInstanceOf(Error);
expect((err as Error).message).toMatch(/authentication failed/i);
// And explicitly *not* the truncation type, because that type is
// what the retry policy keys on: mislabelling corruption as
// truncation would spend the whole attempt budget re-downloading
// a file that will never decrypt.
expect(err).not.toBeInstanceOf(TruncatedStreamError);
expect(existsSync(outPath)).toBe(false); expect(existsSync(outPath)).toBe(false);
expect(readdirSync(dir)).toEqual([]); expect(readdirSync(dir)).toEqual([]);
}); });
@@ -648,8 +749,8 @@ describe.each(entryPoints)(
const outPath = join(dir, "existing.bin"); const outPath = join(dir, "existing.bin");
writeFileSync(outPath, existing); writeFileSync(outPath, existing);
await expect(download(api, file, outPath)).rejects.toThrow( await expect(download(api, file, outPath)).rejects.toBeInstanceOf(
/truncated/i, TruncatedStreamError,
); );
expect(readFileSync(outPath)).toEqual(Buffer.from(existing)); expect(readFileSync(outPath)).toEqual(Buffer.from(existing));
@@ -755,3 +856,208 @@ describe.each(entryPoints)(
}); });
}, },
); );
// ---------------------------------------------------------------------------
// Retries
//
// What is retried here is the whole download — request, stream consumption,
// decryption — because only the first of those three happens inside
// `ApiClient`. Every assertion counts requests; none of them measures time.
// ---------------------------------------------------------------------------
describe.each(entryPoints)("$name retries", ({ name, download }) => {
const freshDir = (): string => mkdtempSync(join(testDir, `${name}-retry-`));
/** A cheap single-chunk fixture: no 4 MiB encryption in the retry tests. */
const smallFixture = (
seed: number,
): {
key: Uint8Array;
header: Uint8Array;
ciphertext: Uint8Array;
plaintext: Uint8Array;
} => {
const plaintext = patternBytes(1024, seed);
const key = sodium.crypto_secretstream_xchacha20poly1305_keygen();
const { header, ciphertext } = encryptFileBody(plaintext, key);
return { key, header, ciphertext, plaintext };
};
const clientFor = (
fetch: typeof globalThis.fetch,
attempts: number,
): ApiClient => new ApiClient({ fetch, retry: { ...noWait, attempts } });
it("retries a connection reset that happened mid-body", async () => {
// The case `ApiClient` cannot see. Its own request succeeded: headers
// arrived, a `ReadableStream` was handed back, and only then did the
// socket die. Retrying the fetch alone would have caught nothing,
// which is why the retry wraps the whole sequence.
const { key, header, ciphertext, plaintext } = smallFixture(41);
const { fetch, requests } = scriptedCdnFetch(
{ kind: "reset", bytes: ciphertext.slice(0, 16) },
{ kind: "body", bytes: ciphertext },
);
const api = clientFor(fetch, 4);
const file = buildMockEnteFile(key, header, header);
const dir = freshDir();
const outPath = join(dir, "reset-then-ok.bin");
const result = await download(api, file, outPath);
expect(requests()).toBe(2);
expect(result.bytesWritten).toBe(plaintext.length);
expectSameBytes(readFileSync(outPath), plaintext);
});
it("stages one temp file for the attempt that succeeded, not one per attempt", async () => {
// The atomic write stays outside the retry loop. A retried download
// must not leave a trail of half-written scratch files, and the
// destination must be touched exactly once — by the attempt that
// produced a complete, authenticated plaintext.
const { key, header, ciphertext } = smallFixture(42);
const { fetch } = scriptedCdnFetch(
{ kind: "reset", bytes: ciphertext.slice(0, 16) },
{ kind: "reset", bytes: ciphertext.slice(0, 16) },
{ kind: "body", bytes: ciphertext },
);
const api = clientFor(fetch, 4);
const file = buildMockEnteFile(key, header, header);
const dir = freshDir();
const outPath = join(dir, "one-stage.bin");
await download(api, file, outPath);
expect(renameHook.calls).toHaveLength(1);
expect(renameHook.calls[0]!.to).toBe(outPath);
expect(readdirSync(dir)).toEqual(["one-stage.bin"]);
});
it("retries a truncated body and gives up after the configured attempts", async () => {
// Truncation is retryable — the file on the server is intact, the
// transfer was not — but it is not retryable forever. Three attempts
// configured, three requests, then the caller gets the error.
const key = sodium.crypto_secretstream_xchacha20poly1305_keygen();
const { header, ciphertext } = encryptNonFinalBody(
patternBytes(256, 43),
key,
);
const { fetch, requests } = scriptedCdnFetch(
{ kind: "body", bytes: ciphertext },
{ kind: "body", bytes: ciphertext },
{ kind: "body", bytes: ciphertext },
{ kind: "body", bytes: ciphertext },
);
const api = clientFor(fetch, 3);
const file = buildMockEnteFile(key, header, header);
const dir = freshDir();
const outPath = join(dir, "always-truncated.bin");
await expect(download(api, file, outPath)).rejects.toBeInstanceOf(
TruncatedStreamError,
);
expect(requests()).toBe(3);
// Every attempt failed before anything was written, so the directory
// is still empty.
expect(readdirSync(dir)).toEqual([]);
});
it("issues exactly one request when the file is gone", async () => {
// A 404 from the CDN is an answer. `runBackup` logs it and moves on;
// spending three more requests and three backoff waits on it would
// slow a large backup down for nothing.
const { key, header } = smallFixture(44);
const { fetch, requests } = scriptedCdnFetch(
{ kind: "status", status: 404 },
{ kind: "status", status: 404 },
{ kind: "status", status: 404 },
{ kind: "status", status: 404 },
);
const api = clientFor(fetch, 4);
const file = buildMockEnteFile(key, header, header);
const outPath = join(freshDir(), "gone.bin");
const err: unknown = await download(api, file, outPath).catch(
(e: unknown) => e,
);
expect(err).toBeInstanceOf(ApiError);
expect((err as ApiError).status).toBe(404);
expect(requests()).toBe(1);
});
it("retries a 503 from the CDN", async () => {
const { key, header, ciphertext, plaintext } = smallFixture(45);
const { fetch, requests } = scriptedCdnFetch(
{ kind: "status", status: 503 },
{ kind: "status", status: 503 },
{ kind: "body", bytes: ciphertext },
);
const api = clientFor(fetch, 4);
const file = buildMockEnteFile(key, header, header);
const outPath = join(freshDir(), "flaky-cdn.bin");
await download(api, file, outPath);
expect(requests()).toBe(3);
expectSameBytes(readFileSync(outPath), plaintext);
});
it("spends one attempt budget, not one per layer", async () => {
// `ApiClient.getFileStream` retries on its own for direct callers.
// The download layer opts out of that and runs its own retry over the
// whole sequence. If it did not, the two budgets would compose: three
// attempts here would become nine requests to the CDN for a single
// file, and the default four would become sixteen.
const { key, header } = smallFixture(46);
const steps: BodyStep[] = Array.from({ length: 12 }, () => ({
kind: "status" as const,
status: 503,
}));
const { fetch, requests } = scriptedCdnFetch(...steps);
const api = clientFor(fetch, 3);
const file = buildMockEnteFile(key, header, header);
const outPath = join(freshDir(), "budget.bin");
await expect(download(api, file, outPath)).rejects.toBeInstanceOf(
ApiError,
);
expect(requests()).toBe(3);
});
});
describe("download retries: corruption is not retried", () => {
it("gives up immediately on a chunk that failed to authenticate", async () => {
// A whole chunk that failed to authenticate while the stream
// continued past it is corruption or a wrong key. Neither is fixed by
// asking again, and a backup run that retried every such file would
// multiply the cost of a genuinely broken file by the attempt count.
//
// This is also the boundary of the single-chunk ambiguity documented
// at the classifier: the split is only achievable because this body
// has more than one chunk.
const corrupted = Uint8Array.from(multiChunk.body);
corrupted[10] ^= 0xff;
const { fetch, requests } = scriptedCdnFetch(
{ kind: "body", bytes: corrupted },
{ kind: "body", bytes: corrupted },
{ kind: "body", bytes: corrupted },
{ kind: "body", bytes: corrupted },
);
const api = new ApiClient({ fetch, retry: { ...noWait, attempts: 4 } });
const file = buildMockEnteFile(
multiChunkKey,
multiChunk.header,
multiChunk.header,
);
const outPath = join(mkdtempSync(join(testDir, "corrupt-")), "c.bin");
await expect(downloadFile(api, file, outPath)).rejects.toThrow(
/authentication failed/i,
);
expect(requests()).toBe(1);
});
});

View File

@@ -0,0 +1,64 @@
// The Docker build context is load-bearing in two directions, and both
// failures are silent.
//
// Excluding too little: a worktree left under `.claude/` is copied into the
// image, vitest globs its `test/` tree as well as the real one, and the
// containerised `make check` runs the whole suite twice over while reporting
// success. A compiled `bin/quak` is ~100 MB of context nobody needs.
//
// Excluding too much: Prettier 3 reads `.gitignore` as a default ignore file,
// so dropping it from the context silently changes which files
// `make fmt-check` looks at inside the image compared to the host.
//
// Neither shows up as a build failure, so they are asserted here.
import { describe, expect, it } from "vitest";
import { existsSync, readFileSync } from "node:fs";
import { fileURLToPath } from "node:url";
import { join } from "node:path";
const repoRoot = fileURLToPath(new URL("../../", import.meta.url));
const patterns = (name: string): string[] =>
readFileSync(join(repoRoot, name), "utf-8")
.split("\n")
.map((line) => line.trim())
.filter((line) => line !== "" && !line.startsWith("#"));
const dockerignore = patterns(".dockerignore");
describe(".dockerignore", () => {
// Everything here is either generated, enormous, or secret. `.claude/` is
// the correctness one: see the header comment and issue #25.
it.each([
".claude/",
".quak/",
"bin/quak",
"node_modules",
"coverage",
"dist",
".vitest-cache/",
".nyc_output/",
"*.tsbuildinfo",
])("keeps %s out of the build context", (pattern) => {
expect(dockerignore).toContain(pattern);
});
it("leaves .gitignore in the build context for prettier", () => {
expect(dockerignore).not.toContain(".gitignore");
});
// Both images are built from this same context, and the lint image runs
// eslint and prettier across it. BuildKit lets a `<dockerfile>.dockerignore`
// shadow the root one for a single build; such a file would silently give
// the lint build a different, unreviewed context — and eslint's flat config
// does not ignore dot-directories, so a stray `.claude/` worktree would be
// linted.
it.each(["Dockerfile", "Dockerfile.lint"])(
"is not shadowed by a per-Dockerfile ignore file for %s",
(name) => {
expect(existsSync(join(repoRoot, `${name}.dockerignore`))).toBe(
false,
);
},
);
});

View File

@@ -0,0 +1,118 @@
// The package manifest promises three files that only exist after a build:
// `main`, `types`, and the `quak` binary. Nothing in the test suite used to
// look at them, and `make check` runs the suite and the lint container but
// never the build, so `tsconfig.json` and `package.json` were free to drift
// apart. (The formatting check is part of the lint container, not a step of
// its own; `test/packaging/lint-once.test.ts` is what holds that shape.) They
// did: `rootDir` was `./src` while `include` also pulled in `bin/**/*`, which
// is TS6059, and no build had succeeded for as long as that was true.
//
// These tests read both files and check the contract between them, without
// running a build, so they stay in the fast unit suite. The complementary
// check — that the files really landed on disk — is in `script/build`, which
// runs after the compiler and is the only place that can honestly answer it.
import { describe, expect, it } from "vitest";
import { readFileSync } from "node:fs";
import { fileURLToPath } from "node:url";
import { join, posix } from "node:path";
interface TsConfig {
compilerOptions: {
outDir: string;
rootDir: string;
};
include: string[];
}
interface PackageJson {
main: string;
types: string;
bin: Record<string, string>;
scripts: Record<string, string>;
}
const repoRoot = fileURLToPath(new URL("../../", import.meta.url));
const readJSON = <T>(name: string): T =>
JSON.parse(readFileSync(join(repoRoot, name), "utf-8")) as T;
const tsconfig = readJSON<TsConfig>("tsconfig.json");
const pkg = readJSON<PackageJson>("package.json");
// Paths in the two manifests are written with a leading "./"; normalize so
// they can be compared and joined. Everything here is POSIX-style because
// that is what both JSON files contain, regardless of the host OS.
const clean = (p: string): string => posix.normalize(p);
const outDir = clean(tsconfig.compilerOptions.outDir);
const rootDir = clean(tsconfig.compilerOptions.rootDir);
// The directory prefix of a glob: the part before the first segment
// containing a wildcard. "src/**/*" -> "src", "bin/**/*" -> "bin".
const globRoot = (pattern: string): string => {
const segments = clean(pattern).split("/");
const wildcard = segments.findIndex((s) => s.includes("*"));
return (
segments
.slice(0, wildcard === -1 ? segments.length : wildcard)
.join("/") || "."
);
};
// Where tsc will write the output for a source file: the path relative to
// rootDir, re-rooted under outDir, with the extension swapped.
const emitted = (source: string, extension: string): string =>
"./" +
posix
.join(outDir, posix.relative(rootDir, clean(source)))
.replace(/\.ts$/, extension);
describe("tsconfig include and rootDir", () => {
// TS6059 is not a style complaint: tsc refuses to emit anything at all
// when a compiled file sits outside rootDir, so this single mismatch
// took out both the library and the CLI artifacts.
it("compiles only files that live under rootDir", () => {
for (const pattern of tsconfig.include) {
const root = globRoot(pattern);
const relative = posix.relative(rootDir, root);
expect(
relative === "" || !relative.startsWith(".."),
`include pattern ${pattern} matches files outside rootDir ` +
`${rootDir}; tsc rejects that with TS6059`,
).toBe(true);
}
});
});
describe("package.json entrypoints", () => {
// Each of these is the path a consumer resolves — `import { Client } from
// "quak"` for main, the editor for types, `npx quak` for the bin — so a
// wrong value is a broken package even when the build itself is green.
it("names the file tsc emits for src/index.ts as main", () => {
expect(pkg.main).toBe(emitted("src/index.ts", ".js"));
});
it("names the declaration tsc emits for src/index.ts as types", () => {
expect(pkg.types).toBe(emitted("src/index.ts", ".d.ts"));
});
it("names the file tsc emits for bin/quak.ts as the quak binary", () => {
expect(pkg.bin.quak).toBe(emitted("bin/quak.ts", ".js"));
});
// The README's Getting Started block tells the reader to run
// `yarn quak login` straight after `yarn build`. That only works if a
// `quak` script exists and points at the built CLI, not at the source.
it("runs the built CLI from the quak script", () => {
expect(pkg.scripts.quak).toBeDefined();
expect(pkg.scripts.quak).toContain(pkg.bin.quak);
});
});
describe("bin/quak.ts", () => {
// tsc copies the shebang into the emitted file, so it has to be in the
// source for the installed binary to be directly executable.
it("starts with a node shebang", () => {
const source = readFileSync(join(repoRoot, "bin/quak.ts"), "utf-8");
expect(source.split("\n")[0]).toBe("#!/usr/bin/env node");
});
});

View File

@@ -0,0 +1,184 @@
// Linting runs in Docker, one way, everywhere: `script/lint` builds
// `Dockerfile.lint`, which COPYs the repo into a digest-pinned image and runs
// eslint and prettier as build steps, so a successful build IS a clean lint.
//
// Three things can quietly undo that, and none of them shows up as a build
// failure, which is why they are asserted here:
//
// 1. Recursion. `script/check` calls `script/lint`, and `script/lint` is now a
// `docker build`. Anything that runs `make check` inside a container is
// therefore asking for Docker inside Docker, and CI breaks. The image built
// from `Dockerfile` runs the suite and the compile only; lint happens once,
// in `Dockerfile.lint`.
// 2. Cache. A lint build over an unchanged tree returns success in well under a
// second having linted nothing. The `LINT_EPOCH` guard is what forces the
// linter layers to execute, and it has to fail closed: an unset build
// argument is the empty string, which is a perfectly stable cache key, so an
// invocation that omits it must be rejected rather than served a cached
// green.
// 3. A host lint path surviving alongside the container one, which would let a
// lint result come from an unpinned local toolchain.
import { describe, expect, it } from "vitest";
import { readFileSync } from "node:fs";
import { fileURLToPath } from "node:url";
import { join } from "node:path";
const repoRoot = fileURLToPath(new URL("../../", import.meta.url));
const read = (name: string): string =>
readFileSync(join(repoRoot, name), "utf-8");
// The executable lines of a shell script or Dockerfile: comments carry the
// reasoning and frequently name the very commands these tests forbid, so they
// would otherwise trigger every assertion below.
const instructions = (name: string): string[] =>
read(name)
.split("\n")
.map((line) => line.trim())
.filter((line) => line !== "" && !line.startsWith("#"));
const lintScript = instructions("script/lint");
const dockerfileLint = instructions("Dockerfile.lint");
const dockerfile = instructions("Dockerfile");
const cibuild = instructions("script/cibuild");
const has = (lines: string[], pattern: RegExp): boolean =>
lines.some((line) => pattern.test(line));
describe("script/lint", () => {
it("lints by building Dockerfile.lint", () => {
expect(has(lintScript, /docker build .*-f Dockerfile\.lint/)).toBe(
true,
);
});
// The whole point of the ruling: no invocation of a linter against the
// working tree survives, so a lint verdict can only come from the pinned
// image.
it("runs no linter on the host", () => {
expect(has(lintScript, /eslint|prettier/)).toBe(false);
});
// Without a fresh epoch the build is served from cache in under a second,
// having linted nothing, and still exits 0.
it("passes a fresh LINT_EPOCH on every run", () => {
expect(
has(lintScript, /--build-arg LINT_EPOCH="\$\(date \+%s\)"/),
).toBe(true);
});
});
describe("Dockerfile.lint", () => {
// Tag references are server-mutable, so they are remote code execution.
it("pins its base image by digest", () => {
expect(has(dockerfileLint, /^FROM \S+@sha256:[0-9a-f]{64}/)).toBe(true);
});
it("runs eslint as a build step", () => {
expect(has(dockerfileLint, /^RUN .*eslint \./)).toBe(true);
});
it("runs prettier as a build step", () => {
expect(has(dockerfileLint, /^RUN .*prettier --check \./)).toBe(true);
});
// An unset ARG is the empty string, and an empty string is a perfectly
// stable cache key. Rejecting it is what stops a bare
// `docker build -f Dockerfile.lint .` from reporting a green it did not
// earn.
it("refuses to build without LINT_EPOCH", () => {
expect(has(dockerfileLint, /^ARG LINT_EPOCH$/)).toBe(true);
expect(
has(dockerfileLint, /^RUN \[ -n "\$LINT_EPOCH" \] \|\| exit 1$/),
).toBe(true);
});
// The guard only forces execution of the layers below it, so both linters
// have to sit after it. Layer order is the mechanism, not a style choice.
it("puts both linters below the epoch guard", () => {
const guard = dockerfileLint.findIndex((line) =>
/^RUN \[ -n "\$LINT_EPOCH" \]/.test(line),
);
const linters = dockerfileLint
.map((line, index) => ({ line, index }))
.filter(({ line }) => /^RUN .*(eslint|prettier)/.test(line));
expect(linters.length).toBeGreaterThan(0);
for (const { line, index } of linters) {
expect(
index,
`${line} must run below the LINT_EPOCH guard`,
).toBeGreaterThan(guard);
}
});
// Dependency installation is the slow layer and has nothing to do with the
// sources, so it caches separately: manifests first, sources afterwards.
it("copies the manifests before the sources", () => {
const manifests = dockerfileLint.findIndex((line) =>
/^COPY package\.json yarn\.lock/.test(line),
);
const sources = dockerfileLint.findIndex((line) =>
/^COPY \. \.$/.test(line),
);
expect(manifests).toBeGreaterThanOrEqual(0);
expect(sources).toBeGreaterThan(manifests);
});
// script/lint is a docker build; a lint step that shelled out to it would
// recurse.
it("does not call script/lint or make lint", () => {
expect(has(dockerfileLint, /make lint|script\/lint/)).toBe(false);
});
});
describe("Dockerfile", () => {
// `make check` runs script/lint, which is a docker build, so an image that
// ran it would need a Docker daemon inside the container.
it("does not run make check, make lint or script/lint", () => {
expect(
has(dockerfile, /make check|make lint|script\/(check|lint)/),
).toBe(false);
});
// The replaced lint stage took a `COPY --from=lint` dependency to order
// itself before the check stage. Dockerfile.lint is that stage now, and
// two definitions of how to lint is one too many.
it("has no lint stage", () => {
expect(has(dockerfile, /AS lint\b|--from=lint\b/)).toBe(false);
});
it("still runs the suite and the build under the epoch guard", () => {
expect(has(dockerfile, /^RUN make test$/)).toBe(true);
expect(has(dockerfile, /^RUN make build$/)).toBe(true);
expect(
has(dockerfile, /^RUN \[ -n "\$CHECK_EPOCH" \] \|\| exit 1$/),
).toBe(true);
});
});
describe("script/cibuild", () => {
// CI has to get both verdicts. Lint goes first so the fast failure is
// reported before the suite runs.
it("builds the lint image before the test and build image", () => {
const lint = cibuild.findIndex((line) => /\/lint"/.test(line));
const check = cibuild.findIndex((line) =>
/docker build .*CHECK_EPOCH/.test(line),
);
expect(lint).toBeGreaterThanOrEqual(0);
expect(check).toBeGreaterThan(lint);
});
});
describe("package.json", () => {
// `yarn lint` was a second, unpinned way to get a lint verdict, from
// whatever eslint the working tree happened to have installed.
it("exposes no host lint script", () => {
const pkg = JSON.parse(read("package.json")) as {
scripts: Record<string, string>;
};
expect(pkg.scripts.lint).toBeUndefined();
});
});

View File

@@ -0,0 +1,503 @@
// `make check` used to run `prettier --check .` twice: once inside the lint
// container (`script/lint` builds `Dockerfile.lint`, which runs eslint and
// prettier as build steps) and once again on the host, because `script/check`
// also called `script/fmt-check`. Two passes, one verdict, and the host one is
// the weaker of the two — its prettier is whatever the working tree happens to
// have installed, while the container's is digest-pinned and installed under
// `--frozen-lockfile`.
//
// The fix was to delete the host call from `script/check` and `script/precommit`.
// Nothing about that fix is self-enforcing: anyone can wire `script/fmt-check`
// back in, or add a prettier step to a Dockerfile, and every build stays green
// while quietly doing the work twice again. So the count is asserted here
// rather than promised in a comment.
//
// The assertion is a static walk of the invocation graph, not a string match
// against one file. Starting from an entrypoint, it follows every edge the repo
// actually uses to reach another command — `run:` steps in the CI workflow,
// `"$SCRIPT_DIR/<name>"` and `script/<name>` into other scripts, `make <target>`
// through the Makefile shims, `yarn run <name>` through the `package.json`
// scripts, and `docker build -f <file>` into that Dockerfile's `RUN` steps — and
// counts the prettier invocations it finds. A prettier call added anywhere in
// that graph is therefore caught, wherever it is added.
//
// Two entrypoints are walked, because they cover different graphs: `make check`
// is what a developer runs, and `.gitea/workflows/check.yml` is what CI runs.
// The CI walk starts at the workflow file rather than at a hand-picked script,
// so "the path CI executes" is read out of the repo instead of assumed; it
// reaches `script/cibuild`, and through it the `Dockerfile` image that `make
// check` never touches. Walking only `make check` is how a duplicate prettier
// pass in `Dockerfile` stayed invisible.
//
// Undercounting is the failure mode that would make this test worthless. Three
// things guard against it: the walk is asserted to have reached the nodes that
// matter, an unresolvable or empty node is a thrown error rather than a quiet
// zero, and prettier is counted per occurrence rather than per line, so two
// invocations chained with `&&` cannot read as one.
import { describe, expect, it } from "vitest";
import { readFileSync } from "node:fs";
import { fileURLToPath } from "node:url";
import { join } from "node:path";
const repoRoot = fileURLToPath(new URL("../../", import.meta.url));
const read = (name: string): string =>
readFileSync(join(repoRoot, name), "utf-8");
// A backslash at end of line continues the command; the resolver has to see the
// whole invocation, since the interesting flags (`-f Dockerfile.lint`) can sit
// on the continuation.
const joinContinuations = (text: string): string[] => {
const joined: string[] = [];
for (const raw of text.split("\n")) {
const line = raw.trim();
const previous = joined[joined.length - 1];
if (previous !== undefined && previous.endsWith("\\")) {
joined[joined.length - 1] =
`${previous.slice(0, -1).trim()} ${line}`;
} else {
joined.push(line);
}
}
return joined;
};
// Comments are stripped everywhere. The headers of these scripts explain the
// duplication this test exists to prevent, and therefore name `prettier` and
// `script/fmt-check` repeatedly; counting them would make the test assert the
// prose instead of the behaviour.
const executable = (text: string): string[] =>
joinContinuations(text).filter(
(line) => line !== "" && !line.startsWith("#"),
);
// Every occurrence, not "does this line mention prettier": a line that reads
// `yarn run prettier --check . && yarn run prettier --check src` is two passes
// over the same tree, which is exactly the bug this file exists to catch, and
// counting it as one would hide it. `.prettierrc` and `.prettierignore` are not
// invocations and do not match, because `\b` requires a non-word character
// after the name.
const countPrettier = (line: string): number =>
(line.match(/\bprettier\b/g) ?? []).length;
// Makefile targets are thin shims (`check:` / tab / `@script/check`), so a
// `make <target>` edge has to resolve through them to keep "per `make check`"
// meaning what it says. Recipe lines are the tab-indented ones.
const makeRecipes = (): Map<string, string[]> => {
const recipes = new Map<string, string[]>();
let current: string | null = null;
for (const raw of read("Makefile").split("\n")) {
if (raw.startsWith("\t")) {
if (current !== null) {
recipes.get(current)?.push(raw.trim().replace(/^[@-]+/, ""));
}
continue;
}
const target = /^([a-z][a-z-]*)\s*:(?!=)/.exec(raw);
current = target === null ? null : target[1];
if (current !== null && !recipes.has(current)) {
recipes.set(current, []);
}
}
return recipes;
};
const recipes = makeRecipes();
const packageScripts = (): Record<string, string> => {
const pkg = JSON.parse(read("package.json")) as {
scripts?: Record<string, string>;
};
return pkg.scripts ?? {};
};
const scripts = packageScripts();
// Node keys: `script/<name>`, `docker:<Dockerfile>`, `make:<target>`,
// `yarn:<package.json script>`, `workflow:<CI workflow file>`.
const resolve = (node: string): string[] => {
if (node.startsWith("script/")) return executable(read(node));
if (node.startsWith("docker:")) {
return executable(read(node.slice("docker:".length)))
.filter((line) => line.startsWith("RUN "))
.map((line) => line.slice("RUN ".length));
}
// The `run:` steps of a workflow, in file order. `uses:` steps are actions,
// not commands, and have no edges into this repo's graph. A `run: |` block
// would resolve to the bare `|`, which reaches nothing and therefore fails
// the count rather than passing quietly.
if (node.startsWith("workflow:")) {
return executable(read(node.slice("workflow:".length)))
.filter((line) => /^-?\s*run:\s*\S/.test(line))
.map((line) => line.replace(/^-?\s*run:\s*/, ""));
}
if (node.startsWith("make:")) {
const target = node.slice("make:".length);
const recipe = recipes.get(target);
// A renamed or deleted target must be a loud failure: silently walking
// an empty recipe would report zero prettier invocations, which reads
// like the tidiest possible result.
if (recipe === undefined) {
throw new Error(`no such Makefile target: ${target}`);
}
return recipe;
}
if (node.startsWith("yarn:")) {
const name = node.slice("yarn:".length);
const script = scripts[name];
if (script === undefined) {
throw new Error(`no such package.json script: ${name}`);
}
return [script];
}
throw new Error(`unresolvable node: ${node}`);
};
// Same reasoning as the missing-target error, applied to every node kind: a
// node that resolves to no commands contributes zero prettier invocations and
// zero edges, which is indistinguishable from a clean result. Fail instead.
const commandsOf = (node: string): string[] => {
const commands = resolve(node);
if (commands.length === 0) {
throw new Error(`node resolved to no commands: ${node}`);
}
return commands;
};
const edgesOf = (line: string): string[] => {
const edges: string[] = [];
// `"$SCRIPT_DIR/lint"`, `"$ROOT/script/lint"` and a bare `script/lint` are
// all the same edge.
for (const match of line.matchAll(
/(?:\$SCRIPT_DIR|\$\{SCRIPT_DIR\}|script)\/([a-z][a-z-]*)/g,
)) {
edges.push(`script/${match[1]}`);
}
// Only real targets: `pkg_install gnumake make make make` in
// script/bootstrap is a package name, not an invocation of this Makefile.
for (const match of line.matchAll(/\bmake\s+([a-z][a-z-]*)/g)) {
if (recipes.has(match[1] ?? "")) edges.push(`make:${match[1]}`);
}
// Same rule for yarn: `yarn run prettier` is the linter itself (counted,
// not followed), `yarn run fmt-check` would be a package.json script that
// runs it indirectly.
for (const match of line.matchAll(/\byarn(?:\s+run)?\s+([a-z][a-z-]*)/g)) {
if ((match[1] ?? "") in scripts) edges.push(`yarn:${match[1]}`);
}
// The container lint pass lives behind a `docker build`; without following
// it the count would miss the one invocation that is supposed to survive.
if (/\bdocker\s+build\b/.test(line)) {
const file = /\s-f\s+(\S+)/.exec(line);
edges.push(`docker:${file === null ? "Dockerfile" : file[1]}`);
}
return edges;
};
interface Walk {
prettier: number;
reached: Set<string>;
}
// Repeated invocations must count repeatedly — running the same script twice is
// exactly the bug — so nodes are not deduplicated. The path stack is only there
// to turn a cycle into a loud failure instead of a hang.
//
// Counting and edge-following both happen for every line: a line that invokes
// prettier can also invoke something else, and skipping the edges of counted
// lines silently truncated the graph.
const walk = (node: string, path: string[] = [], into?: Walk): Walk => {
const result = into ?? { prettier: 0, reached: new Set<string>() };
if (path.includes(node)) {
throw new Error(`invocation cycle: ${[...path, node].join(" -> ")}`);
}
result.reached.add(node);
for (const line of commandsOf(node)) {
result.prettier += countPrettier(line);
for (const edge of edgesOf(line)) {
walk(edge, [...path, node], result);
}
}
return result;
};
describe("prettier runs exactly once per make check", () => {
const check = walk("make:check");
// The headline assertion, and the one the issue is about.
it("invokes prettier once for the whole of make check", () => {
expect(check.prettier).toBe(1);
});
// Guards against the count being 1 (or 0) because the walk never got
// anywhere. `make check` has to reach the suite, the lint script, and the
// Dockerfile whose build IS the lint verdict.
it.each(["script/check", "script/test", "script/lint", "Dockerfile.lint"])(
"reaches %s while counting",
(node) => {
const key = node.startsWith("script/") ? node : `docker:${node}`;
expect([...check.reached]).toContain(key);
},
);
// The one that survives is the container's, not the host's: that is the
// authoritative verdict, since a successful Dockerfile.lint build is what
// CI treats as proof of a clean tree.
it("keeps the surviving invocation inside the lint container", () => {
expect(walk("docker:Dockerfile.lint").prettier).toBe(1);
});
it("does not reach the host formatting check from make check", () => {
expect([...check.reached]).not.toContain("script/fmt-check");
});
});
describe("prettier runs exactly once per CI build", () => {
// Rooted at the workflow file, so this is the graph CI executes rather than
// the graph someone believed CI executes. `make check` cannot stand in for
// it: CI runs script/cibuild, which builds Dockerfile as well as
// Dockerfile.lint, and nothing under `make check` ever reads Dockerfile.
const ci = walk("workflow:.gitea/workflows/check.yml");
it("invokes prettier once for the whole CI build", () => {
expect(ci.prettier).toBe(1);
});
// script/cibuild is here because the workflow is asserted to run it;
// Dockerfile is here because it is the half of the CI graph that the
// `make check` walk cannot see.
it.each([
"script/cibuild",
"script/lint",
"docker:Dockerfile.lint",
"docker:Dockerfile",
])("reaches %s while counting", (node) => {
expect([...ci.reached]).toContain(node);
});
// The test and build image must not lint: linting is Dockerfile.lint's job,
// and a prettier step added here would be a second pass over the same tree
// for the same verdict — on the one path where it matters most.
it("keeps prettier out of the test and build image", () => {
expect(walk("docker:Dockerfile").prettier).toBe(0);
});
});
describe("the standalone entrypoints still do what their names say", () => {
// REPO_POLICIES.md requires both `make lint` and `make fmt-check` to exist
// and mean something. Dropping fmt-check from script/check must not turn it
// into a target nobody can use, and must not leave `make check` passing
// because both halves became no-ops.
it("still checks formatting under make fmt-check", () => {
expect(walk("make:fmt-check").prettier).toBe(1);
});
it("still checks formatting under make lint", () => {
expect(walk("make:lint").prettier).toBe(1);
});
});
describe("script/precommit", () => {
// Same duplication as script/check, same fix. The hook still catches a
// badly formatted tree before the commit lands, because script/lint is the
// container prettier run — that is the whole reason the host call could go.
it("checks formatting exactly once", () => {
expect(walk("script/precommit").prettier).toBe(1);
});
it("gets that check from the lint container", () => {
expect([...walk("script/precommit").reached]).toContain(
"docker:Dockerfile.lint",
);
});
});
// script/bootstrap installs the dependencies, and it has two install sites: one
// for the case where yarn has to be reached through nvm, and one for the case
// where yarn is already on PATH. A substring check against the whole file
// cannot tell them apart, so it reports the first and says nothing about the
// second — which is the one the containers take, because the pinned node image
// ships yarn. Both are resolved separately here.
const installBranches = (): { withoutYarn: string[]; withYarn: string[] } => {
const lines = executable(read("script/bootstrap"));
const open = lines.findIndex((line) =>
/^install_js_deps\s*\(\)/.test(line),
);
if (open === -1) {
throw new Error("script/bootstrap: no install_js_deps function");
}
const close = lines.indexOf("}", open);
const body = lines.slice(open + 1, close === -1 ? undefined : close);
const guard = body.findIndex((line) =>
/^if\b.*\bmissing yarn\b/.test(line),
);
const otherwise = body.indexOf("else", guard);
const end = body.indexOf("fi", otherwise);
if (guard === -1 || otherwise === -1 || end === -1) {
throw new Error(
"script/bootstrap: install_js_deps is not the expected " +
"if missing yarn / else / fi shape",
);
}
return {
withoutYarn: body.slice(guard + 1, otherwise),
withYarn: body.slice(otherwise + 1, end),
};
};
// Every `yarn install` in the given lines, with its flags, so an unpinned
// install cannot hide next to a pinned one.
const yarnInstalls = (lines: string[]): string[] =>
lines.flatMap((line) =>
[...line.matchAll(/\byarn install\b[^"'&|;]*/g)].map((match) =>
match[0].trim(),
),
);
describe("host and container prettier cannot disagree", () => {
// With the host pass gone from `make check`, `make fmt-check` is the only
// host-side formatting check left, and the container is the gate. The two
// must keep producing the same verdict on the same tree, or a developer
// running `make fmt-check` gets a green that CI then rejects.
//
// Three things make them agree, and all three are load-bearing:
it("pins the same prettier for both", () => {
const pkg = JSON.parse(read("package.json")) as {
devDependencies: Record<string, string>;
};
// An exact version, not a range: `^3.8.1` would let the container and
// the host resolve different builds with different formatting.
expect(pkg.devDependencies.prettier).toMatch(/^\d+\.\d+\.\d+$/);
});
it("installs from the lockfile on the branch the container takes", () => {
// Both images are FROM a node image, which ships yarn, so `missing
// yarn` is false and this is the branch that runs in the container.
const installs = yarnInstalls(installBranches().withYarn);
expect(installs).not.toHaveLength(0);
for (const install of installs) {
expect(install).toContain("--frozen-lockfile");
}
});
it("installs from the lockfile on the nvm branch too", () => {
// Not the container's branch, but it is the one a developer without
// yarn on PATH gets, and their prettier has to match the container's.
const installs = yarnInstalls(installBranches().withoutYarn);
expect(installs).not.toHaveLength(0);
for (const install of installs) {
expect(install).toContain("--frozen-lockfile");
}
});
it("runs script/bootstrap inside the lint container", () => {
// Without this the lockfile assertions above would be about a script
// the container never executes.
expect([...walk("docker:Dockerfile.lint").reached]).toContain(
"script/bootstrap",
);
});
it("keeps .gitignore in the build context", () => {
// Prettier 3 reads .gitignore as a default ignore file, so excluding it
// from the context would change which files the container checks.
const dockerignore = read(".dockerignore")
.split("\n")
.map((line) => line.trim());
expect(dockerignore).not.toContain(".gitignore");
});
});
describe("the walk cannot pass vacuously", () => {
// An earlier draft of this file computed a Makefile target as
// `node.slice("make:")` — a string where a number belongs, which coerces to
// NaN and made every target resolve to nothing. The count went to zero and
// an assertion of "not twice" would have been satisfied by a walk that had
// read nothing at all. Every way of reaching nothing is therefore an
// error here, and the ways are tested rather than assumed.
it("reports zero for a subgraph that does not run prettier", () => {
expect(walk("make:clean").prettier).toBe(0);
});
it("refuses a Makefile target that does not exist", () => {
expect(() => walk("make:no-such-target")).toThrow(
/no such Makefile target/,
);
});
it("refuses a package.json script that does not exist", () => {
expect(() => walk("yarn:no-such-script")).toThrow(
/no such package.json script/,
);
});
it("refuses a script that does not exist", () => {
expect(() => walk("script/no-such-script")).toThrow(/ENOENT/);
});
it("refuses a node that resolves to no commands", () => {
// .dockerignore has no RUN steps, standing in for a Dockerfile whose
// steps a restructure moved somewhere the resolver cannot see.
expect(() => walk("docker:.dockerignore")).toThrow(
/resolved to no commands/,
);
});
it("refuses a node kind it does not understand", () => {
expect(() => walk("nonsense")).toThrow(/unresolvable node/);
});
it("refuses to walk in circles", () => {
expect(() => walk("make:check", ["script/check"])).toThrow(
/invocation cycle/,
);
});
});
describe("the resolver reads what the shell would run", () => {
// Counting per line is how `yarn run prettier --check . && yarn run
// prettier --check src` read as a single invocation.
it("counts every prettier invocation on a line", () => {
expect(
countPrettier(
"yarn run prettier --check . && yarn run prettier --check src",
),
).toBe(2);
});
it("does not count the config files as invocations", () => {
expect(countPrettier("COPY .prettierrc .prettierignore ./")).toBe(0);
});
// The counting `continue` also dropped every edge that shared a line with a
// prettier call, so a whole subtree could be hidden behind one `&&`.
it("still follows the edges of a line that invokes prettier", () => {
expect(
edgesOf('yarn run prettier --check . && "$SCRIPT_DIR/lint"'),
).toContain("script/lint");
});
it("resolves every spelling of a script call to one node", () => {
expect(
edgesOf('"$SCRIPT_DIR/lint" "${SCRIPT_DIR}/test" script/fmt'),
).toEqual(["script/lint", "script/test", "script/fmt"]);
});
it("follows a bare docker build to Dockerfile and -f to its file", () => {
expect(edgesOf("docker build .")).toContain("docker:Dockerfile");
expect(edgesOf("docker build -f Dockerfile.lint .")).toContain(
"docker:Dockerfile.lint",
);
});
it("reads the run steps of the CI workflow and not its uses steps", () => {
expect(commandsOf("workflow:.gitea/workflows/check.yml")).toEqual([
"script/cibuild",
]);
});
});

View File

@@ -0,0 +1,30 @@
// `script/projectname` is the single source of the project's name for every
// script that needs one — `script/docker` builds its image tag from it, which
// is the whole reason that file exists. Nothing checked that it agreed with
// `package.json`, and after the repo was renamed it did not: the script still
// said "quack", so `make docker` produced an image tagged after a name this
// project has not used since May.
//
// The script is executed rather than read, because what matters is the string
// it prints, not the source it prints it from.
import { describe, expect, it } from "vitest";
import { execFileSync } from "node:child_process";
import { readFileSync } from "node:fs";
import { fileURLToPath } from "node:url";
import { join } from "node:path";
const repoRoot = fileURLToPath(new URL("../../", import.meta.url));
const pkg = JSON.parse(
readFileSync(join(repoRoot, "package.json"), "utf-8"),
) as { name: string };
describe("script/projectname", () => {
it("prints the name package.json declares", () => {
const printed = execFileSync(join(repoRoot, "script/projectname"), {
cwd: repoRoot,
encoding: "utf-8",
}).trim();
expect(printed).toBe(pkg.name);
});
});

570
test/retry/retry.test.ts Normal file
View File

@@ -0,0 +1,570 @@
/**
* Tests for `src/retry.ts` — the retry policy shared by every network
* operation in quak.
*
* Two things live in that module and they are deliberately separate:
*
* - **`isRetryable(err)`**, a pure classifier. Given an error, is trying
* again capable of producing a different answer? Nothing else about the
* error matters: not how it was logged, not where it came from.
*
* - **`withRetry(fn, opts)`**, the loop. It calls `fn`, and while the error
* is classified retryable and attempts remain, it sleeps and calls `fn`
* again. It never inspects errors itself.
*
* The classifier's default answer is *no*. quak is a backup tool: a wrongly
* retried permanent failure costs a user round trips and delays the rest of
* the run, while a wrongly rejected transient failure costs one file that the
* next run picks up. When in doubt, fail fast.
*
* ## Reading the backoff assertions
*
* `withRetry` takes its `sleep` and its `random` as injected functions. Every
* test here passes a `sleep` that records the delay it was asked for and
* returns immediately, so the suite never waits, and a `random` that returns a
* fixed number, so jitter is exact rather than approximate. **No assertion in
* this file (or anywhere else in the suite) is about elapsed wall-clock time.**
* They are about what `withRetry` asked for, and how many times `fn` ran.
*
* The delay before retry number *n* (1-based) is:
*
* random() * min(maxDelayMs, baseDelayMs * 2 ** (n - 1))
*
* That is exponential backoff with full jitter: the exponential term is the
* *ceiling*, and the actual wait is drawn uniformly below it. Full jitter,
* rather than a fixed delay plus noise, is what stops a client that lost a
* hundred parallel downloads to one CDN blip from re-sending all hundred at
* the same instant.
*/
import { describe, expect, it } from "vitest";
import {
DEFAULT_RETRY_OPTIONS,
isRetryable,
isSafeToReplay,
resolveRetryOptions,
withRetry,
} from "../../src/retry.js";
import { ApiError, TruncatedStreamError } from "../../src/errors.js";
import { ApiError as ApiErrorFromClient } from "../../src/api/client.js";
// ---------------------------------------------------------------------------
// Test helpers
// ---------------------------------------------------------------------------
/**
* A `sleep` that records what it was asked to wait for and returns
* immediately. This is the whole reason `withRetry` takes an injected sleep:
* the retry policy is exercised in full — every branch, every delay — without
* the suite spending a single millisecond waiting.
*/
const recordingSleep = (): {
sleep: (ms: number) => Promise<void>;
delays: number[];
} => {
const delays: number[] = [];
return {
sleep: (ms: number): Promise<void> => {
delays.push(ms);
return Promise.resolve();
},
delays,
};
};
/** An error shaped like a Node transport failure: the errno is on `.code`. */
const errnoError = (code: string, message = code): Error =>
Object.assign(new Error(message), { code });
// ---------------------------------------------------------------------------
// Classification
// ---------------------------------------------------------------------------
describe("isRetryable: HTTP status codes", () => {
it("does not retry ordinary 4xx responses", () => {
// A 4xx is the server saying the request itself is wrong. Repeating
// it verbatim produces the same answer, so retrying only delays the
// failure the caller has to handle. 404 is the load-bearing case:
// `listMissingThumbnails` depends on a 404 arriving promptly and
// exactly once.
for (const status of [400, 401, 403, 404, 409, 410, 422]) {
expect(isRetryable(new ApiError(`HTTP ${status}`, status))).toBe(
false,
);
}
});
it("retries 408 and 429", () => {
// The two 4xx codes that are statements about timing rather than
// about the request. 408 is the server admitting it gave up waiting;
// 429 is it asking for less traffic — which backoff supplies.
expect(isRetryable(new ApiError("timeout", 408))).toBe(true);
expect(isRetryable(new ApiError("slow down", 429))).toBe(true);
});
it("retries every 5xx response", () => {
// A 5xx is the server failing, not the request being wrong. Ente's
// CDN in particular returns 500 and 503 under load.
for (const status of [500, 502, 503, 504, 599]) {
expect(isRetryable(new ApiError(`HTTP ${status}`, status))).toBe(
true,
);
}
});
it("treats a 2xx or 3xx ApiError as not retryable", () => {
// These exist: `getFileStream` raises an ApiError carrying the
// response status when a 200 arrives with a null body. That is a
// malformed response, not a transport failure, and repeating the
// request will produce the same malformed response.
expect(isRetryable(new ApiError("response body is null", 200))).toBe(
false,
);
expect(isRetryable(new ApiError("redirect", 304))).toBe(false);
});
it("uses the same ApiError class that ApiClient exports", () => {
// `ApiError` lives in `src/errors.ts` and is re-exported from
// `src/api/client.ts`, which is where every existing caller and test
// imports it from. If those ever became two separate classes the
// classifier would silently stop recognising errors raised by the
// client, and every 5xx in the wild would be treated as permanent.
expect(ApiErrorFromClient).toBe(ApiError);
expect(new ApiErrorFromClient("boom", 503)).toBeInstanceOf(ApiError);
expect(isRetryable(new ApiErrorFromClient("boom", 503))).toBe(true);
});
});
describe("isRetryable: transport failures", () => {
it("retries a TypeError, which is how fetch reports a failed request", () => {
// Node's fetch rejects with `TypeError: fetch failed` for everything
// below HTTP: DNS failure, refused connection, TLS error, reset
// socket. The real diagnosis is on `cause`, but there is nothing on
// the object that distinguishes it from a TypeError thrown by a bug,
// so this rule is deliberately literal. The cost of the imprecision
// is bounded by the attempt count; the alternative — demanding a
// recognised `cause` — would classify real network failures as
// permanent and fail backups that should have succeeded.
expect(isRetryable(new TypeError("fetch failed"))).toBe(true);
});
it("retries an errno carried on the error itself", () => {
for (const code of [
"ECONNRESET",
"ETIMEDOUT",
"EPIPE",
"ENOTFOUND",
"EAI_AGAIN",
"ECONNREFUSED",
"EHOSTUNREACH",
"ENETUNREACH",
]) {
expect(isRetryable(errnoError(code))).toBe(true);
}
});
it("retries an errno buried in the cause chain", () => {
// undici does not put the errno on the error it throws; it hangs the
// underlying socket error off `cause`, sometimes more than one level
// down. A classifier that only looked at the top-level error would
// see a bare `Error` and call every dropped connection permanent.
const nested = new Error("request to files.ente.io failed", {
cause: new Error("socket hang up", {
cause: errnoError("ECONNRESET", "read ECONNRESET"),
}),
});
expect(isRetryable(nested)).toBe(true);
});
it("accepts a plain object as a cause", () => {
// Not everything in a cause chain is an Error instance.
expect(
isRetryable(new Error("failed", { cause: { code: "ETIMEDOUT" } })),
).toBe(true);
});
it("retries an aborted request", () => {
// `AbortSignal.timeout()` aborts with a `TimeoutError`; an explicit
// `abort()` produces an `AbortError`. quak only ever aborts a request
// on its own deadline, so both mean "this attempt ran out of time",
// which is exactly the condition a later attempt might not hit.
expect(isRetryable(new DOMException("timed out", "TimeoutError"))).toBe(
true,
);
expect(isRetryable(new DOMException("aborted", "AbortError"))).toBe(
true,
);
});
it("does not confuse an unrelated errno with a transport failure", () => {
// A filesystem error surfaces the same way an errno network error
// does. Retrying a full disk or a missing directory is pointless.
expect(isRetryable(errnoError("ENOSPC", "no space left"))).toBe(false);
expect(isRetryable(errnoError("ENOENT", "no such file"))).toBe(false);
expect(isRetryable(errnoError("EACCES", "permission denied"))).toBe(
false,
);
});
it("terminates on a cause chain that points at itself", () => {
// Defensive: `cause` is an arbitrary user-settable property and
// nothing stops it forming a cycle. Without a bound on the walk this
// classifier would hang the process, which is a worse failure than
// any misclassification.
const looped: Error & { cause?: unknown } = new Error("loop");
looped.cause = looped;
expect(isRetryable(looped)).toBe(false);
});
});
describe("isRetryable: stream truncation versus corruption", () => {
it("retries a truncated stream", () => {
// Truncation is a transfer that stopped early. The bytes that did
// arrive are useless, but the file on the server is fine, so asking
// again is exactly right.
expect(isRetryable(new TruncatedStreamError("stream truncated"))).toBe(
true,
);
});
it("retries a truncated stream whose cause is an authentication failure", () => {
// The single-chunk ambiguity, recorded on issue #2 and inherited from
// the truncation work: when a body ends part-way through a chunk,
// Poly1305 fails and carries no framing signal, so a cut connection
// and genuinely corrupt bytes are indistinguishable. That case is
// reported as truncation with the authentication failure preserved as
// `cause`, and it is therefore retried.
//
// Retrying is the deliberate choice. For a multi-chunk body the split
// is real — a corrupt chunk mid-stream stays an authentication
// failure, see the next test — but for a single-chunk body (most
// thumbnails, every small file) a wrong key and a cut connection look
// identical. The cost of guessing wrong is bounded: a few extra round
// trips before the same failure. The cost of guessing the other way
// is a silently truncated file kept forever.
const err = new TruncatedStreamError("stream truncated", {
cause: new Error("secretstream chunk authentication failed"),
});
expect(isRetryable(err)).toBe(true);
});
it("does not retry an authentication failure that is not truncation", () => {
// A whole chunk that failed to authenticate while the stream carried
// on past it cannot be a short transfer. It is corruption or a wrong
// key, and no number of retries fixes either.
expect(
isRetryable(new Error("secretstream chunk authentication failed")),
).toBe(false);
});
});
describe("isRetryable: everything else", () => {
it("does not retry programming errors or unknown values", () => {
expect(isRetryable(new Error("boom"))).toBe(false);
expect(isRetryable(new RangeError("out of range"))).toBe(false);
expect(isRetryable(new SyntaxError("bad JSON"))).toBe(false);
expect(isRetryable("a string")).toBe(false);
expect(isRetryable(undefined)).toBe(false);
expect(isRetryable(null)).toBe(false);
});
});
// ---------------------------------------------------------------------------
// The replay-safety classifier for non-idempotent requests
// ---------------------------------------------------------------------------
describe("isSafeToReplay", () => {
/**
* `isRetryable` answers "could a retry succeed?". For a POST or a PUT
* that is not the whole question: the other half is "could the first
* attempt already have taken effect on the server?".
*
* quak's non-idempotent calls are `/users/srp/create-session`,
* `/users/two-factor/verify` (which consumes one of a limited number of
* 2FA attempts) and `/files/thumbnail`. A blind replay of any of them can
* do real damage, so they retry only on the failures that establish no TCP
* connection to the server ever existed — DNS produced no address, or the
* peer refused the connection — and therefore that no request byte can
* have been transmitted.
*/
it("replays only failures where the connection was never established", () => {
for (const code of ["ENOTFOUND", "EAI_AGAIN", "ECONNREFUSED"]) {
expect(isSafeToReplay(errnoError(code))).toBe(true);
}
// Also when undici has buried it, which is how it actually arrives.
expect(
isSafeToReplay(
new TypeError("fetch failed", {
cause: errnoError("ECONNREFUSED"),
}),
),
).toBe(true);
});
it("does not replay a failure that could have happened after the server acted", () => {
// Every one of these is ambiguous about whether the server processed
// the request. A 5xx proves it did. A reset or a broken pipe can
// arrive after the request was fully sent and handled. A timeout says
// nothing at all about the server's state. A bare `fetch failed` with
// no recognisable cause could be any of them.
expect(isSafeToReplay(new ApiError("HTTP 500", 500))).toBe(false);
expect(isSafeToReplay(new ApiError("HTTP 429", 429))).toBe(false);
expect(isSafeToReplay(errnoError("ECONNRESET"))).toBe(false);
expect(isSafeToReplay(errnoError("EPIPE"))).toBe(false);
expect(isSafeToReplay(errnoError("ETIMEDOUT"))).toBe(false);
// The routing errnos look like connect-time failures but are not. On
// Linux an ICMP destination-unreachable delivered on an established
// connection sets the socket error, and the next read or write returns
// `EHOSTUNREACH` or `ENETUNREACH`; a local interface going down after
// the request was fully written surfaces as `ENETDOWN` the same way.
// In each case the server may already have consumed the request — a
// replayed `/users/two-factor/verify` would burn a second attempt.
// They remain retryable for the idempotent calls; this asserts only
// that they are not replayable.
expect(isSafeToReplay(errnoError("EHOSTUNREACH"))).toBe(false);
expect(isSafeToReplay(errnoError("ENETUNREACH"))).toBe(false);
expect(isSafeToReplay(errnoError("ENETDOWN"))).toBe(false);
// ...and that the narrowing did not make them non-retryable.
expect(isRetryable(errnoError("EHOSTUNREACH"))).toBe(true);
expect(isRetryable(errnoError("ENETUNREACH"))).toBe(true);
expect(isRetryable(errnoError("ENETDOWN"))).toBe(true);
expect(
isSafeToReplay(new DOMException("timed out", "TimeoutError")),
).toBe(false);
expect(isSafeToReplay(new TypeError("fetch failed"))).toBe(false);
});
});
// ---------------------------------------------------------------------------
// The retry loop
// ---------------------------------------------------------------------------
describe("withRetry", () => {
it("calls the function once and does not sleep when it succeeds", async () => {
const { sleep, delays } = recordingSleep();
let calls = 0;
const result = await withRetry(
() => {
calls++;
return Promise.resolve("ok");
},
{ sleep },
);
expect(result).toBe("ok");
expect(calls).toBe(1);
expect(delays).toEqual([]);
});
it("stops at the first success and returns its value", async () => {
const { sleep, delays } = recordingSleep();
let calls = 0;
const result = await withRetry(
() => {
calls++;
if (calls < 3) {
return Promise.reject(new ApiError("HTTP 503", 503));
}
return Promise.resolve(calls);
},
{ attempts: 5, sleep },
);
expect(result).toBe(3);
expect(calls).toBe(3);
// Two failures, so two waits — and none after the attempt that
// succeeded.
expect(delays).toHaveLength(2);
});
it("gives up after `attempts` calls and throws the last error", async () => {
// `attempts` counts calls, not retries: `attempts: 3` means the
// function runs three times in total. The error that escapes is the
// one from the final attempt, because that is the current state of
// the world; a caller logging it is logging what is true now.
const { sleep, delays } = recordingSleep();
let calls = 0;
const failure = withRetry(
() => {
calls++;
return Promise.reject(new ApiError(`attempt ${calls}`, 503));
},
{ attempts: 3, sleep },
);
await expect(failure).rejects.toThrow("attempt 3");
expect(calls).toBe(3);
// Three attempts, two gaps between them. Sleeping after the last
// attempt would delay the caller's failure for nothing.
expect(delays).toHaveLength(2);
});
it("does not retry at all when attempts is 1", async () => {
const { sleep, delays } = recordingSleep();
let calls = 0;
await expect(
withRetry(
() => {
calls++;
return Promise.reject(new ApiError("HTTP 500", 500));
},
{ attempts: 1, sleep },
),
).rejects.toThrow("HTTP 500");
expect(calls).toBe(1);
expect(delays).toEqual([]);
});
it("rethrows a non-retryable error immediately", async () => {
const { sleep, delays } = recordingSleep();
let calls = 0;
await expect(
withRetry(
() => {
calls++;
return Promise.reject(new ApiError("HTTP 404", 404));
},
{ attempts: 5, sleep },
),
).rejects.toBeInstanceOf(ApiError);
expect(calls).toBe(1);
expect(delays).toEqual([]);
});
it("preserves the error object, not just its message", async () => {
// Callers classify what escapes: `listMissingThumbnails` needs the
// `ApiError` and its status to tell a genuine 404 from a transient
// failure. Wrapping the error in a "retries exhausted" error would
// break that.
const original = new ApiError("gone", 410, { code: "GONE" });
const err: unknown = await withRetry(() => Promise.reject(original), {
sleep: () => Promise.resolve(),
}).catch((e: unknown) => e);
expect(err).toBe(original);
});
it("honours a caller-supplied classifier", async () => {
// This is how the non-idempotent call sites narrow the policy: same
// loop, same backoff, stricter question.
const { sleep } = recordingSleep();
let calls = 0;
await expect(
withRetry(
() => {
calls++;
// Retryable under the default policy...
return Promise.reject(new ApiError("HTTP 503", 503));
},
{ attempts: 4, sleep, isRetryable: isSafeToReplay },
),
).rejects.toThrow("HTTP 503");
// ...but not under `isSafeToReplay`, so it ran exactly once.
expect(calls).toBe(1);
});
});
describe("withRetry backoff", () => {
it("doubles the ceiling on each retry and caps it", async () => {
// `random: () => 1` pins the jitter to the top of its range, which
// makes the ceiling itself observable. The sequence is
// base, base*2, base*4, ... clamped at maxDelayMs — so a long outage
// settles into a steady poll instead of growing to hours.
const { sleep, delays } = recordingSleep();
await expect(
withRetry(() => Promise.reject(new ApiError("HTTP 500", 500)), {
attempts: 6,
baseDelayMs: 100,
maxDelayMs: 250,
sleep,
random: () => 1,
}),
).rejects.toThrow();
expect(delays).toEqual([100, 200, 250, 250, 250]);
});
it("draws each delay uniformly below its ceiling", async () => {
// Full jitter. The exponential value is the maximum wait, not the
// wait itself, so a fleet of clients that failed together does not
// come back in lockstep.
const { sleep, delays } = recordingSleep();
await expect(
withRetry(() => Promise.reject(new ApiError("HTTP 500", 500)), {
attempts: 4,
baseDelayMs: 100,
maxDelayMs: 10_000,
sleep,
random: () => 0.25,
}),
).rejects.toThrow();
expect(delays).toEqual([25, 50, 100]);
});
it("never asks to sleep longer than the cap or less than zero", async () => {
// Whatever `random` returns from its [0, 1) contract, the delay stays
// inside the configured envelope.
const draws = [0, 0.999_999, 0.5, 0.1, 0.9];
let i = 0;
const { sleep, delays } = recordingSleep();
await expect(
withRetry(() => Promise.reject(new ApiError("HTTP 500", 500)), {
attempts: 6,
baseDelayMs: 1000,
maxDelayMs: 2000,
sleep,
random: () => draws[i++]!,
}),
).rejects.toThrow();
expect(delays).toHaveLength(5);
for (const d of delays) {
expect(d).toBeGreaterThanOrEqual(0);
expect(d).toBeLessThanOrEqual(2000);
}
expect(delays[0]).toBe(0);
});
});
describe("retry defaults", () => {
it("ships a bounded, documented default policy", () => {
// These are the numbers the README documents. They are asserted here
// so the README and the code cannot drift apart silently. Four
// attempts at a 500ms base put the three ceilings at 500, 1000 and
// 2000ms, so a file that is going to fail gives up after at most
// three and a half seconds of waiting — which keeps a
// several-thousand-file backup moving past a bad file rather than
// stalling on it. The 10s cap only comes into play for a caller that
// raises the attempt count.
expect(DEFAULT_RETRY_OPTIONS.attempts).toBe(4);
expect(DEFAULT_RETRY_OPTIONS.baseDelayMs).toBe(500);
expect(DEFAULT_RETRY_OPTIONS.maxDelayMs).toBe(10_000);
});
it("fills in only the fields the caller left out", () => {
const resolved = resolveRetryOptions({ attempts: 2 });
expect(resolved.attempts).toBe(2);
expect(resolved.baseDelayMs).toBe(DEFAULT_RETRY_OPTIONS.baseDelayMs);
expect(resolved.maxDelayMs).toBe(DEFAULT_RETRY_OPTIONS.maxDelayMs);
expect(typeof resolved.sleep).toBe("function");
expect(typeof resolved.random).toBe("function");
});
it("resolves to the defaults when given nothing", () => {
expect(resolveRetryOptions()).toEqual(DEFAULT_RETRY_OPTIONS);
expect(resolveRetryOptions({})).toEqual(DEFAULT_RETRY_OPTIONS);
});
it("defaults random to a real generator in [0, 1)", () => {
const { random } = resolveRetryOptions();
for (let i = 0; i < 100; i++) {
const r = random();
expect(r).toBeGreaterThanOrEqual(0);
expect(r).toBeLessThan(1);
}
});
});

View File

@@ -32,6 +32,7 @@ import {
fixMissingThumbnails, fixMissingThumbnails,
} from "../../src/thumbnails.js"; } from "../../src/thumbnails.js";
import type { KeyAttributes } from "../../src/auth/types.js"; import type { KeyAttributes } from "../../src/auth/types.js";
import type { RetryOptions } from "../../src/retry.js";
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
// Mock server with controllable thumbnail behavior // Mock server with controllable thumbnail behavior
@@ -51,7 +52,7 @@ interface ThumbMockState {
filesByCollection: Record<number, Record<string, unknown>[]>; filesByCollection: Record<number, Record<string, unknown>[]>;
fileCiphertexts: Record<number, Uint8Array>; fileCiphertexts: Record<number, Uint8Array>;
fileKeys: Record<number, Uint8Array>; fileKeys: Record<number, Uint8Array>;
thumbnailBehavior: Record<number, "ok" | "empty" | "404">; thumbnailBehavior: Record<number, "ok" | "empty" | "404" | "500">;
// Captures from fix operations // Captures from fix operations
uploadedThumbnails: { uploadedThumbnails: {
fileID: number; fileID: number;
@@ -302,6 +303,9 @@ const buildThumbFetch = (m: ThumbMockState) => {
if (behavior === "empty") { if (behavior === "empty") {
return new Response(new Uint8Array(0), { status: 200 }); return new Response(new Uint8Array(0), { status: 200 });
} }
if (behavior === "500") {
return new Response("Internal Server Error", { status: 500 });
}
return new Response("not found", { status: 404 }); return new Response("not found", { status: 404 });
} }
@@ -345,6 +349,38 @@ const buildThumbFetch = (m: ThumbMockState) => {
}) as typeof globalThis.fetch; }) as typeof globalThis.fetch;
}; };
/**
* A retry policy with the waiting removed. `listMissingThumbnails` walks every
* file in the account, so a transient failure is retried; without an injected
* `sleep` these tests would spend real seconds waiting out backoff.
*/
const noWait: RetryOptions = {
sleep: () => Promise.resolve(),
random: () => 0,
};
/** Wrap a fetch so the tests can count how often one endpoint was hit. */
const countingFetch = (
inner: typeof globalThis.fetch,
match: (url: string) => boolean,
): { fetch: typeof globalThis.fetch; matched: () => number } => {
let matched = 0;
const fake = async (
input: RequestInfo | URL,
init?: RequestInit,
): Promise<Response> => {
const url =
typeof input === "string"
? input
: input instanceof URL
? input.href
: input.url;
if (match(url)) matched++;
return inner(input, init);
};
return { fetch: fake as typeof globalThis.fetch, matched: () => matched };
};
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
// Tests // Tests
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
@@ -378,7 +414,77 @@ describe("listMissingThumbnails", () => {
expect(emptyEntry.collection).toBe("Photos"); expect(emptyEntry.collection).toBe("Photos");
const notFoundEntry = missing.find((m) => m.fileID === 102)!; const notFoundEntry = missing.find((m) => m.fileID === 102)!;
expect(notFoundEntry.reason).toContain("fetch failed"); // A 404 is the server stating the thumbnail is not there. That is the
// only network answer that means "missing", and the reason says so
// rather than the older catch-all "fetch failed" — which used to
// cover a 500 and a dropped connection too.
expect(notFoundEntry.reason).toContain("not found");
});
it("does not report a thumbnail as missing when the server is failing", async () => {
// The distinction that matters for `helper fix-missing-thumbnails`.
// Reporting a file here leads to downloading the original,
// regenerating a thumbnail, and uploading it over a thumbnail that
// was fine all along — because the server was briefly returning 500s.
//
// File 102 serves 500 on every attempt, so the retries are genuinely
// exhausted. It must still not be reported.
const failingMock = await buildThumbMock();
failingMock.thumbnailBehavior[102] = "500";
const counted = countingFetch(
buildThumbFetch(failingMock),
(url) => url.includes("thumbnails.ente.io") && url.includes("102"),
);
const client = await Client.login({
email: TEST_EMAIL,
password: TEST_PASSWORD,
apiOptions: { fetch: counted.fetch, retry: { ...noWait } },
});
const missing = await listMissingThumbnails(client);
// Only the genuinely empty thumbnail is reported.
expect(missing.map((m) => m.fileID)).toEqual([101]);
// And the 500 was retried rather than accepted as an answer: four
// attempts is the library default.
expect(counted.matched()).toBe(4);
});
it("does not report a thumbnail as missing when the connection fails", async () => {
// Same rule for a transport failure, which carries no status at all.
const failingMock = await buildThumbMock();
const inner = buildThumbFetch(failingMock);
let thumbRequests = 0;
const fetch = (async (
input: RequestInfo | URL,
init?: RequestInit,
): Promise<Response> => {
const url =
typeof input === "string"
? input
: input instanceof URL
? input.href
: input.url;
if (url.includes("thumbnails.ente.io") && url.includes("102")) {
thumbRequests++;
throw Object.assign(new Error("socket hang up"), {
code: "ECONNRESET",
});
}
return inner(input, init);
}) as typeof globalThis.fetch;
const client = await Client.login({
email: TEST_EMAIL,
password: TEST_PASSWORD,
apiOptions: { fetch, retry: { ...noWait } },
});
const missing = await listMissingThumbnails(client);
expect(missing.map((m) => m.fileID)).toEqual([101]);
expect(thumbRequests).toBe(4);
}); });
it("deduplicates files seen in multiple collections", async () => { it("deduplicates files seen in multiple collections", async () => {

View File

@@ -5,7 +5,8 @@
"moduleResolution": "NodeNext", "moduleResolution": "NodeNext",
"lib": ["ES2022"], "lib": ["ES2022"],
"outDir": "./dist", "outDir": "./dist",
"rootDir": "./src", "rootDir": ".",
"noEmitOnError": true,
"strict": true, "strict": true,
"noImplicitOverride": true, "noImplicitOverride": true,
"noUncheckedIndexedAccess": true, "noUncheckedIndexedAccess": true,