7 Commits

Author SHA1 Message Date
156fe871e8 Make the image build multi-stage and cache-proof (closes #4)
All checks were successful
check / check (push) Successful in 23s
The image build reported a green it had not earned. `script/cibuild` is a
bare `docker build .`, and with `COPY . .` followed by `RUN make check`,
an unchanged tree served that layer from cache: the suite never ran and
the build still exited 0, while the script's header comment asserted the
opposite.

CHECK_EPOCH, passed by `script/cibuild` and `script/docker`, changes the
cache key of the check and build layers on every invocation. It is
guarded, because an unset ARG is the empty string and therefore a stable
key: without the guard a plain `docker build .` — the command the policy
names, and the one anyone debugging types — would still get the false
green. A missing argument is now a hard failure rather than a silent
degradation to the behaviour the epoch was added to prevent.

The Dockerfile is now two stages: `fmt-check` and `lint` run first, and
the check stage takes a `COPY --from=lint` dependency on them, so a
formatting mistake fails the build in seconds instead of racing the
suite to the finish. Both stages stay pinned to the same digest.

The remaining fixes are one-liners that had made the target unusable:
`script/projectname` still printed the pre-rename name, so `make docker`
tagged its image after a name this project dropped in May;
`script/bootstrap` installed without fetching apt's package lists, which
cannot work on a Debian base; and `.dockerignore` had drifted far enough
from `.gitignore` to ship a ~100 MB compiled binary and any agent
worktree under `.claude/` into the build context. The second of those is
a correctness problem, not a size one — vitest globs a copied worktree's
tests alongside the real ones and runs the suite twice over. `.gitignore`
itself stays in the context, because prettier reads it as a default
ignore file and dropping it would change what `make fmt-check` sees.
2026-08-09 14:38:32 +00:00
118f8e22c3 Add failing tests for the project name and build context
script/projectname still prints the pre-rename name, so make docker tags
its image quack, and .dockerignore has drifted from .gitignore: a compiled
bin/quak, the vitest and tsc caches, the CLI's runtime directory and the
local worktree directory all reach the build context. Neither failure is
visible in a build that exits 0.

Red until the fixes land.
2026-08-09 14:27:50 +00:00
69bd6d1539 Compile bin/ alongside src/ so the package can be built (closes #3)
All checks were successful
check / check (push) Successful in 5s
rootDir was ./src while include also matched bin/**/*, which is TS6059: tsc
refuses to emit at all when a compiled file sits outside rootDir. rootDir is
now the repository root, which is the smallest change that makes the two
agree and leaves the source layout the README documents alone. Output keeps
the shape of the source tree, so main and types move to dist/src/index.js and
dist/src/index.d.ts while bin.quak stays at dist/bin/quak.js. The alternative,
moving the CLI body into src/ behind a shim in bin/, would hold main at
dist/index.js at the cost of churning the CLI and contradicting the layout
diagram in the README.

Clearing TS6059 exposed two type errors that had never been reached, because
the config error aborts before checking: StateAddress was read as a namespace
member off the default import, and the secretstream pull was called without
the additional-data argument, which libsodium does not make optional. The type
is now taken from the module's named export and the pull passes null for ad,
matching the null already passed on the push side in encryptBlob. Neither
changes what runs.

noEmitOnError stops a failed build from leaving output behind. It emitted
despite the error before, which is how a stale bin/quak.js came to sit next to
bin/quak.ts in a working tree, where eslint then read it and failed make check
on a generated file.

script/build compiles and then checks that the files package.json advertises
are among the ones the compiler wrote, since tsc knows nothing about the
manifest and a green build could still ship a package whose main resolves to
nothing. It also sets the executable bit on the bin entries, which tsc does
not carry over from the source even though it does copy the shebang. The
Makefile target is now a shim over it, as the other targets are, and
package.json's build script points at it so yarn build gets the same checks.

The Dockerfile runs make build after make check, so a branch that does not
compile cannot reach main. What make check itself runs is unchanged.

package.json gains a quak script, so the yarn quak commands the README's
Getting Started block has always listed resolve to the built CLI.
2026-08-09 10:06:31 +00:00
d79ed83f4d Add failing tests for the package entrypoint contract
The manifests disagreed and nothing noticed. tsconfig.json set rootDir to
./src while include also matched bin/**/*, which is TS6059, so no build had
succeeded; package.json meanwhile advertised main, types and a bin that a
successful build would have to produce. make check runs test, lint and
fmt-check, so neither half was ever exercised.

These tests read tsconfig.json and package.json and assert the contract
between them without invoking a compiler, which keeps them in the fast unit
suite: every include pattern must root under rootDir, and main, types and
bin.quak must equal the paths tsc will emit for src/index.ts and bin/quak.ts.
They also require a quak script pointing at the built CLI, which the README's
Getting Started block has always told the reader to run.

Three of them fail at this commit.
2026-08-09 10:01:32 +00:00
348f23bac9 Narrow the replay errno set to failures that prove no connection existed
All checks were successful
check / check (push) Successful in 4s
`CONNECT_CODES` drives `isSafeToReplay`, which is the only thing standing
between a transport failure and a replayed `POST /users/two-factor/verify`.
It included `EHOSTUNREACH`, `ENETUNREACH` and `ENETDOWN` on the stated
grounds that those errnos can only be reported before any request byte was
written. That is not true on Linux: an ICMP destination-unreachable
delivered on an already-established connection sets the socket error and
the next read or write returns `EHOSTUNREACH` or `ENETUNREACH`, and a local
interface going down after the request was fully written surfaces as
`ENETDOWN` the same way. In each case the server may already have received
and acted on the request -- exactly the ambiguity the rule exists to
exclude, on the paths that consume a second-factor attempt or register a
thumbnail.

The three are dropped from `CONNECT_CODES` and stay in `TRANSPORT_CODES`,
so they remain retryable for the idempotent calls; only replay eligibility
narrows. What is left -- `ENOTFOUND`, `EAI_AGAIN`, `ECONNREFUSED` -- means
no TCP connection to the server ever existed, so no request byte can have
been transmitted.

The justification is corrected everywhere it was stated: the comment on
`CONNECT_CODES`, the one on `isSafeToReplay`, the `postJSON` call site, the
README's idempotency section and the `client.test.ts` docblock. All of them
now describe what the narrowed set actually establishes rather than
claiming a proof it did not support.

The narrowing is enforced by the suite rather than asserted in a comment:
the three errnos join `ECONNRESET`/`EPIPE`/`ETIMEDOUT` in the
`isSafeToReplay`-returns-false test, with companion `isRetryable` assertions
so a future edit cannot make them non-retryable by accident. Putting the
three back into `CONNECT_CODES` turns that test red (1 failure, verified).
2026-08-09 05:41:59 +00:00
f3cf4af833 Retry transient network failures with exponential backoff (closes #2)
All checks were successful
check / check (push) Successful in 20s
No retry on 4xx, backoff on 5xx and transport failures, and a deadline on
every request. Before this, one transient 503 or TCP reset failed a file
for good, and a CDN connection that went quiet after accepting the request
blocked `quak backup` forever, because there was no timeout anywhere.

src/retry.ts holds the policy: a classifier that decides whether another
attempt could produce a different answer, and a loop that acts on it with
exponential backoff and full jitter. Retried: 5xx, 408, 429, transport
failures (the errno is read out of the cause chain, which is where Node's
fetch puts it), deadline aborts, and truncated transfers. Not retried:
every other 4xx, and anything unrecognised — a wrongly retried permanent
failure delays every remaining file, while a wrongly abandoned transient
one costs a single file the next run picks up. Attempt count, delays,
sleep and jitter source are all configurable through ApiClientOptions;
sleep being injectable is what lets the suite exercise the policy without
waiting.

Truncation needed a type before it could be classified. streamDecrypt
threw plain Errors whose messages began "download: stream truncated", and
classifying on message text would mean the next reword silently turned
every truncated download into a permanent failure. It now throws
TruncatedStreamError, which lives in src/errors.ts alongside ApiError so
the classifier can recognise both without importing the modules that
import it; api/client.ts re-exports ApiError, so it stays one class and
every existing import path still resolves.

Downloads retry the request, the stream consumption and the decryption
together. Only the first of those happens inside ApiClient: a socket
reset after the headers arrived throws in streamDecrypt, and retrying the
request alone would never see it. The client's own retry is switched off
for those two calls so the budgets do not multiply into sixteen requests
per file, and the atomic write stays outside the loop so a download that
took three attempts still performs one write and one rename.

Non-idempotent requests are not blindly replayed. postJSON and putJSON
reach create-session, two-factor/verify — which burns one of a few
second-factor attempts — and files/thumbnail, so they retry only when the
connection was never established and the server provably never saw the
request. putFile is exempt and retries fully: a presigned PUT stores one
whole object at one key, with no partial state to damage. It now throws
ApiError with the status, as do the two null-body paths, which previously
threw bare Errors that nothing could classify.

Timeouts come from AbortSignal.timeout(), renewed per attempt: 30s for
JSON and upload calls, 10 minutes for file bodies, since a value short
enough to keep a hung API call from stalling a backup would cancel a
legitimate multi-gigabyte download. The download deadline is enforced
over the body rather than only the headers, by racing each read against
the signal, so the guarantee does not depend on the fetch implementation
tearing down a stream it already handed over.

listMissingThumbnails now separates a genuine 404 from an exhausted
retry. Its bare catch reported both as missing, which after this change
would have let a few minutes of 500s talk fix-missing-thumbnails into
regenerating and re-uploading thumbnails that were fine. runBackup and
runMetadataBackup are untouched: the retry sits below them and their
per-file resilience is unchanged.
2026-08-09 05:21:31 +00:00
0cbe338b58 Add failing tests for the download retry policy
Tests only; the retry module they import does not exist yet, so the
branch is red at this commit.

New test/retry/retry.test.ts documents the classifier and the backoff:
which errors are worth another attempt, which are not, and how the delay
before each retry is derived. It asserts on the arguments handed to an
injected sleep function rather than on elapsed time, so the suite never
waits and the numbers are exact.

test/api/client.test.ts gains the request-count contract for each of the
six call sites, the deadline behaviour, the ApiError typing that the
presigned PUT and the null-body paths need in order to be classified at
all, and the replay rule for the two non-idempotent methods.

test/download/download.test.ts gains the case that motivates the whole
design: a socket reset after the response headers arrived, which happens
below ApiClient and can only be caught by retrying the request, the
stream consumption and the decryption together. It also pins that a
retried download stages exactly one temp file, and that the two retry
layers do not compose into a multiplied request budget. The existing
truncation tests now assert on the error type rather than its wording,
since that type is what the classifier reads.

test/thumbnails/thumbnails.test.ts separates a genuine 404 from an
exhausted retry, so a failing server can no longer make
fix-missing-thumbnails re-upload thumbnails that already exist.
2026-08-09 05:11:07 +00:00
28 changed files with 2606 additions and 134 deletions

View File

@@ -1,8 +1,50 @@
# Mirrors .gitignore, with one deliberate exception: .gitignore itself stays
# in the build context, because prettier 3 reads it as a default ignore file
# and dropping it would change what `make fmt-check` sees inside the image.
# VCS
.git .git
# OS
.DS_Store
Thumbs.db
# Editors
*.swp
*.swo
*~
*.bak
.idea/
.vscode/
*.sublime-*
# Node
node_modules node_modules
# TypeScript / build artifacts
dist dist
build build
*.tsbuildinfo
coverage coverage
.DS_Store .nyc_output/
# Vitest
.vitest-cache/
# Environment / secrets
.env .env
.env.* .env.*
*.pem
*.key
# Compiled binary (built by make build-bin); around 100 MB
bin/quak
# quak runtime data (in case anyone runs the CLI from inside the repo)
.quak/
# Local per-developer tool state, including agent worktrees. Correctness,
# not context size: a worktree copied in here has its own test/ tree, which
# vitest globs alongside the real one, so the containerised suite runs N+1
# times over and still reports success.
.claude/

3
.gitignore vendored
View File

@@ -36,5 +36,6 @@ bin/quak
# quak runtime data (in case anyone runs the CLI from inside the repo) # quak runtime data (in case anyone runs the CLI from inside the repo)
.quak/ .quak/
# Local Claude Code settings (per-developer) # Local per-developer tool settings and scratch state, including the
# worktrees agents check out under this directory
.claude/ .claude/

View File

@@ -1,10 +1,37 @@
# node 22-alpine, 2026-02-22 # Lint stage — fast feedback on formatting and lint issues
FROM node@sha256:e4bf2a82ad0a4037d28035ae71529873c069b13eb0455466ae0bc13363826e34 # node 22.22.0 on Alpine 3.23.3 (node:22-alpine), 2026-08-09
FROM node@sha256:e4bf2a82ad0a4037d28035ae71529873c069b13eb0455466ae0bc13363826e34 AS lint
WORKDIR /app WORKDIR /app
COPY script/ script/ COPY script/ script/
COPY package.json yarn.lock ./ COPY package.json yarn.lock ./
RUN script/bootstrap RUN script/bootstrap
COPY . . COPY . .
RUN make fmt-check
RUN make lint
# Check stage — the full suite and the build
# node 22.22.0 on Alpine 3.23.3 (node:22-alpine), 2026-08-09
FROM node@sha256:e4bf2a82ad0a4037d28035ae71529873c069b13eb0455466ae0bc13363826e34 AS check
WORKDIR /app
# Force BuildKit to run the lint stage before proceeding. Without this the
# two stages run in parallel and a lint failure can lose the race.
COPY --from=lint /app/yarn.lock /dev/null
COPY script/ script/
COPY package.json yarn.lock ./
RUN script/bootstrap
COPY . .
# CHECK_EPOCH is a cache buster: without it Docker serves `make check` from
# cache on an unchanged tree, the suite never executes, and the build still
# exits 0. The guard makes an absent argument a hard failure — an unset ARG
# is the empty string, which is a perfectly stable cache key, so a plain
# `docker build .` would otherwise still get the false green. Fail closed.
ARG CHECK_EPOCH
RUN [ -n "$CHECK_EPOCH" ] || exit 1
RUN make check RUN make check
ARG CHECK_EPOCH
RUN [ -n "$CHECK_EPOCH" ] || exit 1
RUN make build

View File

@@ -24,7 +24,7 @@ check:
@script/check @script/check
build: build:
@$(YARN) tsc @script/build
build-bin: build-bin:
nix-shell -p bun --run "bun build bin/quak.ts --compile --outfile bin/quak" nix-shell -p bun --run "bun build bin/quak.ts --compile --outfile bin/quak"

131
README.md
View File

@@ -81,6 +81,9 @@ alpine. We provide:
`script/bootstrap`, then `script/install-precommit` `script/bootstrap`, then `script/install-precommit`
- `script/projectname` — output the project name (our own extension); used by - `script/projectname` — output the project name (our own extension); used by
`script/docker` for the image tag `script/docker` for the image tag
- `script/build` — compile the TypeScript sources into `dist/`, then verify that
the entrypoints `package.json` declares (`main`, `types`, `bin`) are among the
files the compiler wrote, and make the CLI executable (our own extension)
- `script/test` — run the test suite (vitest, hard-capped at 30s where `timeout` - `script/test` — run the test suite (vitest, hard-capped at 30s where `timeout`
is available, verbose rerun on failure) is available, verbose rerun on failure)
- `script/lint` — run eslint and a prettier check - `script/lint` — run eslint and a prettier check
@@ -89,9 +92,9 @@ alpine. We provide:
- `script/check` — run all checks: `test`, `lint`, `fmt-check` (our own - `script/check` — run all checks: `test`, `lint`, `fmt-check` (our own
extension) extension)
- `script/docker` — build the Docker image, tagged via `script/projectname` - `script/docker` — build the Docker image, tagged via `script/projectname`
(byte-identical across repos) - `script/cibuild` — cd to the repo root and run the image build (what CI runs;
- `script/cibuild` — cd to the repo root and `docker build .` (what CI runs; the the build runs `make fmt-check` and `make lint` in a first stage, then
image build runs `make check`) `make check` and `make build` in a second)
- `script/precommit` — run by the git pre-commit hook (our own extension); runs - `script/precommit` — run by the git pre-commit hook (our own extension); runs
`script/lint` and `script/fmt-check` but deliberately not the tests, so the `script/lint` and `script/fmt-check` but deliberately not the tests, so the
TDD red-phase commit can land TDD red-phase commit can land
@@ -100,6 +103,15 @@ alpine. We provide:
`make hooks` installs the pre-commit hook that runs `script/precommit`. `make hooks` installs the pre-commit hook that runs `script/precommit`.
Both `script/docker` and `script/cibuild` pass
`--build-arg CHECK_EPOCH="$(date +%s)"`. The Dockerfile refuses to build without
it. This is deliberate: on an unchanged tree Docker would otherwise serve the
`make check` layer from cache, so the suite would never run and the build would
still exit 0. A changing epoch invalidates the check and build layers on every
invocation while leaving the dependency layers below them cached, and the
missing-argument guard means a bare `docker build .` fails loudly instead of
quietly reporting a green it did not earn.
## Rationale ## Rationale
Ente is one of very few photo services with a credible end-to-end encryption Ente is one of very few photo services with a credible end-to-end encryption
@@ -133,8 +145,8 @@ All work on quak is test-driven. No exceptions.
3. Subsequent commits add the implementation and any refactors needed to make 3. Subsequent commits add the implementation and any refactors needed to make
the tests pass. the tests pass.
4. A feature branch can only be merged into `main` when `make check` is green. 4. A feature branch can only be merged into `main` when `make check` is green.
`main` is always green. The Dockerfile runs `make check`, so a red branch `main` is always green. The Dockerfile runs `make check` and `make build`, so
cannot pass CI. neither a red branch nor one that does not compile can pass CI.
5. Tests are the canonical API documentation for this library. Every test file 5. Tests are the canonical API documentation for this library. Every test file
is commented thoroughly enough that a reader who has never seen quak can is commented thoroughly enough that a reader who has never seen quak can
learn how to use it from the tests alone. Comments explain why a behavior learn how to use it from the tests alone. Comments explain why a behavior
@@ -150,8 +162,9 @@ All work on quak is test-driven. No exceptions.
8. The pre-commit hook installed by `make hooks` runs `script/precommit`, which 8. The pre-commit hook installed by `make hooks` runs `script/precommit`, which
runs the lint and format checks but not the full `make check`. This is runs the lint and format checks but not the full `make check`. This is
deliberate so the TDD red-phase commit (failing tests, no implementation yet) deliberate so the TDD red-phase commit (failing tests, no implementation yet)
can land. The full `make check` runs as part of `docker build .`, which is can land. The full `make check` runs as part of the image build, which is
what CI executes, so a red branch still cannot reach `main`. what CI executes via `script/cibuild`, so a red branch still cannot reach
`main`.
## Design ## Design
@@ -169,6 +182,8 @@ quak/
model/ decrypted Collection, File, Metadata types + decrypt fns model/ decrypted Collection, File, Metadata types + decrypt fns
download/ streaming file/thumbnail download + decryption download/ streaming file/thumbnail download + decryption
backup.ts resilient full-account backup with dedup backup.ts resilient full-account backup with dedup
errors.ts error types shared across layers
retry.ts retry classifier + exponential backoff with jitter
thumbnails.ts detect + regenerate missing thumbnails thumbnails.ts detect + regenerate missing thumbnails
client.ts high-level Client class assembled from the above client.ts high-level Client class assembled from the above
index.ts public library exports index.ts public library exports
@@ -181,6 +196,12 @@ quak/
tsconfig.json tsconfig.json
``` ```
`make build` compiles that tree into `dist/`, preserving its shape: the library
lands in `dist/src/` and the CLI in `dist/bin/quak.js`, which is what
`package.json` points `main`, `types` and `bin` at. The compiler's `rootDir` is
the repository root rather than `src/`, because `bin/` is compiled too and
`rootDir` has to contain everything that is compiled.
### Cryptography ### Cryptography
All cryptography is done by `libsodium-wrappers-sumo` (the "sumo" build is All cryptography is done by `libsodium-wrappers-sumo` (the "sumo" build is
@@ -250,6 +271,93 @@ Endpoints used:
- `POST /files/upload-url`: mint a presigned upload URL (for thumbnail repair). - `POST /files/upload-url`: mint a presigned upload URL (for thumbnail repair).
- `PUT /files/thumbnail`: register an uploaded thumbnail's object key. - `PUT /files/thumbnail`: register an uploaded thumbnail's object key.
### Retries and timeouts
Every request in the library goes through one policy, in `src/retry.ts`. A
request is repeated only when repeating it could produce a different answer:
- `ApiError` with a 5xx status: retried. So are `408` and `429`, the two 4xx
codes that are statements about timing rather than about the request.
- Every other 4xx: not retried. A 404 in particular is an answer, and
`listMissingThumbnails` depends on getting it promptly and once.
- Transport failures — a `fetch` rejection, `ECONNRESET`, `ETIMEDOUT`, a DNS or
TLS failure — and deadline aborts: retried. The errno is looked for in the
error's `cause` chain, because that is where Node's `fetch` puts it.
- A truncated download: retried.
- Anything else, including a secretstream authentication failure that is not
truncation: not retried. The default answer is no. For a backup tool, retrying
a permanent failure spends round trips and delays every remaining file, while
declining to retry a transient one costs a single file that the next run picks
up.
Backoff is exponential with full jitter: the delay before retry _n_ is
`random() * min(maxDelayMs, baseDelayMs * 2 ** (n - 1))`. The exponential term
is the ceiling and the wait is drawn below it, so a client that lost many
parallel downloads to one CDN blip does not send them all again at the same
instant. Defaults, configurable through `ApiClientOptions.retry`:
| Option | Default | Meaning |
| ------------- | ------- | ----------------------------------- |
| `attempts` | `4` | total calls, not retries |
| `baseDelayMs` | `500` | ceiling for the first retry's delay |
| `maxDelayMs` | `10000` | upper bound on that ceiling |
With those defaults a file that is going to fail gives up after at most three
and a half seconds of waiting. `sleep` and `random` are injectable through the
same option, which is how the test suite exercises the whole policy without
waiting.
Two deadlines, applied with `AbortSignal.timeout()` and renewed for each
attempt:
| Option | Default | Applies to |
| ------------------- | -------- | ------------------------------------------- |
| `requestTimeoutMs` | `30000` | `getJSON`, `postJSON`, `putJSON`, `putFile` |
| `downloadTimeoutMs` | `600000` | file and thumbnail body transfers |
They are separate because one number cannot serve both: a value short enough to
keep a hung API call from stalling a backup would cancel a legitimate
multi-gigabyte download. The download deadline covers the body, not just the
headers — `getFileStream` returns as soon as headers arrive, so a deadline that
only guarded the initial request would leave the same hang one layer down.
**Non-idempotent requests are not blindly replayed.** `postJSON` and `putJSON`
reach `/users/srp/create-session`, `/users/two-factor/verify` — which consumes
one of a small number of second-factor attempts — and `/files/thumbnail`. They
are retried only on the three failures that establish no TCP connection to the
server ever existed, so no request byte can have been transmitted: `ENOTFOUND`
and `EAI_AGAIN` (name resolution produced no address) and `ECONNREFUSED` (the
peer refused the connection). A 5xx, a mid-flight reset and a deadline are all
left to the caller, because each of them can happen after the server has already
acted. The routing errnos `EHOSTUNREACH`, `ENETUNREACH` and `ENETDOWN` are
excluded for the same reason, despite looking like connect-time failures: on
Linux an ICMP unreachable arriving mid-flight, or a local interface going down
after the request was written, delivers them on an already-established socket.
They stay retryable for the idempotent calls. `putFile` is exempt: a presigned
PUT stores one whole object at one key in one request, so replaying it has no
partial state to damage.
A download is retried as a whole — request, stream consumption, and decryption —
because a socket reset after the response headers have arrived surfaces in the
download layer rather than in `ApiClient`, and that is the common failure for
multi-megabyte photos over a CDN. The secretstream pull state is not resumable
and these endpoints have no Range support, so a retry starts the file over. The
atomic write stays outside the retry, so a download that needed three attempts
still performs exactly one write and one rename. `runBackup` and
`runMetadataBackup` are unchanged: the retry sits below them, and a file that
fails after exhausting it is still logged, counted, and stepped over.
One imprecision is deliberate and worth knowing about. When a body ends part-way
through a secretstream chunk, Poly1305 fails and carries no framing signal, so a
cut connection and genuinely corrupt bytes are indistinguishable. quak reports
that as truncation, which means it is retried. For a body of more than one chunk
the distinction is real — a chunk that failed while the stream carried on past
it stays an authentication failure and is not retried — but for a single-chunk
body, which is most thumbnails and every small file, a wrong key, server-side
corruption and a mid-chunk cutoff all present alike and all get retried. The
cost is bounded by the attempt count, and it buys never silently keeping a
truncated file.
### Session handling ### Session handling
The `Client` class holds the auth token, master key, secret key, and public key The `Client` class holds the auth token, master key, secret key, and public key
@@ -308,10 +416,10 @@ code is non-zero if any files failed.
## TODO ## TODO
- [ ] Retry policy: no retry on 4xx, exponential backoff on 5xx and network - [x] Retry policy: no retry on 4xx, exponential backoff on 5xx and network
errors errors
- [ ] Update the API reference section below to match the current implementation - [ ] Update the API reference section below to match the current implementation
- [ ] `make docker` green - [x] `make docker` green
- [ ] Tag `v1.0.0` - [ ] Tag `v1.0.0`
Future (desktop client, separate repo): Future (desktop client, separate repo):
@@ -333,7 +441,10 @@ are correct.
The key types and their actual signatures can be found in: The key types and their actual signatures can be found in:
- `src/client.ts`: `Client`, `LoginOptions`, `ClientSnapshot` - `src/client.ts`: `Client`, `LoginOptions`, `ClientSnapshot`
- `src/api/client.ts`: `ApiClient`, `ApiClientOptions`, `ApiError` - `src/api/client.ts`: `ApiClient`, `ApiClientOptions`, `ApiError`,
`StreamOptions`
- `src/errors.ts`: `ApiError`, `TruncatedStreamError`
- `src/retry.ts`: `withRetry`, `isRetryable`, `isSafeToReplay`, `RetryOptions`
- `src/auth/types.ts`: `KeyAttributes`, `SRPAttributes`, - `src/auth/types.ts`: `KeyAttributes`, `SRPAttributes`,
`AuthorizationResponse`, `LoginChallenge` `AuthorizationResponse`, `LoginChallenge`
- `src/model/types.ts`: `Collection`, `EnteFile`, `FileMetadata`, `FileBlob`, - `src/model/types.ts`: `Collection`, `EnteFile`, `FileMetadata`, `FileBlob`,

31
TODO.md
View File

@@ -14,12 +14,33 @@ pre-1.0
# Next Step # Next Step
Implement the download retry policy from the README TODO: no retry on 4xx, Update the README API reference section to match the current implementation.
exponential backoff on 5xx and network errors. Apply it to file and thumbnail
downloads, cover it with mock-server tests, and update the README TODO checkbox.
# Completed Steps # Completed Steps
- 2026-08-09: Made `make docker` green and policy-conformant. Multi-stage
Dockerfile: a lint stage runs `make fmt-check` and `make lint`, and the check
stage takes a `COPY --from=lint` dependency on it before running `make check`
and `make build`. `CHECK_EPOCH` and a fail-closed guard stop Docker serving
those two layers from cache, which is what let a build report success without
running the suite. `script/projectname` says `quak`, so the image is tagged
`quak`; `script/bootstrap` updates apt lists before installing, so a Debian
base works; `.dockerignore` no longer ships the compiled binary, the caches or
agent worktrees into the build context, and keeps `.gitignore` in it for
prettier.
- 2026-08-09: Fixed the TypeScript build. `rootDir` is the repo root, so `bin/`
compiles alongside `src/` instead of failing with TS6059; output is
`dist/src/` and `dist/bin/`, which is where `main`, `types` and `bin.quak` now
point. `script/build` verifies the declared entrypoints exist after the
compiler runs and makes the CLI executable, the Dockerfile runs `make build`
as well as `make check`, and a `quak` script makes the README's
`yarn quak <command>` examples work.
- 2026-08-09: Retry policy: no retry on 4xx (except `408` and `429`),
exponential backoff with full jitter on 5xx, transport failures and truncated
transfers, under per-attempt deadlines that cover the response body as well as
the request. Downloads retry request, stream consumption and decryption as one
unit; `postJSON` and `putJSON` are replayed only when the connection was never
established.
- 2026-08-09: Downloads verify the secretstream terminated on `TAG_FINAL` and - 2026-08-09: Downloads verify the secretstream terminated on `TAG_FINAL` and
write output atomically: a truncated body is rejected instead of landing on write output atomically: a truncated body is rejected instead of landing on
disk as a short file, and plaintext is staged in a sibling temp file and disk as a short file, and plaintext is staged in a sibling temp file and
@@ -45,10 +66,6 @@ downloads, cover it with mock-server tests, and update the README TODO checkbox.
# Future Steps # Future Steps
- Retry policy: no retry on 4xx, exponential backoff on 5xx and network errors
(the Next Step).
- Update the README API reference section to match the current implementation.
- Make `make docker` green.
- Tag v1.0.0. - Tag v1.0.0.
- Future desktop client, separate repo: - Future desktop client, separate repo:
- Electron app skeleton consuming this library. - Electron app skeleton consuming this library.

View File

@@ -10,8 +10,8 @@
"url": "https://git.eeqj.de/sneak/quak.git" "url": "https://git.eeqj.de/sneak/quak.git"
}, },
"type": "module", "type": "module",
"main": "./dist/index.js", "main": "./dist/src/index.js",
"types": "./dist/index.d.ts", "types": "./dist/src/index.d.ts",
"bin": { "bin": {
"quak": "./dist/bin/quak.js" "quak": "./dist/bin/quak.js"
}, },
@@ -21,7 +21,8 @@
"LICENSE" "LICENSE"
], ],
"scripts": { "scripts": {
"build": "tsc", "build": "script/build",
"quak": "node ./dist/bin/quak.js",
"test": "vitest run", "test": "vitest run",
"lint": "eslint .", "lint": "eslint .",
"fmt": "prettier --write .", "fmt": "prettier --write .",

View File

@@ -19,6 +19,7 @@ YARN_VERSION="1.22.22"
PKGMGR="" PKGMGR=""
SUDO="" SUDO=""
APT_UPDATED=""
detect_pkgmgr() { detect_pkgmgr() {
[ -n "$PKGMGR" ] && return 0 [ -n "$PKGMGR" ] && return 0
@@ -42,12 +43,24 @@ detect_pkgmgr() {
fi fi
} }
# A fresh Debian image ships no package lists at all, so apt-get install
# fails with "E: Unable to locate package make" until they are fetched.
# Done once per run, since the lists do not go stale mid-bootstrap.
apt_update_once() {
[ -n "$APT_UPDATED" ] && return 0
$SUDO env DEBIAN_FRONTEND=noninteractive apt-get update
APT_UPDATED="yes"
}
# pkg_install <nix-attr> <apt-pkg> <brew-formula> <apk-pkg> # pkg_install <nix-attr> <apt-pkg> <brew-formula> <apk-pkg>
pkg_install() { pkg_install() {
detect_pkgmgr detect_pkgmgr
case "$PKGMGR" in case "$PKGMGR" in
nix) nix-env -iA "nixpkgs.$1" ;; nix) nix-env -iA "nixpkgs.$1" ;;
apt) $SUDO env DEBIAN_FRONTEND=noninteractive apt-get install -y "$2" ;; apt)
apt_update_once
$SUDO env DEBIAN_FRONTEND=noninteractive apt-get install -y "$2"
;;
brew) brew install "$3" ;; brew) brew install "$3" ;;
apk) apk add --no-cache "$4" ;; apk) apk add --no-cache "$4" ;;
esac esac

54
script/build Executable file
View File

@@ -0,0 +1,54 @@
#!/bin/sh
# script/build: compile the TypeScript sources into dist/, then verify that
# the artifacts package.json advertises are among the files the compiler
# actually wrote. tsc reports success by exit status alone and knows nothing
# about the manifest, so without this step a green build can still ship a
# package whose main, types or bin resolve to nothing. Our own extension to
# scripts-to-rule-them-all.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
# Reads package.json, requires every declared entrypoint to exist, requires
# each bin entry to have kept its shebang, and makes the bin entries
# executable: tsc copies the shebang through but not the mode bits, and an
# installed CLI has to be runnable.
verify_entrypoints() {
node -e '
const { readFileSync, statSync, chmodSync } = require("node:fs");
const pkg = JSON.parse(readFileSync("package.json", "utf-8"));
const bins = Object.values(pkg.bin ?? {});
const fail = (message) => {
console.error("build: " + message);
process.exit(1);
};
for (const declared of [pkg.main, pkg.types, ...bins]) {
if (!declared) continue;
try {
statSync(declared);
} catch {
fail("package.json declares " + declared + ", which the build did not produce");
}
console.log("build: verified " + declared);
}
for (const bin of bins) {
const firstLine = readFileSync(bin, "utf-8").split("\n")[0];
if (!firstLine.startsWith("#!")) {
fail(bin + " lost its shebang, so it cannot be executed directly");
}
chmodSync(bin, 0o755);
console.log("build: " + bin + " is executable (" + firstLine + ")");
}
'
}
main() {
cd "$ROOT"
yarn run tsc
verify_entrypoints
}
main "$@"

View File

@@ -1,13 +1,17 @@
#!/bin/sh #!/bin/sh
# script/cibuild: run the CI build. The Dockerfile runs script/check, so # script/cibuild: run the CI build. The Dockerfile runs script/check and
# a successful build implies all checks pass. # script/build, and CHECK_EPOCH differs on every invocation, so those two
# layers cannot be served from Docker's cache: a green build here means
# the checks ran now, not that a previous run was remembered. The layers
# below the epoch (bootstrap, yarn install) are unaffected and stay
# cached. A build that omits the argument fails by design.
set -eu set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)" ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
main() { main() {
cd "$ROOT" cd "$ROOT"
docker build . docker build --build-arg CHECK_EPOCH="$(date +%s)" .
} }
main "$@" main "$@"

View File

@@ -1,6 +1,9 @@
#!/bin/sh #!/bin/sh
# script/docker: build the Docker image tagged with the project name. # script/docker: build the Docker image tagged with the project name.
# Identical in all repos; the tag comes from script/projectname. # Identical in all repos; the tag comes from script/projectname.
# CHECK_EPOCH is passed for the same reason script/cibuild passes it: the
# Dockerfile refuses to build without it, so that no path to an image can
# quietly serve the check and build layers from cache.
set -eu set -eu
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd -P)" SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd -P)"
@@ -8,7 +11,8 @@ ROOT="$(cd "$SCRIPT_DIR/.." && pwd -P)"
main() { main() {
cd "$ROOT" cd "$ROOT"
docker build -t "$("$SCRIPT_DIR/projectname")" . docker build --build-arg CHECK_EPOCH="$(date +%s)" \
-t "$("$SCRIPT_DIR/projectname")" .
} }
main "$@" main "$@"

View File

@@ -6,7 +6,7 @@
set -eu set -eu
main() { main() {
echo "quack" echo "quak"
} }
main "$@" main "$@"

View File

@@ -1,8 +1,33 @@
import { ApiError } from "../errors.js";
import {
isSafeToReplay,
resolveRetryOptions,
withRetry,
type ResolvedRetryOptions,
type RetryOptions,
} from "../retry.js";
// `ApiError` is defined in `src/errors.ts` so that the retry classifier can
// recognise it without importing this module, which imports the classifier.
// It is re-exported here because this is where callers have always imported it
// from, and it must remain one class: a second copy would make `instanceof`
// fail in the classifier and every 5xx would look permanent.
export { ApiError };
const DEFAULT_API_ORIGIN = "https://api.ente.io"; const DEFAULT_API_ORIGIN = "https://api.ente.io";
const DEFAULT_FILES_ORIGIN = "https://files.ente.io"; const DEFAULT_FILES_ORIGIN = "https://files.ente.io";
const DEFAULT_THUMBS_ORIGIN = "https://thumbnails.ente.io"; const DEFAULT_THUMBS_ORIGIN = "https://thumbnails.ente.io";
const CLIENT_PACKAGE = "berlin.sneak.quak"; const CLIENT_PACKAGE = "berlin.sneak.quak";
// Two deadlines rather than one, because a single number cannot serve both
// jobs. Thirty seconds is generous for a JSON call and short enough that a
// hung API connection cannot stall a backup for long. A file body is a
// different shape of problem: the deadline has to cover the whole transfer,
// which for a large video on a slow link is minutes, so a value sane for JSON
// would cancel legitimate downloads.
export const DEFAULT_REQUEST_TIMEOUT_MS = 30_000;
export const DEFAULT_DOWNLOAD_TIMEOUT_MS = 600_000;
export interface ApiClientOptions { export interface ApiClientOptions {
apiOrigin?: string; apiOrigin?: string;
filesOrigin?: string; filesOrigin?: string;
@@ -10,33 +35,78 @@ export interface ApiClientOptions {
authToken?: string; authToken?: string;
fetch?: typeof globalThis.fetch; fetch?: typeof globalThis.fetch;
userAgent?: string; userAgent?: string;
retry?: RetryOptions;
requestTimeoutMs?: number;
downloadTimeoutMs?: number;
} }
export class ApiError extends Error { export interface StreamOptions {
readonly status: number; // Opt out of this client's own retry. Exactly one caller wants that: the
readonly code?: string; // download layer, which retries the request, the stream consumption and
readonly requestID?: string; // the decryption as one unit. Leaving both layers enabled would multiply
readonly body?: unknown; // the budgets — four attempts each becoming sixteen requests per file.
constructor( retry?: boolean;
message: string,
status: number,
opts?: { code?: string; requestID?: string; body?: unknown },
) {
super(message);
this.name = "ApiError";
this.status = status;
this.code = opts?.code;
this.requestID = opts?.requestID;
this.body = opts?.body;
}
} }
// Enforce a deadline over a response body, not merely over its headers.
//
// `getFileStream` returns as soon as headers arrive; the bytes are pulled
// later, in the download layer. Whether the signal passed to `fetch` also
// tears down the body afterwards is up to the fetch implementation, so this
// wrapper makes it a property of quak instead: every read races the signal,
// and an abort errors the stream with the abort reason — which the retry
// classifier recognises.
const deadlineStream = (
body: ReadableStream<Uint8Array>,
signal: AbortSignal,
): ReadableStream<Uint8Array> => {
const reader = body.getReader();
let rejectOnAbort: (reason: unknown) => void = () => undefined;
const aborted = new Promise<never>((_resolve, reject) => {
rejectOnAbort = reject;
});
// The abort may fire when nothing is awaiting `aborted` — after the body
// has been read in full, say. Without this, that rejection would surface
// as an unhandled rejection and take the process down.
void aborted.catch(() => undefined);
const onAbort = (): void => rejectOnAbort(signal.reason);
if (signal.aborted) onAbort();
else signal.addEventListener("abort", onAbort, { once: true });
const release = (): void => signal.removeEventListener("abort", onAbort);
return new ReadableStream<Uint8Array>({
async pull(controller) {
try {
const next = await Promise.race([reader.read(), aborted]);
if (next.done) {
release();
controller.close();
return;
}
controller.enqueue(next.value);
} catch (err) {
release();
await reader.cancel(err).catch(() => undefined);
throw err;
}
},
async cancel(reason) {
release();
await reader.cancel(reason);
},
});
};
export class ApiClient { export class ApiClient {
private readonly apiOrigin: string; private readonly apiOrigin: string;
private readonly isCustomOrigin: boolean; private readonly isCustomOrigin: boolean;
private readonly filesOrigin: string; private readonly filesOrigin: string;
private readonly thumbsOrigin: string; private readonly thumbsOrigin: string;
private readonly _fetch: typeof globalThis.fetch; private readonly _fetch: typeof globalThis.fetch;
private readonly retry: ResolvedRetryOptions;
private readonly requestTimeoutMs: number;
private readonly downloadTimeoutMs: number;
private token: string | undefined; private token: string | undefined;
constructor(opts?: ApiClientOptions) { constructor(opts?: ApiClientOptions) {
@@ -53,6 +123,11 @@ export class ApiClient {
opts?.thumbsOrigin ?? DEFAULT_THUMBS_ORIGIN opts?.thumbsOrigin ?? DEFAULT_THUMBS_ORIGIN
).replace(/\/+$/, ""); ).replace(/\/+$/, "");
this._fetch = opts?.fetch ?? globalThis.fetch; this._fetch = opts?.fetch ?? globalThis.fetch;
this.retry = resolveRetryOptions(opts?.retry);
this.requestTimeoutMs =
opts?.requestTimeoutMs ?? DEFAULT_REQUEST_TIMEOUT_MS;
this.downloadTimeoutMs =
opts?.downloadTimeoutMs ?? DEFAULT_DOWNLOAD_TIMEOUT_MS;
this.token = opts?.authToken; this.token = opts?.authToken;
} }
@@ -64,6 +139,13 @@ export class ApiClient {
this.token = undefined; this.token = undefined;
} }
// The policy this client was configured with, so that a caller wrapping a
// whole operation in its own `withRetry` — the download layer — runs under
// the same settings rather than under the library defaults.
getRetryOptions(): ResolvedRetryOptions {
return this.retry;
}
private headers(extra?: Record<string, string>): Record<string, string> { private headers(extra?: Record<string, string>): Record<string, string> {
const h: Record<string, string> = { const h: Record<string, string> = {
"X-Client-Package": CLIENT_PACKAGE, "X-Client-Package": CLIENT_PACKAGE,
@@ -128,38 +210,54 @@ export class ApiClient {
} }
} }
} }
// A GET changes nothing, so it is retried under the full policy.
return withRetry(async () => {
const resp = await this._fetch(url.href, { const resp = await this._fetch(url.href, {
method: "GET", method: "GET",
headers: this.headers(), headers: this.headers(),
signal: AbortSignal.timeout(this.requestTimeoutMs),
}); });
await this.throwIfError(resp); await this.throwIfError(resp);
return (await resp.json()) as T; return (await resp.json()) as T;
}, this.retry);
} }
async postJSON<T>(path: string, body: unknown): Promise<T> { async postJSON<T>(path: string, body: unknown): Promise<T> {
const url = `${this.apiOrigin}${path}`; const url = `${this.apiOrigin}${path}`;
// Idempotency: this reaches `/users/srp/create-session`,
// `/users/two-factor/verify` and `/users/ott`, all of which change
// server state — verifying a second factor consumes one of a small
// number of attempts. So a POST is replayed only on a failure that
// establishes no TCP connection to the server ever existed: DNS
// produced no address, or the peer refused the connection. A 5xx, a
// mid-flight reset, a routing errno (which Linux also delivers on an
// established socket) and a timeout are all left to the caller,
// because each of them can occur after the server has already acted.
return withRetry(
async () => {
const resp = await this._fetch(url, { const resp = await this._fetch(url, {
method: "POST", method: "POST",
headers: this.headers({ "Content-Type": "application/json" }), headers: this.headers({
"Content-Type": "application/json",
}),
body: JSON.stringify(body), body: JSON.stringify(body),
signal: AbortSignal.timeout(this.requestTimeoutMs),
}); });
await this.throwIfError(resp); await this.throwIfError(resp);
return (await resp.json()) as T; return (await resp.json()) as T;
},
{ ...this.retry, isRetryable: isSafeToReplay },
);
} }
async getFileStream(fileID: number): Promise<ReadableStream<Uint8Array>> { async getFileStream(
fileID: number,
opts?: StreamOptions,
): Promise<ReadableStream<Uint8Array>> {
const url = this.isCustomOrigin const url = this.isCustomOrigin
? `${this.apiOrigin}/files/download/${fileID}` ? `${this.apiOrigin}/files/download/${fileID}`
: `${this.filesOrigin}/?fileID=${fileID}`; : `${this.filesOrigin}/?fileID=${fileID}`;
const resp = await this._fetch(url, { return this.streamRequest(url, opts);
method: "GET",
headers: this.headers(),
});
await this.throwIfError(resp);
if (!resp.body) {
throw new Error("response body is null");
}
return resp.body;
} }
async getUploadURL( async getUploadURL(
@@ -173,6 +271,11 @@ export class ApiClient {
} }
async putFile(presignedURL: string, data: Uint8Array): Promise<void> { async putFile(presignedURL: string, data: Uint8Array): Promise<void> {
// Idempotent despite being a write: a presigned PUT stores one whole
// object at one key in one request, so replaying it either overwrites
// the same bytes or lands them for the first time. There is no partial
// state to protect, hence the full policy rather than the POST rule.
await withRetry(async () => {
const resp = await this._fetch(presignedURL, { const resp = await this._fetch(presignedURL, {
method: "PUT", method: "PUT",
headers: { headers: {
@@ -180,21 +283,40 @@ export class ApiClient {
"Content-Length": String(data.length), "Content-Length": String(data.length),
}, },
body: data, body: data,
signal: AbortSignal.timeout(this.requestTimeoutMs),
}); });
if (!resp.ok) { if (!resp.ok) {
throw new Error(`PUT to presigned URL failed: HTTP ${resp.status}`); // An ApiError, not a bare Error: without the status on the
// error the upload path cannot be classified at all, and a
// 503 from S3 would be indistinguishable from a bug.
throw new ApiError(
`PUT to presigned URL failed: HTTP ${resp.status}`,
resp.status,
);
} }
}, this.retry);
} }
async putJSON<T>(path: string, body: unknown): Promise<T> { async putJSON<T>(path: string, body: unknown): Promise<T> {
const url = `${this.apiOrigin}${path}`; const url = `${this.apiOrigin}${path}`;
// Same idempotency rule as `postJSON`, for the same reason: this
// reaches `/files/thumbnail`, which registers an uploaded thumbnail
// against a file.
return withRetry(
async () => {
const resp = await this._fetch(url, { const resp = await this._fetch(url, {
method: "PUT", method: "PUT",
headers: this.headers({ "Content-Type": "application/json" }), headers: this.headers({
"Content-Type": "application/json",
}),
body: JSON.stringify(body), body: JSON.stringify(body),
signal: AbortSignal.timeout(this.requestTimeoutMs),
}); });
await this.throwIfError(resp); await this.throwIfError(resp);
return (await resp.json()) as T; return (await resp.json()) as T;
},
{ ...this.retry, isRetryable: isSafeToReplay },
);
} }
async updateThumbnail( async updateThumbnail(
@@ -210,18 +332,36 @@ export class ApiClient {
async getThumbnailStream( async getThumbnailStream(
fileID: number, fileID: number,
opts?: StreamOptions,
): Promise<ReadableStream<Uint8Array>> { ): Promise<ReadableStream<Uint8Array>> {
const url = this.isCustomOrigin const url = this.isCustomOrigin
? `${this.apiOrigin}/files/preview/${fileID}` ? `${this.apiOrigin}/files/preview/${fileID}`
: `${this.thumbsOrigin}/?fileID=${fileID}`; : `${this.thumbsOrigin}/?fileID=${fileID}`;
return this.streamRequest(url, opts);
}
private async streamRequest(
url: string,
opts?: StreamOptions,
): Promise<ReadableStream<Uint8Array>> {
const once = async (): Promise<ReadableStream<Uint8Array>> => {
// A fresh deadline per attempt, so a retry gets the whole budget
// rather than the remainder of the one that just expired.
const signal = AbortSignal.timeout(this.downloadTimeoutMs);
const resp = await this._fetch(url, { const resp = await this._fetch(url, {
method: "GET", method: "GET",
headers: this.headers(), headers: this.headers(),
signal,
}); });
await this.throwIfError(resp); await this.throwIfError(resp);
if (!resp.body) { if (!resp.body) {
throw new Error("response body is null"); // Carries the status, and is not retryable: a response that
// arrived without a body is malformed, and asking again
// produces the same malformed response.
throw new ApiError("response body is null", resp.status);
} }
return resp.body; return deadlineStream(resp.body, signal);
};
return opts?.retry === false ? once() : withRetry(once, this.retry);
} }
} }

View File

@@ -1,4 +1,4 @@
import sodium from "libsodium-wrappers-sumo"; import sodium, { type StateAddress } from "libsodium-wrappers-sumo";
// Plaintext chunk size used by Ente for file content streams. Hard-coded by // Plaintext chunk size used by Ente for file content streams. Hard-coded by
// the server; clients must match. // the server; clients must match.
@@ -47,8 +47,10 @@ export const encryptBlob = (
}; };
// Opaque handle to libsodium's secretstream pull state. Threaded through // Opaque handle to libsodium's secretstream pull state. Threaded through
// successive pullStreamChunk calls. // successive pullStreamChunk calls. The type comes from the named export;
export type StreamPullState = sodium.StateAddress; // the default import is the module's value side and has no type namespace
// under it.
export type StreamPullState = StateAddress;
// Initialise a pull stream from the per-file decryption header and the // Initialise a pull stream from the per-file decryption header and the
// per-file key. // per-file key.
@@ -84,9 +86,13 @@ export const pullStreamChunk = (
state: StreamPullState, state: StreamPullState,
ciphertext: Uint8Array, ciphertext: Uint8Array,
): { plaintext: Uint8Array; tag: number } => { ): { plaintext: Uint8Array; tag: number } => {
// The additional-data argument is not optional in libsodium's signature.
// null is "no additional data", matching the null passed on the push side
// in encryptBlob; Ente's file streams carry none.
const result = sodium.crypto_secretstream_xchacha20poly1305_pull( const result = sodium.crypto_secretstream_xchacha20poly1305_pull(
state, state,
ciphertext, ciphertext,
null,
); );
if (result === false) { if (result === false) {
throw new Error("secretstream chunk authentication failed"); throw new Error("secretstream chunk authentication failed");

View File

@@ -9,6 +9,8 @@ import {
STREAM_CHUNK_SIZE, STREAM_CHUNK_SIZE,
streamTagFinal, streamTagFinal,
} from "../crypto/index.js"; } from "../crypto/index.js";
import { TruncatedStreamError } from "../errors.js";
import { withRetry } from "../retry.js";
import type { ApiClient } from "../api/client.js"; import type { ApiClient } from "../api/client.js";
import type { EnteFile } from "../model/types.js"; import type { EnteFile } from "../model/types.js";
@@ -65,7 +67,7 @@ const streamDecrypt = async (
try { try {
pulled = pullStreamChunk(state, buffer); pulled = pullStreamChunk(state, buffer);
} catch (err) { } catch (err) {
throw new Error( throw new TruncatedStreamError(
`download: stream truncated: response body ended with ${buffer.length} trailing bytes that did not authenticate as a final chunk (transfer stopped mid-chunk, or the data is corrupt)`, `download: stream truncated: response body ended with ${buffer.length} trailing bytes that did not authenticate as a final chunk (transfer stopped mid-chunk, or the data is corrupt)`,
{ cause: err }, { cause: err },
); );
@@ -85,13 +87,13 @@ const streamDecrypt = async (
// Returning a short plaintext here would put a corrupt file on disk that // Returning a short plaintext here would put a corrupt file on disk that
// later backup runs would treat as complete. // later backup runs would treat as complete.
if (chunksPulled === 0) { if (chunksPulled === 0) {
throw new Error( throw new TruncatedStreamError(
"download: stream truncated: response body contained no secretstream chunks", "download: stream truncated: response body contained no secretstream chunks",
); );
} }
const tagFinal = streamTagFinal(); const tagFinal = streamTagFinal();
if (lastTag !== tagFinal) { if (lastTag !== tagFinal) {
throw new Error( throw new TruncatedStreamError(
`download: stream truncated: last chunk tag ${lastTag}, expected TAG_FINAL (${tagFinal})`, `download: stream truncated: last chunk tag ${lastTag}, expected TAG_FINAL (${tagFinal})`,
); );
} }
@@ -128,15 +130,49 @@ const writeAtomic = async (
} }
}; };
// Fetch a stream and decrypt it, retrying the whole sequence.
//
// The request is only the first third of a download. `getXStream` returns as
// soon as headers arrive, and the bytes are pulled here, so a socket reset
// mid-body — the dominant failure mode for multi-megabyte photos over a CDN —
// throws in `streamDecrypt` and never reaches `ApiClient` at all. Retrying the
// request alone would miss it entirely.
//
// The client's own retry is therefore switched off for these two calls: with
// both layers active the budgets would multiply, and the library default of
// four attempts would mean sixteen requests for one file. The policy comes
// from the client so a caller that configured one gets it here too.
//
// A retry starts the file over from byte zero: the secretstream pull state is
// not resumable and there is no Range support on these endpoints.
const fetchAndDecrypt = async (
api: ApiClient,
openStream: () => Promise<ReadableStream<Uint8Array>>,
header: Uint8Array,
key: Uint8Array,
): Promise<Uint8Array> =>
withRetry(async () => {
const stream = await openStream();
return streamDecrypt(stream, header, key);
}, api.getRetryOptions());
export const downloadFile = async ( export const downloadFile = async (
api: ApiClient, api: ApiClient,
file: EnteFile, file: EnteFile,
outPath?: string, outPath?: string,
): Promise<DownloadResult> => { ): Promise<DownloadResult> => {
const resolvedPath = outPath ?? file.metadata.title; const resolvedPath = outPath ?? file.metadata.title;
const stream = await api.getFileStream(file.id);
const header = fromBase64(file.file.decryptionHeader); const header = fromBase64(file.file.decryptionHeader);
const plaintext = await streamDecrypt(stream, header, file.key); const plaintext = await fetchAndDecrypt(
api,
() => api.getFileStream(file.id, { retry: false }),
header,
file.key,
);
// Outside the retry, deliberately: only the attempt that produced a
// complete, authenticated plaintext gets to stage a temporary file, so a
// download that needed three tries still performs exactly one write and
// one rename.
await writeAtomic(resolvedPath, plaintext); await writeAtomic(resolvedPath, plaintext);
return { path: resolvedPath, bytesWritten: plaintext.length }; return { path: resolvedPath, bytesWritten: plaintext.length };
}; };
@@ -147,9 +183,13 @@ export const downloadThumbnail = async (
outPath?: string, outPath?: string,
): Promise<DownloadResult> => { ): Promise<DownloadResult> => {
const resolvedPath = outPath ?? `thumb_${file.metadata.title}`; const resolvedPath = outPath ?? `thumb_${file.metadata.title}`;
const stream = await api.getThumbnailStream(file.id);
const header = fromBase64(file.thumbnail.decryptionHeader); const header = fromBase64(file.thumbnail.decryptionHeader);
const plaintext = await streamDecrypt(stream, header, file.key); const plaintext = await fetchAndDecrypt(
api,
() => api.getThumbnailStream(file.id, { retry: false }),
header,
file.key,
);
await writeAtomic(resolvedPath, plaintext); await writeAtomic(resolvedPath, plaintext);
return { path: resolvedPath, bytesWritten: plaintext.length }; return { path: resolvedPath, bytesWritten: plaintext.length };
}; };

45
src/errors.ts Normal file
View File

@@ -0,0 +1,45 @@
// Error types that more than one layer of quak needs to recognise.
//
// They live here rather than beside the code that throws them so that the
// retry classifier can identify them without importing the HTTP client or the
// download layer — both of which import the classifier. `ApiError` is
// re-exported from `src/api/client.ts`, which is where callers have always
// imported it from and where it still belongs conceptually.
export class ApiError extends Error {
readonly status: number;
readonly code?: string;
readonly requestID?: string;
readonly body?: unknown;
constructor(
message: string,
status: number,
opts?: { code?: string; requestID?: string; body?: unknown },
) {
super(message);
this.name = "ApiError";
this.status = status;
this.code = opts?.code;
this.requestID = opts?.requestID;
this.body = opts?.body;
}
}
// A response body that ended before the secretstream did.
//
// This is a type rather than a message prefix because it is a decision, not a
// diagnostic: the retry policy asks "was this a short transfer?" and acts on
// the answer. Matching on the wording of an error message would make the next
// person to reword a diagnostic silently turn every truncated download into a
// permanent failure, and the failure would look like a corrupt file rather
// than like a bug.
//
// `cause` carries the underlying authentication failure on the one path where
// there is one — a body that stopped part-way through a chunk, which Poly1305
// cannot distinguish from corruption.
export class TruncatedStreamError extends Error {
constructor(message: string, opts?: ErrorOptions) {
super(message, opts);
this.name = "TruncatedStreamError";
}
}

View File

@@ -1,7 +1,25 @@
export const VERSION = "0.0.0"; export const VERSION = "0.0.0";
export { Client, type LoginOptions, type ClientSnapshot } from "./client.js"; export { Client, type LoginOptions, type ClientSnapshot } from "./client.js";
export { ApiClient, ApiError, type ApiClientOptions } from "./api/client.js"; export {
ApiClient,
ApiError,
DEFAULT_DOWNLOAD_TIMEOUT_MS,
DEFAULT_REQUEST_TIMEOUT_MS,
type ApiClientOptions,
type StreamOptions,
} from "./api/client.js";
export { TruncatedStreamError } from "./errors.js";
export {
DEFAULT_RETRY_OPTIONS,
isRetryable,
isSafeToReplay,
resolveRetryOptions,
withRetry,
type ResolvedRetryOptions,
type RetryOptions,
type WithRetryOptions,
} from "./retry.js";
export { unwrapAuth, type UnwrapResult } from "./auth/unwrap.js"; export { unwrapAuth, type UnwrapResult } from "./auth/unwrap.js";
export { export {
beginLogin, beginLogin,

201
src/retry.ts Normal file
View File

@@ -0,0 +1,201 @@
// The retry policy shared by every network operation in quak.
//
// Two independent pieces: a classifier that decides whether an error is worth
// another attempt, and a loop that acts on that decision with exponential
// backoff. Keeping them apart is what lets the non-idempotent call sites reuse
// the loop under a stricter question (see `isSafeToReplay`).
//
// The classifier's default answer is no. For a backup tool, retrying a
// permanent failure spends round trips and delays every remaining file, while
// declining to retry a transient one costs a single file that the next run
// picks up anyway. When the evidence is ambiguous, fail fast.
import { ApiError, TruncatedStreamError } from "./errors.js";
export interface RetryOptions {
// Total calls, not retries: `attempts: 1` disables retrying.
attempts?: number;
// Ceiling for the first retry's delay; doubles with each retry.
baseDelayMs?: number;
// Upper bound on that ceiling, so a long outage settles into a steady
// poll instead of growing without limit.
maxDelayMs?: number;
// Injected so tests exercise the whole policy without waiting.
sleep?: (ms: number) => Promise<void>;
// Injected so the jitter is reproducible under test.
random?: () => number;
}
export type ResolvedRetryOptions = Required<RetryOptions>;
export const DEFAULT_RETRY_OPTIONS: ResolvedRetryOptions = {
attempts: 4,
baseDelayMs: 500,
maxDelayMs: 10_000,
sleep: (ms: number): Promise<void> =>
new Promise((resolve) => setTimeout(resolve, ms)),
random: Math.random,
};
export const resolveRetryOptions = (
opts?: RetryOptions,
): ResolvedRetryOptions => ({
attempts: opts?.attempts ?? DEFAULT_RETRY_OPTIONS.attempts,
baseDelayMs: opts?.baseDelayMs ?? DEFAULT_RETRY_OPTIONS.baseDelayMs,
maxDelayMs: opts?.maxDelayMs ?? DEFAULT_RETRY_OPTIONS.maxDelayMs,
sleep: opts?.sleep ?? DEFAULT_RETRY_OPTIONS.sleep,
random: opts?.random ?? DEFAULT_RETRY_OPTIONS.random,
});
// Transport failures: the request did not complete, for reasons below HTTP.
const TRANSPORT_CODES = new Set([
"ECONNRESET",
"ECONNABORTED",
"ETIMEDOUT",
"EPIPE",
"ENOTFOUND",
"EAI_AGAIN",
"ECONNREFUSED",
"EHOSTUNREACH",
"ENETUNREACH",
"ENETRESET",
"ENETDOWN",
]);
// The subset of the above that can only be reported before a TCP connection
// exists, and therefore before any request byte could have been written: name
// resolution produced no address (`ENOTFOUND`, `EAI_AGAIN`) or the peer
// refused the connection with an RST to the SYN (`ECONNREFUSED`).
//
// The routing errnos — `EHOSTUNREACH`, `ENETUNREACH`, `ENETDOWN` — are
// deliberately absent even though they look like connect-time failures. On
// Linux they are also delivered on an already-established socket: an ICMP
// destination-unreachable arriving mid-flight sets the socket error and the
// next read or write returns it, and a local interface going down after the
// request was fully written surfaces the same way. In those cases the server
// may already have received and acted on the request, which is exactly the
// ambiguity this set exists to exclude. They stay in `TRANSPORT_CODES`, so
// they remain retryable for idempotent calls; only replay eligibility is
// narrowed. See `isSafeToReplay`.
const CONNECT_CODES = new Set(["ENOTFOUND", "EAI_AGAIN", "ECONNREFUSED"]);
// `cause` is an arbitrary user-settable property and nothing prevents it from
// forming a cycle, so the walk is bounded. Hanging the process would be a
// worse outcome than any misclassification.
const MAX_CAUSE_DEPTH = 8;
// Collect every `code` in an error's cause chain. undici does not put the
// errno on the error it throws — it hangs the underlying socket error off
// `cause`, sometimes more than one level down — so a classifier that only read
// the top-level error would see a bare `Error` and call every dropped
// connection permanent.
const causeCodes = (err: unknown): string[] => {
const codes: string[] = [];
let current: unknown = err;
for (let depth = 0; depth < MAX_CAUSE_DEPTH; depth++) {
if (current === null || typeof current !== "object") break;
const { code, cause } = current as { code?: unknown; cause?: unknown };
if (typeof code === "string") codes.push(code);
if (cause === current) break;
current = cause;
}
return codes;
};
const isAbort = (err: unknown): boolean => {
if (err === null || typeof err !== "object") return false;
const { name } = err as { name?: unknown };
return name === "AbortError" || name === "TimeoutError";
};
// Is another attempt capable of producing a different answer?
//
// A note on the truncation case, recorded on issue #2. `TruncatedStreamError`
// is retried, and for a body of more than one chunk that is exactly right: a
// chunk that failed to authenticate while the stream carried on past it is
// corruption, stays an ordinary authentication failure, and is not retried.
//
// For a single-chunk body — most thumbnails, every small file — the split is
// not achievable. A wrong file key, server-side corruption and a connection
// cut mid-chunk are cryptographically identical: Poly1305 fails and carries no
// framing signal. All three are reported as truncation and therefore retried.
// That imprecision is deliberate and bounded by the attempt count: one wasted
// round trip on a genuinely corrupt file is a fair price for never silently
// keeping a truncated one, and the alternative — treating a single-chunk
// authentication failure as permanent — would reintroduce exactly the
// silent-corruption risk the truncation check exists to remove.
export const isRetryable = (err: unknown): boolean => {
if (err instanceof ApiError) {
// 408 and 429 are the two 4xx codes that are statements about timing
// rather than about the request, and backoff is the right answer to
// both. Every other 4xx will answer the same way however often it is
// asked.
if (err.status === 408 || err.status === 429) return true;
return err.status >= 500 && err.status <= 599;
}
if (err instanceof TruncatedStreamError) return true;
// quak only ever aborts a request on its own deadline, so an abort means
// this attempt ran out of time — which a later one may not.
if (isAbort(err)) return true;
// Node's fetch rejects with `TypeError: fetch failed` for everything below
// HTTP: DNS failure, refused connection, TLS error, reset socket. Nothing
// on the object separates it from a TypeError thrown by a bug, so this is
// deliberately literal. Demanding a recognised `cause` instead would
// classify real network failures as permanent and fail backups that should
// have succeeded; the cost of the imprecision is bounded by the attempt
// count.
if (err instanceof TypeError) return true;
return causeCodes(err).some((code) => TRANSPORT_CODES.has(code));
};
// Could the first attempt already have taken effect on the server?
//
// `isRetryable` is the wrong question for a request that changes state.
// quak's non-idempotent calls are `/users/srp/create-session`,
// `/users/two-factor/verify` — which consumes one of a small number of 2FA
// attempts — and `/files/thumbnail`. They are replayed only on the failures in
// `CONNECT_CODES`, which establish that no TCP connection to the server ever
// existed: there was no address to connect to, or the peer refused the
// connection outright. A request byte cannot have been transmitted, so the
// server cannot have acted.
//
// Everything else is ambiguous. A 5xx proves the server did process the
// request. A reset or a broken pipe can arrive after it was fully sent and
// acted on. A routing errno can be delivered on an established socket. A
// deadline says nothing at all about the server's state.
export const isSafeToReplay = (err: unknown): boolean =>
isRetryable(err) && causeCodes(err).some((code) => CONNECT_CODES.has(code));
export interface WithRetryOptions extends RetryOptions {
isRetryable?: (err: unknown) => boolean;
}
// Exponential backoff with full jitter: the exponential term is the ceiling,
// and the actual wait is drawn uniformly below it. Full jitter, rather than a
// fixed delay plus noise, is what stops a client that lost a hundred parallel
// downloads to one CDN blip from re-sending all hundred at the same instant.
const backoffMs = (retryNumber: number, policy: ResolvedRetryOptions): number =>
policy.random() *
Math.min(policy.maxDelayMs, policy.baseDelayMs * 2 ** (retryNumber - 1));
export const withRetry = async <T>(
fn: () => Promise<T>,
opts?: WithRetryOptions,
): Promise<T> => {
const policy = resolveRetryOptions(opts);
const retryable = opts?.isRetryable ?? isRetryable;
for (let attempt = 1; ; attempt++) {
try {
return await fn();
} catch (err) {
if (attempt >= policy.attempts || !retryable(err)) {
// The error escapes unwrapped, and it is the one from the
// final attempt: callers classify what they catch, and
// `listMissingThumbnails` in particular needs the `ApiError`
// and its status intact.
throw err;
}
await policy.sleep(backoffMs(attempt, policy));
}
}
};

View File

@@ -2,6 +2,7 @@ import { createHash } from "node:crypto";
import { readFileSync } from "node:fs"; import { readFileSync } from "node:fs";
import * as jpeg from "jpeg-js"; import * as jpeg from "jpeg-js";
import type { Client } from "./client.js"; import type { Client } from "./client.js";
import { ApiError } from "./api/client.js";
import { encryptBlob, toBase64 } from "./crypto/index.js"; import { encryptBlob, toBase64 } from "./crypto/index.js";
import { downloadFile } from "./download/index.js"; import { downloadFile } from "./download/index.js";
import type { EnteFile } from "./model/types.js"; import type { EnteFile } from "./model/types.js";
@@ -62,13 +63,31 @@ export const listMissingThumbnails = async (
reason: "empty thumbnail (0 bytes)", reason: "empty thumbnail (0 bytes)",
}); });
} }
} catch { } catch (err) {
// A 404 is the server stating the thumbnail is not there:
// that, and an empty body, are the only two answers that mean
// "missing". Anything else reaching this point is a failure
// that already exhausted its retries — a failing server, a
// dropped connection, a deadline — and says nothing about
// whether the thumbnail exists.
//
// The distinction is what stops `helper
// fix-missing-thumbnails` from downloading originals,
// regenerating thumbnails and uploading them over thumbnails
// that were fine all along, because the CDN was briefly
// returning 500s while this ran.
if (err instanceof ApiError && err.status === 404) {
missing.push({ missing.push({
fileID: file.id, fileID: file.id,
title: file.metadata.title, title: file.metadata.title,
collection: col.name, collection: col.name,
reason: "thumbnail fetch failed", reason: "thumbnail not found (HTTP 404)",
}); });
} else {
log(
`[${col.name}] Could not check ${file.metadata.title}: ${err instanceof Error ? err.message : String(err)} (not reported as missing)`,
);
}
} }
} }
} }

View File

@@ -29,12 +29,23 @@
* return a `ReadableStream<Uint8Array>` from the appropriate CDN * return a `ReadableStream<Uint8Array>` from the appropriate CDN
* (or the self-hosted fallback path). * (or the self-hosted fallback path).
* *
* - Retries and timeouts. Every request is issued under a deadline and,
* where it is safe to do so, retried with exponential backoff. The
* policy is `src/retry.ts`; what the last section of this file
* documents is which requests get it and which deliberately do not.
*
* All tests inject a fake `fetch` via the constructor so nothing touches * All tests inject a fake `fetch` via the constructor so nothing touches
* the network. The fake records every call for assertion. * the network. The fake records every call for assertion.
*/ */
import { describe, expect, it } from "vitest"; import { describe, expect, it } from "vitest";
import { ApiClient, ApiError } from "../../src/api/client.js"; import {
ApiClient,
ApiError,
DEFAULT_DOWNLOAD_TIMEOUT_MS,
DEFAULT_REQUEST_TIMEOUT_MS,
} from "../../src/api/client.js";
import type { RetryOptions } from "../../src/retry.js";
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
// Test helpers // Test helpers
@@ -102,6 +113,92 @@ const recordingFetch = (
return { fetch: fake as typeof globalThis.fetch, calls }; return { fetch: fake as typeof globalThis.fetch, calls };
}; };
/**
* One scripted outcome for a single `fetch` call:
*
* - a `Response`, returned as-is;
* - an `Error`, thrown — this is how `fetch` reports a network failure;
* - `HANG`, a request that never answers until its own deadline aborts it.
*
* `HANG` is what makes the timeout tests honest. A fake that resolved after a
* delay would be testing the clock; this one resolves *only* when the signal
* quak attached fires. If no signal is attached, or the signal is not wired to
* the body, the promise never settles and the test fails on its own timeout
* rather than passing by accident.
*/
const HANG = Symbol("hang until aborted");
type FetchStep = Response | Error | typeof HANG;
const scriptedFetch = (
...steps: FetchStep[]
): {
fetch: typeof globalThis.fetch;
calls: { url: string; init: RequestInit | undefined }[];
} => {
const calls: { url: string; init: RequestInit | undefined }[] = [];
let i = 0;
const fake = async (
input: RequestInfo | URL,
init?: RequestInit,
): Promise<Response> => {
const url =
typeof input === "string"
? input
: input instanceof URL
? input.href
: input.url;
calls.push({ url, init });
const step = steps[i++];
if (step === undefined) {
throw new Error(`scriptedFetch: no step for call #${i - 1}`);
}
if (step === HANG) {
return new Promise<Response>((_resolve, reject) => {
const signal = init?.signal;
if (!signal) return;
if (signal.aborted) {
reject(signal.reason as Error);
return;
}
signal.addEventListener(
"abort",
() => reject(signal.reason as Error),
{ once: true },
);
});
}
if (step instanceof Error) throw step;
return step;
};
return { fetch: fake as typeof globalThis.fetch, calls };
};
/** An error shaped like a Node transport failure: the errno is on `.code`. */
const errnoError = (code: string, message = code): Error =>
Object.assign(new Error(message), { code });
/**
* A retry policy with the waiting removed. Backoff arithmetic is covered in
* `test/retry/retry.test.ts`; what the tests below are about is *how many
* requests* each call site issues, so they inject a `sleep` that returns
* immediately. Nothing in this file waits.
*/
const noWait: RetryOptions = {
sleep: () => Promise.resolve(),
random: () => 0,
};
/** Drain a stream and return the bytes, so body-level failures surface. */
const readAll = async (stream: ReadableStream<Uint8Array>): Promise<number> => {
const reader = stream.getReader();
let total = 0;
for (;;) {
const { done, value } = await reader.read();
if (value) total += value.length;
if (done) return total;
}
};
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
// Tests // Tests
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
@@ -383,3 +480,446 @@ describe("ApiClient.getFileStream / getThumbnailStream", () => {
expect(headers.get("X-Auth-Token")).toBe("tk"); expect(headers.get("X-Auth-Token")).toBe("tk");
}); });
}); });
// ---------------------------------------------------------------------------
// Retries
//
// The rule the whole section turns on: a request is repeated only when
// repeating it could produce a different answer, and only when repeating it
// cannot do harm. Those are two separate questions and the second one is why
// `postJSON` and `putJSON` behave differently from everything else here.
//
// Every assertion below counts requests. None of them measures how long
// anything took: the retry policy's `sleep` is injected and returns
// immediately, so a machine under load and an idle one produce identical
// results.
// ---------------------------------------------------------------------------
describe("ApiClient retries", () => {
it("issues exactly one request for a 404", async () => {
// A 404 is an answer, not a failure to get one. Repeating it wastes
// a round trip and — for `listMissingThumbnails`, which reads a 404
// as "this thumbnail really is missing" — delays a correct result.
// The script holds five responses; only the first may be consumed.
const { fetch, calls } = scriptedFetch(
textResponse("Not Found", 404),
textResponse("Not Found", 404),
textResponse("Not Found", 404),
textResponse("Not Found", 404),
textResponse("Not Found", 404),
);
const client = new ApiClient({ fetch, retry: noWait });
await expect(client.getJSON("/missing")).rejects.toBeInstanceOf(
ApiError,
);
expect(calls).toHaveLength(1);
});
it("retries a 500 up to the configured attempt count, then throws", async () => {
// `attempts` is a total, not a number of retries: three attempts mean
// three requests. The script offers five responses so that a client
// which ignored the limit would be visible as a count of 4 or 5
// rather than as a crash.
const { fetch, calls } = scriptedFetch(
textResponse("boom", 500),
textResponse("boom", 500),
textResponse("boom", 500),
textResponse("boom", 500),
textResponse("boom", 500),
);
const client = new ApiClient({
fetch,
retry: { ...noWait, attempts: 3 },
});
const err: unknown = await client
.getJSON("/flaky")
.catch((e: unknown) => e);
expect(err).toBeInstanceOf(ApiError);
expect((err as ApiError).status).toBe(500);
expect(calls).toHaveLength(3);
});
it("returns the first successful response after a 503", async () => {
const { fetch, calls } = scriptedFetch(
textResponse("unavailable", 503),
jsonResponse({ ok: true }),
);
const client = new ApiClient({ fetch, retry: noWait });
await expect(
client.getJSON<{ ok: boolean }>("/health"),
).resolves.toEqual({ ok: true });
expect(calls).toHaveLength(2);
});
it("retries 408 and 429", async () => {
// The two 4xx codes that are about timing rather than about the
// request. Backoff is precisely the right response to both.
for (const status of [408, 429]) {
const { fetch, calls } = scriptedFetch(
textResponse("wait", status),
jsonResponse({ ok: true }),
);
const client = new ApiClient({ fetch, retry: noWait });
await expect(
client.getJSON("/rate-limited"),
).resolves.toBeDefined();
expect(calls).toHaveLength(2);
}
});
it("retries a fetch rejection and succeeds on a later attempt", async () => {
// A dropped or refused connection is the failure this policy exists
// for: the request never got an answer, so asking again is free of
// consequence and likely to work.
const { fetch, calls } = scriptedFetch(
new TypeError("fetch failed"),
errnoError("ECONNRESET", "read ECONNRESET"),
jsonResponse({ collections: [] }),
);
const client = new ApiClient({ fetch, retry: noWait });
await expect(client.getJSON("/collections/v2")).resolves.toEqual({
collections: [],
});
expect(calls).toHaveLength(3);
});
it("applies the policy to file and thumbnail streams", async () => {
// Both CDN endpoints go through the same wrapper. A 503 from
// files.ente.io during a large backup is common enough that not
// retrying it would fail files for no reason.
for (const get of ["getFileStream", "getThumbnailStream"] as const) {
const { fetch, calls } = scriptedFetch(
textResponse("unavailable", 503),
streamResponse(new Uint8Array([1, 2, 3])),
);
const client = new ApiClient({ fetch, retry: noWait });
const stream = await client[get](7);
expect(await readAll(stream)).toBe(3);
expect(calls).toHaveLength(2);
}
});
it("does not retry when the caller opts out", async () => {
// `{ retry: false }` exists for one caller: the download layer, which
// wraps request *and* body consumption *and* decryption in a single
// retry of its own. Without the opt-out the two budgets would
// multiply — four attempts each becoming sixteen requests for one
// file.
const { fetch, calls } = scriptedFetch(
textResponse("unavailable", 503),
streamResponse(new Uint8Array([1, 2, 3])),
);
const client = new ApiClient({ fetch, retry: noWait });
await expect(
client.getFileStream(7, { retry: false }),
).rejects.toBeInstanceOf(ApiError);
expect(calls).toHaveLength(1);
});
it("exposes its resolved retry policy to the download layer", async () => {
// The download layer runs its own `withRetry` and must run it under
// the same policy the client was configured with, not under the
// library defaults.
const { fetch } = scriptedFetch(jsonResponse({}));
const client = new ApiClient({
fetch,
retry: { ...noWait, attempts: 9, baseDelayMs: 7, maxDelayMs: 11 },
});
const policy = client.getRetryOptions();
expect(policy.attempts).toBe(9);
expect(policy.baseDelayMs).toBe(7);
expect(policy.maxDelayMs).toBe(11);
});
});
describe("ApiClient timeouts", () => {
it("ships bounded default deadlines", () => {
// Asserted here so the README and the code cannot drift. Two numbers
// rather than one, because a deadline that is sane for a JSON call is
// nowhere near enough for a multi-gigabyte body, and a deadline long
// enough for that body would let a hung API call stall a backup for
// ten minutes.
expect(DEFAULT_REQUEST_TIMEOUT_MS).toBe(30_000);
expect(DEFAULT_DOWNLOAD_TIMEOUT_MS).toBe(600_000);
});
it("attaches an abort signal to every request", async () => {
const { fetch, calls } = recordingFetch(
jsonResponse({}),
jsonResponse({}),
new Response(null, { status: 200 }),
jsonResponse({}),
);
const client = new ApiClient({ fetch, retry: noWait });
await client.getJSON("/a");
await client.postJSON("/b", {});
await client.putFile("https://s3.example/x", new Uint8Array([1]));
await client.putJSON("/c", {});
for (const call of calls) {
expect(call.init?.signal).toBeInstanceOf(AbortSignal);
}
});
it("gives up on a request that never answers, and retries it", async () => {
// Before this policy existed there was no timeout anywhere in quak: a
// CDN connection that accepted the request and then went quiet would
// hang `quak backup` forever. `HANG` reproduces exactly that — the
// fake never answers, so the only thing that can end the call is the
// deadline quak attached.
const { fetch, calls } = scriptedFetch(HANG, HANG, HANG);
const client = new ApiClient({
fetch,
requestTimeoutMs: 20,
retry: { ...noWait, attempts: 3 },
});
await expect(client.getJSON("/black-hole")).rejects.toThrow();
// A timeout is retryable, so all three attempts were spent...
expect(calls).toHaveLength(3);
// ...each under its own fresh deadline, not one shared one that
// expired during the first attempt.
const signals = calls.map((c) => c.init?.signal);
expect(new Set(signals).size).toBe(3);
}, 5000);
it("recovers when a later attempt answers in time", async () => {
const { fetch, calls } = scriptedFetch(HANG, jsonResponse({ ok: 1 }));
const client = new ApiClient({
fetch,
requestTimeoutMs: 20,
retry: noWait,
});
await expect(client.getJSON("/slow-then-fast")).resolves.toEqual({
ok: 1,
});
expect(calls).toHaveLength(2);
}, 5000);
it("aborts a body that stalls after the headers arrived", async () => {
// The failure mode that a naive timeout misses. `getFileStream`
// returns as soon as headers arrive; the bytes are pulled later, in
// the download layer. A deadline that only guarded the initial fetch
// would leave the identical hang one layer down — which is where
// multi-megabyte photo downloads actually stall.
//
// This fake resolves its headers immediately and then serves a body
// that never produces a chunk and never observes the signal, so the
// only thing that can unblock the read is quak's own enforcement of
// the deadline over the stream it hands out.
const stalling = new Response(
new ReadableStream<Uint8Array>({
pull: () => new Promise<void>(() => {}),
}),
{ status: 200 },
);
const { fetch } = scriptedFetch(stalling);
const client = new ApiClient({
fetch,
downloadTimeoutMs: 20,
retry: { ...noWait, attempts: 1 },
});
const stream = await client.getFileStream(42);
const err: unknown = await readAll(stream).catch((e: unknown) => e);
expect(err).toBeInstanceOf(Error);
expect((err as Error).name).toBe("TimeoutError");
}, 5000);
it("lets a body that arrives in time through untouched", async () => {
// The counterpart to the previous test: enforcing the deadline over
// the stream must not corrupt or truncate a body that is simply being
// read normally.
const payload = new Uint8Array([9, 8, 7, 6, 5]);
const { fetch } = scriptedFetch(streamResponse(payload));
const client = new ApiClient({ fetch, retry: noWait });
const stream = await client.getFileStream(42);
const reader = stream.getReader();
const chunks: Uint8Array[] = [];
for (;;) {
const { done, value } = await reader.read();
if (done) break;
chunks.push(value);
}
const joined = new Uint8Array(chunks.reduce((n, c) => n + c.length, 0));
let offset = 0;
for (const c of chunks) {
joined.set(c, offset);
offset += c.length;
}
expect(joined).toEqual(payload);
});
});
describe("ApiClient error typing", () => {
it("throws ApiError with the status when a presigned PUT fails", async () => {
// `putFile` used to throw a bare Error with the status baked into a
// string. Nothing downstream could classify it, so a 500 from S3 was
// indistinguishable from a bug and could never be retried.
const { fetch, calls } = scriptedFetch(
new Response("Forbidden", { status: 403 }),
);
const client = new ApiClient({ fetch, retry: noWait });
const err: unknown = await client
.putFile("https://s3.example/obj", new Uint8Array([1, 2]))
.catch((e: unknown) => e);
expect(err).toBeInstanceOf(ApiError);
expect((err as ApiError).status).toBe(403);
expect(calls).toHaveLength(1);
});
it("retries a presigned PUT on a 5xx", async () => {
// A presigned PUT writes the whole object at one key in one request,
// so repeating it either overwrites the same bytes or lands them for
// the first time. There is no partial state to protect.
const { fetch, calls } = scriptedFetch(
new Response("slow down", { status: 503 }),
new Response(null, { status: 200 }),
);
const client = new ApiClient({ fetch, retry: noWait });
await client.putFile("https://s3.example/obj", new Uint8Array([1, 2]));
expect(calls).toHaveLength(2);
});
it("throws ApiError when a download response has no body", async () => {
// Also previously a bare Error. It carries the response status so a
// caller can see what arrived — and it is *not* retried: a 200 with
// no body is a malformed response, and asking again produces the same
// malformed response.
for (const get of ["getFileStream", "getThumbnailStream"] as const) {
const { fetch, calls } = scriptedFetch(
new Response(null, { status: 200 }),
new Response(null, { status: 200 }),
);
const client = new ApiClient({ fetch, retry: noWait });
const err: unknown = await client[get](5).catch((e: unknown) => e);
expect(err).toBeInstanceOf(ApiError);
expect((err as ApiError).status).toBe(200);
expect(calls).toHaveLength(1);
}
});
});
describe("ApiClient non-idempotent requests", () => {
/**
* `postJSON` and `putJSON` carry quak's only requests that change server
* state: `/users/srp/create-session`, `/users/two-factor/verify` — which
* consumes one of a small number of 2FA attempts — and `/files/thumbnail`.
*
* They are retried only on a failure that establishes no TCP connection to
* the server ever existed — DNS produced no address, or the peer refused
* the connection — so no request byte can have been transmitted.
* Everything else is ambiguous: a 5xx proves the server did process the
* request, and a reset or a timeout can arrive after it did.
* Replaying under that ambiguity can burn a 2FA attempt or register a
* thumbnail twice, and neither is worth the round trip it saves.
*/
it("does not replay a POST after a 5xx", async () => {
const { fetch, calls } = scriptedFetch(
textResponse("boom", 500),
jsonResponse({ sessionID: "second" }),
);
const client = new ApiClient({ fetch, retry: noWait });
await expect(
client.postJSON("/users/two-factor/verify", { code: "123456" }),
).rejects.toBeInstanceOf(ApiError);
expect(calls).toHaveLength(1);
});
it("does not replay a POST after a mid-flight connection reset", async () => {
// A reset can happen after the request was fully sent and acted on.
// It is retryable in general — `getJSON` retries it — but it is not
// replay-safe.
const { fetch, calls } = scriptedFetch(
errnoError("ECONNRESET", "socket hang up"),
jsonResponse({ sessionID: "second" }),
);
const client = new ApiClient({ fetch, retry: noWait });
await expect(
client.postJSON("/users/srp/create-session", {}),
).rejects.toThrow(/ECONNRESET|socket hang up/);
expect(calls).toHaveLength(1);
});
it("does not replay a POST after a timeout", async () => {
// A deadline says nothing about whether the server acted.
const { fetch, calls } = scriptedFetch(HANG, jsonResponse({}));
const client = new ApiClient({
fetch,
requestTimeoutMs: 20,
retry: noWait,
});
await expect(client.postJSON("/users/ott", {})).rejects.toThrow();
expect(calls).toHaveLength(1);
}, 5000);
it("replays a POST when the connection was never established", async () => {
// A refused connection or a DNS failure happens before any request
// byte is written, so the server cannot have seen it. This is the one
// case where replaying is provably harmless.
const { fetch, calls } = scriptedFetch(
new TypeError("fetch failed", {
cause: errnoError("ECONNREFUSED", "connect ECONNREFUSED"),
}),
errnoError("EAI_AGAIN", "getaddrinfo EAI_AGAIN api.ente.io"),
jsonResponse({ sessionID: "third" }),
);
const client = new ApiClient({ fetch, retry: noWait });
await expect(
client.postJSON<{ sessionID: string }>(
"/users/srp/create-session",
{},
),
).resolves.toEqual({ sessionID: "third" });
expect(calls).toHaveLength(3);
});
it("applies the same rule to PUT", async () => {
// `/files/thumbnail` is reached through `putJSON`.
const failing = scriptedFetch(
textResponse("boom", 503),
jsonResponse({}),
);
const failingClient = new ApiClient({
fetch: failing.fetch,
retry: noWait,
});
await expect(
failingClient.updateThumbnail(1, "key", "header"),
).rejects.toBeInstanceOf(ApiError);
expect(failing.calls).toHaveLength(1);
const refused = scriptedFetch(
errnoError("ECONNREFUSED", "connect ECONNREFUSED"),
jsonResponse({}),
);
const refusedClient = new ApiClient({
fetch: refused.fetch,
retry: noWait,
});
await refusedClient.updateThumbnail(1, "key", "header");
expect(refused.calls).toHaveLength(2);
});
});

View File

@@ -411,11 +411,22 @@ describe("quak backup", () => {
// File 101 (sunset.jpg) will return HTTP 500. The other two // File 101 (sunset.jpg) will return HTTP 500. The other two
// files must still download. The result must report the failure // files must still download. The result must report the failure
// without throwing. // without throwing.
//
// A 500 is retryable, so this file now costs several requests before
// it is given up on — that is the point of the retry policy, and
// `runBackup`'s own resilience is unchanged by it: the retry lives
// strictly below this loop, and an exhausted file is still logged,
// counted, and stepped over rather than aborting the run. The
// injected `sleep` is what keeps the suite from actually waiting out
// the backoff.
const outDir = join(testDir, "partial-failure"); const outDir = join(testDir, "partial-failure");
const client = await Client.login({ const client = await Client.login({
email: TEST_EMAIL, email: TEST_EMAIL,
password: TEST_PASSWORD, password: TEST_PASSWORD,
apiOptions: { fetch: buildMockFetch(mock, { failFileID: 101 }) }, apiOptions: {
fetch: buildMockFetch(mock, { failFileID: 101 }),
retry: { sleep: () => Promise.resolve(), random: () => 0 },
},
}); });
const result = await runBackup(client, outDir); const result = await runBackup(client, outDir);

View File

@@ -34,6 +34,15 @@
* forever after. The staging file and the rename are observed directly (see * forever after. The staging file and the rename are observed directly (see
* the `rename` hook below), not inferred from an empty directory. * the `rename` hook below), not inferred from an empty directory.
* *
* 3. **A failed transfer is retried as a whole.** A download is a request, a
* stream consumption, and a decryption, and only the first of those three
* happens inside `ApiClient`. A socket reset after the response headers
* have arrived therefore surfaces here, in the download layer — and that is
* the dominant failure mode for multi-megabyte photos over a CDN. So the
* entire sequence is retried as one unit, not just the request. The
* secretstream pull state is not resumable and there is no Range support,
* so a retry starts the file over from byte zero.
*
* These tests build synthetic encrypted files using sodium's push API, * These tests build synthetic encrypted files using sodium's push API,
* serve them from a mock fetch, and verify the decrypted output on disk. * serve them from a mock fetch, and verify the decrypted output on disk.
*/ */
@@ -61,6 +70,8 @@ import {
} from "vitest"; } from "vitest";
import { init, toBase64, STREAM_CHUNK_SIZE } from "../../src/crypto/index.js"; import { init, toBase64, STREAM_CHUNK_SIZE } from "../../src/crypto/index.js";
import { ApiClient } from "../../src/api/client.js"; import { ApiClient } from "../../src/api/client.js";
import { ApiError, TruncatedStreamError } from "../../src/errors.js";
import type { RetryOptions } from "../../src/retry.js";
import { downloadFile, downloadThumbnail } from "../../src/download/index.js"; import { downloadFile, downloadThumbnail } from "../../src/download/index.js";
import type { EnteFile, FileMetadata } from "../../src/model/types.js"; import type { EnteFile, FileMetadata } from "../../src/model/types.js";
@@ -326,18 +337,92 @@ beforeAll(() => {
multiChunk = encryptMultiChunkBody(multiChunkKey, 1, 1024); multiChunk = encryptMultiChunkBody(multiChunkKey, 1, 1024);
}); });
/** An error shaped like a Node transport failure: the errno is on `.code`. */
const errnoError = (code: string, message = code): Error =>
Object.assign(new Error(message), { code });
/**
* A retry policy with the waiting removed, used by every fixture in this
* file. Backoff arithmetic belongs to `test/retry/retry.test.ts`; here the
* only interesting quantity is how many requests a download issued, so the
* injected `sleep` returns immediately and nothing in this file waits.
*/
const noWait: RetryOptions = {
sleep: () => Promise.resolve(),
random: () => 0,
};
/**
* One scripted outcome for a single request to the CDN.
*
* - `body` — a complete response body.
* - `status` — an HTTP error response.
* - `reset` — a response whose headers arrive, whose body delivers `bytes`,
* and which then dies with a socket reset. This is the failure that
* motivates retrying the download rather than the request: by the time it
* happens `ApiClient` has already returned successfully.
*/
type BodyStep =
| { kind: "body"; bytes: Uint8Array }
| { kind: "status"; status: number }
| { kind: "reset"; bytes: Uint8Array };
/**
* A fetch that serves one scripted step per call and counts the calls. It
* deliberately refuses to serve more requests than it was given steps for, so
* a retry loop that ran away is a test failure rather than a silent success.
*/
const scriptedCdnFetch = (
...steps: BodyStep[]
): { fetch: typeof globalThis.fetch; requests: () => number } => {
let calls = 0;
const fake = async (): Promise<Response> => {
const step = steps[calls++];
if (step === undefined) {
throw new Error(`scriptedCdnFetch: no step for request #${calls}`);
}
if (step.kind === "status") {
return new Response("error", { status: step.status });
}
if (step.kind === "body") {
return new Response(step.bytes, { status: 200 });
}
const bytes = step.bytes;
return new Response(
new ReadableStream<Uint8Array>({
start(controller) {
controller.enqueue(bytes);
controller.error(
errnoError("ECONNRESET", "aborted by peer"),
);
},
}),
{ status: 200 },
);
};
return { fetch: fake as typeof globalThis.fetch, requests: () => calls };
};
/** /**
* Build an EnteFile plus ApiClient whose file *and* thumbnail streams both * Build an EnteFile plus ApiClient whose file *and* thumbnail streams both
* serve `body` under `header`. The download path under test is otherwise * serve `body` under `header`. The download path under test is otherwise
* identical for the two, so every truncation/atomicity case below runs * identical for the two, so every truncation/atomicity case below runs
* against both entry points from a single fixture. * against both entry points from a single fixture.
*
* The default policy here is a single attempt. The failure-contract tests are
* about what the caller and the filesystem are left with, not about how many
* times quak asked; pinning attempts to one keeps them saying exactly that,
* and keeps them from re-decrypting a 4 MiB fixture four times over. The
* retry counts have their own tests at the bottom of this file, which set the
* attempt count explicitly.
*/ */
const fixtureFor = ( const fixtureFor = (
key: Uint8Array, key: Uint8Array,
header: Uint8Array, header: Uint8Array,
body: Uint8Array, body: Uint8Array,
retry: RetryOptions = { ...noWait, attempts: 1 },
): { api: ApiClient; file: EnteFile } => ({ ): { api: ApiClient; file: EnteFile } => ({
api: new ApiClient({ fetch: mockFetchForBody(body) }), api: new ApiClient({ fetch: mockFetchForBody(body), retry }),
file: buildMockEnteFile(key, header, header), file: buildMockEnteFile(key, header, header),
}); });
@@ -502,9 +587,17 @@ describe.each(entryPoints)(
); );
const outPath = join(freshDir(), "truncated.bin"); const outPath = join(freshDir(), "truncated.bin");
await expect(download(api, file, outPath)).rejects.toThrow( const err: unknown = await download(api, file, outPath).catch(
/truncated/i, (e: unknown) => e,
); );
// The type, not the wording, is the contract. The retry policy
// classifies truncation as worth another attempt, and it decides
// that with `instanceof`: matching on message text would make
// rewording a diagnostic silently turn every truncated download
// into a permanent failure.
expect(err).toBeInstanceOf(TruncatedStreamError);
expect((err as Error).message).toMatch(/truncated/i);
}); });
it("rejects a body whose final chunk arrived only in part", async () => { it("rejects a body whose final chunk arrived only in part", async () => {
@@ -537,7 +630,7 @@ describe.each(entryPoints)(
(e: unknown) => e, (e: unknown) => e,
); );
expect(err).toBeInstanceOf(Error); expect(err).toBeInstanceOf(TruncatedStreamError);
expect((err as Error).message).toMatch(/truncated/i); expect((err as Error).message).toMatch(/truncated/i);
expect((err as Error).cause).toBeInstanceOf(Error); expect((err as Error).cause).toBeInstanceOf(Error);
expect(((err as Error).cause as Error).message).toMatch( expect(((err as Error).cause as Error).message).toMatch(
@@ -564,8 +657,8 @@ describe.each(entryPoints)(
); );
const outPath = join(freshDir(), "empty.bin"); const outPath = join(freshDir(), "empty.bin");
await expect(download(api, file, outPath)).rejects.toThrow( await expect(download(api, file, outPath)).rejects.toBeInstanceOf(
/truncated/i, TruncatedStreamError,
); );
}); });
@@ -588,8 +681,8 @@ describe.each(entryPoints)(
const dir = freshDir(); const dir = freshDir();
const outPath = join(dir, "absent.bin"); const outPath = join(dir, "absent.bin");
await expect(download(api, file, outPath)).rejects.toThrow( await expect(download(api, file, outPath)).rejects.toBeInstanceOf(
/truncated/i, TruncatedStreamError,
); );
expect(existsSync(outPath)).toBe(false); expect(existsSync(outPath)).toBe(false);
@@ -621,10 +714,18 @@ describe.each(entryPoints)(
const dir = freshDir(); const dir = freshDir();
const outPath = join(dir, "corrupt.bin"); const outPath = join(dir, "corrupt.bin");
await expect(download(api, file, outPath)).rejects.toThrow( const err: unknown = await download(api, file, outPath).catch(
/authentication failed/i, (e: unknown) => e,
); );
expect(err).toBeInstanceOf(Error);
expect((err as Error).message).toMatch(/authentication failed/i);
// And explicitly *not* the truncation type, because that type is
// what the retry policy keys on: mislabelling corruption as
// truncation would spend the whole attempt budget re-downloading
// a file that will never decrypt.
expect(err).not.toBeInstanceOf(TruncatedStreamError);
expect(existsSync(outPath)).toBe(false); expect(existsSync(outPath)).toBe(false);
expect(readdirSync(dir)).toEqual([]); expect(readdirSync(dir)).toEqual([]);
}); });
@@ -648,8 +749,8 @@ describe.each(entryPoints)(
const outPath = join(dir, "existing.bin"); const outPath = join(dir, "existing.bin");
writeFileSync(outPath, existing); writeFileSync(outPath, existing);
await expect(download(api, file, outPath)).rejects.toThrow( await expect(download(api, file, outPath)).rejects.toBeInstanceOf(
/truncated/i, TruncatedStreamError,
); );
expect(readFileSync(outPath)).toEqual(Buffer.from(existing)); expect(readFileSync(outPath)).toEqual(Buffer.from(existing));
@@ -755,3 +856,208 @@ describe.each(entryPoints)(
}); });
}, },
); );
// ---------------------------------------------------------------------------
// Retries
//
// What is retried here is the whole download — request, stream consumption,
// decryption — because only the first of those three happens inside
// `ApiClient`. Every assertion counts requests; none of them measures time.
// ---------------------------------------------------------------------------
describe.each(entryPoints)("$name retries", ({ name, download }) => {
const freshDir = (): string => mkdtempSync(join(testDir, `${name}-retry-`));
/** A cheap single-chunk fixture: no 4 MiB encryption in the retry tests. */
const smallFixture = (
seed: number,
): {
key: Uint8Array;
header: Uint8Array;
ciphertext: Uint8Array;
plaintext: Uint8Array;
} => {
const plaintext = patternBytes(1024, seed);
const key = sodium.crypto_secretstream_xchacha20poly1305_keygen();
const { header, ciphertext } = encryptFileBody(plaintext, key);
return { key, header, ciphertext, plaintext };
};
const clientFor = (
fetch: typeof globalThis.fetch,
attempts: number,
): ApiClient => new ApiClient({ fetch, retry: { ...noWait, attempts } });
it("retries a connection reset that happened mid-body", async () => {
// The case `ApiClient` cannot see. Its own request succeeded: headers
// arrived, a `ReadableStream` was handed back, and only then did the
// socket die. Retrying the fetch alone would have caught nothing,
// which is why the retry wraps the whole sequence.
const { key, header, ciphertext, plaintext } = smallFixture(41);
const { fetch, requests } = scriptedCdnFetch(
{ kind: "reset", bytes: ciphertext.slice(0, 16) },
{ kind: "body", bytes: ciphertext },
);
const api = clientFor(fetch, 4);
const file = buildMockEnteFile(key, header, header);
const dir = freshDir();
const outPath = join(dir, "reset-then-ok.bin");
const result = await download(api, file, outPath);
expect(requests()).toBe(2);
expect(result.bytesWritten).toBe(plaintext.length);
expectSameBytes(readFileSync(outPath), plaintext);
});
it("stages one temp file for the attempt that succeeded, not one per attempt", async () => {
// The atomic write stays outside the retry loop. A retried download
// must not leave a trail of half-written scratch files, and the
// destination must be touched exactly once — by the attempt that
// produced a complete, authenticated plaintext.
const { key, header, ciphertext } = smallFixture(42);
const { fetch } = scriptedCdnFetch(
{ kind: "reset", bytes: ciphertext.slice(0, 16) },
{ kind: "reset", bytes: ciphertext.slice(0, 16) },
{ kind: "body", bytes: ciphertext },
);
const api = clientFor(fetch, 4);
const file = buildMockEnteFile(key, header, header);
const dir = freshDir();
const outPath = join(dir, "one-stage.bin");
await download(api, file, outPath);
expect(renameHook.calls).toHaveLength(1);
expect(renameHook.calls[0]!.to).toBe(outPath);
expect(readdirSync(dir)).toEqual(["one-stage.bin"]);
});
it("retries a truncated body and gives up after the configured attempts", async () => {
// Truncation is retryable — the file on the server is intact, the
// transfer was not — but it is not retryable forever. Three attempts
// configured, three requests, then the caller gets the error.
const key = sodium.crypto_secretstream_xchacha20poly1305_keygen();
const { header, ciphertext } = encryptNonFinalBody(
patternBytes(256, 43),
key,
);
const { fetch, requests } = scriptedCdnFetch(
{ kind: "body", bytes: ciphertext },
{ kind: "body", bytes: ciphertext },
{ kind: "body", bytes: ciphertext },
{ kind: "body", bytes: ciphertext },
);
const api = clientFor(fetch, 3);
const file = buildMockEnteFile(key, header, header);
const dir = freshDir();
const outPath = join(dir, "always-truncated.bin");
await expect(download(api, file, outPath)).rejects.toBeInstanceOf(
TruncatedStreamError,
);
expect(requests()).toBe(3);
// Every attempt failed before anything was written, so the directory
// is still empty.
expect(readdirSync(dir)).toEqual([]);
});
it("issues exactly one request when the file is gone", async () => {
// A 404 from the CDN is an answer. `runBackup` logs it and moves on;
// spending three more requests and three backoff waits on it would
// slow a large backup down for nothing.
const { key, header } = smallFixture(44);
const { fetch, requests } = scriptedCdnFetch(
{ kind: "status", status: 404 },
{ kind: "status", status: 404 },
{ kind: "status", status: 404 },
{ kind: "status", status: 404 },
);
const api = clientFor(fetch, 4);
const file = buildMockEnteFile(key, header, header);
const outPath = join(freshDir(), "gone.bin");
const err: unknown = await download(api, file, outPath).catch(
(e: unknown) => e,
);
expect(err).toBeInstanceOf(ApiError);
expect((err as ApiError).status).toBe(404);
expect(requests()).toBe(1);
});
it("retries a 503 from the CDN", async () => {
const { key, header, ciphertext, plaintext } = smallFixture(45);
const { fetch, requests } = scriptedCdnFetch(
{ kind: "status", status: 503 },
{ kind: "status", status: 503 },
{ kind: "body", bytes: ciphertext },
);
const api = clientFor(fetch, 4);
const file = buildMockEnteFile(key, header, header);
const outPath = join(freshDir(), "flaky-cdn.bin");
await download(api, file, outPath);
expect(requests()).toBe(3);
expectSameBytes(readFileSync(outPath), plaintext);
});
it("spends one attempt budget, not one per layer", async () => {
// `ApiClient.getFileStream` retries on its own for direct callers.
// The download layer opts out of that and runs its own retry over the
// whole sequence. If it did not, the two budgets would compose: three
// attempts here would become nine requests to the CDN for a single
// file, and the default four would become sixteen.
const { key, header } = smallFixture(46);
const steps: BodyStep[] = Array.from({ length: 12 }, () => ({
kind: "status" as const,
status: 503,
}));
const { fetch, requests } = scriptedCdnFetch(...steps);
const api = clientFor(fetch, 3);
const file = buildMockEnteFile(key, header, header);
const outPath = join(freshDir(), "budget.bin");
await expect(download(api, file, outPath)).rejects.toBeInstanceOf(
ApiError,
);
expect(requests()).toBe(3);
});
});
describe("download retries: corruption is not retried", () => {
it("gives up immediately on a chunk that failed to authenticate", async () => {
// A whole chunk that failed to authenticate while the stream
// continued past it is corruption or a wrong key. Neither is fixed by
// asking again, and a backup run that retried every such file would
// multiply the cost of a genuinely broken file by the attempt count.
//
// This is also the boundary of the single-chunk ambiguity documented
// at the classifier: the split is only achievable because this body
// has more than one chunk.
const corrupted = Uint8Array.from(multiChunk.body);
corrupted[10] ^= 0xff;
const { fetch, requests } = scriptedCdnFetch(
{ kind: "body", bytes: corrupted },
{ kind: "body", bytes: corrupted },
{ kind: "body", bytes: corrupted },
{ kind: "body", bytes: corrupted },
);
const api = new ApiClient({ fetch, retry: { ...noWait, attempts: 4 } });
const file = buildMockEnteFile(
multiChunkKey,
multiChunk.header,
multiChunk.header,
);
const outPath = join(mkdtempSync(join(testDir, "corrupt-")), "c.bin");
await expect(downloadFile(api, file, outPath)).rejects.toThrow(
/authentication failed/i,
);
expect(requests()).toBe(1);
});
});

View File

@@ -0,0 +1,49 @@
// The Docker build context is load-bearing in two directions, and both
// failures are silent.
//
// Excluding too little: a worktree left under `.claude/` is copied into the
// image, vitest globs its `test/` tree as well as the real one, and the
// containerised `make check` runs the whole suite twice over while reporting
// success. A compiled `bin/quak` is ~100 MB of context nobody needs.
//
// Excluding too much: Prettier 3 reads `.gitignore` as a default ignore file,
// so dropping it from the context silently changes which files
// `make fmt-check` looks at inside the image compared to the host.
//
// Neither shows up as a build failure, so they are asserted here.
import { describe, expect, it } from "vitest";
import { readFileSync } from "node:fs";
import { fileURLToPath } from "node:url";
import { join } from "node:path";
const repoRoot = fileURLToPath(new URL("../../", import.meta.url));
const patterns = (name: string): string[] =>
readFileSync(join(repoRoot, name), "utf-8")
.split("\n")
.map((line) => line.trim())
.filter((line) => line !== "" && !line.startsWith("#"));
const dockerignore = patterns(".dockerignore");
describe(".dockerignore", () => {
// Everything here is either generated, enormous, or secret. `.claude/` is
// the correctness one: see the header comment and issue #25.
it.each([
".claude/",
".quak/",
"bin/quak",
"node_modules",
"coverage",
"dist",
".vitest-cache/",
".nyc_output/",
"*.tsbuildinfo",
])("keeps %s out of the build context", (pattern) => {
expect(dockerignore).toContain(pattern);
});
it("leaves .gitignore in the build context for prettier", () => {
expect(dockerignore).not.toContain(".gitignore");
});
});

View File

@@ -0,0 +1,116 @@
// The package manifest promises three files that only exist after a build:
// `main`, `types`, and the `quak` binary. Nothing in the test suite used to
// look at them, and `make check` runs test, lint and fmt-check but never the
// build, so `tsconfig.json` and `package.json` were free to drift apart. They
// did: `rootDir` was `./src` while `include` also pulled in `bin/**/*`, which
// is TS6059, and no build had succeeded for as long as that was true.
//
// These tests read both files and check the contract between them, without
// running a build, so they stay in the fast unit suite. The complementary
// check — that the files really landed on disk — is in `script/build`, which
// runs after the compiler and is the only place that can honestly answer it.
import { describe, expect, it } from "vitest";
import { readFileSync } from "node:fs";
import { fileURLToPath } from "node:url";
import { join, posix } from "node:path";
interface TsConfig {
compilerOptions: {
outDir: string;
rootDir: string;
};
include: string[];
}
interface PackageJson {
main: string;
types: string;
bin: Record<string, string>;
scripts: Record<string, string>;
}
const repoRoot = fileURLToPath(new URL("../../", import.meta.url));
const readJSON = <T>(name: string): T =>
JSON.parse(readFileSync(join(repoRoot, name), "utf-8")) as T;
const tsconfig = readJSON<TsConfig>("tsconfig.json");
const pkg = readJSON<PackageJson>("package.json");
// Paths in the two manifests are written with a leading "./"; normalize so
// they can be compared and joined. Everything here is POSIX-style because
// that is what both JSON files contain, regardless of the host OS.
const clean = (p: string): string => posix.normalize(p);
const outDir = clean(tsconfig.compilerOptions.outDir);
const rootDir = clean(tsconfig.compilerOptions.rootDir);
// The directory prefix of a glob: the part before the first segment
// containing a wildcard. "src/**/*" -> "src", "bin/**/*" -> "bin".
const globRoot = (pattern: string): string => {
const segments = clean(pattern).split("/");
const wildcard = segments.findIndex((s) => s.includes("*"));
return (
segments
.slice(0, wildcard === -1 ? segments.length : wildcard)
.join("/") || "."
);
};
// Where tsc will write the output for a source file: the path relative to
// rootDir, re-rooted under outDir, with the extension swapped.
const emitted = (source: string, extension: string): string =>
"./" +
posix
.join(outDir, posix.relative(rootDir, clean(source)))
.replace(/\.ts$/, extension);
describe("tsconfig include and rootDir", () => {
// TS6059 is not a style complaint: tsc refuses to emit anything at all
// when a compiled file sits outside rootDir, so this single mismatch
// took out both the library and the CLI artifacts.
it("compiles only files that live under rootDir", () => {
for (const pattern of tsconfig.include) {
const root = globRoot(pattern);
const relative = posix.relative(rootDir, root);
expect(
relative === "" || !relative.startsWith(".."),
`include pattern ${pattern} matches files outside rootDir ` +
`${rootDir}; tsc rejects that with TS6059`,
).toBe(true);
}
});
});
describe("package.json entrypoints", () => {
// Each of these is the path a consumer resolves — `import { Client } from
// "quak"` for main, the editor for types, `npx quak` for the bin — so a
// wrong value is a broken package even when the build itself is green.
it("names the file tsc emits for src/index.ts as main", () => {
expect(pkg.main).toBe(emitted("src/index.ts", ".js"));
});
it("names the declaration tsc emits for src/index.ts as types", () => {
expect(pkg.types).toBe(emitted("src/index.ts", ".d.ts"));
});
it("names the file tsc emits for bin/quak.ts as the quak binary", () => {
expect(pkg.bin.quak).toBe(emitted("bin/quak.ts", ".js"));
});
// The README's Getting Started block tells the reader to run
// `yarn quak login` straight after `yarn build`. That only works if a
// `quak` script exists and points at the built CLI, not at the source.
it("runs the built CLI from the quak script", () => {
expect(pkg.scripts.quak).toBeDefined();
expect(pkg.scripts.quak).toContain(pkg.bin.quak);
});
});
describe("bin/quak.ts", () => {
// tsc copies the shebang into the emitted file, so it has to be in the
// source for the installed binary to be directly executable.
it("starts with a node shebang", () => {
const source = readFileSync(join(repoRoot, "bin/quak.ts"), "utf-8");
expect(source.split("\n")[0]).toBe("#!/usr/bin/env node");
});
});

View File

@@ -0,0 +1,30 @@
// `script/projectname` is the single source of the project's name for every
// script that needs one — `script/docker` builds its image tag from it, which
// is the whole reason that file exists. Nothing checked that it agreed with
// `package.json`, and after the repo was renamed it did not: the script still
// said "quack", so `make docker` produced an image tagged after a name this
// project has not used since May.
//
// The script is executed rather than read, because what matters is the string
// it prints, not the source it prints it from.
import { describe, expect, it } from "vitest";
import { execFileSync } from "node:child_process";
import { readFileSync } from "node:fs";
import { fileURLToPath } from "node:url";
import { join } from "node:path";
const repoRoot = fileURLToPath(new URL("../../", import.meta.url));
const pkg = JSON.parse(
readFileSync(join(repoRoot, "package.json"), "utf-8"),
) as { name: string };
describe("script/projectname", () => {
it("prints the name package.json declares", () => {
const printed = execFileSync(join(repoRoot, "script/projectname"), {
cwd: repoRoot,
encoding: "utf-8",
}).trim();
expect(printed).toBe(pkg.name);
});
});

570
test/retry/retry.test.ts Normal file
View File

@@ -0,0 +1,570 @@
/**
* Tests for `src/retry.ts` — the retry policy shared by every network
* operation in quak.
*
* Two things live in that module and they are deliberately separate:
*
* - **`isRetryable(err)`**, a pure classifier. Given an error, is trying
* again capable of producing a different answer? Nothing else about the
* error matters: not how it was logged, not where it came from.
*
* - **`withRetry(fn, opts)`**, the loop. It calls `fn`, and while the error
* is classified retryable and attempts remain, it sleeps and calls `fn`
* again. It never inspects errors itself.
*
* The classifier's default answer is *no*. quak is a backup tool: a wrongly
* retried permanent failure costs a user round trips and delays the rest of
* the run, while a wrongly rejected transient failure costs one file that the
* next run picks up. When in doubt, fail fast.
*
* ## Reading the backoff assertions
*
* `withRetry` takes its `sleep` and its `random` as injected functions. Every
* test here passes a `sleep` that records the delay it was asked for and
* returns immediately, so the suite never waits, and a `random` that returns a
* fixed number, so jitter is exact rather than approximate. **No assertion in
* this file (or anywhere else in the suite) is about elapsed wall-clock time.**
* They are about what `withRetry` asked for, and how many times `fn` ran.
*
* The delay before retry number *n* (1-based) is:
*
* random() * min(maxDelayMs, baseDelayMs * 2 ** (n - 1))
*
* That is exponential backoff with full jitter: the exponential term is the
* *ceiling*, and the actual wait is drawn uniformly below it. Full jitter,
* rather than a fixed delay plus noise, is what stops a client that lost a
* hundred parallel downloads to one CDN blip from re-sending all hundred at
* the same instant.
*/
import { describe, expect, it } from "vitest";
import {
DEFAULT_RETRY_OPTIONS,
isRetryable,
isSafeToReplay,
resolveRetryOptions,
withRetry,
} from "../../src/retry.js";
import { ApiError, TruncatedStreamError } from "../../src/errors.js";
import { ApiError as ApiErrorFromClient } from "../../src/api/client.js";
// ---------------------------------------------------------------------------
// Test helpers
// ---------------------------------------------------------------------------
/**
* A `sleep` that records what it was asked to wait for and returns
* immediately. This is the whole reason `withRetry` takes an injected sleep:
* the retry policy is exercised in full — every branch, every delay — without
* the suite spending a single millisecond waiting.
*/
const recordingSleep = (): {
sleep: (ms: number) => Promise<void>;
delays: number[];
} => {
const delays: number[] = [];
return {
sleep: (ms: number): Promise<void> => {
delays.push(ms);
return Promise.resolve();
},
delays,
};
};
/** An error shaped like a Node transport failure: the errno is on `.code`. */
const errnoError = (code: string, message = code): Error =>
Object.assign(new Error(message), { code });
// ---------------------------------------------------------------------------
// Classification
// ---------------------------------------------------------------------------
describe("isRetryable: HTTP status codes", () => {
it("does not retry ordinary 4xx responses", () => {
// A 4xx is the server saying the request itself is wrong. Repeating
// it verbatim produces the same answer, so retrying only delays the
// failure the caller has to handle. 404 is the load-bearing case:
// `listMissingThumbnails` depends on a 404 arriving promptly and
// exactly once.
for (const status of [400, 401, 403, 404, 409, 410, 422]) {
expect(isRetryable(new ApiError(`HTTP ${status}`, status))).toBe(
false,
);
}
});
it("retries 408 and 429", () => {
// The two 4xx codes that are statements about timing rather than
// about the request. 408 is the server admitting it gave up waiting;
// 429 is it asking for less traffic — which backoff supplies.
expect(isRetryable(new ApiError("timeout", 408))).toBe(true);
expect(isRetryable(new ApiError("slow down", 429))).toBe(true);
});
it("retries every 5xx response", () => {
// A 5xx is the server failing, not the request being wrong. Ente's
// CDN in particular returns 500 and 503 under load.
for (const status of [500, 502, 503, 504, 599]) {
expect(isRetryable(new ApiError(`HTTP ${status}`, status))).toBe(
true,
);
}
});
it("treats a 2xx or 3xx ApiError as not retryable", () => {
// These exist: `getFileStream` raises an ApiError carrying the
// response status when a 200 arrives with a null body. That is a
// malformed response, not a transport failure, and repeating the
// request will produce the same malformed response.
expect(isRetryable(new ApiError("response body is null", 200))).toBe(
false,
);
expect(isRetryable(new ApiError("redirect", 304))).toBe(false);
});
it("uses the same ApiError class that ApiClient exports", () => {
// `ApiError` lives in `src/errors.ts` and is re-exported from
// `src/api/client.ts`, which is where every existing caller and test
// imports it from. If those ever became two separate classes the
// classifier would silently stop recognising errors raised by the
// client, and every 5xx in the wild would be treated as permanent.
expect(ApiErrorFromClient).toBe(ApiError);
expect(new ApiErrorFromClient("boom", 503)).toBeInstanceOf(ApiError);
expect(isRetryable(new ApiErrorFromClient("boom", 503))).toBe(true);
});
});
describe("isRetryable: transport failures", () => {
it("retries a TypeError, which is how fetch reports a failed request", () => {
// Node's fetch rejects with `TypeError: fetch failed` for everything
// below HTTP: DNS failure, refused connection, TLS error, reset
// socket. The real diagnosis is on `cause`, but there is nothing on
// the object that distinguishes it from a TypeError thrown by a bug,
// so this rule is deliberately literal. The cost of the imprecision
// is bounded by the attempt count; the alternative — demanding a
// recognised `cause` — would classify real network failures as
// permanent and fail backups that should have succeeded.
expect(isRetryable(new TypeError("fetch failed"))).toBe(true);
});
it("retries an errno carried on the error itself", () => {
for (const code of [
"ECONNRESET",
"ETIMEDOUT",
"EPIPE",
"ENOTFOUND",
"EAI_AGAIN",
"ECONNREFUSED",
"EHOSTUNREACH",
"ENETUNREACH",
]) {
expect(isRetryable(errnoError(code))).toBe(true);
}
});
it("retries an errno buried in the cause chain", () => {
// undici does not put the errno on the error it throws; it hangs the
// underlying socket error off `cause`, sometimes more than one level
// down. A classifier that only looked at the top-level error would
// see a bare `Error` and call every dropped connection permanent.
const nested = new Error("request to files.ente.io failed", {
cause: new Error("socket hang up", {
cause: errnoError("ECONNRESET", "read ECONNRESET"),
}),
});
expect(isRetryable(nested)).toBe(true);
});
it("accepts a plain object as a cause", () => {
// Not everything in a cause chain is an Error instance.
expect(
isRetryable(new Error("failed", { cause: { code: "ETIMEDOUT" } })),
).toBe(true);
});
it("retries an aborted request", () => {
// `AbortSignal.timeout()` aborts with a `TimeoutError`; an explicit
// `abort()` produces an `AbortError`. quak only ever aborts a request
// on its own deadline, so both mean "this attempt ran out of time",
// which is exactly the condition a later attempt might not hit.
expect(isRetryable(new DOMException("timed out", "TimeoutError"))).toBe(
true,
);
expect(isRetryable(new DOMException("aborted", "AbortError"))).toBe(
true,
);
});
it("does not confuse an unrelated errno with a transport failure", () => {
// A filesystem error surfaces the same way an errno network error
// does. Retrying a full disk or a missing directory is pointless.
expect(isRetryable(errnoError("ENOSPC", "no space left"))).toBe(false);
expect(isRetryable(errnoError("ENOENT", "no such file"))).toBe(false);
expect(isRetryable(errnoError("EACCES", "permission denied"))).toBe(
false,
);
});
it("terminates on a cause chain that points at itself", () => {
// Defensive: `cause` is an arbitrary user-settable property and
// nothing stops it forming a cycle. Without a bound on the walk this
// classifier would hang the process, which is a worse failure than
// any misclassification.
const looped: Error & { cause?: unknown } = new Error("loop");
looped.cause = looped;
expect(isRetryable(looped)).toBe(false);
});
});
describe("isRetryable: stream truncation versus corruption", () => {
it("retries a truncated stream", () => {
// Truncation is a transfer that stopped early. The bytes that did
// arrive are useless, but the file on the server is fine, so asking
// again is exactly right.
expect(isRetryable(new TruncatedStreamError("stream truncated"))).toBe(
true,
);
});
it("retries a truncated stream whose cause is an authentication failure", () => {
// The single-chunk ambiguity, recorded on issue #2 and inherited from
// the truncation work: when a body ends part-way through a chunk,
// Poly1305 fails and carries no framing signal, so a cut connection
// and genuinely corrupt bytes are indistinguishable. That case is
// reported as truncation with the authentication failure preserved as
// `cause`, and it is therefore retried.
//
// Retrying is the deliberate choice. For a multi-chunk body the split
// is real — a corrupt chunk mid-stream stays an authentication
// failure, see the next test — but for a single-chunk body (most
// thumbnails, every small file) a wrong key and a cut connection look
// identical. The cost of guessing wrong is bounded: a few extra round
// trips before the same failure. The cost of guessing the other way
// is a silently truncated file kept forever.
const err = new TruncatedStreamError("stream truncated", {
cause: new Error("secretstream chunk authentication failed"),
});
expect(isRetryable(err)).toBe(true);
});
it("does not retry an authentication failure that is not truncation", () => {
// A whole chunk that failed to authenticate while the stream carried
// on past it cannot be a short transfer. It is corruption or a wrong
// key, and no number of retries fixes either.
expect(
isRetryable(new Error("secretstream chunk authentication failed")),
).toBe(false);
});
});
describe("isRetryable: everything else", () => {
it("does not retry programming errors or unknown values", () => {
expect(isRetryable(new Error("boom"))).toBe(false);
expect(isRetryable(new RangeError("out of range"))).toBe(false);
expect(isRetryable(new SyntaxError("bad JSON"))).toBe(false);
expect(isRetryable("a string")).toBe(false);
expect(isRetryable(undefined)).toBe(false);
expect(isRetryable(null)).toBe(false);
});
});
// ---------------------------------------------------------------------------
// The replay-safety classifier for non-idempotent requests
// ---------------------------------------------------------------------------
describe("isSafeToReplay", () => {
/**
* `isRetryable` answers "could a retry succeed?". For a POST or a PUT
* that is not the whole question: the other half is "could the first
* attempt already have taken effect on the server?".
*
* quak's non-idempotent calls are `/users/srp/create-session`,
* `/users/two-factor/verify` (which consumes one of a limited number of
* 2FA attempts) and `/files/thumbnail`. A blind replay of any of them can
* do real damage, so they retry only on the failures that establish no TCP
* connection to the server ever existed — DNS produced no address, or the
* peer refused the connection — and therefore that no request byte can
* have been transmitted.
*/
it("replays only failures where the connection was never established", () => {
for (const code of ["ENOTFOUND", "EAI_AGAIN", "ECONNREFUSED"]) {
expect(isSafeToReplay(errnoError(code))).toBe(true);
}
// Also when undici has buried it, which is how it actually arrives.
expect(
isSafeToReplay(
new TypeError("fetch failed", {
cause: errnoError("ECONNREFUSED"),
}),
),
).toBe(true);
});
it("does not replay a failure that could have happened after the server acted", () => {
// Every one of these is ambiguous about whether the server processed
// the request. A 5xx proves it did. A reset or a broken pipe can
// arrive after the request was fully sent and handled. A timeout says
// nothing at all about the server's state. A bare `fetch failed` with
// no recognisable cause could be any of them.
expect(isSafeToReplay(new ApiError("HTTP 500", 500))).toBe(false);
expect(isSafeToReplay(new ApiError("HTTP 429", 429))).toBe(false);
expect(isSafeToReplay(errnoError("ECONNRESET"))).toBe(false);
expect(isSafeToReplay(errnoError("EPIPE"))).toBe(false);
expect(isSafeToReplay(errnoError("ETIMEDOUT"))).toBe(false);
// The routing errnos look like connect-time failures but are not. On
// Linux an ICMP destination-unreachable delivered on an established
// connection sets the socket error, and the next read or write returns
// `EHOSTUNREACH` or `ENETUNREACH`; a local interface going down after
// the request was fully written surfaces as `ENETDOWN` the same way.
// In each case the server may already have consumed the request — a
// replayed `/users/two-factor/verify` would burn a second attempt.
// They remain retryable for the idempotent calls; this asserts only
// that they are not replayable.
expect(isSafeToReplay(errnoError("EHOSTUNREACH"))).toBe(false);
expect(isSafeToReplay(errnoError("ENETUNREACH"))).toBe(false);
expect(isSafeToReplay(errnoError("ENETDOWN"))).toBe(false);
// ...and that the narrowing did not make them non-retryable.
expect(isRetryable(errnoError("EHOSTUNREACH"))).toBe(true);
expect(isRetryable(errnoError("ENETUNREACH"))).toBe(true);
expect(isRetryable(errnoError("ENETDOWN"))).toBe(true);
expect(
isSafeToReplay(new DOMException("timed out", "TimeoutError")),
).toBe(false);
expect(isSafeToReplay(new TypeError("fetch failed"))).toBe(false);
});
});
// ---------------------------------------------------------------------------
// The retry loop
// ---------------------------------------------------------------------------
describe("withRetry", () => {
it("calls the function once and does not sleep when it succeeds", async () => {
const { sleep, delays } = recordingSleep();
let calls = 0;
const result = await withRetry(
() => {
calls++;
return Promise.resolve("ok");
},
{ sleep },
);
expect(result).toBe("ok");
expect(calls).toBe(1);
expect(delays).toEqual([]);
});
it("stops at the first success and returns its value", async () => {
const { sleep, delays } = recordingSleep();
let calls = 0;
const result = await withRetry(
() => {
calls++;
if (calls < 3) {
return Promise.reject(new ApiError("HTTP 503", 503));
}
return Promise.resolve(calls);
},
{ attempts: 5, sleep },
);
expect(result).toBe(3);
expect(calls).toBe(3);
// Two failures, so two waits — and none after the attempt that
// succeeded.
expect(delays).toHaveLength(2);
});
it("gives up after `attempts` calls and throws the last error", async () => {
// `attempts` counts calls, not retries: `attempts: 3` means the
// function runs three times in total. The error that escapes is the
// one from the final attempt, because that is the current state of
// the world; a caller logging it is logging what is true now.
const { sleep, delays } = recordingSleep();
let calls = 0;
const failure = withRetry(
() => {
calls++;
return Promise.reject(new ApiError(`attempt ${calls}`, 503));
},
{ attempts: 3, sleep },
);
await expect(failure).rejects.toThrow("attempt 3");
expect(calls).toBe(3);
// Three attempts, two gaps between them. Sleeping after the last
// attempt would delay the caller's failure for nothing.
expect(delays).toHaveLength(2);
});
it("does not retry at all when attempts is 1", async () => {
const { sleep, delays } = recordingSleep();
let calls = 0;
await expect(
withRetry(
() => {
calls++;
return Promise.reject(new ApiError("HTTP 500", 500));
},
{ attempts: 1, sleep },
),
).rejects.toThrow("HTTP 500");
expect(calls).toBe(1);
expect(delays).toEqual([]);
});
it("rethrows a non-retryable error immediately", async () => {
const { sleep, delays } = recordingSleep();
let calls = 0;
await expect(
withRetry(
() => {
calls++;
return Promise.reject(new ApiError("HTTP 404", 404));
},
{ attempts: 5, sleep },
),
).rejects.toBeInstanceOf(ApiError);
expect(calls).toBe(1);
expect(delays).toEqual([]);
});
it("preserves the error object, not just its message", async () => {
// Callers classify what escapes: `listMissingThumbnails` needs the
// `ApiError` and its status to tell a genuine 404 from a transient
// failure. Wrapping the error in a "retries exhausted" error would
// break that.
const original = new ApiError("gone", 410, { code: "GONE" });
const err: unknown = await withRetry(() => Promise.reject(original), {
sleep: () => Promise.resolve(),
}).catch((e: unknown) => e);
expect(err).toBe(original);
});
it("honours a caller-supplied classifier", async () => {
// This is how the non-idempotent call sites narrow the policy: same
// loop, same backoff, stricter question.
const { sleep } = recordingSleep();
let calls = 0;
await expect(
withRetry(
() => {
calls++;
// Retryable under the default policy...
return Promise.reject(new ApiError("HTTP 503", 503));
},
{ attempts: 4, sleep, isRetryable: isSafeToReplay },
),
).rejects.toThrow("HTTP 503");
// ...but not under `isSafeToReplay`, so it ran exactly once.
expect(calls).toBe(1);
});
});
describe("withRetry backoff", () => {
it("doubles the ceiling on each retry and caps it", async () => {
// `random: () => 1` pins the jitter to the top of its range, which
// makes the ceiling itself observable. The sequence is
// base, base*2, base*4, ... clamped at maxDelayMs — so a long outage
// settles into a steady poll instead of growing to hours.
const { sleep, delays } = recordingSleep();
await expect(
withRetry(() => Promise.reject(new ApiError("HTTP 500", 500)), {
attempts: 6,
baseDelayMs: 100,
maxDelayMs: 250,
sleep,
random: () => 1,
}),
).rejects.toThrow();
expect(delays).toEqual([100, 200, 250, 250, 250]);
});
it("draws each delay uniformly below its ceiling", async () => {
// Full jitter. The exponential value is the maximum wait, not the
// wait itself, so a fleet of clients that failed together does not
// come back in lockstep.
const { sleep, delays } = recordingSleep();
await expect(
withRetry(() => Promise.reject(new ApiError("HTTP 500", 500)), {
attempts: 4,
baseDelayMs: 100,
maxDelayMs: 10_000,
sleep,
random: () => 0.25,
}),
).rejects.toThrow();
expect(delays).toEqual([25, 50, 100]);
});
it("never asks to sleep longer than the cap or less than zero", async () => {
// Whatever `random` returns from its [0, 1) contract, the delay stays
// inside the configured envelope.
const draws = [0, 0.999_999, 0.5, 0.1, 0.9];
let i = 0;
const { sleep, delays } = recordingSleep();
await expect(
withRetry(() => Promise.reject(new ApiError("HTTP 500", 500)), {
attempts: 6,
baseDelayMs: 1000,
maxDelayMs: 2000,
sleep,
random: () => draws[i++]!,
}),
).rejects.toThrow();
expect(delays).toHaveLength(5);
for (const d of delays) {
expect(d).toBeGreaterThanOrEqual(0);
expect(d).toBeLessThanOrEqual(2000);
}
expect(delays[0]).toBe(0);
});
});
describe("retry defaults", () => {
it("ships a bounded, documented default policy", () => {
// These are the numbers the README documents. They are asserted here
// so the README and the code cannot drift apart silently. Four
// attempts at a 500ms base put the three ceilings at 500, 1000 and
// 2000ms, so a file that is going to fail gives up after at most
// three and a half seconds of waiting — which keeps a
// several-thousand-file backup moving past a bad file rather than
// stalling on it. The 10s cap only comes into play for a caller that
// raises the attempt count.
expect(DEFAULT_RETRY_OPTIONS.attempts).toBe(4);
expect(DEFAULT_RETRY_OPTIONS.baseDelayMs).toBe(500);
expect(DEFAULT_RETRY_OPTIONS.maxDelayMs).toBe(10_000);
});
it("fills in only the fields the caller left out", () => {
const resolved = resolveRetryOptions({ attempts: 2 });
expect(resolved.attempts).toBe(2);
expect(resolved.baseDelayMs).toBe(DEFAULT_RETRY_OPTIONS.baseDelayMs);
expect(resolved.maxDelayMs).toBe(DEFAULT_RETRY_OPTIONS.maxDelayMs);
expect(typeof resolved.sleep).toBe("function");
expect(typeof resolved.random).toBe("function");
});
it("resolves to the defaults when given nothing", () => {
expect(resolveRetryOptions()).toEqual(DEFAULT_RETRY_OPTIONS);
expect(resolveRetryOptions({})).toEqual(DEFAULT_RETRY_OPTIONS);
});
it("defaults random to a real generator in [0, 1)", () => {
const { random } = resolveRetryOptions();
for (let i = 0; i < 100; i++) {
const r = random();
expect(r).toBeGreaterThanOrEqual(0);
expect(r).toBeLessThan(1);
}
});
});

View File

@@ -32,6 +32,7 @@ import {
fixMissingThumbnails, fixMissingThumbnails,
} from "../../src/thumbnails.js"; } from "../../src/thumbnails.js";
import type { KeyAttributes } from "../../src/auth/types.js"; import type { KeyAttributes } from "../../src/auth/types.js";
import type { RetryOptions } from "../../src/retry.js";
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
// Mock server with controllable thumbnail behavior // Mock server with controllable thumbnail behavior
@@ -51,7 +52,7 @@ interface ThumbMockState {
filesByCollection: Record<number, Record<string, unknown>[]>; filesByCollection: Record<number, Record<string, unknown>[]>;
fileCiphertexts: Record<number, Uint8Array>; fileCiphertexts: Record<number, Uint8Array>;
fileKeys: Record<number, Uint8Array>; fileKeys: Record<number, Uint8Array>;
thumbnailBehavior: Record<number, "ok" | "empty" | "404">; thumbnailBehavior: Record<number, "ok" | "empty" | "404" | "500">;
// Captures from fix operations // Captures from fix operations
uploadedThumbnails: { uploadedThumbnails: {
fileID: number; fileID: number;
@@ -302,6 +303,9 @@ const buildThumbFetch = (m: ThumbMockState) => {
if (behavior === "empty") { if (behavior === "empty") {
return new Response(new Uint8Array(0), { status: 200 }); return new Response(new Uint8Array(0), { status: 200 });
} }
if (behavior === "500") {
return new Response("Internal Server Error", { status: 500 });
}
return new Response("not found", { status: 404 }); return new Response("not found", { status: 404 });
} }
@@ -345,6 +349,38 @@ const buildThumbFetch = (m: ThumbMockState) => {
}) as typeof globalThis.fetch; }) as typeof globalThis.fetch;
}; };
/**
* A retry policy with the waiting removed. `listMissingThumbnails` walks every
* file in the account, so a transient failure is retried; without an injected
* `sleep` these tests would spend real seconds waiting out backoff.
*/
const noWait: RetryOptions = {
sleep: () => Promise.resolve(),
random: () => 0,
};
/** Wrap a fetch so the tests can count how often one endpoint was hit. */
const countingFetch = (
inner: typeof globalThis.fetch,
match: (url: string) => boolean,
): { fetch: typeof globalThis.fetch; matched: () => number } => {
let matched = 0;
const fake = async (
input: RequestInfo | URL,
init?: RequestInit,
): Promise<Response> => {
const url =
typeof input === "string"
? input
: input instanceof URL
? input.href
: input.url;
if (match(url)) matched++;
return inner(input, init);
};
return { fetch: fake as typeof globalThis.fetch, matched: () => matched };
};
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
// Tests // Tests
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
@@ -378,7 +414,77 @@ describe("listMissingThumbnails", () => {
expect(emptyEntry.collection).toBe("Photos"); expect(emptyEntry.collection).toBe("Photos");
const notFoundEntry = missing.find((m) => m.fileID === 102)!; const notFoundEntry = missing.find((m) => m.fileID === 102)!;
expect(notFoundEntry.reason).toContain("fetch failed"); // A 404 is the server stating the thumbnail is not there. That is the
// only network answer that means "missing", and the reason says so
// rather than the older catch-all "fetch failed" — which used to
// cover a 500 and a dropped connection too.
expect(notFoundEntry.reason).toContain("not found");
});
it("does not report a thumbnail as missing when the server is failing", async () => {
// The distinction that matters for `helper fix-missing-thumbnails`.
// Reporting a file here leads to downloading the original,
// regenerating a thumbnail, and uploading it over a thumbnail that
// was fine all along — because the server was briefly returning 500s.
//
// File 102 serves 500 on every attempt, so the retries are genuinely
// exhausted. It must still not be reported.
const failingMock = await buildThumbMock();
failingMock.thumbnailBehavior[102] = "500";
const counted = countingFetch(
buildThumbFetch(failingMock),
(url) => url.includes("thumbnails.ente.io") && url.includes("102"),
);
const client = await Client.login({
email: TEST_EMAIL,
password: TEST_PASSWORD,
apiOptions: { fetch: counted.fetch, retry: { ...noWait } },
});
const missing = await listMissingThumbnails(client);
// Only the genuinely empty thumbnail is reported.
expect(missing.map((m) => m.fileID)).toEqual([101]);
// And the 500 was retried rather than accepted as an answer: four
// attempts is the library default.
expect(counted.matched()).toBe(4);
});
it("does not report a thumbnail as missing when the connection fails", async () => {
// Same rule for a transport failure, which carries no status at all.
const failingMock = await buildThumbMock();
const inner = buildThumbFetch(failingMock);
let thumbRequests = 0;
const fetch = (async (
input: RequestInfo | URL,
init?: RequestInit,
): Promise<Response> => {
const url =
typeof input === "string"
? input
: input instanceof URL
? input.href
: input.url;
if (url.includes("thumbnails.ente.io") && url.includes("102")) {
thumbRequests++;
throw Object.assign(new Error("socket hang up"), {
code: "ECONNRESET",
});
}
return inner(input, init);
}) as typeof globalThis.fetch;
const client = await Client.login({
email: TEST_EMAIL,
password: TEST_PASSWORD,
apiOptions: { fetch, retry: { ...noWait } },
});
const missing = await listMissingThumbnails(client);
expect(missing.map((m) => m.fileID)).toEqual([101]);
expect(thumbRequests).toBe(4);
}); });
it("deduplicates files seen in multiple collections", async () => { it("deduplicates files seen in multiple collections", async () => {

View File

@@ -5,7 +5,8 @@
"moduleResolution": "NodeNext", "moduleResolution": "NodeNext",
"lib": ["ES2022"], "lib": ["ES2022"],
"outDir": "./dist", "outDir": "./dist",
"rootDir": "./src", "rootDir": ".",
"noEmitOnError": true,
"strict": true, "strict": true,
"noImplicitOverride": true, "noImplicitOverride": true,
"noUncheckedIndexedAccess": true, "noUncheckedIndexedAccess": true,